ON TIME BUSINESS MANAGEMENT A+ WORK, ON TIME, NO PLAGARIZING; ON TIME
Online sentiment analysis in marketing research: a review
Meena Rambocas and Barney G. Pacheco Department of Management Studies, The University of the West Indies,
St. Augustine, Trinidad and Tobago
Abstract Purpose – The explosion of internet-generated content, coupled with methodologies such as sentiment analysis, present exciting opportunities for marketers to generate market intelligence on consumer attitudes and brand opinions. The purpose of this paper is to review the marketing literature on online sentiment analysis and examines the application of sentiment analysis from three main perspectives: the unit of analysis, sampling design and methods used in sentiment detection and statistical analysis. Design/methodology/approach – The paper reviews the prior literature on the application of online sentiment analysis published in marketing journals over the period 2008-2016. Findings – The findings highlight the uniqueness of online sentiment analysis in action-oriented marketing research and examine the technical, practical and ethical challenges faced by researchers. Practical implications – The paper discusses the application of sentiment analysis in marketing research and offers recommendations to address the challenges researchers confront in using this technique. Originality/value – This study provides academics and practitioners with a comprehensive review of the application of online sentiment analysis within the marketing discipline. The paper focuses attention on the limitations surrounding the utilization of this technique and provides suggestions for mitigating these challenges.
Keywords Online marketing, Qualitative research, Quantitative research, Methodology, Text mining, High technology marketing
Paper type Literature review
Introduction In recent years, consumers’ willingness to share consumption experiences online, coupled with the technology to analyze “big data”, offer marketing managers an unprecedented opportunity to collect market intelligence (Erevelles et al., 2016). Through online sentiment analysis, hereafter referred to as sentiment analysis, researchers can systematically extract and classify consumer emotions about products and services expressed in social network discussions and online postings to track brand attitudes and emerging market trends. While sentiment analysis presents tremendous opportunities to interpret a large body of data collected in a naturalistic setting, concerns have been expressed about the technique’s accuracy and practicality (Gonçalves et al., 2013). Moreover, apprehensions over online data volume, fragmented data sources, content bias and user exploitation have exposed the technique to critical scrutiny.
In light of these challenges, it is somewhat surprising that researchers have not devoted more attention to evaluating the overall feasibility of using sentiment analysis as a tool for online marketing research. Our study serves to fill this gap by reviewing the literature on the application of sentiment analysis in the marketing discipline. The review is specific to the literature published in scholarly peer-reviewed marketing journals between 2008 and 2016, which coincides with the technique’s general usage within the marketing discipline.
JRIM 12,2
146
Received 9 May 2017 Revised 19 October 2017 Accepted 14 December 2017
Journal of Research in Interactive Marketing Vol. 12 No. 2, 2018 pp. 146-163 © Emerald Publishing Limited 2040-7122 DOI 10.1108/JRIM-05-2017-0030
The current issue and full text archive of this journal is available on Emerald Insight at: www.emeraldinsight.com/2040-7122.htm
The current study makes two unique contributions to the field of marketing research in an interactive environment. First, it is one of the few papers to review the application of sentiment analysis in marketing research comprehensively. Second, the paper focuses attention on the limitations surrounding the utilization of this technique for marketing research and provides suggestions for more effective use.
Overview of sentiment analysis While reviewing the literature, it is apparent that a misunderstanding often exists about what constitutes sentiment analysis. To provide conceptual clarity, sentiment analysis first needs to be distinguished from the broader literature on online text mining. With text-mining applications, researchers structure a large body of data from various online sources into numerous topics or themes which emerge from the body of textual data. In this regard, text mining is similar to traditional content analysis, since it allows researchers to efficiently extract, classify and manage a large body of data to identify hidden patterns or trends (He et al., 2013).
In contrast, sentiment analysis refers to the application of machine learning techniques to evaluate and classify attitudes and opinions on a specific topic of interest (Rambocas and Gama, 2013). Sentiment analysis focuses on extracting emotions from the online text but classifies specific problem areas into predefined mutually exclusive categories (Liu, 2012). These categories imply bi-polar classifications of emotions (positive and negative) and are typically represented by numeric codes for subsequent statistical analyses.
A major advantage of sentiment analysis is that it collects and analyzes online comments in real time. This is especially significant to researchers, given the exponential updating of user-generated content on social media platforms. With sentiment analysis, researchers can also automatically extract high-quality data on emotional expressions that are measurable, objective and consistent. Given these advantages, it is no surprise that sentiment analysis has attracted the interest of both academics and practitioners. However, previous research on this analytical tool has primarily focused on the technique’s methodological properties by identifying efficient methods for extracting textual content from a body of online data via natural language processing, computational linguistics and text analytics (Günther and Furrer, 2013). From a marketing research perspective, however, extracting and classifying online text through sentiment analysis remains a relatively new area with burgeoning potential.
The scope and approach of the review The objective of this review is to examine the use of sentiment analysis in the marketing literature published between 2008 and 2016. The studies identified in this paper were sourced using a combination of computerized and manual search methods. We first surveyed several online scientific databases including Business Source Complete, Proquest and Emerald, and conducted an issue by issue search of the top-ranked marketing journals. We also searched for articles using Google Scholar with the search term “sentiment analysis”. Finally, we used a snowballing procedure where the references of each article on sentiment analysis were examined to identify additional studies, a search technique consistent with Babi�c Rosario et al. (2016). This broad search yielded a total of 21,456 articles.
Next, the title, abstract and keywords of each article were examined to determine whether it was relevant to the application of online sentiment analysis or simply contained keywords such as “sentiments” or “emotions”. Articles not related to the application of online sentiment analysis, defined as “uses natural language processing, computational
Online sentiment
analysis
147
linguistics and text analytics to identify and extract the content of interest from a body of textual data” (Rambocas and Gama, 2013, p. 4), were excluded from the data set. Articles in conference papers and non-peer-reviewed journals were also eliminated. This resulted in a reduced list of 257 articles.
As Table I indicates, these articles were published across disciplines in nine subject matters, namely, Communications, Computer, Education, Engineering, Finance, Health, Marketing, Political Science and General literature review. A comparison over the review period 2008-2016 shows the preponderance of sentiment analysis articles were concentrated in the computer-related discipline (72 per cent of published work). The remaining 28 per cent of publications were distributed across several other disciplines, with Marketing accounting for 22 articles or roughly 9 per cent of this total. Figure 1 shows the publication trends over the review period.
Characteristics of the marketing articles reviewed The authors selected the 22 marketing articles and evaluated them on four criteria:
(1) utilized sentiment analysis and not text mining or other social media analytics; (2) applied sentiment analysis in the study of marketing-related issue/s from the
perspective of the consumer, business or both; (3) published in a peer-reviewed academic journal; and (4) empirical in nature, with a large body of data and utilized statistical tests for data
analysis.
Non-English publications were not included in this review. Only 12 articles met all four criteria (Table II), which is perhaps indicative of how recent the technique is within the marketing discipline. Many of the studies included in our sample were published in relatively high-ranking marketing journals such as Journal of Marketing, Journal of Marketing Research, Marketing Science and Journal of the Academy of Marketing Science. Of the 12 articles reviewed, the majority of studies were led by researchers from the USA (five), followed by Germany (two), The Netherlands (one), England (one), South Korea (one), Taiwan (one) and Denmark (one). The relatively low number of sentiment analysis articles published in marketing journals (average 1.33 per year) sharply contrasts with the emergence of the internet as an important tool for both consumers and marketers and is much lower than expected. We suspect that this trend might be attributable to the challenges in using the technique by researchers in the marketing discipline.
Table I. Comparison of published work on sentiment analysis from 2008 to 2016
Discipline Frequency (%)
Communication 26 10.12 Computer 185 71.98 Education 1 0.39 Engineering 2 0.78 Finance 5 1.95 Health 3 1.17 Marketing 22 8.56 Political Science 1 0.39 Review 12 4.67 Total 257 100
JRIM 12,2
148
Figure 1. Sentiment analysis
publication trend by discipline from 2008
to 2016
Table II. Summary of the
sentiment analysis publications published in
marketing journals
Journal/year 2008 2009 2010 2011 2012 2013 2014 2015 2016 Grand
total
Journal of Marketing 1 1 2 Journal of Marketing Research 1 1 2 Marketing Science 1 1 2 International Journal of Electronic commerce 1 1 2 Academy of Marketing Studies Journal 1 1 Academy of Marketing Science 1 1 Corporate Communications: An International Journal 1 1 The Journal of Consumer Affairs 1 1 Total 0 0 0 1 3 2 2 3 1 12
Online sentiment
analysis
149
Categorization of the articles The contents of each article were categorized into three predetermined themes, which form the basis for subsequent discussions. The themes followed the categorization scheme adopted by Sousa et al. (2008) and sought to summarize the research design of the articles reviewed. These themes are:
(1) unit of analysis; (2) sampling design (size, venue and sample characteristics) and (3) methods used in sentiment detection and statistical analysis.
Following the methods used in Cummins et al. (2013) and Rodriguez et al. (2014), we prepared a data file that listed each article’s title, authors, journal, volume, issue, abstract, unit of analysis, sampling characteristics, data venues, sample size, extraction methods, statistical analysis, challenges and key findings. This technique allowed us to summarize the information into a smaller set of categories and identify patterns across the articles reviewed. The authors examined and discussed each category and then proceeded with independent coding. Only minor discrepancies in coding were discovered, which were resolved through discussions. The properties of the 12 articles are presented in Table III.
Unit of analysis The majority of studies (nine) focused on comments generated by individual users who purchased goods and services for personal and household consumption. The textual content focal object of interest related to products within the consumer market. These studies surveyed a broad range of tangible categories including books, cars, personal computers, phones, footwear, toys, clothing and accessories and personal care items. Service-related industries were also evident with analyses conducted on movies, telecommunications, mobile applications, music downloads, video games, hotels, airlines and retail stores. Two studies focused on business-related cases. The first investigated the accuracy of automated Twitter sentiment coding by third-party research companies and the second evaluated sentiments toward an organization’s corporate communication strategies. Finally, there was one study that used sentiment analysis to examine boycott messages and quantify the intensity of emotions expressed.
Our review revealed that the majority of the studies focused on user-generated comments within a single industry or product category. For instance, Schweidel and Moe (2014) evaluated two leading brands in a single sector – the enterprise and development software sector. Likewise, Ludwig et al. (2013) extracted user-generated content on books from the fiction and non-fiction categories such as academia, religion, philosophy, humor, mystery and crime, romance, horror, science fictions, among others. Similarly, Tang et al. (2014) surveyed the automotive industry and extracted user-generated comments on 39 brands, while Hennig-Thurau et al. (2015) focused on the movie industry. Focusing on a single product category may be viewed as a pragmatic decision by researchers, given the sheer volume of user-generated content on social media.
There were, however, a few notable exceptions to the single-category approach among the studies included in this review. Homburg, Ehm and Artz (2015), for instance, extracted reviews from both a firm’s sponsored online community of do-it-yourself individuals and online travel-related forums. Tirunillai and Tellis (2012), on the other hand, analyzed an extensive range of reviews on personal computing, phones, personal digital assistants, footwear, toys and data storage. Similarly, Baek et al. (2012) collected reviews on 28 product- related categories, and Sonnier et al. (2011) collected comments from a technology firm who
JRIM 12,2
150
Jo ur
na l
T itl
e of
a rt
ic le
A
ut ho
rs
D at
a so
ur ce
Si
ze o
f s am
pl e
W ha
t w as
s en
tim en
t a na
ly si
s us
ed fo
r K
ey fi
nd in
gs
Jo ur
na l o
f M ar
ke tin
g M
or e
th an
w or
ds : T
he
in flu
en ce
o f a
ff ec
tiv e
co nt
en t
an d
lin gu
is tic
s ty
le m
at ch
es
in o
nl in
e re
vi ew
s on
co
nv er
si on
ra te
s
Lu dw
ig e
t a l.
(2 01
3)
A m
az on
.c om
18
,6 82
In
ve st
ig at
ed th
e in
flu en
ce o
f t ex
tu al
pr
op er
tie s
of c
on su
m er
re vi
ew s
on o
nl in
e re
ta ile
r’s c
on ve
rs io
n ra
te s
N on
lin ea
r r el
at io
n be
tw ee
n po
si tiv
e af
fe ct
iv e
co nt
en t a
nd c
on ve
rs io
n ra
te s.
L in
gu is
tic s
ty le
an
d co
nt en
t o f t
he re
vi ew
in flu
en ce
s ch
an ge
s in
co
ns um
er s’
c on
ve rs
io n
be ha
vi or
Jo ur
na l o
f M ar
ke tin
g Is
N eu
tr al
R ea
lly N
eu tr
al ?
T he
E ff
ec ts
o f N
eu tr
al U
se r-
G
en er
at ed
C on
te nt
(U G
C) o
n Pr
od uc
t S al
es
T an
g et
a l.
(2 01
4)
Y ou
tu be
a nd
F ac
eb oo
k 59
2, 31
0 Sp
ec ify
th e
pe rf
or m
an ce
im pl
ic at
io ns
o f
ne ut
ra l U
G C
on p
ro du
ct s
al es
b y
di ff
er en
tia tin
g m
ix ed
-n eu
tr al
U G
C,
w hi
ch c
on ta
in s
an e
qu al
a m
ou nt
o f
po si
tiv e
an d
ne ga
tiv e
cl ai
m s,
fr om
in
di ff
er en
t-n eu
tr al
U G
C, w
hi ch
in cl
ud es
ne
ith er
p os
iti ve
n or
n eg
at iv
e cl
ai m
s
Po si
tiv e
an d
ne ga
tiv e
U G
C pr
ov id
e op
po rt
un iti
es fo
r c on
su m
er s
to p
ro ce
ss
pr od
uc t-r
el at
ed in
fo rm
at io
n, w
he re
as b
ot h
m ix
ed - a
nd in
di ff
er en
t-n eu
tr al
U G
C af
fe ct
co
ns um
er s’
m ot
iv at
io n
an d
ab ili
ty to
p ro
ce ss
po
si tiv
e an
d ne
ga tiv
e U
G C.
U
G C
am pl
ifi es
th e
ef fe
ct s
of p
os iti
ve a
nd
ne ga
tiv e
U G
C, w
he re
as in
di ff
er en
t-n eu
tr al
U
G C
at te
nu at
es th
em . T
he e
ff ec
ts o
f n eu
tr al
U
G C
on p
ro du
ct s
al es
th us
a re
n ot
tr ul
y ne
ut ra
l, an
d th
e di
re ct
io n
of th
e bi
as d
ep en
ds
on b
ot h
th e
ty pe
o f U
G C
an d
th e
di st
ri bu
tio n
of
po si
tiv e
an d
ne ga
tiv e
U G
C Jo
ur na
l o f M
ar ke
tin g
R es
ea rc
h Li
st en
in g
in o
n so
ci al
m ed
ia :
A jo
in t m
od el
o f s
en tim
en t
an d
ve nu
e fo
rm at
c ho
ic e
Sc hw
ei de
l a nd
M
oe (2
01 4)
Co
nv er
se on
e xt
ra ct
ed
da ta
fr om
b lo
gs ,
pr od
uc t r
ev ie
w s,
di
sc us
si on
s tr
ea m
s,
so ci
al n
et w
or k
si te
s an
d m
ic ro
b lo
gs
7, 56
5 M
od el
ed th
e se
nt im
en t e
xp re
ss ed
in u
se r-
ge
ne ra
te d
co m
m en
ts a
cr os
s ve
nu es
Si
gn ifi
ca nt
v ar
ia tio
ns in
s en
tim en
ts e
xp re
ss ed
ac
ro ss
v en
ue s
Jo ur
na l o
f M ar
ke tin
g R
es ea
rc h
M ea
su ri
ng a
nd m
an ag
in g
co ns
um er
s en
tim en
t i n
an
on lin
e co
m m
un ity
en
vi ro
nm en
t
H om
bu rg
e t a
l. (2
01 5)
O
nl in
e fo
ru m
s 11
5, 00
0 E
xa m
in ed
h ow
c on
su m
er s
re ac
t t o
fir m
’s
ac tiv
e pa
rt ic
ip at
io n
in c
on su
m er
-to -
co ns
um er
c on
ve rs
at io
ns in
a n
on lin
e co
m m
un ity
s et
tin g
R ea
ct io
ns w
er e
si gn
ifi ca
nt ly
d iff
er en
t a nd
ba
se d
on b
ot h
cu st
om er
s fu
nc tio
na l a
nd . s
oc ia
l ne
ed s.
C us
to m
er s
ex pr
es se
d di
m in
is hi
ng
re tu
rn s
as fi
rm g
en er
at ed
in fo
rm at
io n
in cr
ea se
s M
ar ke
tin g
Sc ie
nc e
A d
yn am
ic m
od el
o f t
he
ef fe
ct o
f o nl
in e
co m
m un
ic at
io ns
o n
fir m
sa
le s
So nn
ie r e
t a l.
(2 01
1)
Pr op
ri et
ar y
W eb
cr aw
le r
te ch
no lo
gy
A ve
ra ge
2 ,5
72
pe r d
ay
In ve
st ig
at e
th e
ef fe
ct o
f t he
v ol
um e
of
po si
tiv e,
n eg
at iv
e an
d ne
ut ra
l o nl
in e
co m
m en
ts o
n da
ily s
al es
p er
fo rm
an ce
Si gn
ifi ca
nt e
ff ec
t o f p
os iti
ve , n
eg at
iv e
an d
ne ut
ra l o
nl in
e co
m m
un ic
at io
ns o
n da
ily s
al es
pe
rf or
m an
ce
(c on
tin ue
d)
Table III. Characteristics of the
published work on sentiment analysis
0
Online sentiment
analysis
151
Jo ur
na l
T itl
e of
a rt
ic le
A
ut ho
rs
D at
a so
ur ce
Si
ze o
f s am
pl e
W ha
t w as
s en
tim en
t a na
ly si
s us
ed fo
r K
ey fi
nd in
gs
M ar
ke tin
g Sc
ie nc
e D
oe s
ch at
te r r
ea lly
m at
te r?
D
yn am
ic s
of u
se r-
ge ne
ra te
d co
nt en
t a nd
s to
ck
pe rf
or m
an ce
T ir
un ill
ai a
nd
T el
lis (2
01 2)
am
az on
.c om
; E pi
ni on
s.
co m
; Y ah
oo S
ho pp
in g
(y ah
oo .c
om )
In ve
st ig
at e
w he
th er
th er
e is
a
re la
tio ns
hi p
be tw
ee n
ch at
te r a
nd th
e st
oc k
m ar
ke t p
er fo
rm an
ce o
f t he
fi rm
?
Ch at
te r v
ol um
e ha
s th
e st
ro ng
es t r
el at
io ns
hi p
w ith
re tu
rn s
an d
tr ad
in g
vo lu
m e.
T hi
s is
fo
llo w
ed b
y ne
ga tiv
e ch
at te
r w hi
ch in
cr ea
se s
vo la
til ity
(r is
k) in
re tu
rn s
In te
rn at
io na
l J ou
rn al
o f
E le
ct ro
ni c
co m
m er
ce
H el
pf ul
ne ss
o f o
nl in
e co
ns um
er re
vi ew
s: re
ad er
s’
ob je
ct iv
es a
nd re
vi ew
c ue
s
B ae
k et
a l.
(2 01
2)
A m
az on
.c om
75
,2 26
In
ve st
ig at
e th
e fa
ct or
s th
at d
et er
m in
e th
e he
lp fu
ln es
s of
o nl
in e
re vi
ew s
an d
w hi
ch
fa ct
or s
ar e
m or
e im
po rt
an t f
or h
el pf
ul
on lin
e re
vi ew
s
T he
h el
pf ul
ne ss
o f o
nl in
e re
vi ew
s is
de
te rm
in ed
b y
bo th
p er
ip he
ra l c
ue s
(r ev
ie w
ra
tin g
an d
re vi
ew er
’s c
re di
bi lit
y) a
nd c
en tr
al
cu es
(c on
te nt
o f r
ev ie
w s,
in flu
en ce
th e
he lp
fu ln
es s
of re
vi ew
s)
In te
rn at
io na
l J ou
rn al
o f
E le
ct ro
ni c
co m
m er
ce
W ha
t i n
co ns
um er
re vi
ew s
af fe
ct s
th e
sa le
s of
m ob
ile
ap ps
: A M
ul tif
ac et
S en
tim en
t A
na ly
si s
A pp
ro ac
h
Li an
g et
a l.
(2 01
5)
Se ve
nt y-
ni ne
p ai
d an
d se
ve nt
y fr
ee a
pp s
fr om
an
iO S
ap p
st or
e
79 p
ai d
an d
70
fr ee
a pp
s E
xa m
in e
th e
ef fe
ct o
f t ex
tu al
c on
su m
er
re vi
ew s
on th
e sa
le s
of m
ob ile
a pp
s T
he s
tu dy
fo un
d th
at a
lth ou
gh c
on su
m er
s’
op in
io ns
o n
pr od
uc t q
ua lit
y oc
cu pi
es a
la rg
er
po rt
io n
of c
on su
m er
re vi
ew s,
th ei
r c om
m en
ts
on s
er vi
ce q
ua lit
y ha
ve a
s tr
on ge
r u ni
t e ff
ec t
on s
al es
ra nk
in gs
A
ca de
m y
of M
ar ke
tin g
St ud
ie s
Jo ur
na l
A ss
es si
ng th
e ac
cu ra
cy o
f au
to m
at ed
T w
itt er
s en
tim en
t co
di ng
D av
is a
nd
O ’F
la he
rt y
(2 01
2)
T w
itt er
76
7 E
va lu
at ed
th e
au to
m at
ed s
en tim
en t
co di
ng a
cc ur
ac y
an d
m is
cl as
si fic
at io
n er
ro rs
o f s
ix le
ad in
g th
ir d-
pa rt
y co
m pa
ni es
T he
s tu
dy s
ho w
ed th
at a
ut om
at ed
s en
tim en
t co
di ng
h ad
li m
ite d
re lia
bi lit
y an
d ap
pe ar
s w
ith
lim its
fo r s
im pl
e st
at em
en ts
w he
re a
k ey
w or
d is
u se
d Jo
ur na
l o f t
he A
ca de
m y
of M
ar ke
tin g
Sc ie
nc e
D oe
s T
w itt
er m
at te
r? T
he
im pa
ct o
f m ic
ro bl
og gi
ng
w or
d of
m ou
th o
n co
ns um
er s’
a do
pt io
n of
n ew
m
ov ie
s
H en
ni g-
T hu
ra u
et a
l. (2
01 5)
T
w itt
er
4, 04
5, 35
0 tw
ee ts
a bo
ut
th e
10 5
m ov
ie s
E m
pi ri
ca lly
te st
th e
im pa
ct o
f m
ic ro
bl og
gi ng
w or
d- of
-m ou
th (M
W O
M )
on c
on su
m er
s ad
op tio
n of
n ew
m ov
ie s
N eg
at iv
e M
W O
M re
vi ew
s im
pa ct
o n
ea rl
y ad
op tio
n bu
t n ot
p os
iti ve
M W
O M
re vi
ew s,
w
hi ch
in di
ca te
s a
ne ga
tiv ity
b ia
s fo
r M W
O M
C or
po ra
te
C om
m un
ic at
io ns
: A n
In te
rn at
io na
l J ou
rn al
CS R
c om
m un
ic at
io n
st ra
te gi
es fo
r o rg
an iz
at io
na l
le gi
tim ac
y in
s oc
ia l m
ed ia
Co lle
on i (
20 13
) T
w itt
er
N et
w or
k of
9,
58 9
us er
s an
d 32
6, 00
0 CS
R -r
el at
ed
tw ee
ts
In ve
st ig
at e
th e
ef fe
ct iv
en es
s of
w hi
ch
on lin
e so
ci al
m ed
ia c
or po
ra te
co
m m
un ic
at io
n st
ra te
gi es
is m
or e
ef fe
ct iv
e to
c re
at e
co nv
er ge
nc e
be tw
ee n
co rp
or at
io ns
’ c or
po ra
te s
oc ia
l re
sp on
si bi
lit y
(C SR
) a ge
nd a
an d
st ak
eh ol
de rs
’ s oc
ia l e
xp ec
ta tio
ns
N ei
th er
th e
en ga
gi ng
n or
th e
in fo
rm at
io n
st ra
te gi
es le
ad s
to a
lig nm
en t.
T he
fi nd
in gs
s ho
w th
at , e
ve n
w he
n en
ga gi
ng
in a
d ia
lo gu
e, c
om m
un ic
at io
n in
s oc
ia l m
ed ia
is
st ill
c on
ce iv
ed a
s a
m ar
ke tin
g pr
ac tic
e to
co
nv ey
m es
sa ge
s ab
ou t c
om pa
ni es
T he
Jo ur
na l o
f C
on su
m er
A ff
ai rs
Co
ns um
er b
oy co
tt b
eh av
io r:
A n
ex pl
or at
or y
an al
ys is
o f
T w
itt er
fe ed
s
M ak
ar em
a nd
Ja e
(2 01
6)
T w
itt er
E
xt ra
ct ed
14
,6 85
a nd
ra
nd om
ly
se le
ct ed
2 ,0
00
E xa
m in
e an
d qu
an tif
y th
e em
ot io
na l
in te
ns ity
o f b
oy co
tt m
es sa
ge s
Co ns
um er
b oy
co tt
m es
sa ge
s ar
e m
or e
co m
m on
ly m
ot iv
at ed
b y
in st
ru m
en ta
l m ot
iv es
. H
ow ev
er , n
on -in
st ru
m en
ta l m
ot iv
es h
av e
hi gh
er e
m ot
io na
l i nt
en si
ty
Table III. 0
JRIM 12,2
152
sold a variety of goods on the online market. These studies adopted a wider approach to the analysis and broadened the scope of the analysis to include the impact of product type on the nature of customer sentiments.
In relation to the reasons for selecting a specific product category, the majority of authors based their arguments on the specificity of the research context. For instance, Tang et al. (2014) indicated that the automobile industry contributed significantly to the US economy. Also, the US automobile sector has invested a considerable amount of resources into social media marketing aimed at building long-term relationships with customers and promoting sales initiatives. A similar rationale was advanced by Hennig-Thurau et al. (2015) who selected movies as the product of interest. On the other hand, Sonnier et al. (2011) explained that their selection was based on the firm’s unique interest in social media monitoring and data collection. Additionally, Liang et al. (2015) justified the significance of studying mobile apps by highlighting the rapid development of smartphones and business opportunities for mobile applications.
However, some studies were more purposive in their approach and matched the selection strategy with the research designs. For instance, Baek et al. (2012) drew on the scholarly contribution of Nelson (1974) to classify 28 products as either search or experience goods. Davis and O’Flaherty (2012), on the other hand, used textual data from Twitter to assess the accuracy of automated coding and misclassification errors for positive, negative and neutral comments about a fictitious product (beer). The purposive approach allowed the researchers flexibility and control over the tweets extracted.
Sampling design Sentiment analysis provides marketers with the opportunity to collect a vast amount of textual content from large samples of participants, an opportunity fully capitalized on by the 12 papers reviewed. More specifically, the data set of the 12 articles ranged from 767 to 4,045,350 cases. The median number of cases analyzed with sentiment analysis was 66,841. Apart from reducing sampling error, these large samples give analysts the options to use sophisticated statistical data analysis techniques.
Despite the capacity of sentiment analysis to extract user-generated content from multiple online sources, our review revealed that the majority of studies utilized a single- venue approach to extract text, with Amazon and Twitter being the most popular. Amazon is reportedly the preferred venue for data collection because of the easy accessibility of information and the high frequency of updates by the retailer. The unique advantages of amazon.com are discussed by Ludwig et al. (2013) who identified the site’s unique ability to trace customers reviews and consumers conversion, which facilitates insight into the impact of consumer sentiments on actual purchase behavior. The information generated from the site also allowed the authors to control for the effects of extraneous variables such as price, review volume and review helpfulness.
Focusing on the automobile industry, Tang et al. (2014) investigated the performance implications of neutral user-generated extracted content on product sales data from YouTube and Facebook. The authors drew on the Stelzner (2012) social media marketing report, which confirmed that 92 per cent of the automobile companies use Facebook and 57 per cent use YouTube to share information with users, to justify the relevance of the sampling strategy employed. Also, Facebook and YouTube are two of the largest social media and video sharing sites with a significant amount of user-generated content.
Unlike previous studies, Sonnier et al. (2011) used proprietary Web Crawler technology to generate a daily log of online comments relevant to the company of interest and its products and services. This approach allowed the authors to canvass more internet platforms for a
Online sentiment
analysis
153
wider spectrum of reviews. Likewise, Schweidel and Moe (2014) acknowledged the impact of venue on social media monitoring. Using a social media data set provided by a leading online social media listening platform (Converseon), the authors extracted 7,565 comments across a wide sample of website domains and demonstrated how sentiments from different social media venues could vary in both strength and subject matter.
Methods used in sentiment detection and statistical analysis Sentiment detection requires appraising and extracting only the emotionally laden content such as personal expressions, opinions and feelings from the textual data set. The studies reviewed in this paper employed either manual or automated detection mechanisms. Manual sentiment detection requires human input into the analysis and has the advantage of accommodating emoticons, abbreviations, sarcasm and slangs. Makarem and Jae (2016) note that emotional intensity may be detected from peripheral cues hidden in messages such as message length (long messages are typically associated with greater consumer engagements and emotions); capital letter words (an expression for shouting) and the number of profanities or insult words that may be undetected by automatic detection.
Manual coding can also accommodate language idiosyncrasies. For instance, Liang et al. (2015) noted that the Part-of-Speech system in Taiwan is different from mainland China and English where many words that would be considered adjectives in both languages would be transitive verbs laden with sentiments. Although advantageous in many ways, manual coding can reflect individual subjectivity, bias and misinterpretations. Also, manual coding is very time-consuming and costly, with researchers spending several weeks coding and processing data into categories (Makarem and Jae, 2016).
Conversely, automated sentiment analysis is more commonly used by researchers. The method relies on algorithms to process the volume of text. For instance, Tang et al. (2014) adopted a dictionary of affective words from SentiStrength2, while Ludwig et al. (2013) used the Linguistic Inquiry and Word Count program (LIWC), which calculates the proportion of words in text reviews that match predefined dictionaries. LIWC is flexible in nature and facilitates the analysis of files in multiple languages quickly and efficiently (Scholand et al., 2010). These programs, however, require training documents or data corpus, which can be generated from the original data set.
Regarding statistical analysis, the majority of studies complemented sentiment analysis with econometric linear and non-linear modeling, panel data modeling and time series analysis. However, while linear regression is the most popular analytical approach, other multivariate analysis techniques such as multiple discriminant analysis and structural equation modeling are notably absent.
The application of sentiment analysis in marketing research In reviewing the application of sentiment analysis, we found the majority of articles focused on quantifying the effect of online textual comments on corporate financial performance as measured by sales, preferential consumer behavior and corporate stock performance. In almost all cases, the research models were causal and driven by a strong theoretical underpinning.
For instance, Sonnier et al. (2011) considered the effect of the volume of positive, negative and neutral user-generated comments on sales. The authors’ study was among the first in the marketing literature to model the dynamic effects of online communication and found positive feedback has the greatest effect on sales followed by negative and neutral comments. In an extension of that work, Tang et al. (2014) showed that mixed-neutral comments intensify the impact of positive and negative comments while indifferent-neutral comments diminish the
JRIM 12,2
154
effect. The findings suggest that mixed-neutral comments are associated with credibility and honesty and aid in consumers deliberate processing of product-related information.
Additional evidence of the influence of online comments on sales of a specific product is provided by Liang et al. (2015) who used sentiment analysis to categorize customer feedback and model the influence of electronic word of mouth on the sale of mobile phone applications. The authors concluded that sentiments on overall product’s quality and service attributes would effectively predict overall sales but comments on service have a stronger effect.
Drawing on human communication theories, Ludwig et al. (2013) used dynamic panel data modeling to dissect text reviews into extreme positive and negative changes and advanced the notion that affect and communication style can increase consumers’ product conversion rates via better rapport, credibility and shared perceptions among online users. This effect existed even after controlling for customers rating, weekly review quantity, affective content variation, price discounts and reviewer expertise. Similarly, Hennig- Thurau et al. (2015) found that unlike positive reviews, negative reviews impacted on consumers’ early adoption of movies – a relationship rooted in the theory of information diagnosticity and prospect theory. The authors also included a series of control variables, namely, movie hype, production budget and studio.
In a departure from the focus on sales as the primary variable of interest, Makarem and Jae (2016) used sentiment analysis in conjunction with content analysis to show that boycott messages driven by non-instrumental motives have higher emotional intensity. Additionally, Tirunillai and Tellis (2012) found that negative product reviews and ratings (online chatter) increased volatility in returns and significantly impacted on traded volume. Finally, in a related stream of research, Schweidel and Moe (2014) provided evidence that the effect of online sentiments on a brand’s stock market performance may vary depending on the social media venue where the comments are posted.
In summary, most of the articles reviewed used rigorous, comprehensive research designs that were primarily driven by pre-defined research frameworks. In this regard, sentiment analysis was primarily used in explanatory or causal research designs rather than in descriptive or exploratory studies.
Challenges of applying sentiment analysis in marketing research Despite the technique’s potential benefits, it is still relatively new to the marketing discipline with evolving methodological properties. Given the methodology’s embryonic stage, questions on its relevance and applicability to marketing research are at the forefront, and there is growing recognition of the various application challenges. This section groups these challenges into three broad categories (technical, practical and ethical) and discusses the implications of each.
Technical limitations Liu (2012) described the technical limitations of sentiment analysis as “multifaceted” based on object identification, feature extraction and opinion grouping. In object identification and feature extraction, sentiment analysis only extracts and classifies comments related to the study’s focal object. However, the views expressed in textual logs often refer to many different issues, sometimes having an indirect link to the problem at hand. These are typically excluded in the analysis which compromises the accuracy and validity of the sentiment analysis. The CEO and President of Beyond the Arc, Steven Ramirez, attributed these misclassifications to the early development of machine learning as well as the “cultural factors, linguistic nuances and differing contexts which make it extremely difficult to turn a
Online sentiment
analysis
155
string of written text into a simple pro and con statement” (Mullich, 2012). However, the classification hurdles may be more pronounced in some product categories than others. Venkat Viswanathan – the CEO and founder of LatentView (a data analytics company that works with Fortune 500 companies) – suggested that the classification accuracy is higher among consumer electronics products, primarily because their distinct features and relatively limited number of main features simplifies classification.
Finally, technical limitations relating to opinion groupings refer to lower accuracy and high response errors in classifying textual data. Online texts are notorious for unstructured textual content laden with grammatical and spelling errors, which make sentiment detection more difficult. The problems are exacerbated by the use of multiple phrases, nouns or linguistic variances (slangs) to describe features and attributes. Even the context of the text should be taken into account since it can affect the accuracy of classification. Conversations are typically domain-specific, and opinions can be misclassified from positive to negative depending on the conversation context. A detailed discussion of errors in classifications can be found in the study conducted by Davis and O’Flaherty (2012).
Criticisms are also rooted in the classification approach given that the technique classifies data into mutually exclusive groups. While the limited groups are simple and easy to present, the classification ignores the diversity and richness of the online comments in addition to excluding data that straddle multiple categories or are less readily assigned and thus not captured by the analysis. Davis and O’Flaherty (2012) report that statements without keywords, or statements in which keywords are reversed through negation or context, are more likely to be miscoded. Also, neutral statements are problematic for coders and brand managers are urged to independently verify the level of accuracy across a broad range of statement types.
Researchers also face challenges in assessing brand sentiments across social media venues, given that media platforms appeal to different audiences and have different usage patterns. Very often, the data are single sourced (drawn from a single type of site), which fails to consider the diversity of media appeal, interest and attention (Ordenes et al., 2017). For instance, sites such as Yelp.com and Epinons.com host reviews on consumer-related product and services including restaurants, nightlife, shopping and even medical providers. Other sites like Twitter and Facebook are more generic and host a broad range of data on an unlimited number of subject matters. Additionally, in their investigations of sentiments toward brands, Schweidel and Moe (2014) found that blogs have the highest percentage of positive comments, while forums have the highest percentage of negative comments. Moreover, the authors noted that the quantity and nature of discussions on blogs and microblogs exhibit temporal inconsistency, while forums tend to be more consistent over time.
The demographic profile of users on the various social media platforms also differs. For instance, Instagram is particularly appealing to a younger (18-29 years) African American and Latino population, while LinkedIn usage rates are higher in a more mature population who have graduated from college and Pinterest appeals to middle-class professional women between the ages 25 and 34 (Duggan, 2015). These differences are likely to influence the nature of the comments found on these sites and thus the sentiments that are extracted. This challenge is especially significant given that most marketing studies using sentiment analysis have extracted data from a single online source and those that draw data from multiple sources usually do so in a fragmented manner. Even if data are extracted from multiple sites, sampling bias may still exist given that most researchers usually consolidate responses across platforms and ignore the response variances based on where the data come from (Schweidel and Moe 2014). Therefore, the efficiency gains from gathering and
JRIM 12,2
156
classifying information from a single internet source often masks population differences across diverse media platforms that can challenge the validity of the research findings.
Practical limitations The issues of research cost and accuracy are real concerns for market researchers. Organizations planning to use sentiment analysis either acquire expensive software and infrastructure or pay significant consulting fees to external companies. This could be a disincentive to using this technique for smaller firms who cannot afford expensive third- party services. In addition to the high research cost, the burgeoning volume of online data creates administration and data storage challenges. Moreover, although the abundance of user-generated content can provide useful insights, they are usually laden with short and irregular phrases, which hinder fast and effective sentiment classification (Saif et al., 2012).
Additional challenges in using sentiment analysis in real-life applications include problems with coding text using automatic coding systems (Davis and O’Flaherty, 2012). The authors noted that the misclassification is especially important when lengthy sentences are extracted or when sentences do not contain keywords. Furthermore, misclassification and errors in coding are also common when neutral opinions are expressed. Misclassification therefore presents difficulties in measuring and tracking brand-related sentiments and may distort a study’s conclusions.
In addition to misclassifications and inaccuracies, the evidence points to variations in both textual style and content based on gender (Thelwall et al., 2010; Otterbacher, 2013). Boldon and Carter (2013) also described how age differences can bias the findings of sentiment analysis, given that older customers are less likely to share information on social media as opposed to a younger more digital-centered group. However, despite these potential confounds, it is clear from the articles reviewed that authors have often not considered the influence of demographics on their analysis and findings.
The dramatic increase in the number of deceptive or fake reviews has also emerged as another practical issue that challenges the trustworthiness of sentiment analysis (Heydari et al., 2015). Deceptive reviews are created sinister motives designed to mislead businesses and consumers. To the extent that the opinions expressed publicly on websites and other social media websites do not represent authentic sentiments, they threaten the accuracy of the conclusions derived from sentiment analysis.
The inefficacy of sentiment analysis in cross-cultural research is another reality that researchers confront when using sentiment analysis as a research tool. It is instructive, for instance, that all the articles in our review, with the notable exception of Liang et al. (2015), relied on English language text for their analysis. This imbalance may perhaps be attributed to the widely available lexicons in English, most of which offer strong and reliable convergence between human and automatic coding (Ludwig et al., 2013). Additionally, most of the tools and classification documents used are in English, which impedes analysts’ abilities in building multi-language classifiers and conducting cross- cultural analyses (Liu, 2012).
Although improvements in automated technology make language translation easier, automatic translators can be unreliable. Research shows that automated translation often fails to identify and adapt to varying cultural idiosyncrasies and can compromise the validity of the results (Lotz and Van Rensburg, 2016). Specifically, automated translators fail to consider cultural differences in expressions and can overlook subtle nuances in the way opinions are communicated. These inefficiencies widen the gap between actual meaning and translated meaning and may cause misclassifications. In this regard, researchers should exercise caution in using automated translation when collecting opinions on foreign
Online sentiment
analysis
157
customers or conducting cross-cultural comparisons based solely on online linguistic expressions.
Ethical limitations Although there is ample literature available on ethical standards for online research, sentiment analysis raises several fundamental ethical concerns, which to a large extent have been under-reported in the articles reviewed. One main concern hinges on the right to user privacy since the technology enables researchers to surreptitiously collate a comprehensive set of personal and confidential information about an individual through their online activity that may not have otherwise been made public (Nunan and Di Domenico, 2013). Although this data are useful to marketers interested in building a more accurate customer profile, questions on commercial exploitation arise, especially if third-party companies contact users via unsolicited direct marketing initiatives. The technique therefore fails to adequately consider users’ right of refusal to participate in or withdraw from market research programs.
Further, the trustworthiness of the companies responsible for data extracting can also be an important issue, given the sensitive nature of personal data, which, if exposed, may cause irreparable damage to consumers (Dumas et al., 2014; Hasan et al., 2013). In this regard, researchers have a duty to maintain the highest ethical standards in their research activities to protect participants from harm and safeguard their privacy wherever possible.
Recommendations for sentiment analysis application in marketing research This review highlighted several constraints of sentiment analysis that can compromise the validity of the information generated. For instance, our findings indicated that researchers often rely on a single venue for data extraction rather than using a combination of sources such as blogs, discussion boards and social network sites. This practice increases the potential for sampling bias to emerge since it does not account for user heterogeneity across social media platforms. Marketers and users of sentiment analysis should therefore view data extracted from these narrow sources with skepticism and insist on extracting data from a wider cross-section of users across multiple venues who display varying demographic and psychographic characteristics. Schweidel and Moe (2014) warn, however, that merely aggregating data from multiple forums will not automatically reduce sampling bias and may continue to yield misleading results, especially if comments are simply combined. It is therefore necessary for future researchers to consider using robust statistical analysis that can either control or explicitly account for the impact of the data source on their results.
In addition to sampling considerations, the high cost of contracting specialized vendors to track brand-specific comments and code consumer sentiments across a wide range of social media platform can be prohibitive, especially for smaller companies. Smaller companies (who are under-staffed and under-resourced) may thus want to consider open- source text-analytics tools that can be used for sentiment analysis (e.g. Watson Natural Language; Python NLTK; RapidMiner). Although many of these online interfaces have restricted functions, it allows users to run common queries on any topic of interest and analyze text-based sentiments.
Sentiment analysis in a multilingual environment also remains a challenge, primarily because of the inadequacy of online language translators and the fact that publicly available lexicons are usually only available in English. While online translators help bridge the language divide, they cannot understand language context and consequently often fail to provide an accurate rendition of the original text. However, improvements in machine learning technology are now showing dramatic improvements in translation. Grimes (2014)
JRIM 12,2
158
suggested that cross-cultural analysis is imminent given the recent advancements in natural language processing, stylistic analysis and profile extraction within the academic and commercial environments. However, to date, the application of such technologies remains sparse. Efforts still need to be devoted to developing training documents and foreign language data corpus which can be used to enhance the applicability of sentiment analysis to a broader range of material.
As a more short-term measure, Gopaldas (2014) suggested that the cultural divide can be bridged by employing new types of skills in the data environment. The author suggested that big data companies like Facebook and Google should consider hiring graduates from clinical psychology and cultural anthropology to complement scientists and statisticians, to absorb multimodal data and recognize linguistic variances (humor, irony and sarcasm) and capture more holistic market sentiments. Also, researchers have started to address the lexicon gaps. For instance, Chen and Skiena (2014) integrated a variety of linguistic resources to produce “high-quality” lexicons for 136 different languages and concurred that similar work is being done in natural language processing tasks with specific dictionaries and seed words. However, the authors admit that more work must be done in the technical areas of learning modifiers, negations terms and sentiment attributions.
Researchers can also address some of the methodological limitations of automated text processing by integrating human analysis to classify sentiments according to the language context as well as interpret the valence of a sentiment from the text. The benefits of human coders in an automated sentiment analysis environment are outlined by Venkat Viswanathan, CEO and founder of LatentView (a data analytics company that works with Fortune 500 companies). According to Viswanathan, “Some topics and conversations are easy to classify, some are complex [. . .] In any case, you always need humans to provide the context”.
The review also presented implications from an academic and research perspective. For instance, incorporating sentiment analysis into marketing doctoral programs as a mainstream methodology may not only increase usage of the technique in academic research but will produce a cadre of trained users who can provide this service at lower cost for marketing practitioners. Expanding the pool of users will also increase competition among service providers which should further reduce costs. The analytical rigor with which the technique is applied also stands to benefit once a critical mass of users is attained, since there will be a wider understanding of the standards which should be used to judge the quality of findings produced by this technique.
Additionally, marketers and researchers should be mindful of the devastating impact of deceptive reviews on the validity of the findings. Consequently, through spam detection techniques (automated or manual), marketers should vigorously identify and isolate these predatory comments from the analysis. Admittedly, these detection methods could be very complex and may require considerable resources to develop and implement. Nevertheless, the purification of content could assist in improving the accuracy of the sentiment analysis and the overall results.
The ethical issues surrounding the execution, analysis and presentation of data remain a contentious area that requires a balance between market research goals and an individual’s right to privacy, especially if the data are being extracted from private sources. Holmes (2009) suggests that posting a message on online forums or social media explaining the nature of the research and inviting volunteers can help minimize the potential damage to some online communities. This ensures that users are reminded of how and when data will be collected as well as the reasons why it is being collected. The risk of following this
Online sentiment
analysis
159
recommendation, however, is that it exposes the sample to self-selection bias, which is a limitation of this approach.
In the absence of specific ethical rules regarding the application of sentiment analysis, academic researchers may consider following the general ethical guidelines outlined by the institutional review board (IRB) in their institution. In relation to privacy, IRB principles dictate that the researcher should ensure that there are sufficient allowances within the research design to protect the privacy of the participants and maintain the confidentiality of the data collected. In the case of sentiment analysis, this can be done through a combination of steps that may include collecting data in anonymous environments, purging identifying information from the data set and restricting the number of personnel with access to the data.
Another critical ethical area for sentiment analysis researchers is informed consent, which guarantees voluntary participation based on participants’ knowledge of the intended study. Lunnay et al. (2015) suggested that social media researchers (including users of sentiment analysis) should give participants the right to participate or not participate before the research starts. In the case of minors (less than 18 years), researchers should make conscientious efforts to identify minors and seek parental consent before extracting data, although this may be difficult, given that there may not be any practical way of discovering the true age of online participants.
Table IV outlines the main limitations we have identified that are experienced by researchers when using sentiment analysis in marketing research and summarizes our suggestions to mitigate their effects.
Conclusions It is clear that the online environment provides rich and valuable information about consumer opinions, though harnessing and analyzing that data can be difficult. Improvements in computer modeling and techniques like sentiment analysis provide
Table IV. A summary of the challenges and recommendations in applying sentiment analysis
Challenges Aspects Recommendations
Technical limitations
Accuracy, reliability and validity
Include a cross-section of product categories in the analysis Extract data from multiple venues that appeal to different consumer demographical characteristics Triangulate results with more traditional research methods
Practical limitations
Cost concerns Free services are available, but the brand managers should be cautioned when relying on these free services for strategic decisions Invest in programs that will produce a cadre of trained users who can provide this service at lower cost for marketing practitioners
Miscoding Integrate human analysis with automated text processing to classify opinions
Cross-cultural research
Cross-cultural analysis is imminent given the recent advancements in natural language processing, stylistic analysis and profile extraction within the academic and commercial environments
Deceptive reviews Use spam detection techniques (automated or manual) to identify and isolate these predatory comments
Ethical concerns
The right to user privacy
IRB ethical guidelines for online research should be followed to preserve the rights of consent, privacy, confidentiality and anonymity, assurances of voluntary participation and protection from harm Special effort should be taken to remove children and other vulnerable groups from the analysis
Exploitation
JRIM 12,2
160
powerful mechanisms by which this information can be converted into deep insights about the attitudes held by a brand’s target market. The recommendations for use provided in the current research are thus the first step to guide marketers and academics who wish to adopt this nascent technology. Hopefully, by integrating these recommendations into their research designs, academics and marketers will be better equipped to produce more enriching, meaningful and rigorous analyses. It is anticipated, therefore, that as sentiment analysis emerges as a powerful tool to understand consumer opinion, the technique will attain the methodological rigor associated with other more widely used analytic techniques and feature more prominently in future marketing research.
References Babi�c Rosario, A., Sotgiu, F., De Valck, K. and Bijmolt, T.H. (2016), “The effect of electronic word of
mouth on sales: a meta-analytic review of platform, product, and metric factors”, Journal of Marketing Research, Vol. 53 No. 3, pp. 297-318.
Baek, H., Ahn, J. and Choi, Y. (2012), “Helpfulness of online consumer reviews: readers’ objectives and review cues”, International Journal of Electronic Commerce, Vol. 17 No. 2, pp. 99-126.
Boldon, R. and Carter, R. (2013), “Lost in translation/managing multi-lingual A/V and metadata in the digital supply chain”, Journal of Digital Media Management, Vol. 1 No. 4, pp. pp. 330-335.
Chen, Y. and Skiena, S. (2014), “Building sentiment lexicons for all major languages”, ACL, Vol. 2, pp. 383-389, available at: http://ai2-s2dfs.s3.amazonaws.com/c5e3/b065e352a93d8754b86 baaf8ec20bf81a5c3.pdf (accessed 20 March 2017).
Colleoni, E. (2013), “CSR communication strategies for organizational legitimacy in social media”, Corporate Communications: An International Journal, Vol. 18 No. 2, pp. 228-248.
Cummins, S., Peltier, J.W., Erffmeyer, R. and Whalen, J. (2013), “A critical review of the literature for sales educators”, Journal of Marketing Education, Vol. 35 No. 1, pp. 68-78.
Davis, J.J. and O’Flaherty, S. (2012), “Assessing the accuracy of automated Twitter sentiment coding”, Academy of Marketing Studies Journal, Vol. 16, pp. 35-50.
Duggan, M. (2015), “The demographics of social media users”, Pew Research Center: Internet, Science & Tech, available at: www.pewinternet.org/2015/08/19/the-demographics-of-social-media-users (assessed 20 March 2017).
Dumas, G., Serfass, D.G., Brown, N.A. and Sherman, R.A. (2014), “The evolving nature of social network research: a commentary to Gleibs”, Analyses of Social Issues and Public Policy, Vol. 14 No. 1, pp. 374-378.
Erevelles, S., Fukawa, N. and Swayne, L. (2016), “Big data consumer analytics and the transformation of marketing”, Journal of Business Research, Vol. 69 No. 2, pp. 897-904.
Gonçalves, P. Araújo, M. Benevenuto, F. and and Cha, M. (2013), “Comparing and combining sentiment analysis methods”, paper presented at the ACM Conference on Online Social Networks, Boston, MA, 7-8 October 2013. available at: https://arxiv.org/pdf/1406.0032.pdf (accessed 3 February 2017).
Gopaldas, A. (2014), “Marketplace sentiments”, Journal of Consumer Research, Vol. 41 No. 4, pp. 995-1014.
Grimes, S. (2014), “Can sentiment analysis decode cross-cultural social media”, Breakthrough Analysis, available at https://breakthroughanalysis.com/2014/03/18/can-sentiment-analysis-decode-cross- cultural-social-media (accessed 4 May 2016).
Günther, T. and and Furrer, L. (2013), “GU-MLT-LT: Sentiment analysis of short messages using linguistic features and stochastic gradient descent”, paper presented at the Second Joint Conference on Lexical and Computational Semantic (SemEval 2013), Atlanta, Georgia, 14-15 June, available at: www.aclweb.org/anthology/S/S13/S13-2.pdf#page=364 (accessed 1 May 2017).
Online sentiment
analysis
161
Hasan, O., Habegger, B., Brunie, L., Bennani, N., and and Damiani, E. (2013), “A discussion of privacy challenges in user profiling with big data techniques: the eexcess use case”, 2013 IEEE International Congress on Big Data (BigData Congress), pp. 25-30, available at https://pdfs. semanticscholar.org/b5fb/e425c94c8e6477f1e4abd2af47bb1cac5f71.pdf (accessed 16 February 2017).
He, W., Zha, S. and Li, L. (2013), “Social media competitive analysis and text mining: a case study in the pizza industry”, International Journal of Information Management, Vol. 33 No. 3, pp. 464-472.
Hennig-Thurau, T., Wiertz, C. and Feldhaus, F. (2015), “Does twitter matter? the impact of microblogging word of mouth on consumers’ adoption of new movies”, Journal of the Academy of Marketing Science, Vol. 43 No. 3, pp. 375-394.
Heydari, A., Ali Tavakoli, M., Salim, N. and Heydari, Z. (2015), “Detection of review spam: a survey”, Expert Systems with Applications, Vol. 42 No. 7, pp. 3634-3642.
Holmes, S. (2009), “Methodological and ethical considerations in designing an Internet study of quality of life: a discussion paper”, International Journal of Nursing Studies, Vol. 46 No. 3, pp. 394-405.
Homburg, C., Ehm, L. and Artz, M. (2015), “Measuring and managing consumer sentiment in an online community environment”, Journal of Marketing Research, Vol. 52 No. 5, pp. 629-641.
Liang, T.P., Li, X., Yang, C.T. and Wang, M. (2015), “What in consumer reviews affects the sales of mobile apps: a multifacet sentiment analysis approach”, International Journal of Electronic Commerce, Vol. 20 No. 2, pp. 236-260.
Liu, B. (2012), “Sentiment analysis and opinion mining”, Synthesis Lectures on Human Language Technologies, Vol. 5 No. 1, pp. 1-167.
Lotz, S. and Van Rensburg, A. (2016), “Omission and other sins: tracking the quality of online machine translation output over four years”, Stellenbosch Papers in Linguistics, Vol. 46 No. 0, pp. 77-97.
Ludwig, S., De Ruyter, K., Friedman, M., Brüggen, E.C., Wetzels, M. and Pfann, G. (2013), “More than words: the influence of affective content and linguistic style matches in online reviews on conversion rates”, Journal of Marketing, Vol. 77 No. 1, pp. 87-103.
Lunnay, B., Borlagdan, J., McNaughton, D. and Ward, P. (2015), “Ethical use of social media to facilitate qualitative research”, Qualitative Health Research, Vol. 25 No. 1, pp. 99-109.
Makarem, S.C. and Jae, H. (2016), “Consumer boycott behavior: an exploratory analysis of twitter feeds”, Journal of Consumer Affairs, Vol. 50 No. 1, pp. 193-223.
Mullich, J. (2012), “Improving the effectiveness of customer sentiment analysis”, Data Informed, available at: http://data-informed.com/improving-effectiveness-of-customer-sentiment-analysis/ (assessed 23 February, 2017).
Nelson, P. (1974), “Advertising as information”, Journal of Political Economy, Vol. 82 No. 4, pp. 729-754.
Nunan, D. and Di Domenico, M. (2013), “Market research and the ethics of big data”, International Journal of Market Research, Vol. 55 No. 4, pp. 2-13.
Ordenes, F.V., Ludwig, S., De Ruyter, K., Grewal, D. and Wetzels, M. (2017), “Unveiling what is written in the stars: analyzing explicit, implicit, and discourse patterns of sentiment in social media”, Journal of Consumer Research, available at: http://openaccess.city.ac.uk/16047/1/Villa%20Roel %20et%20al.%202017.pdf (accessed 3 May 2017).
Otterbacher, J. (2013), “Gender, writing and ranking in review forums: a case study of the IMDb”, Knowledge and Information Systems, Vol. 35 No. 3, pp. 645-664.
Rambocas, M. and Gama, J. (2013), “The role of sentiment analysis”, Working Paper [489], FEP- UP, University of Porto, April, available at: https://pdfs.semanticscholar.org/acd0/c9f75152 acd2a622be442d20f96b0a3225d4.pdf (accessed 17 January 2017).
Rodriguez, M.L., Dixon, A.W. and Peltier, J. (2014), “A review of the interactive marketing literature in the context of personal selling and sales management: a research agenda”, Journal of Research in Interactive Marketing, Vol. 8 No. 4, pp. 294-308.
JRIM 12,2
162
Saif, H., He, Y., and Alani, H. (2012), “Semantic sentiment analysis of twitter”, 11th International Semantic Web Conference (ISWC 2012), Boston, MA, 11-15 November, available at: http://oro. open.ac.uk/34929/1/76490497.pdf (accessed 7 January 2017).
Scholand, A.J., Tausczik, Y.R. and Pennebaker, J.W. (2010), “Assessing group interaction with social language network analysis”, International Conference on Social Computing, Behavioral Modeling, and Prediction, Springer, Berlin, pp. 248-255.
Schweidel, D.A. and Moe, W.W. (2014), “Listening in on social media: a joint model of sentiment and venue format choice”, Journal of Marketing Research, Vol. 51 No. 4, pp. 387-402.
Sonnier, G.P., McAlister, L. and Rutz, O.J. (2011), “A dynamic model of the effect of online communications on firm sales”, Marketing Science, Vol. 30 No. 4, pp. 702-716.
Sousa, C.M., Martínez-L�opez, F.J. and Coelho, F. (2008), “The determinants of export performance: a review of the research in the literature between 1998 and 2005”, International Journal of Management Reviews, Vol. 10 No. 4, pp. 343-374.
Stelzner, M.A. (2012), “Social media marketing industry report”, Social Media Examiner, available at: www.socialmediaexaminer.com/SocialMediaMarketingIndustryReport2012.pdf (accessed 1 May 2017).
Tang, T., Fang, E. and Wang, F. (2014), “Is neutral really neutral? The effects of neutral user-generated content on product sales”, Journal of Marketing, Vol. 78 No. 4, pp. 41-58.
Thelwall, M., Wilkinson, D. and Uppal, S. (2010), “Data mining emotion in social network communication: gender differences in MySpace”, Journal of the American Society for Information Science and Technology, Vol. 61 No. 1, pp. 190-199.
Tirunillai, S. and Tellis, G.J. (2012), “Does chatter really matter? Dynamics of user-generated content and stock performance”, Marketing Science, Vol. 31 No. 2, pp. 198-215.
Corresponding author Meena Rambocas can be contacted at: [email protected]
For instructions on how to order reprints of this article, please visit our website: www.emeraldgrouppublishing.com/licensing/reprints.htm Or contact us for further details: [email protected]
Online sentiment
analysis
163
Reproduced with permission of copyright owner. Further reproduction prohibited without permission.
- Online sentiment analysis in marketing research: a review
- Introduction
- Overview of sentiment analysis
- The scope and approach of the review
- Characteristics of the marketing articles reviewed
- Categorization of the articles
- Unit of analysis
- Sampling design
- Methods used in sentiment detection and statistical analysis
- The application of sentiment analysis in marketing research
- Challenges of applying sentiment analysis in marketing research
- Technical limitations
- Practical limitations
- Ethical limitations
- Recommendations for sentiment analysis application in marketing research
- Conclusions
- References