help

profilebcs
A_Literature_Analysis_for_the_Identification_of_Machine_Learning_and_Feature_Extraction_Methods_for_Sentiment_Analysis.pdf

A Literature Analysis for the Identification of Machine Learning and Feature Extraction Methods

for Sentiment Analysis

Haberzettl, Markus and Markscheffel, Bernd

Chair for Information and Knowledge Management

Technische Universität Ilmenau

Ilmenau, Germany

{markus.haberzettl & bernd.markscheffel}@tu-ilmenau.de

Abstract - The increase in daily emails sent to the customer service of companies is creating new challenges. Sentiment

analysis, i.e. the automated recognition of mood and polarity in

texts, is a solution to this problem, but the sentiment analysis of

German emails is still an open research problem. With the help

of a literature analysis we identify and analyze the most relevant

machine learning methods and the corresponding feature

extraction methods.

Keywords – sentiment analysis. literature analysis, machine learning, feature extraction methods.

I. INTRODUCTION

Email is one of the preferred communication channels in

the field of customer service [1]. Therefore, the increasing

number of emails arriving daily in customer service

represents a challenge for the prompt processing of customer

requests in companies [2, 3, 4,]. Automated prioritization is

necessary in order to identify and give priority to critical

requests, for example complaints. Otherwise, there is a risk of

negative effects on the perception of companies and the

associated financial effects, e.g. due to customers leaving the

company.

One form of prioritization is sentiment, i.e. the

emotionally annotated mood and opinion in an email [5].

Sentiment is also an approach for solving further problems

such as the analysis of the course of customer contacts, email

marketing and campaign activities and the identification of

critical topics [6]. Approaches of linguistic data processing

(LDP) are used to automatically capture the sentiment [7].

One approach is the functional bundling of procedures and

methods. Methods are sets of rules for the targeted use of one

or more methods, methods being systematically executed

procedures geared to defined goals [8]. LDP is thus a

framework for computer-supported methods and procedures

for language processing as close to humans as possible [7, 9,

10]. The sentiment analysis deals with the automated

recording of sentiment in texts, sentences and words as part of

the LDP [6, 10, 11].

Although the number of published research projects is

increasing, sentiment analysis continues to be an open

research problem [12, 13]. In particular, there is a lack of

approaches specifically for the German language, whereby

the automated classification of polarity is of particular interest

[14, 15, 16]. Polarity denotes the expression of the sentiment

into the categories positive, negative and neutral [11, 12]. In

addition, much research work has flowed into the sentiment

analysis of microblogs such as Twitter in recent years; the

further development of approaches for emails at document

level, i.e. the classification of an entire document and not only

its sub-areas, has been neglected [17, 18].

In research, methods of machine learning have prevailed

over knowledge- and dictionary-based methods to determine

the polarity [19, 20, 21, 22]. The reason for this is that

machine learning methods approach human accuracy and are

not subject to some limitations of the other two, for example

lack of dynamics in relation to informal language [20, 23].

While knowledge- and dictionary-based methods are manual

rule definitions, machine learning represents the fully

automated inductive detection of such rules using algorithms

developed for this purpose [23] So far, no machine learning

method has been identified as dominant - another reason why

sentiment analysis is still an unsolved research problem today

[24, 25, 3]. Ohana and Tierney see a solution for the

classification of polarity in the combination of sentiment

lexica and machine learning methods [26, 14]. Sentiment

lexica are dictionaries in which words are assigned to a

polarity index [27, 28] In addition, Ohana and Tierney

suspect further potential in linking such lexica and learning

methods with further methods of feature extraction [26].

Feature extraction methods generate or extract features from

data which are the input for machine learning methods [29,

30, 31]. Features designate numerically measurable attributes

and properties of data [31].

The main objective of this work is to identify the major

approaches of machine learning and the corresponding

feature extraction methods for a sentiment analysis with the

help of a literature analysis. In chapter II we will present the

related work for the context of sentiment analysis. In section

III we will outline the methodology of the literature analysis

and finally, in section IV we will show the results and close

with an outlook on future work.

978-1-5386-5244-2/18/$31.00 ©2018 IEEE

!"#$%%&$!'(&$%#&)$"*&)+',*&-%#%&.%'*&'/"0"$)+'(&-*#1)$"*&'2)&)0%1%&$'3(,/(2'45678

6

II. RELATED WORK

Automated text categorization has been a research

problem since the 1960s, but only towards the end of the last

century did the leap in computer computing capacity allow

for more extensive research [23]. Therefore, research on

sentiment analysis only began at the beginning of this

millennium and above all with the groundbreaking work of

Pang, Lee and Vaithyanathan [21].

In this paper, the authors point out the distinction

between sentiment analysis and classical text categorization

procedures. They justify their differentiation by the

insufficiently functioning recognition of sentimental sentences

by keywords. Until then, the categorization of texts by means

of keywords or manually derived rules was the recognized

procedure for text categorization - the so-called knowledge

engineering. Instead, they relied on the monitored machine

learning methods that became popular in the 1990 [21, 23].

Nevertheless, both approaches were pursued in parallel

in research. Although knowledge engineering lost its

importance due to its inferiority with regard to informal

language. Nevertheless, the sentiment dictionaries still used

today for feature extraction are based on this approach. The

most groundbreaking lexicon, SentiWord-Net, was created in

2006 by Esuli and Sebastiani to recognize the polarity of

English words [27]. It is a model for numerous English and

other-language sentiment encyclopedias, for example for the

German version SentiWS by Remus, Quasthoff and Heyer

[32].

With regard to machine learning methods, there is no

consensus as to which method is dominant. A problem in this

context is the comparability of the results of the various

research activities. Authors usually use individually and

differently generated corpora of different size and content

[33]. In addition, the multitude of adjustment possibilities of

the individual learning and feature extraction methods is

complex and extensive. Therefore, not every combination of

characteristics and learning methods can be tested or used in

every work. After all, hybrids are the result of competing

approaches. These try to combine the respective strengths,

like conducted by Prabowo and Thelwall [34] and Ohana and

Tierney [26], for example but here too there is no consensus

on a dominant method; rather, research has entered a search

cycle for combinations of already known methods.

Only in recent years has the increasing computing power

allowed a further approach to research: Following its success,

deep learning is now regarded as a source of hope for new

findings in sentiment analysis [35, 36]. Its strength is above

all in the automation of feature extraction and selection,

whereby significantly more features can be processed at the

same time than with other approaches [10]. However, since

there is no consensus to date regarding the dominance of deep

learning over their approaches, further experiments are being

conducted with hybrids [20, 37].

Meanwhile, research on the German language is, as

mentioned, less frequent. The most important works are the

SentiWS by Remus, Quasthoff and Heyer [32] and the

production of the GermanPolarityClues by Waltinger [16]

which is another German-language sentiment encyclopedia.

Furthermore, the use of SentiWS by Dollmann and Geierhos

[38] and the attempt by Momtazi [39] to create a SentiWS

and GermanPolarityClues exceeding dictionary are worth

mentioning. Momtazi's results should be viewed critically due

to a small corpus and the intransparent use of machine

learning methods [39]. Dollmanns and Geierhos' approach

also did not achieve a breakthrough due to moderate precision

and recall values, but clearly shows the use and specialization

of SentiWS and thus permits a transfer of their findings to

new approaches around the encyclopedia [38]. In addition,

Clematide's attempt to develop a German-language reference

corpus is also worth mentioning, in response to the criticism

of incomparable corpora expressed above [40]. German

works based on an email corpus cannot be found. Finally, the

efforts of the Interest Group on German Sentiment Analysis,

founded in 2012, should be mentioned, which counteracts the

lack of German-language research with its 30 publications (as

of November 2017).

III. METHODOLOGY

A. Determinants of literature research

The aim of the literature research, which is oriented

towards Webster and Watson [41], is twofold. First,

monitored machine learning methods must be identified that

can be used in the context of sentiment analysis. The focus is

not on finding any possible learning method, but on finding

the most relevant ones. The premise applies: The most

frequently cited research includes the most relevant methods.

Secondly, methods of feature extraction must be identified on

the same premise. The premise must be specified to the effect

that the characteristic extraction methods used for the most

relevant machine learning methods are the most relevant

extraction methods, i.e. the learning and extraction methods

sought must occur together, since both are interdependent

components of a Knowledge Discovery in Databases process.

Due to the dependency, the methods can be identified in a

single, comprehensive search.

The research identified in the course of the research is

prepared in accordance with the concept-centric approach

required by Webster and Watson [41]. concept-centric means

that each research work is to be assigned in tabular form to

the concepts it contains [41]. Accordingly, the work must be

reviewed for contained concepts. A separate column must be

created for each concept and the table must thus be

successively expanded during the course of the search [41].

Each research work must therefore be listed in a new line for

the assignment. The existence of the concept is then recorded

in the columns of the respective row. Accordingly, concept-

centric tables contain little information for evaluation and

discussion of the concepts. Therefore, the table of this work

contains the following additional dimensions oriented to

Prabowo and Thelwall [34]: Author, subject matter, corpus

and quality (Similar dimensions use for example Vinodhini

7

and Chandrasekaran [42]. The first three are to be recorded

for subsequent discussion and comparability of the work.

Since the documentation level is of interest in this work, the

level used is specified in addition to the object of

investigation. For better comparability, the origin of the data

and the polarity scale and corpus size used shall also be

recorded for each corpus. The data origin provides

information about the domains contained in the corpus. The

quality criteria are used to evaluate of the concepts. Similar to

Prabowo and Thelwall, it comprises the following four

quality criteria: Accuracy, Precision, Recall and F1-Measure

[34]. It is also recorded whether and which validation

procedure was used to better assess the quality of the quality

criteria. Only the quality criteria of the best performing

concept in the respective research work (or the concept

combination of a machine learning and feature extraction

concept) will be noted. In the respective work, this is the most

relevant concept in the sense of the premise mentioned at the

beginning of the chapter. These concepts must be marked in

the table. The previous premise must therefore be further

specified: Only concepts that are the best concept or part of

the best combination of concepts in at least one research

project are relevant for literature research.

Finally, notes should be added to increase the

transparency of the decisions made during the search. As

recommended by Webster and Watson, the Web of Science is

used for research [41]. Due to the limitation of this approach,

further framework conditions must be defined for literature

research. Only articles from the Web of Science Core

Collection published between 2002 and 2018 are to be used

for research. The time limit is based on the publication date

of Pangs, Lees and Vaithyanathan's work [21], which was

identified as the starting point of modern research. In addition,

only articles and ongoing work are to be used. The relevant

research work shall be selected in accordance with the

premise defined in the course of this chapter as follows: Sort

the search results in descending order of their citation rank

(sum of their citations, highest sum first) and select the first

thirty results.

B. Phases

Webster and Watson mainly describe the steps in their

explanations how to find relevant literature [41]. They treat

the subsequent iterative procedure for preparation, structuring

and coding only superficially. The process must therefore be

further specified. In phase 1 (Generate Query), the query is

generated with which the Web of Science Core Collection is

searched for relevant articles. Naive search terms are first

defined and the result is then analyzed. Analyze means the

explorative examination of the found articles with regard to

their contextual affiliation. For reasons of efficiency, only the

abstract is used; only in cases of uncertainty is the text itself

considered in more detail. Afterwards the key terms, topics

assigned to the article as well as headings and the abstract are

to be searched for key search terms. If the article belongs to

the context, missing search terms found must be added to the

query. If the article is context-independent, terms identified as

context-independent must be eliminated from the query or

introduced in the query as exclusion criteria. Elimination or

exclusion cannot be carried out for ambiguous terms, since

articles with potential relevance are no longer recorded.

Articles found faulty in this way are subsequently excluded in

phase 2. This is more time-consuming but improves the

quality of the research. After the exclusion of a work, the next

lower cited work will follow (the required 30 work must be

observed in this way). The search query must then be

repeated iteratively until no new search terms are found.

(((TS=(sentiment) OR TI=(sentiment))

AND (TS=(classification) OR TI=(classification)

OR TS=(extraction) OR TI=(extraction)

OR TS=(analysis) OR TI=(analysis)

OR TS=(mining) OR TI=(mining)

OR TS=(polarity) OR TI=(polarity)))

OR (TS=(“opinion mining”) OR TI=(“opinion mining”)))

AND (TS=(machine learning) OR TI=(machine learning)

OR TS=(artificial) OR TI=(artificial)

OR TS=(supervised) OR TI=(supervised))

AND PY=(2002-2018) AND DOCUMENT TYPES:

(Article OR Proceedings Paper) Fig. 1 Query in Web of Science Syntax

In phase 2 starts the content analysis of the articles found

with the help of the search query. Despite the preparatory

work in phase 1, you can recognize the need to adjust the

query again. Phase 2 must then be restarted. The analysis

includes the concept-centric coding of the articles and filling

of the additional dimensions. Encoding decisions made must

be documented in an annotation field for traceability. This is

necessary, because approaches, definitions and contents in the

articles are often heterogeneous. Therefore, the final

structuring and coding of the table in phase 2 is also not

possible - an overall impression of all relevant work is

required first. In addition, it is necessary to maintain a to-do

list in which discrepancies and arising questions are to be

noted. This is important for phase 3 in order to better identify

concepts to be merged. Furthermore, non-context articles are

to be excluded. In step 3 all articles have been analyzed at

least once and a final structuring is explored. After labeling

each work, the best learning method and the best extraction

methods were selected according to the above quality criteria.

The Accuracy represents the greatest intersection between all

works; only two works do not have an accuracy specification.

Therefore, the classification result with the highest Accuracy

(test data only) was selected and the associated Precision,

Recall and F1 values were determined. If these or the

Accuracy itself are distributed over several classification

results, for example when using several corpora, the mean

value of all relevant results is calculated. The identification of

machine learning methods is trouble free. The methods are

mathematically justified and therefore have no scope for

interpretation. In contrast, feature extraction methods are not

always clearly differentiated or defined. It should be

explicitly noted that e.g. preprocessing tasks like stemming

and lemmatization are not to be interpreted as characteristic

extraction methods. As a result, 11 feature extraction methods

and 12 machine learning methods were identified.

8

IV. RESULTS

The 30 research papers [12, 21, 22, 43, 44, 45, 46, 47,

48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63,

64, 65, 66, 67, 68, 69] were determined on the basis of their

citation rank in the Web of Science Core Collection. Twelve

monitored machine learning methods and eleven feature

extraction methods were identified. The premise was that the

most relevant research work, measured by citation rank,

included the most relevant methods. This premise was

specified to the effect that the most relevant machine learning

methods use the most relevant feature extraction methods.

Furthermore, it was determined that in the context of

literature research, methods are only relevant if they are most

relevant in at least one of the research projects compared to

the other methods. This assumption is fulfilled by the

following five of the twelve machine learning methods:

Support Vector Machine (SVM), Artificial Neural Network

(ANN), Naive Bayes (NB), Logistic Regression (LR) or

Maximum Entropy (ME) and k-nN nearest neighbour (k-nN).

The SVM is used in most of the papers (28) and in over half

(53.33%) of the papers as the best machine learning method.

ANNs are only used in seven papers but dominate the SVM in

five of them. Table I shows the evaluation of the identified

relevant monitored machine learning methods.

TABLE I EVALUATION OF THE IDENTIFIED RELEVANT

MACHINE LEARNING METHODS Identified machine

learning method

Number of papers

in which the

method appears

Number of times

that a method is

the best method

Naive Bayes 17 4

Rule Based Classifier 3 0

Maximum Entropy 10 4

Regularized Least Squares 1 0

Winnow Classifier 1 0

n-Gram Model 1 0

Support Vector Machine 28 16

Artificial Neural Network 7 5

Decision Tree 6 0

k-nearest Neighbour 5 1

Conditional Random Field 1 0

Nearest Centroid Classifier 1 0

Table II illustrates the different subjects of investigation.

TABLE II SOURCE DOMAIN Source domain Number of papers

Film Reviews 15

Product Reviews 10

Other Reviews 4

MySpace 2

Tweets 9

Facebook 1

Other Social Media 1

News 2

Forum Posts 1

Blog Posts 1

Comments 3

SMS 1

Table III illustrates the relevant identified feature extraction

methods.

Table III EVALUATION OF THE IDENTIFIED RELEVANT

FEATURE EXTRACTION METHODS

Identified feature

extraction method

Number of papers

in which the

method appears

Number of times

that a method is

part of the best

combination

n-Gramm 26 26

Term frequency 8 3

Term presence 19 17

Term frequency -

Inverse document frequency

1 6

Part of speech tagging 11 8

Modification feature 3 3

Negation 7 7

Pointwise Mutual

Information

3 3

Sentiment Dictionary 10 9

Category 5 5

Corpus specific 8 8

Already at the beginning of the content analysis the

heterogeneity of the found works became clear. Heterogeneity

here means the differences in the objectives of the study, the

approach, the data basis used and the quality criteria and,

accordingly, the quality of the work. As a result, the contents

of the papers had to be standardized. This carries the risk of

distortion in the interpretation of the results of individual

works but increases their comparability. Nevertheless, the

effects of homogenisation must be discussed critically against

the background of the above-mentioned differences in the

papers. For example, the quality criteria collected for the

work in the literature search illustrate the problem of direct

comparability of the papers. Half of the work contains only

one quality criterion, the Accuracy. The low expressiveness

of Accuracy alone, e.g. if the corpus is dominated by one

class, for example if 90% of a corpus of sentiment analysis

consists of documents of the class "neutral" and consequently

an accuracy of 90% is already achieved with a continuous

classification of all documents with this class, must be

accepted due to the lack of availability of other criteria. Also,

a 10-fold cross-validation is used as the best validation

procedure in only 57% of all work. Thus, the quality criterion

of 43% of the works is to be regarded as not sufficiently

valid.

We must also consider the dependency of the quality

criteria of the corpus. A total of 31 different corpora are used

in the 30 papers (one to four corpora per work), 16 of which

are created by the authors themselves and are therefore

generally not comparable (because not available). Moreover,

the corpora differ significantly in their size (500 - 1.6 million

documents, median 8,000). Regarding the corpus size should

be mentioned that the corpora |D| ≫ 8,000 are often silver

standard corpora. Unlike gold standard corpora, which are

manually coded and of high quality, silver corpora consist of

automatically acquired data, which possibly simply learn the

heuristics used for the automatically generation of the corpus

instead of real patterns.

9

A further difference can be seen in the scale level used.

73% of the examined works use a binary classification into

"positive" and "negative". This is relevant because binary

classifications sometimes require different algorithms for a

machine learning method than the multi-class case. However,

since the collection of algorithms in the context of literature

research is negligible, or since various algorithms do not

change the use of the method itself, no limitation for the

findings of the research arises from this. But it should be

noticed that quality criteria are not directly comparable for

scales of different granularity, as the complexity for

classifiers increases with increasing the scale width.

V. SUMMARY AND FUTURE WORK

This work can be seen as a first step towards an advanced

research for an effective implementation of a solution for a

sentiment analysis of German emails. In this further work,

various machine learning methods are combined with

sentiment lexicons and, in a further approach, with

corresponding feature extraction methods identified by this

literature analysis. A third approach is a Deep Learning

approach based on Word Embeddings and a Convolutional

Neural Network to ultimately answer research questions such

as: Do machine learning methods based on sentiment lexicons

generate better results in the context of sentiment analysis

when the lexicon is combined with additional methods of

feature extraction or will the results of the Deep Learning

approach be better than the first two approaches mentioned?

The experiments to answer these questions are conducted

according to the KDD (Knowledge Discovery in Databases)

process [70]. It requires the solution of a number of further

problems, like the acquisition and coding of an own corpus,

which meets gold standard requirements. The experiments

themselves are implemented with the help of the Konstanz

Information Miner (KNIME) version 3.5.2. [71] due to the

many methods available and the integration of other common

tools.

VI. REFERENCES

[1] N. Gupta, M. Gilbert, and G. Di Fabbrizio, “Emotion Detection in Email Customer Care,“ in D. Inkpen, and C. Strapparava, Eds. Proceedings of the NAACL HLT 2010 Workshop on Computational Approaches to Analysis and Generation of Emotion in Text. Los Angeles 2010, pp. 10-16.

[2] J. Maier, “Kfz-Versicherer: Schwächen im E-Mail-Verkehr,” vb Magazin, nr. 2, 2015, pp. 20-23.

[3] novomind AG, “Ein Ohr für den Kunden: schneller und zuverlässiger Kundenservice bei Teufel Lautsprecher,” https://www.novomind.com/de/newsevents/news/de-tail/ein-ohr-fuer- den-kunden-schneller-und-zuverlaessiger-kundenservice-bei-teu-fel- lautsprecher, 2016, retrieved: 2017-11-02.

[4] The Radicati Group, “Email Statistics Report, 2018-2022 – Executive Summary,” https://www.radicati.com/wp/wp-content/uploads/2018/01/ Email_Statistics_Report,_2018-2022_Executive_Summary.pdf, 2018, retrieved: 2018-04-01.

[5] P. Borele, and D. A. Borikar, “An Approach to Sentiment Analysis using Artificial Neural Network with Comparative Analysis of Different Techniques,” IOSR Journal of Computer Engineering, vol. 18, nr. 2, 2016, pp. 64-69.

[6] T. Nasukawa, and J. Yi, “Sentiment Analysis: Capturing Favorability Using Natural Language Processing,” in J. Gennari, B. Porter, and Y.

Gil, Eds. Proceedings of the 2nd International Conference on Knowledge Capture (K-CAP'03). Sanibel Island 2003, pp. 70-77.

[7] A. Agarwal, et al, “Sentiment Analysis of Twitter Data,” in Proceedings of the Workshop on Language in Social Media LSM 2011, Portland 2011, pp. 30-38.

[8] W. Hesse, G. Merbeth, and Rainer Frölich, “Software-Entwicklung: Vorgehensmodelle, Projektführung, Produktverwaltung,” Oldenburg 1992.

[9] E. D. Liddy,“Natural Language Processing,” http://surface.syr.edu/cgi/ view-content.cgi?article=1043& context=istpub, 2001, retrieved: 2017- 06-17.

[10] D. Stojanovski, G. Strezoski, G. Madjarov, and I. Dimitrovski, “Twitter Sentiment Analysis Using Deep Convolutional Neural Network,” in E. Onieva, et al, Eds. Hybrid Artificial Intelligent Systems, HAIS 2015. Lecture Notes in Computer Science. vol 9121, Cham 2015, pp. 726-737.

[11] H. Yanagimoto, M. Shimada, and A. Yoshimura, “Word Classification for Sentiment Polarity Estimation Using Neural Network,” in S. Yamamoto, Eds. Human Interface and the Management of Information. Information and Interaction Design. Part I, LNCS 8016. Las Vegas 2013, pp. 669-677.

[12] F. Bravo-Marquez, M. Mendoza, and B. Poblete, “Meta-level sentiment models for big social data analysis,” Knowledge-Based Systems, vol. 69, 2014, pp. 86-99.

[13] K. Ravi, and V. Ravi, “A survey on opinion mining and sentiment analysis: Tasks, approaches and applications,” Knowledge-Based Systems, vol. 89, 2015, pp. 14-46.

[14] T. Scholz, S. Conrad, and L. Hillekamps, “Opinion Mining on a German Corpus of a Media Response Analysis,” in P. Sojka, et al ,. Text, Speech and Dialogue, 15th International Conference, TSD 2012. Proceedings. Brno 2012, pp. 39-46.

[15] F. Steinbauer, M. Kröll, “Sentiment Analysis for German Facebook Pages”, in E. Métaiset al, Eds. Natural Language Processing and Information Systems, 21st International Conference on Applications of Natural Language to Information Systems, NLDB 2016. Proceedings. Salford 2016, pp. 427-432.

[16] U. Waltinger, “GERMANPOLARITYCLUES: A Lexical Resource for German Sentiment Analysis,” in N. Calzolari, et al, Eds. Proceedings of the Seventh Conference on International Language Resources and Evaluation (LREC-10), Valletta 2010, pp. 1638-1642.

[17] S. Alhojely, “Sentiment Analysis and Opinion Mining: A Survey,” International Journal of Computer Applications, vol. 150, nr. 6, 2016, pp. 22-25.

[18] A. Tripathya, A. Agrawalb, and S. Kumar Rath, “Classification of Sentimental Reviews Using Machine Learning Techniques,” Procedia Computer Science, vol. 57, 2015, pp. 821-829.

[19] E. Cambria, et al, “New Avenues in Opinion Mining and Sentiment Analysis,” IEEE Intelligent Systems, vol. 28, nr. 2, 2013, pp. 15-21.

[20] Y. Cao, R. Xu, T. Chen, “Combining Convolutional Neural Network and Support Vector Machine for Sentiment Classification,” in X. Zhang, et al, Eds. Social Media Processing. SMP 2015. Communications in Computer and Information Science, vol. 568, Guangzhou 2015, pp. 144-155.

[21] B. Pang, L. Lee, and S. Vaithyanathan, “Thumbs up? Sentiment Classification using Machine Learning Techniques,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), vol. 10, Philadelphia 2002, pp. 79-86.

[22] S. Poria, E. Cambria, G. Winterstein, G.-B. Huang: Sentic Patterns, “Dependency-Based Rules for Concept-Level Sentiment Analysis,” in Knowledge-Based Systems, 2014, pp. 1-32.

[23] F. Sebastiani, “Machine Learning in Automated Text Categorization,” ACM computing surveys (CSUR), vol. 34, nr. 1, 2002, pp. 1-47.

[24] G. Vinodhini, and RM. Chandrasekaran, “Sentiment Analysis and Opinion Mining: A Survey,” International Journal of Advanced Research in Computer Science and Software Engineering. vol.2, nr. 6, June 2012, pp. 282-292.

[25] S. Argamonet al, “Automatically Determining Attitude Type and Force for Sentiment Analysis,”in Z. Vetulani and H. Uszkoreit, Eds. Human Language Technology. Challenges of the Information Society. Third Language and Technology Conference, LTC 2007. LNAI 5603. Poznan 2007, pp. 218-231.

[26] B. Ohana, and B. Tierney, “Sentiment Classification of Reviews Using SentiWordNet,”. in: 9th. IT&T Conference, Dublin Oktober 2009, pp. 1-9.

10

[27] A. Esuli, and F. Sebastiani,”SENTIWORDNET: A Publicly Available Lexical Resource for Opinion Mining,”https://www.researchgate.net/ publica-tion/200044289_SentiWordNet_A_Publicly_Available_ Lexical_Resource_for_Opinion_Mining, retrieved: 2017-06-17.

[28] G. Qiuet et al, “Opinion Word Expansion and Target Extraction through Double Propagation,” Computational Linguistics, vol. 37, nr. 1, 2011, pp. 9-27.

[29] M. Bramer, “Principles of Data Mining,” London, 2013.

[30] I. Guyon, and A. Elisseeff, “An Introduction to Feature Extraction,” in I. Guyon, S. Gunn, M. Nikravesh, and L. A. Zadeh Eds. Feature Extraction. Foundations and Applications, Berlin Heidelberg, 2006, pp. 1-28.

[31] T. A. Runkler, “Data Mining. Modelle und Algorithmen intelligenter Datenanalyse,” Wiesbaden, 2015.

[32] R. Remus, U. Quasthoff, and G. Heye, “SentiWS – a Publicly Available German-language Resource for Sentiment Analysis” in International Conference on Language Resources and Evaluation, 2010, pp. 1168-1171.

[33] F. Bütow, F. Schultze and L. Strauch, “Semantic Search: Sentiment Analysis with Machine Learning Algorithms on German News Articles,” http://www.dai-Labor.de/fileadmin/Files/Publikationen/ Buchdatei/BuetowEtAl--SentimentAnalysisOnGermanNews.pdf, retrieved 2017-12-29.

[34] R. Prabowo, and M. Thelwall, “Sentiment analysis: A combined approach,” Journal of Informetrics, vol. 3, 2009, pp. 143-157.

[35] G. Cai, and B. Xia, “Convolutional Neural Networks for Multimedia Sentiment Analysis,” in J. Li, Heng Ji, et al, Eds. Natural Language Processing and Chinese ComputinG; 4th CCF Conference, NLPCC 2015, Nanchang 2015, pp. 159-167.

[36] X. Gu, Y. Gu, H. Wu, ”Cascaded Convolutional Neural Networks for Aspect-Based Opinion Summary,” Neural Processing Letters, vol. 46, nr. 2, 2017, pp. 581-594.

[37] S. Ebert, N. T. Vu, and H. Schütze, “CIS-positive: Combining Con- volutional Neural Networks and SVMs for Sentiment Analysis in Twitter,” in P. Nakov, T. Zesch, D. Cer, and D. Jurgens, Eds. Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015), Denver 2015, pp. 527-532.

[38] M. Dollmann, and M. Geierhos, “SentiBA: Lexicon-based Sentiment Analysis on German Product Reviews,” in J. Ruppenhofer, and G. Faaß, Eds. Workshop Proceedings of the 12th Edition of the Konvens Conference. Hildesheim 2014, pp. 185-191.

[39] S. Momtazi, “Fine-grained German Sentiment Analysis on Social Media,” in N. Calzolari, et al Eds. Proceedings of the 8th International Conference on Language Ressources and Evaluation (LREC-2012), Istanbul 2012, pp. 1215-1220.

[40] S. Clematide, et al, “MLSA – A Multi-layered Reference Corpus for German Sentiment Analysis,” in N. Calzolari, et al Eds. Proceedings of the 8th International Conference on Language Ressources and Evaluation (LREC-2012), Istanbul 2012, pp. 3551-3556.

[41] J. Webster, and R. T. Watson, “Analyzing the Past to Prepare for the Future: Writing a Literatur Review,“ MIS Quarterly, vol. 26, nr. 2, 2002, pp. 13-23.

[42] G. Vinodhini, and RM. Chandrasekaran, “Sentiment Analysis and Opinion Mining: A Survey,” International Journal of Advanced Research in Computer Science and Software Engineering, vol. 2, nr. 6, June 2012, pp. 282-292.

[43] M. Thelwall, et al, “Sentiment Strength Detection in Short Informal Text,” Journal of the American Society for Information Science and Technology, vol. 61, nr. 12, 2010, pp. 2544-2558.

[44] M. Thelwall, K. Buckley, and G. Paltoglou, “Sentiment Strength Detection for the Social Web,” Journal of the American Society for Information Science and Technology, vol. 63, nr. 1, 2012, pp. 163-173.

[45] A. Kennedy, and D. Inkpen, “Sentiment Classification of Movie Reviews using Contextual Valence Shifters,” Computational Intelligence, vol. 22, nr. 2, 2006, pp. 110-125.

[46] T. Wilson, J. Wiebe, and P. Hoffmann, “Recognizing Contextual Polarity: An Exploration of Features for Phrase-Level Sentiment Analysis,” Association for Computational Linguistics, vol. 35, nr. 3, 2009, pp. 399-433.

[47] R. Xia, C. Zong, and S. Li: Ensemble of feature sets and classifi-cation algorithms for sentiment classification. In: Information Sciences. Vol. 181, 2011, S. 1138-1152.

[48] Q. Ye, Z. Zhang, and R. Law, “Sentiment classification of online reviews to travel destinations by supervised machine learning

approaches,” Expert Systems with Applications, vol. 36, 2009, pp. 6527-6535.

[49] S. Tan, and J. Zhang, “An empirical study of sentiment analysis for chinese documents,” Expert Systems with Applications, vol. 34, 2008, pp. 2622-2629.

[50] R. Moraes, et al, “Document-level sentiment classification. An empirical comparison between SVM and ANN,” Expert Systems with Applications. vol. 40, 2013, pp. 621-633.

[51] M. Ghiassi, J. Skinner, and David Zimbra, “Twitter brand sentiment analysis: A hybrid system using n-gram analysis and dynamic artificial neural network,” Expert Systems with Applications, vol. 40, 2013, pp. 6266-6282.

[52] M. Myslín, et al, “Using Twitter to Examine Smoking Behavior and Perceptions of Emerging Tobacco Products,” Journal of Medical Internet Research, vol. 15, nr. 8, e174, 2013, pp. 1-16.

[53] C. Zhang, et al, “Sentiment Analysis of Chinese Documents: From Sentence to Document Level,” Journal of the American Society for Information Science and Technology, vol. 60, nr. 12, 2009, pp. 2474- 2487.

[54] A. Ortigosa, J. M. Martín, and R. M. Carro, “Sentiment analysis in Face-book and ist application to e-learning,” Computers in Human Behavior, vol. 31, 2014, pp. 527-541.

[55] S. Kiritchenko, X. Zhu, and S. M. Mohammad, “Sentiment Analysis of Short Informal Texts,” Journal of Artificial Intelligence Research, vol. 50, 2014, pp. 723-762.

[56] G. Wang, et al, “Sentiment classification: The contribution of ensemble learning,” Decision Support Systems, vol. 57, 2014, pp. 77- 93.

[57] M. Rushdi-Saleh, et al, “OCA: Opinion Corpus for Arabic,” Journal of the American Society for Information Science and Technology, Vol. 62, Nr. 10, 2011, S. 2045-2054.

[58] Y. He, and D. Zhou, “Self-training from labeled features for sentiment analysis,” Information Processing and Management, vol. 47, 2011, pp. 606-616.

[59] M. Rushdi-Saleh, et al, “Experiments with SVM to classify opinions in different domains,” Expert Systems with Applications, vol. 38, 2011, pp. 14799-14804.

[60] E. Fersini, E. Messina, and F. A. Pozzi, “Sentiment analysis: Bayesian Ensemble Learning,” Decision Support Systems, vol. 68, 2014, pp. 26- 38.

[61] S. Poria, et al, “EmoSenticSpace: A novel framework for affective common-sense reasoning,” Knowledge-Based Systems, vol. 69, 2014, S. 108-123.

[62] F. Greaves, et al, “Use of Sentiment Analysis for Capturing Patient Experience From Free-Text Comments Posted Online,” Journal of Medical Internet Research, vol. 15, nr. 11, e239, 2013, pp. 1-9.

[63] J. Smailović, et al, “Stream-based active learning for sentiment analysis in the financial domain,“ Information Sciences, vol. 285, 2014, pp. 181-203.

[64] A. Balahur, and M. Turchi, “Comparative experiments using supervised learning and machine translation for multilingual sentiment analysis,“ Computer Speech and Language, vol. 28, 2014, pp. 56-75.

[65] D. Bollegala, D. Weir, and J. Carroll, “Cross-domain sentiment classification using a sentiment sensitive thesaurus,” IEEE Transactions on Knowledge and Data Engineering, vol. 25, Nr. 8, 2013, S. 1719-1731.

[66] S. Poria, et al, “Sentiment Data Flow Analysis by Means of Dynamic Linguistic Patterns,” IEEE Computational Intelligence Magazine, vol. 10, nr. 4, 2015, pp. 26-36.

[67] A. Montejo-Ráez, et al, “Ranked WordNet graph for Sentiment Polarity Classification in Twitter,“ Computer Speech and Language, vol. 28, 2014, pp. 93-107.

[68] A. Duric, and F. Song, “Feature Selection for Sentiment Analysis Based on Content and Syntax Models,” Decision Support Systems, vol. 53, 2012, pp. 704-711.

[69] V. Sindhwani, and P. Melville, “Document-Word Co-Regularization for Semisupervised Sentiment Analysis,” Eighth IEEE International Conference on Data Mining, Pisa 2008, pp. 1025-1030.

[70] U. Fayyad, G. Piatetsky-Shapiro, and P. Smyth: From Data Mining to Knowledge Discovery in Databases. In: AI Magazine. vol 17, nr. 3, 1996, pp. 37-54.

[71] KNIME AG, “Konstanz Information Miner“. https://www.knime.com/ 2018, retrieved: 2018-09-02.

11