helpfn

profilebcs
A_Survey_of_Sentiment_Analysis_from_Social_Media_Data.pdf

450 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

A Survey of Sentiment Analysis from Social Media Data

Koyel Chakraborty , Siddhartha Bhattacharyya , Senior Member, IEEE, and Rajib Bag

Abstract— In the current era of automation, machines are constantly being channelized to provide accurate interpretations of what people express on social media. The human race nowadays is submerged in the idea of what and how people think and the decisions taken thereafter are mostly based on the drift of the masses on social platforms. This article provides a multifaceted insight into the evolution of sentiment analysis into the limelight through the sudden explosion of plethora of data on the internet. This article also addresses the process of capturing data from social media over the years along with the similarity detection based on similar choices of the users in social networks. The techniques of communalizing user data have also been surveyed in this article. Data, in its different forms, have also been analyzed and presented as a part of survey in this article. Other than this, the methods of evaluating sentiments have been studied, categorized, and compared, and the limitations exposed in the hope that this shall provide scope for better research in the future.

Index Terms— Clustering, community, sentiment analysis, social media, social networks.

I. INTRODUCTION

AS HUMANS, we always tend to get attracted to like-minded people. Even studies suggest that we are com- fortable in socializing with people with similar beliefs, with people on whom we can trust and who can facilitate to help achieve our aspirations. Etymologically, people have a tendency to be associated with similar-minded communities. Multiple clusters make a community. Modularity is one of the prime mechanisms considered while determining the quantity of communities [1]. If the characteristics of the clusters are minutely analyzed, then it can be instrumental in helping to identify the specific character set of individual clusters or like- minded people groups.

Putting it in the other way, it can also be said that the presence of a common connect between a set of individuals ensures that there lie similar principles and purposes between that set of people.

To be more specific, there is the availability of two types of social media; Social Networks and Online Communities.

Manuscript received July 8, 2019; revised October 15, 2019; accepted November 18, 2019. Date of publication January 7, 2020; date of current version April 3, 2020. (Corresponding author: Siddhartha Bhattacharyya.)

K. Chakraborty and R. Bag are with the Computer Science and Engi- neering Department, Supreme Knowledge Foundation Group of Institu- tions, Chandannagar 712139, India (e-mail: [email protected]; [email protected]).

S. Bhattacharyya is with the Department of Computer Science and Engineering, Christ University, Bengaluru 560029, India (e-mail: [email protected]).

This article has supplementary downloadable material available at http://ieeexplore.ieee.org, provided by the authors.

Digital Object Identifier 10.1109/TCSS.2019.2956957

Social networks are formed of people who are interconnected through some previous personal relationships, retain the same socially and further would prefer connecting to new associ- ations to enlarge their personal contacts. It associates people who have a straight connect to the other. Compared to the former, communities comprise people from multiple fields having less or no connection between them. The main connect between individuals in a community lies in the fondness toward a familiar interest. Apparently, people stay within a community for varied reasons, it might be the liking for a special thing, or it might be that the person feels that he/she should be associated with that community or he/she might achieve something by adhering to that community. Clearly, social networks contain organized arrangement whereas com- munities contain arrangements which overlap and are nested amidst them.

Social Media is the method of sharing data with a huge and vast audience. It can be addressed as a medium of propagating information through an interface. Social media in tandem with social networks helps individuals cater their content to a wider society and reach out to more people for sharing or promotion [2].

Sentiment analysis is the procedure of categorizing the views expressed over a particular object. With the advent of varied technological tools, it has become an important measure to be aware of the mass view in business, prod- ucts or in matters of common like and dislike. Tracking down the emotion behind the posts on social media can help relate the context in which the user shall react and progress.

This article provides the flow of social network analysis (SNA) from over 200 papers and presents the research that has been performed in social network and its related fields. This article is organized in the following manner—Section II deals with the inception of social networks in the research vicinity. Section III mentions the motivation for this arti- cle. Section IV contains the detailed methods implied to find clusters as well as communities from over 40 papers. Section V deals with the review done on 45 papers in the corresponding field and Section VI consists of the variety of the efforts that have been meted out for social media data. Section VIII deals with the varied techniques applied to detect accurate sentiments from data on social media. Section VIII concludes the paper with some future scope for research in this field. A crawler, tool to evaluate seman- tics, an engine that allows language preprocessing and a classifier are the main components of a sentiment analysis system [3].

2329-924X © 2020 IEEE. Personal use is permitted, but republication/redistribution requires IEEE permission. See https://www.ieee.org/publications/rights/index.html for more information.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

CHAKRABORTY et al.: SURVEY OF SENTIMENT ANALYSIS FROM SOCIAL MEDIA DATA 451

II. ONSET OF SOCIAL NETWORK AND ITS ANALYSIS

In SNA, the interconnectivity of humans in a social network is called as cliques, which can be defined as a structure where every member of the collection of people is directly and cohesively tied to each other. Bron and Kerbosch [4] suggested that there are ways to find maximal total subgraphs of graphs which are not directed, especially through back- tracking algorithms utilizing branch and bound to slice off the branches that do not lead to the formation of clique. Research related to develop algorithms to cluster correlated data from different applications of social networks has been initiated way back in 1975 [5], where the converging nature of product-moment corelational matrices was used in the CONCOR algorithm to cluster in a hierarchical manner. It is to be considered also, that with the latest splurge in the amount of social media data, it must have been a humongous task for the search engines to index the billions of data that they store in the web pages. Accordingly, the framework of Google was explained in detail [6] along with its incorporated features of scalability and sturdiness so that they yield perfect search results. As providing faultless web page results on the internet has been a concern since inception of the social network and its analysis context, researchers proposed various techniques to trace efficient relatable data when information relating to a broad subject was searched for in the web [7]. Certain factors were kept in mind such as the vast amount of data that was increasing at an exponential rate and the corresponding results that were shown had to be of supe- rior quality. Also, in cases of broad topics, the quality of pages was important as they would be of most relevance to the user and finally finding the hub pages that were densely linked to the set of correlated authoritative pages through which the outline of the social association could be judged.

The contribution of this article lies in the fact that the evolution of digital data which is extremely necessary to understand the opinion of the masses has been studied since inception. The path initiating from website data to the final analysis of sentiments has been portrayed. The diverse areas in which data have been collected for sentiments to be analyzed have been made ready in a single place. Ample options have been provided which can be applied to solve real-life problems collecting data from social media and finally evaluating them.

III. MOTIVATION

With the plethora of data being amassed on the internet, it is high time that matters relating to social media and its data be given utmost importance. What others think has to be followed by all, is a trend being perceived among the masses. This article provides an almost chronological order in the development of the social networks, acquiring data on social media and analyzing them and finally prediction of feelings within these data have been discussed in detail. To the best of our knowledge, this article is the first attempt to amalgamate this huge amount of information from over an assortment of 200 papers mentioning their contributions to this

field. It is a known fact that qualitative study comprises how people observe the reality that is taking place around a human being; hence the problems are learned in their natural setting. This article provides a qualitative insight into the dealings that can be made with social media data and how they are analyzed to help the world to understand the emotions and behavioral patterns of their contemporaries. The authors strongly feel that this one-of-a-kind paper will help experienced researchers get access to a vast variety of meaningful work at the same place and can also judge the amount of work that has already been done in this field. The future scopes for each of the papers mentioned in the tables act as an easy reference from which the idea of further work can be considered. For naïve researchers, this article can act as a base upon which they can start their work pertaining to social media and its analysis. As we have become extremely technology dependant, this article will prove as a benchmark to fellow researchers who further wish to work on the mentioned challenges faced in social media data.

As almost all of the conversations take place online and not in offline mode, it is important that the social media and its components be understood properly in order to analyze sentiments. It has been observed that the focus on sentiment analysis has been more since 2004 [207], and that is why the authors have considered papers of social media and it ancillaries from the year 2008 in this review. Reviewing only the work that has been done on sentiments does not meet the requirement to properly get a grip of the methods in which the semantics of sentiments is to be judged. Hence, it was a strategic decision to provide a detailed description of social media, its components and finally jump into the main topic of the most trending sentiment analysis. In a nutshell, the main objective of the paper is to present the work done in this particular field and to address its limitation making scope of ample research for scholars.

IV. DATA ACQUISITION

Acquiring data from social networks demands a huge atten- tion in order to accurately predict the ideology of the user behind posting that data on social media. Popular social media sites like Facebook, LinkedIn and Twitter provide media for the community to post, share, like and comment along with other friends within the same network. On the other hand, sites like Youtube, Flickr and Digg also are gearing up in providing various facilities for enhancing the connection with social friends. Data collected from social media have lots of factors associated with it, it may or may not be noisy, it may be homogeneous or heterogeneous, it may be of diverse range, etc.

The main techniques through which data are obtained from social media include: 1) collection of fresh data; 2) reuse of previously available data; 3) reuse of data not belonging to the specific person; 4) procuring data; and 5) data obtained from the internet (social media, texts and photosets).

The data after being obtained are initially processed. Data processing comprises multiple actions like checking the authenticity, understanding the outline, changing and assimi- lating into a suitable format for further use. After this step,

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

452 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

the processed data are analyzed to deduce the outcome of the actual underlying emotion of the users behind those data.

Technically, three main methods widely used to acquire data are [8].

1) Network Traffic Analysis—This is the method in which packet streams are collected from a network connection, which further helps in tracking the browsing information in the network. Due to security concerns, this method is very rarely used specially in private groups.

2) Ad hoc Applications—This is a set of air position indicators (APIs) which give the information regarding the account holder in a specific site and can further track the vibrant behavior of the user.

3) Crawling—This is the most popular method used to acquire data from social media. Public information is provided on asking for a specific type of data through queries. Crawling also helps to achieve data through the APIs available in some of the social media sites.

V. SOCIAL NETWORKS

Social Networks can be defined as the utilization of social media to connect to known or sometimes unknown acquain- tances. It might be to bond with friends or relatives, contem- poraries or colleagues and maybe for bonding with users for business purposes as well. The entire history leading to the progress of SNA can be found in [9]. Early research in this area leads back to [10], where movie data have been collected to construct models from relational statistics. Data in huge volumes were considered along with a language to extract related information and an algorithm to produce relational probability theory to produce a classifier for relational data. In [11], a small sample of emails showed that algorithms sup- ported by graphs were more efficient to identify who-knows- what within the organization compared to content-driven algo- rithms. Relational dependence networks have been proposed in [12] emphasizing the fact that models competent of finding dependencies result in improved categorization of data.

VI. CLUSTER AND COMMUNITY IN SOCIAL NETWORKS

Though clustering is one of the most sought after methods utilized to distinguish and estimate community structures in social networks, there lie differences in both the terms. While varied type of characteristics is considered while dealing with clusters, community generally adheres to any one type of attribute while it is detected. There lies distinction in consideration of connections in both clusters and communities; cluster discovery can be done easily on crowded connections, but the discovery of the latter presumes that there is very little connection in the network.

A. Clustering and Its Applications

Clustering can be defined as the technique to identify normal assembly within a group of units [13]. The probing nature of different types of traditional clustering methods along with the comparatively new spectral clustering methods helps to successfully solve the problem of finding similarity

in conduct of a set of data [14]. On social media, specific clustering mechanisms like hierarchical clustering [15] have been applied to a great extent for folksonomy, which helps in understanding the intended interest of the client as well as the specific matter of the resource [16]. Studies show that social graphs have also been used to designate clusters of connections which are of importance to the user [17], where factors like the regularity, intimacy and recentness with the contact help in identifying a substantially significant online association. Social cohesion has been a matter of concern even before the advent of the computer age and hence intense constituent detection and its scrutiny are necessary for the proper understanding of social networks [18]. Participation of users in sharing opinions on different issues on social media and later distinguishing those themes is not an easy task which urges the requirement of clustering web views for intelligent detection of security threatening cases. It is in this regard that the scalable distance-based algorithm [19] shows high precision regarding the mining of essential issues while eliminating noise. Noisy link detection and its reduced effect in clustering can also be eliminated utilizing assorted arbitrary areas combining related data and social indicators to cluster information [20]. Clustering has also been applied to combine both socially and geographically to find the proximity of visits to the related clusters [21]. This method outdoes the traditional spatial clustering [22] in grouping a plethora of spaces in a fraction of time. With the recent outburst of social data, problem of finding recurrent itemsets in large data has been solved by MapReduce model [23], which deploys k-means clustering algorithm [24] to preprocess the data and frequent data sets are mined through a priori [25] and Eclat [26] algorithms [27]. The experiment proves that this method yields excellent results on big data at a very elevated pace.

B. Community Detection in Social Networks

To manage the incessant rise while indexing web pages and to preserve the stability of precision and recall it was devised that identifying a cohesive community and linking them to relevant links solves the problem [28]. The goal of the community is to identify intensely joined associations of individuals in social networks. Traditional approaches like Kernighan–Lin were based on particular problems. Hence, many algorithms have been proposed to identify community models within networks, some of which may differ from traditional methods of detecting communities but ultimately yields similar qualitative results and that too involving sur- plus of vertices [29], while the other makes use of ethereal techniques, cluster scrutiny and modularity perception [30]. Some techniques have also made use of traditional approaches imposing some shortcuts, resulting in linear time execution of the algorithms [31]. In [32], the Lennard-Jones clusters have been studied to identify the potential energy landscapes (PEL) utilizing the network topology and the community structure was detected. Considering the centrality measure of an edge in a network [33], algorithms were also designed to find out the most vital edge by truncating the less important edges succes-

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

CHAKRABORTY et al.: SURVEY OF SENTIMENT ANALYSIS FROM SOCIAL MEDIA DATA 453

sively to ultimately form disconnected groups. Other variations of quick yet general divisive algorithms were designed in the same year [34], [35], one recounting the betweenness count after elimination of each edge while the other refining the computationally expensive GN algorithm which required nontopological data to clarify the branch details which bears significance to the structure. Initial researches concentrated on graphs from which the complete configuration was identified. It was next to which Clauset [36] proposed an approach of agglomerative algorithm which worked on dynamic and too hefty graphs. This approach took into account one vertex at a time and also proved that the straightforward application of that algorithm which would be appropriate to implement crawler programs helping unearth neighboring communities on the internet. Out of the algorithms that were implemented on large networks, study in [37] showed methods to explore extremely overlapping, nested and linked association of nodes in binary networks. A stringent version of the web community detection is mentioned in [38], which incorporated to the build- ing of a Gomory–Hu tree [39] yielding in computationally proficient results. Apart from the varied approaches based on the works by GN, one technique has been mentioned in [40], which maps the community identification dilemma into the discovery of the ground situation of an unbound scoped Potts whirl glass through ansatz coalescing data from in cooperation with current and absent associates. Though finding communities in networks has been well considered since its inception, a prescribed definition of the same was absent, taking clue of which, researchers in [41] made use of benchmark techniques and expressed community detection as an inference or maximum likelihood problem. A notable example of detecting communities utilizing the eigenvectors of matrices has been portrayed in [42], where the modularity function has been revisited in terms of matrices leading to representation of the optimization job as a spectral quandary. In another case [43], the degree concept from a solitary vertex was extended to subgraphs which reduce the overall complexity and the same was applied to produce a tool referred to as ModuleNetwork (MoNet) [43], to find commu- nity formations with outsized networks. Random walks were used to compute resemblance in structures within the ver- tex spaces, which if implemented in hierarchical algorithms, provided efficient results [44]. Also instances of methods taking into account the smallest loops passing through a particular node were also cited in [45], which is claimed as an adaptation of the neighboring competence introduced by Latora and Marchiori [46]. Raghavan et al. [47] probed into a mechanism which considered only the structure of the network and implemented simple label broadcast algorithm where every node has been provided with a unique label and at every step individual nodes acquire the label that majority of its neighbors currently possess. Other notable works that need mention in the community detection field are that of techniques in which heuristics were used to optimize the modularity [48], genetics supported approach used to discover communities by optimizing a simple yet successful fitness function as mentioned in [49]. While the concept in [47] was used to detect real-time communities in

large networks [50] and utilized to include the information about as many communities as there are parameters to handle bipartite graphs [51]; optical data mining methodologies were being used to identify overlapping communities to facilitate proper constraint selection post viewing the preliminary data visualizations [52]. The narrowing of the distance between humans and social media was kept in mind to opt for a coclustering construction, which makes use of the intercon- nected data among the users and the tags applied on social media to understand the preference of the group by an indi- vidual [53]. Other applications of community-related services extend to the maintenance of interpretational significance between question–answer duos along with data sets, for which a deep belief inspired structure via barely phrase characteris- tics was projected within the social community as illustrated in [54]. An overall survey of the methods to detect communi- ties especially in social networks can be found in [55]–[58]. Other than these, edges have been considered to yield opti- mum community detection results [59], multiple data sources have been combined to enhance the performance of distin- guishing communities [60] and prediction made to find how trendy information within communities will be broadcasted contagiously [61].

VII. CENTRALITY FACTORS TO MEASURE THE INFLUENCE OF A NODE IN SOCIAL NETWORKS

Another important aspect of networks is the centrality factor where the network with high centrality is conquered by the entity controlling the passage flow of the network. In [62], the centrality measure calculation methods have been revisited and two deviations have been suggested, one, to find the centrality measure through the course of traffic and the other through the extension of the network. Upon finding the suitable flow, outcomes relating to the significance or contribution of the node may also be obtained. Observation of centrality and its decreasing effect in the presence of error have been mentioned in [63]. Design of an effective algorithm that yields the uppermost graded intimate centrality vertex in a network in less time has been described in [64]. Zhou et al. [65] have demonstrated through parameterized measures how terrorism is being broadened utilizing the social networks by connecting to like-minded clusters, and hence their usage should be monitored. As social networks comprise innumerable links and nodes, it becomes difficult to survey them in an organized man- ner. To maintain an equilibrium between methodical yet elastic discovery of social networks, a system called SocialAction was offered in [66], to efficiently consider several geometric and optical network analysis parameters. The parameter ranking methodology in the paper not only allows receiving summary, strain nodes, locate outliers but also helps to integrate nodes to decrease complexity, find consistent subgroups and pay atten- tion on communities of attention. Social networks have been utilized to foretell psychological fitness as portrayed in [67], where dissimilar networks were found, the nonfamily and nonfriends network which helped to predict the symptoms of depression within these isolated individuals. With the passage of time, the increasing prominence of large networks led to

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

454 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

a major problem of handling the real-time data generated from those networks to understand the basic nuances of those groups. Backstrom et al. [68] devised methods which predicted how the large groups within the networks showed growth in terms of members and substance. It also explored the networks at an individual level to ascertain the joining in a certain overlapping community. Wikipedia [69], being one of the biggest repositories of information doubts the trustworthiness of the material provided to it. To take charge of this situation, Sicilia et al. [70] proposed that this problem should be confronted with in the beginning as reliability issues shall increase with the increase in the growth of the network. The people contributing to Wiki and the references that are linked to external sources were mapped through a graph and certain metrics were imbibed to extort data from the network. On the commercial front, social networks started playing an important role in the decision-making process of entrepre- neurs. It is in this regard, Aldrich and Kim [71] defined three models of networks—accidental, small world and truncated scale free. Two types of association formation models were found, one which was formed logically and the other based on individual relationships. Though many variations of the system were proposed, it was suggested that instead of banking on a specific type, it was recommended to envelop the entire entrepreneurial series in the aim of their business to grow. Businessmen promoting their brands utilized social networks to reach out to a maximum number of customers through social networks. Further availability of data corresponding to a specific consumer network could help to identify the controller using the concept of centrality; though ascertaining the parameters to be considered still remains an issue. In [72], a real-time network data were considered to evaluate varied centrality parameters to spread out messages within a network. SenderRank [72] is a novel centrality measure introduced in the paper which outperformed many existing measures, but the category of message and its contents along with the type of networks affect the prediction of favorable customers. As social networks seem to form the entire habit of an individual, it should also depict the mental state of a person along with geographical preferences in the social networks which ultimately shall manipulate the traveling inclination of that person. But challenges faced in collecting data for this survey showed that many factors like economic stability, mentality, future plans affect the prediction [73]. On the academic front, retrieving information profiles of researchers through the architecture of ArnetMiner structure has been explored in [74], where a united tagging approach has been used to haul out profiles of researchers from the internet without human intervention. This tool will also be helping in assimilating the existing information related to publications from online libraries into the network, helping to mold the scholarly network in its entirety and provide probing facilities to the academicians. The construction of the tool also provides a probabilistic solution to tackle redundancy problems raised due to similar names. With the advent of mobiles into our daily lives, Eagle et al. [75] guided us into the way through which information could be collected from mobile phones instead of self-reports. This process resulted in a vivid representation

of the actual dynamics between individuals and also allowed to study the progression of these associations with time. But privacy and security were to be considered as the most important criteria while using data with these types. It is in this context that breaching the security can be a matter of serious concern for even private profile users. Hence, users should be aware of the connections and participation in groups as they can offer to release private information intended to be kept covered [76]. A social networking middleware MobiClique has been presented in [77], that associates with other devices when they meet avoiding the requirement of a machine in between to exchange the messages. Messages can be easily dispersed within separate networks through this middleware. To develop a better analysis of social networks, a fresh idea of ModulLand was presented in [78], where linked parts in a community who have a selected centrality threshold have been termed as modules. ModuLand techniques comprise four steps—establishing the functions which control a node or link, followed by the creation of a community landscape, formatting the hills in the community and determination of higher level and resolving a hierarchy of superior intensity networks.

VIII. SPAM DETECTION IN SOCIAL NETWORKS

Based on multiple assumptions, an experiment was per- formed and results were demonstrated in [79], where dissem- ination of a fitness conduct was implemented through social networks. Results showed that more the span and number of clusters in the network, the more effect it had on helping to extend the health conduct. It was observed that clustered networks performed better in accepting the behavior and also in less time than random networks. As users may contribute any content or make use of the social platform to spread any unwanted content through the help of social networks, Gao et al. [80] proposed a method to measure and distin- guish spam promotions performed through pseudo accounts in online networks. Based on a data set of the messages from “Facebook” wall, it was observed that around 97% of the accounts were formed for the sole purpose of spreading spam in the networks and the spam messages are usually activated in the wee hours of the morning. Hence, it was more or less established that social networks were the target to spread spam and malware and exposure methods should be devised to detect online social spam. Another framework was projected in [81] that could be adapted by available social networking sites to restrict spams. This proposed framework had many advantages like identifying spam in the network and spreading information about the same through the entire network. It was observed that the model worked well in terms of accuracy for a large amount of data. Hence, comparing it to the rise of humongous data through social sites, it would prove helpful to efficiently detect spam in the sites. It was also anticipated that new social networking sites could prevent spam at an early age if this method was imbibed. Multiple classifiers were used whose results were passed through AND, OR, Majority voting and Bayesian strategies to detect spams. As every entity has a positive and negative aspect to it, social networks bear no

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

CHAKRABORTY et al.: SURVEY OF SENTIMENT ANALYSIS FROM SOCIAL MEDIA DATA 455

exception. An algorithm was also proposed for the same reason but which can be applied to small scale networks in [82], combining graph theory and machine learning concepts. The specialty of this method is that it requires only the graph topology construction to detect spams. On one hand, when social networking sites like “Facebook” are being utilized for malicious activities, on the other hand, it is also used for educational purposes as well [83]. As the teaching paradigm had started to shift from the conventional methods to more technologically enriched techniques, it was found that back in 2010, that 73% of the teachers had an account on Facebook compared to students, where 93% of whom had an account. It was also seen that college-goers had the same regularity of checking Facebook as well as emails, but faculties checked their emails more than checking the social networking site. But it was found out that both college students and faculties did not consider Facebook as a noteworthy means of sharing instructions, hence adhering to its name of being a social site rather than an educational site. Contrary to this, another work [84] employs a priori algorithm along with association rules to understand the involvement of Facebook in connecting students with each other versus students with teachers. It was found out that considering a few aspects like the number of times the students check their social media accounts and the time given to Facebook, students believe that Facebook is the best medium to access affluent information.

IX. INFLUENCE ANALYSIS IN SOCIAL NETWORKS

Along with online social networks, mobile social networks were slowly making its mark in the digital world. Accordingly, to spread information and to make people convince them to adapt those, significant individuals were to be found out. A greedy approach was incorporated in communities to mine the most significant nodes in a mobile social network [85]. Initially, the communities were selected and found out which will be used as a medium to disseminate the information followed by a dynamic programming methodology to consider dominant nodes. This method shows better results than the traditional greedy approach with minimum errors. To facilitate inter public system procedures, a framework was designed in [86] to find out the most number of user profiles that are being used by a person. To accomplish the task, parameters considered to find out whether two profiles belong to the same person, were allotted weights both physically and without human intervention, semantic comparisons were made and collective functions were used to take decisions. The results performed well compared to the existing traditional tech- niques. Another interesting service to take active participation in social networks is that of the micro-blogging site Twitter. Twitter seems to perform the opposite of normal networks of individuals. As there is a provision of 140 characters to be written on its platform, analysis of the tweets shows that most of the posts on Twitter are based on heading or trending news [87]. Like other social networking sites, prag- matic investigation was also performed on Twitter in search of prevalent computer-generated illicit systems. The internal social associations reveal that illicit accounts are publicly

linked and the influential hubs of criminal system have a tendency to follow the accounts more. Incorporation of the Mr. SPA [88] results in categorizing the accounts in divisions, namely, social butterflies, which haphazardly follow back-and- forth through any account on Twitter. But it is to be kept in mind that these accounts are very rare in nature. Another algorithm, Criminal Account Inference [88] is also applied in this work to gather more details in the field by analyzing their social accounts and by calculating the semantic harmo- nization amid accounts. A case study has been demonstrated in [89] where data were collected from various sites including homepages, blogs and Twitter against the backdrop of the National Assembly elections of Korea in 2010. The aim was to check if this means of social network was used purely to con- verse with fellow members and citizens or was it intentional. It was found out after experimentation that the politicians were more connected to their contemporaries on Twitter than with their citizens. A study conducted on 15 students who were gaining knowledge of imbibing social networks into data collected from the internet showed that beginners can easily grasp the concepts of SNA using NodeXL [90] within a short span of time. The intention was to facilitate beginners with prime quality material to learn online communities as per their preference [91]. Initially, the network analysis and visualization (NAV) tool was used and the SNA equipment was made available to the beginners. Another paper presents a methodical system which helps individuals visually surf in cases of hefty dimensional networks [92]. This approach helps to view social communities as well as their involved exchanges in an amalgamated view. The network structure is accumulated in the form of a social graph and their corresponding intercon- nected subgraphs are presented visually. As already discussed, identifying potential users in clusters or communities has been a scope for research always, and continuing with the same, [93] presents a technique to identify the significant person based on the post exchanges made upon a given theme. The first step involves preparing a graph that exhibits the relationship between the posts on the specific theme followed by the next step where a user graph is prepared which represents the influential users. Based on attributes collected from both the graphs, the most significant one is selected. Not only from the significant users, but the manners among online clients are also prejudiced by their peers using entropy assessment techniques as mentioned in [94]. To understand the nature of links present in social networks, a clustering along with a combined sorting algorithm has been used to differentiate between positive and negative links [95]. The method ensures that there is a social equilibrium between the clusters and preprocessing is done on the network before calculating the indication of the link. But against this concept, a study mentioned in [96] shows that as most of the previous works are based on supervised learning, the concealed ethics that initiate a specific behavior of social members are not con- sidered. Hence, a deep belief network-based technique based on unsupervised learning has been proposed to determine links in the network. Other than machine learning models, deep learning mechanisms have also been used for sculpting individual deeds as well as envisage health social networks.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

456 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

A study of how social networks could influence personal con- duct and how assumptions could be made to select parameters to predict the behavior of a person has been studied in detail in [97]. Factors such as personality inspiration, implied and precise social persuasion and ecological measures have been considered to foresee the expected movement of the individual. Methods have also been presented to efficiently fragment ego network by a genetic algorithm based K-means clustering structure combined with the information received from the behavior of the social circles in [98]. Hopeful results were achieved while studying the developing of associations with time on social networking sites like Facebook and Twitter. Considering the fact that Twitter data are extremely noisy and comes with loaded additional information, a Twitter-Network structure has been proposed to entirely model the network employing the hierarchical Poisson–Dirichlet processes for wordings and a Gaussian arbitrary task for social network modeling [99]. The role of social media sites like Twitter in cases of emergency is extremely accepted, but it has to be understood that the insufficiency of categorized data at the onset of emergency postpones the learning method and hence takes time to predict. This problem has been addressed in [100], where Convolution Neural Network has been utilized to detect and swiftly categorize important tweets at the time of disaster.

Social networks have resulted in a major impact on the socio-economic facets of the world and will continue to do so in the upcoming years. Out of the many crucial factors determining the concept of social networks, some of them are mentioned here along with a reference to the surveys conducted on these areas. To identify the significant user in a social network is a matter of serious concern to researchers worldwide. A survey of all the techniques used based on the construction of the network as well as the substance available in the network has been considered in [101]. The trend toward which the social data are pointing to, termed as opinion propagation, is also a matter of concern to the social world analytics. Cercel and Trausan-Matu [102] present a survey of the judgment broadcasting process in detail along with the scope of research in that field.

With the recent surge of data on the internet and the requirement of gaining meaningful insight from the huge amounts of data needs automatic processes for its analysis. Data mining techniques are appropriate to help with these issues. A list of possible procedures followed to understand social data based on various facets is available in [103].

X. SOCIAL MEDIA DATA USED IN SOCIAL NETWORKS

Social media facilitates the vast distribution of information through the virtual medium. The content on social media may vary from being personal data to documents, photos and official data. It was initially imbibed as a means of communicating to each other but as days are going by it has become entirely business based as it has the advantage to reach out to the entire population concurrently. A detailed study of the contributions made in the field of social media has been presented in Supplementary Table I [104]–[174].

XI. ANALYZING SENTIMENTS ON SOCIAL MEDIA

As it has been observed from Supplementary Table I, that social media mainly deals with comments from its users, it is very important to efficiently analyze them to help understand the emotions and opinions of the mass in general. The proxim- ity in which humans use social media platforms to present their views against each and every event leads to the necessity of exploring sentiments and try to resolve ways to evaluate them optimally. The prioritization of including the sentiments of public through online platforms has risen in the last few years.

Sentiment analysis is termed as the method in which invol- untary procedures are formulated to infer the sentiment of a text. The individual data which have been identified through computation help in forming planned insights to be utilized by judgment manufacturers. Owing to the rapid technological advancement and unbounded access to social media, sentiment analysis is constantly gaining popularity in the current business scenario. Sentiment analysis necessitates the use of handling natural language processing and its varied responsibilities like analysis of micro texts, detection of irony, anaphora detection, situation as well as feature identification.

Social media data from several spheres are collected from multiple means to extract the sentiments from texts. Pres- ence of multiple languages within the texts on social media, informal spellings due to message size constraints, spelling mistakes, grammatical and logical errors make the task of analyzing sentiments difficult on social media. In some cases, file demonstration in the form of N-gram graphs has been pre- sented to evaluate substance-based sentiment analysis [175]. It ably captures the sentiments of words by matching with a part of the string and by making no conjectures to the primary language.

A. Sentiment Analysis Techniques

There are two main methods of extracting sentiments, viz., lexicon-based approach and classification-based approach. Uti- lizing the former one, [176] shows the performance of Seman- tic Orientation Calculator (SO-CAL). Initially, emotion-related expressions (comprising different parts of speech) are used to evaluate valence shifters that are responsible in communicating the attitude of the text organization and finally the senti- ment is calculated. Two theories have been considered while calculating sentiments, first, that sentiments are independent of contexts and second that sentiments can be articulated through numbers. This work emphasizes on adding a hint of contradiction that reallocates the rate of the word in the presence of a negator. Considering health as one of the prime issues of our lives, a study [177] shows how analyzing small text messages collected from Twitter could help in understand- ing the emotions of people against respiratory tract infection A(HINI) vaccination. All the messages that were collected were connected to vaccinations as well as they provided the geographical position of the person behind the tweets. The accuracy of mining the sentiment was 84.29%, incorporating Naïve Bayes classifier [178] to identify the optimistic and pessimistic tweets and maximum entropy classifier to identify unbiased and inappropriate tweets [179]. Similarly, for the

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

CHAKRABORTY et al.: SURVEY OF SENTIMENT ANALYSIS FROM SOCIAL MEDIA DATA 457

case of disasters, a method has been presented in [180], in which visual analysis has been used to exploit location related tweets hearting on emotions of public in general. The outbreak of Ebola was considered in this case and queries like whether dissimilarity between several feeling categorizers can be exposed through the model and whether there remains optimistic attitude during disasters were addressed in the paper.

Another instance shows the usage of hybrid unsupervised approach along with language processing, dictionary-based methods and ontology procedures to create classifiers for emo- tions expressed on social media [181]. The Palavras software is used for the experimentation to determine the sentiments from Portuguese texts in social networks. The process comprises collecting all the corpus relating to a particular topic, stan- dardizing the texts, the relevant entity recognition, uncovering of the circumstance in which the entity was mined, selection of identifying characteristics, sentiment revealing, culminating of sentiment rates, accumulation of the retrieved data and finally analyzing them.

B. Opinion Mining

While sentiment analysis deals with judging the feeling within texts, Opinion Mining is said to be the process of judging the attitude of people about an entity. A detailed study of Opinion Mining on social media, including its problem definition in detail, categorization of sentiments in varied aspects, regulations related to designating an opinion, mining of features, extracting relative opinions and ultimately spam detection within opinions is available in [182]. Sentiments regarding the transportation service in Milan on Twitter were analyzed to provide enhanced itinerary services and also the sentiments could be used to adjust the services as per the preference of the traveler [183]. Tweets were collected both commencing and concluding from the respective travel agency, and the model designed in the paper analyzed the contents to categorize the occurrences as well as the categorization of attitudes about the transportation service. One of the straight- forward rule-based sentiment analysis methods is mentioned in [184], named VADER which results in a 0.96 accuracy compared to other methods. Factors relating to value and magnitude were considered to design a universal, valence based, manually crafted gold standard word list suited for microblog texts of limited characters.

C. Optical Sentiment Analysis

Not only in texts, but analysis for human emotions also travels through pictures as well. An optical sentiment analysis categorization approach [185] relying on deep convolution neural networks had been applied over a million labeled images collected from Flicker. This method, implemented on a novel deep learning framework Caffe, helped to determine the emotions portrayed in the images through the use of adjective- noun expressions involuntarily extracted from the images. This approach proved to perform well against conventional methods like support vector machine (SVM) categorization techniques. Another implementation of deep convolution neural frame- work has also been employed in [186], where the automatic

yet precise feature detection characteristic of deep learning has been utilized to find out the inadequately characterized images among half a million images from Flicker. These feeble labels are fine-tuned with the help of an advanced and domain shift approach which in turn hone the neural network. Other than this, a huge physically marked visual sentiment factual data via Amazon Mechanical Turk was created in this work. From both the works, it could be inferred that well-trained convolution neural networks outperformed existing optical sentiment analysis methods on social media. Although, Yuan et al. [187] are of the opinion that both texts and images are helpful while detecting the emotion of a user on social media. Primarily low-rated traits are dug out from the SUN database [188] and classified to produce 102 intermediately rated characteristics which are further utilized to envisage sentiments. On the other hand, the look wise emotions are predicted using eigenfaces. This entire method performs well to identify well-built positive and negative sentiments showing 82% accuracy after the complete execution. An example of this type of application is also found for microblogs as mentioned in [189], which along with the proposal of a framework, helps to acquire a minuscule and large-scale view of the details of the emotions extracted. Sentiment analysis on social media has been explored in many languages like Czech language as mentioned in [174], where administered machine learning techniques have been employed on document level emotion detection on 10 000 Facebook posts.

To facilitate industries facing a strong challenge against each other, a comparative analysis of what people are saying on social media has been modeled in [190], which provides an option to devise methods required for product-specific promotion strategies. The projected model has also been implemented into an investigative tool VOZIQ and further tested on five trading concerns to produce significant business reports. This structure aims to find out the most important companies related to a specific business type and provides a detailed report of their performances focusing on the vital features and ultimately paving the way for clever judgment building. Dynamic Architecture for Artificial Neural Networks (DAN2) [191] addresses such an issue, where the product efficacy is tested on a Starbucks related tweet data set with above 80% accuracy in all test cases. The feature in this case was evaluated through administered characteristic manufac- turing resulting in a feature with precisely seven dimensions. Three-class and five-class categorization of emotions were applied on the data set to provide effective insights to placid sentiments which might be required for crucial brand market- ing strategies. As mentioned earlier, emotion detection and its analysis have been performed on almost all languages of the world, one such application of segregating multiple languages has been projected in [192], where adjective-noun duo has been used to create a huge multiple language optical emotion system considering data of 12 languages from assorted origins.

D. Multiple Facets of Sentiments and Its Analysis

Other aspects of social media, where texts are posted along with images gained attention in [193], which facilitated both

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

458 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

solitary and multifaceted views on analyzing sentiments on social media. The multi-view sentiment analysis (MVSA) data set can be considered as a threshold which yielded positive relations between textual and optical data. Data from social media can also be utilized to find out the unfinished errands of an area of Sejong City, as explained in [194], where the sentiment analysis model has been assessed using the Naïve Bayes classifier at an accuracy of 75%. Mere availability of lexical source for sentiment analysis, for example, SentiWord- Net, is not enough, accuracy also matters when issues relating to trends of the public opinion are involved. In [195], SentiMI has been created which separates the individual examples from the object-oriented ones in SentiWordNet, and which extracts the parts of speech and evaluates the joint data for both optimistic and pessimistic terms. Out of all the methods men- tioned above, a novel approach was concentrated on in [196], where verbal communication processes were considered to detect the overall sentiment on a topic. Several disparities and unsymmetrical properties of unambiguous and implied expressions along with the straight effect that discourse pat- terns make on attitude power lead to the implementation of the above concept. Initially, the level of stir, enhancers and attenuators on emotion-laden words are studied followed by ways devised as to how the emotions can be expressed barred the use of emotion-filled words and finally it shows how the cohesiveness between phrases decide the total nature of a review. The first work on irony recognition was reported in [197], where sentiments were explored to find the sarcasm within texts. The use of personalized features and pretrained models for trait mining yielded high-performance results. The classification was done by applying CNN first followed by SVM [198].

An overall descriptive advancement in the field of sentiment analysis in the last decade has been described in the previous sections. Supplementary Table II [199]–[218] provides a tab- ular format comprising some major details about sentiment analysis.

In recent times, constant urge to compute accurate sentiment identification, researchers are extensively working on varied aspects. Not only that individual technique is explored daily, but hybridization techniques are also followed which turn out to be better result producers in case of perfect senti- ment detection. A visual sentiment analysis framework along with amalgamation of low and middle level characteristics of images is proposed in [219], resulting in a 9% rise in accuracy. Supervised learning methodologies like K-nearest neighbor (KNN) [220] and SVM [198] have been used to extract the emotions within images. Features are designed using the singular value decomposition (SVD) [221] and hue saturation intensity (HSI) [222] methods. Sentiment extraction from social media to create health-related awareness without the influence of medicines was found out in [223]. The most popular medicine category with respect to handling, value, cost and regularity of procurement could be derived due to a high precision acquired in the process and also it was also perceived that people pay attention to health problems pertaining to eye, skin and sexual well-being in the current age.

XII. CONCLUSION AND FUTURE SCOPE

It is high time that humans prioritize the unremitting rise of data from social networks. As almost all real-life complex problems ranging from biological to technological types can be represented by means of social networks, its challenges should also be addressed. Rumor detection [224], echoing of opinions, trends of online conversations leading to chaotic situations and community shaming [225], bring a change to preconceived notions, to understand that social prevalence in form of quantity of likes, shares and retweets. Features like finding the appropriate content and the right time to post are some of the important issues that need to be addressed in social networks before imbibing into the lives of humans completely. Even the detection of false comments should be addressed at the micro-level of social sites like Twitter to avoid unnecessary harassment from spams [226], [227]. Health issues of serious concern should be addressed in further research so that they make a strong impact on social media users. It would be fitting at this era if a unified linguistic model be prepared that understands the sentiments of the users while she/he is posting comments on the social media. To make the brain think like humans, the theme of object perception must be concentrated on to correctly understand the look and feel of any object as humans and simultaneously their behavioral patterns be studied from their responses to certain happenings [228]. Video analysis is a major research field that might gain popularity in the upcoming years. Influ- ential nodes which are responsible for sharing appropriate information must be constrained by some feature mining so that irrelevant information may not become viral within a fraction of second. Last but not least, personalization in terms of content portrayal on social media and social networks should be given utmost importance to enhance the quality of the web content. Effective methods to rate the comments of users in social sites for recommendation systems should be trodden upon [229]. Further, reducing ambiguity in the gallon of data being generated daily in these networks always provides ample scope for research. Importance should also be given into the amalgamation of literature and technology where consistency between the adaptation of original novels and its visual counterparts are being dealt with recently [230].

Authors of this article have recently trodden into the field of sentiment analysis making minor contributions in finding the near semantic meaning of a word in sentences as well as in documents [231], [232] and has presented a survey of all the works performed in native Indian languages [233].

This one-of-a-kind paper presents a detailed survey of social networks and its related terms. The works that have been accomplished relating to cluster, community and social networks have been described in its scope. This article mainly aims to bring out the shortfalls of the wide variety of papers making it easy for researches to apply sentiment analysis meth- ods after accumulating data from social media. The novelties have also been mentioned for papers in sentiment analysis to help scholars think of innovative ideas to train machines more efficiently in recognizing the opinion of the masses. Papers from the 20th century have been considered mainly following

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

CHAKRABORTY et al.: SURVEY OF SENTIMENT ANALYSIS FROM SOCIAL MEDIA DATA 459

the rise in the trend of social media data and its corresponding analysis. It is anticipated that more of the omnipresent deep learning mechanisms can be employed for social networks as they automatically detect features from patterns and hence will provide more structure to unstructured information without the minimum amount of human intervention.

REFERENCES

[1] M. Hoffman, D. Steinley, K. M. Gates, M. J. Prinstein, and M. J. Brusco, “Detecting clusters/communities in social networks,” Multivariate Behav. Res., vol. 53, no. 1, pp. 57–73, 2018, doi: 10.1080/ 00273171.2017.1391682.

[2] J. Leskovec, “Social media analytics: Tracking, modeling and predict- ing the flow of information through networks,” in Proc. 20th Int. Conf. Companion World Wide Web, Mar. 2011, pp. 277–278.

[3] F. Neri, C. Aliprandi, F. Capeci, M. Cuadros, and T. By, “Sentiment analysis on social media,” in Proc. IEEE/ACM Int. Conf. Adv. Social Netw. Anal. Mining, Aug. 2012, pp. 919–926.

[4] C. Bron and J. Kerbosch, “Algorithm 457: Finding all cliques of an undirected graph,” Commun. ACM, vol. 16, no. 9, pp. 575–577, 1973.

[5] R. L. Breiger, S. A. Boorman, and P. Arabie, “An algorithm for clustering relational data with applications to social network analysis and comparison with multidimensional scaling,” J. Math. Psychol., vol. 12, no. 3, pp. 328–383, 1975.

[6] S. Brin and L. Page, “The anatomy of a large-scale hypertextual Web search engine,” Comput. Netw. ISDN Syst., vol. 30, nos. 1–7, pp. 107–117, 1998.

[7] J. M. Kleinberg, “Authoritative sources in a hyperlinked environment,” J. ACM, vol. 46, no. 5, pp. 604–632, 1999.

[8] C. Canali, M. Colajanni, and R. Lancellotti, “Data acquisition in social networks: Issues and proposals,” in Proc. Int. Workshop Services Open Sources (SOS), Jun. 2011, pp. 1–12.

[9] B. Wellman, “The development of social network analysis: A study in the sociology of science,” Contemp. Sociol., vol. 37, no. 3, p. 221, 2008.

[10] D. Jensen and J. Neville, “Data mining in social networks,” in Dynamic Social Network Modeling and Analysis: Workshop Summary and Papers (Computer Science Department Faculty Publication Series). Amherst, MA, USA: Univ. of Massachusetts, 2003, pp. 287–302.

[11] C. S. Campbell, P. P. Maglio, A. Cozzi, and B. Dom, “Expertise identification using email communications,” in Proc. 12th Int. Conf. Inf. Knowl. Manage., Nov. 2003, pp. 528–531.

[12] J. Neville and D. Jensen, “Collective classification with relational dependency networks,” in Proc. Workshop Multi-Relational Data Min- ing (MRDM), 2003, p. 77.

[13] S. M. Van Dongen, “Graph clustering by flow simulation,” Ph.D. dissertation, Dept. Center Math. Comput. Sci., Univ. Utrecht, Utrecht, The Netherlands, 2000.

[14] U. Von Luxburg, “A tutorial on spectral clustering,” Statist. Comput., vol. 17, no. 4, pp. 395–416, 2007.

[15] S. C. Johnson, “Hierarchical clustering schemes,” Psychometrika, vol. 32, no. 3, pp. 241–254, 1967, doi: 10.1007/BF02289588.

[16] A. Shepitsen, J. Gemmell, B. Mobasher, and R. Burke, “Personalized recommendation in social tagging systems using hierarchical cluster- ing,” in Proc. ACM Conf. Recommender Syst., Oct. 2008, pp. 259–266.

[17] M. Roth et al., “Suggesting friends using the implicit social graph,” in Proc. 16th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, Jul. 2010, pp. 233–242.

[18] V. E. Lee, N. Ruan, R. Jin, and C. Aggarwal, “A survey of algorithms for dense subgraph discovery,” in Managing and Mining Graph Data. Boston, MA, USA: Springer, 2010, pp. 303–336.

[19] C. C. Yang and T. D. Ng, “Analyzing and visualizing Web opinion development and social interactions with density-based clustering,” IEEE Trans. Syst., Man, Cybern. A, Syst. Humans, vol. 41, no. 6, pp. 1144–1155, Mar. 2011.

[20] G.-J. Qi, C. C. Aggarwal, and T. S. Huang, “On clustering heteroge- neous social media objects with outlier links,” in Proc. 5th ACM Int. Conf. Web Search Data Mining, Feb. 2012, pp. 553–562.

[21] J. Shi, N. Mamoulis, D. Wu, and D. W. Cheung, “Density-based place clustering in geo-social networks,” in Proc. ACM SIGMOD Int. Conf. Manage. Data, Jun. 2014, pp. 99–110.

[22] S. Scellato, A. Noulas, R. Lambiotte, and C. Mascolo, “Socio-spatial properties of online location-based social networks,” in Proc. 5th Int. AAAI Conf. Weblogs Social Media, Jul. 2011, pp. 1–8.

[23] J. Tang, J. Sun, C. Wang, and Z. Yang, “Social influence analysis in large-scale networks,” in Proc. 15th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, Jun. 2009, pp. 807–816.

[24] J. A. Hartigan and M. A. Wong, “Algorithm AS 136: A k-means clustering algorithm,” J. Roy. Stat. Soc. C, Appl. Statist., vol. 28, no. 1, pp. 100–108, 1979.

[25] G. Schwarz, “Estimating the dimension of a model,” Ann. Statist., vol. 6, no. 2, pp. 461–464, 1978.

[26] C. Borgelt, “Efficient implementations of apriori and eclat,” in Proc. IEEE ICDM Workshop Frequent Itemset Mining Implement. (FIMI), Nov. 2003, pp. 1–10.

[27] S. Gole and B. Tidke, “Frequent itemset mining for big data in social media using ClustBigFIM algorithm,” in Proc. Int. Conf. Pervasive Comput. (ICPC), Jan. 2015, pp. 1–6.

[28] G. W. Flake, S. Lawrence, and C. L. Giles, “Efficient identification of Web communities,” in Proc. KDD, Aug. 2000, pp. 150–160.

[29] M. E. J. Newman, “Fast algorithm for detecting community structure in networks,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 69, no. 6, pp. 66133–66138, 2004.

[30] L. Donetti and M. A. Munoz, “Detecting network communities: A new systematic and efficient algorithm,” J. Stat. Mech., Theory Exp., vol. 2004, no. 10, 2004, Art. no. P10012.

[31] A. Clauset, M. E. Newman, and C. Moore, “Finding community structure in very large networks,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 70, no. 6, 2004, Art. no. 066111.

[32] C. P. Massen and J. P. K. Doye, “Identifying communities within energy landscapes,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 71, no. 4, 2005, Art. no. 046101.

[33] S. Fortunato, V. Latora, and M. Marchiori, “Method to find commu- nity structures based on information centrality,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 70, no. 5, 2004, Art. no. 056104.

[34] M. E. J. Newman and M. Girvan, “Finding and evaluating community structure in networks,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 69, no. 2, 2004, Art. no. 026113.

[35] F. Radicchi, C. Castellano, F. Cecconi, V. Loreto, and D. Parisi, “Defining and identifying communities in networks,” Proc. Nat. Acad. Sci. USA, vol. 101, no. 9, pp. 2658–2663, 2004.

[36] A. Clauset, “Finding local community structure in networks,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 72, no. 2, 2005, Art. no. 026132.

[37] G. Palla, I. Derényi, I. Farkas, and T. Vicsek, “Uncovering the overlap- ping community structure of complex networks in nature and society,” Nature, vol. 435, no. 7043, p. 814, 2005.

[38] H. Ino, M. Kudo, and A. Nakamura, “Partitioning of Web graphs by community topology,” in Proc. 14th Int. Conf. World Wide Web, May 2005, pp. 661–669.

[39] R. E. Gomory and T. C. Hu, “Multi-terminal network flows,” J. Soc. Ind. Appl. Math., vol. 9, no. 4, pp. 551–570, 1961.

[40] J. Reichardt and S. Bornholdt, “Statistical mechanics of community detection,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 74, no. 1, 2006, Art. no. 016110.

[41] M. B. Hastings, “Community detection as an inference problem,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 74, no. 3, 2006, Art. no. 035102.

[42] M. E. J. Newman, “Finding community structure in networks using the eigenvectors of matrices,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 74, no. 3, 2006, Art. no. 036104.

[43] F. Luo, J. Z. Wang, and E. Promislow, “Exploring local community structures in large networks,” Web Intell. Agent Syst., Int. J., vol. 6, no. 4, pp. 387–400, 2008.

[44] P. Pons and M. Latapy, “Computing communities in large networks using random walks,” in Proc. Int. Symp. Comput. Inf. Sci. Berlin, Germany: Springer, Oct. 2005, pp. 284–293.

[45] I. Vragović and E. Louis, “Network community structure and loop coefficient method,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 74, no. 1, 2006, Art. no. 016105.

[46] V. Latora and M. Marchiori, “Efficient behavior of small-world net- works,” Phys. Rev. Lett., vol. 87, Oct. 2001, Art. no. 198701.

[47] U. N. Raghavan, R. Albert, and S. Kumara, “Near linear time algorithm to detect community structures in large-scale networks,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 76, no. 3, 2007, Art. no. 036106.

[48] V. D. Blondel, J. L. Guillaume, R. Lambiotte, and E. Lefebvre, “Fast unfolding of communities in large networks,” J. Stat. Mech., Theory Exp., vol. 2008, no. 10, 2008, Art. no. P10008.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

460 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

[49] C. Pizzuti, “Ga-net: A genetic algorithm for community detection in social networks,” in Proc. Int. Conf. Parallel Problem Solving Nature. Berlin, Germany: Springer, Sep. 2008, pp. 1081–1090.

[50] I. X. Y. Leung, P. Hui, P. Lio, and J. Crowcroft, “Towards real- time community detection in large networks,” Phys. Rev. E, Stat. Phys. Plasmas Fluids Relat. Interdiscip. Top., vol. 79, no. 6, 2009, Art. no. 066107.

[51] S. Gregory, “Finding overlapping communities in networks by label propagation,” New J. Phys., vol. 12, no. 10, 2010, Art. no. 103018.

[52] J. Chen, O. Zaïane, and R. Goebel, “A visual data mining approach to find overlapping communities in networks,” in Proc. Int. Conf. Adv. Social Netw. Anal. Mining, Jul. 2009, pp. 338–343.

[53] X. Wang, L. Tang, H. Gao, and H. Liu, “Discovering overlapping groups in social media,” in Proc. IEEE Int. Conf. Data Mining, Dec. 2010, pp. 569–578.

[54] B. Wang, X. Wang, C. Sun, B. Liu, and L. Sun, “Modeling semantic relevance for question-answer pairs in Web social communities,” in Proc. 48th Annu. Meeting Assoc. Comput. Linguistics, Jul. 2010, pp. 1230–1238.

[55] S. Parthasarathy, Y. Ruan, and V. Satuluri, “Community discovery in social networks: Applications, methods and emerging trends,” in Social Network Data Analytics. Boston, MA, USA: Springer, 2011, pp. 79–113.

[56] S. Papadopoulos, Y. Kompatsiaris, A. Vakali, and P. Spyridonos, “Com- munity detection in social media,” Data Mining Knowl. Discovery, vol. 24, no. 3, pp. 515–554, May 2012.

[57] M. Plantié and M. Crampes, “Survey on social community detection,” in Social Media Retrieval. London, U.K.: Springer, 2013, pp. 65–85.

[58] J. Xie, S. Kelley, and B. K. Szymanski, “Overlapping community detection in networks: The state-of-the-art and comparative study,” ACM Comput. Surv., vol. 45, no. 4, p. 43, 2013.

[59] G. J. Qi, C. C. Aggarwal, and T. Huang, “Community detection with edge content in social media networks,” in Proc. IEEE 28th Int. Conf. Data Eng., Apr. 2012, pp. 534–545.

[60] J. Tang, X. Wang, and H. Liu, “Integrating social media data for community detection,” in Modeling and Mining Ubiquitous Social Media. Berlin, Germany: Springer, 2011, pp. 1–20.

[61] L. Weng, F. Menczer, and Y. Y. Ahn, “Virality prediction and com- munity structure in social networks,” Sci. Rep., vol. 3, Aug. 2013, Art. no. 2522.

[62] S. P. Borgatti, “Centrality and network flow,” Social Netw., vol. 27, no. 1, pp. 55–71, 2005.

[63] S. P. Borgatti, K. M. Carley, and D. Krackhardt, “On the robustness of centrality measures under conditions of imperfect data,” Social Netw., vol. 28, no. 2, pp. 124–136, 2006.

[64] K. W. Axhausen, “Social networks, mobility biographies, and travel: Survey challenges,” Environ. Planning B, Planning Des., vol. 35, no. 6, pp. 981–996, 2008.

[65] Y. Zhou, E. Reid, J. Qin, H. Chen, and G. Lai, “US domestic extremist groups on the Web: Link and content analysis,” IEEE Intell. Syst., vol. 20, no. 5, pp. 44–51, Sep. 2005.

[66] A. Perer and B. Shneiderman, “Balancing systematic and flexible exploration of social networks,” IEEE Trans. Vis. Comput. Graphics, vol. 12, no. 5, pp. 693–700, Nov. 2006.

[67] K. L. Fiori, T. C. Antonucci, and K. S. Cortina, “Social network typologies and mental health among older adults,” J. Gerontol. B, Psychol. Sci. Social Sci., vol. 61, no. 1, pp. P25–P32, 2006.

[68] L. Backstrom, D. Huttenlocher, J. Kleinberg, and L. X. , “Group for- mation in large social networks: Membership, growth, and evolution,” in Proc. 12th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, Aug. 2006, pp. 44–54.

[69] J. Voß, “Measuring wikipedia,” in Proc. ISSI 10th Int. Conf. Int. Soc. Scientometrics Informetrics, 2005.

[70] M. A. Sicilia, N. T. Korfiatis, M. Poulos, and G. Bokos, “Evalu- ating authoritative sources using social networks: An insight from Wikipedia,” Online Inf. Rev., vol. 30, no. 3, pp. 252–262, 2006.

[71] H. E. Aldrich and P. H. Kim, “Small worlds, infinite possibilities? How social networks affect entrepreneurial team formation and search,” Strategic Entrepreneurship J., vol. 1, nos. 1–2, pp. 147–165, 2007.

[72] C. Kiss and M. Bichler, “Identification of influencers—Measuring influence in customer networks,” Decis. Support Syst., vol. 46, no. 1, pp. 233–253, 2008.

[73] K. Okamoto, W. Chen, and X.-Y. Li, “Ranking of closeness centrality for large-scale social networks,” in Proc. Int. Workshop Frontiers Algorithmics. Berlin, Germany: Springer, Jun. 2008, pp. 186–195.

[74] J. Tang, J. Zhang, L. Yao, J. Li, L. Zhang, and Z. Su, “Arnetminer: Extraction and mining of academic social networks,” in Proc. 14th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, Aug. 2008, pp. 990–998.

[75] N. Eagle, A. S. Pentland, and D. Lazer, “Inferring friendship network structure by using mobile phone data,” Proc. Nat. Acad. Sci. USA, vol. 106, no. 36, pp. 15274–15278, 2009.

[76] E. Zheleva and L. Getoor, “To join or not to join: The illusion of privacy in social networks with mixed public and private user profiles,” in Proc. 18th Int. Conf. World Wide Web, Apr. 2009, pp. 531–540.

[77] A. K. Pietiläinen, E. Oliver, J. LeBrun, G. Varghese, and C. Diot, “MobiClique: Middleware for mobile social networking,” in Proc. 2nd ACM Workshop Online Social Netw., Aug. 2009, pp. 49–54.

[78] I. A. Kovács, R. Palotai, M. S. Szalay, and P. Csermely, “Community landscapes: An integrative approach to determine overlapping network module hierarchy, identify key nodes and predict network dynamics,” PLoS ONE, vol. 5, no. 9, 2010, Art. no. e12528.

[79] D. Centola, “The spread of behavior in an online social network experiment,” Science, vol. 329, no. 5996, pp. 1194–1197, 2010.

[80] H. Gao, J. Hu, C. Wilson, Z. Li, Y. Chen, and B. Y. Zhao, “Detect- ing and characterizing social spam campaigns,” in Proc. 10th ACM SIGCOMM Conf. Internet Meas., Nov. 2010, pp. 35–47.

[81] D. Wang, D. Irani, and C. Pu, “A social-spam detection framework,” in Proc. 8th Annu. Collaboration, Electron. Messaging, Anti-Abuse Spam Conf., Sep. 2011, pp. 46–54.

[82] M. Fire, G. Katz, and Y. Elovici, “Strangers intrusion detection- detecting spammers and fake profiles in social networks based on topology anomalies,” Hum. J., vol. 1, no. 1, pp. 26–39, 2012.

[83] M. D. Roblyer, M. McDaniel, M. Webb, J. Herman, and J. V. Witty, “Findings on Facebook in higher education: A comparison of college faculty and student uses and perceptions of social networking sites,” Internet Higher Educ., vol. 13, no. 3, pp. 134–140, 2010.

[84] A. S. Bozkır, S. G. Mazman, and E. A. Sezer, “Identification of user patterns in social networks by data mining techniques: Facebook case,” in Proc. Int. Symp. Inf. Manage. Changing World. Berlin, Germany: Springer, Sep. 2010, pp. 145–153.

[85] Y. Wang, G. Cong, G. Song, and K. Xie, “Community-based greedy algorithm for mining top-k influential nodes in mobile social networks,” in Proc. 16th ACM SIGKDD Int. Conf. Knowl. Discovery Data Mining, Jul. 2010, pp. 1039–1048.

[86] E. Raad, R. Chbeir, and A. Dipanda, “User profile matching in social networks,” in Proc. 13th Int. Conf. Netw.-Based Inf. Syst., Sep. 2010, pp. 297–304.

[87] H. Kwak, C. Lee, H. Park, and S. Moon, “What is Twitter, a social network or a news media?” in Proc. 19th Int. Conf. World Wide Web, Apr. 2010, pp. 591–600.

[88] C. Yang, R. Harkreader, J. Zhang, S. Shin, and G. Gu, “Analyzing spammers’ social networks for fun and profit: A case study of cyber criminal ecosystem on Twitter,” in Proc. 21st Int. Conf. World Wide Web, Apr. 2012, pp. 71–80.

[89] C.-L. Hsu and H. W. Park, “Mapping online social networks of Korean politicians,” Government Inf. Quart., vol. 29, no. 2, pp. 169–181, 2012.

[90] D. Hansen, B. Shneiderman, and M. A. Smith, Analyzing Social Media Networks With NodeXL: Insights From a Connected World. San Mateo, CA, USA: Morgan Kaufmann, 2010.

[91] D. L. Hansen et al., “Do you know the way to SNA?: A process model for analyzing and visualizing social media network data,” in Proc. Int. Conf. Social Inf., Dec. 2012, pp. 304–313.

[92] F. Zhao and A. K. H. Tung, “Large scale cohesive subgraphs discovery for social network visual analysis,” Proc. VLDB Endowment, vol. 6, no. 2, pp. 85–96, 2012.

[93] B. Sun and V. T. Ng, “Identifying influential users by their postings in social networks,” in Ubiquitous Social Media Analysis. Berlin, Germany: Springer, 2012, pp. 128–151.

[94] S. He, X. Zheng, D. Zeng, K. Cui, Z. Zhang, and C. Luo, “Identifying peer influence in online social networks using transfer entropy,” in Proc. Pacific-Asia Workshop Intell. Security Inform. Berlin, Germany: Springer, Aug. 2013, pp. 47–61.

[95] A. Javari and M. Jalili, “Cluster-based collaborative filtering for sign prediction in social networks with positive and negative links,” ACM Trans. Intell. Syst. Technol., vol. 5, no. 2, p. 24, 2014.

[96] F. Liu, B. Liu, C. Sun, M. Liu, and X. Wang, “Deep belief network- based approaches for link prediction in signed social networks,” Entropy, vol. 17, no. 4, pp. 2140–2169, 2015.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

CHAKRABORTY et al.: SURVEY OF SENTIMENT ANALYSIS FROM SOCIAL MEDIA DATA 461

[97] N. Phan, D. Dou, B. Piniewski, and D. Kil, “Social restricted boltzmann machine: Human behavior prediction in health social networks,” in Proc. IEEE/ACM Int. Conf. Adv. Social Netw. Anal. Mining, Aug. 2015, pp. 424–431.

[98] V. Agarwal and K. K. Bharadwaj, “Predicting the dynamics of social circles in ego networks using pattern analysis and GA K-means clustering,” Wiley Interdiscipl. Rev., Data Mining Knowl. Discovery, vol. 5, no. 3, pp. 113–141, 2015.

[99] K. W. Lim, C. Chen, and W. Buntine, “Twitter-network topic model: A full Bayesian treatment for social network and text mod- eling,” 2016, arXiv:1609.06791. [Online]. Available: https://arxiv.org/ abs/1609.06791

[100] D. T. Nguyen, K. A. A. Mannai, S. Joty, H. Sajjad, M. Imran, and P. Mitra, “Robust classification of crisis-related data on social networks using convolutional neural networks,” in Proc. 11th Int. AAAI Conf. Web Social Media, May 2017, pp. 1–4.

[101] R. Rabade, N. Mishra, and S. Sharma, “Survey of influential user identification techniques in online social networks,” in Recent Advances in Intelligent Informatics. Cham, Switzerland: Springer, 2014, pp. 359–370.

[102] D. C. Cercel and T.-M. Stefan, “Opinion propagation in online social networks: A survey,” in Proc. 4th Int. Conf. Web Intell., Mining, Semantics (WIMS), Jun. 2014, p. 11.

[103] M. Adedoyin-Olowe, M. M. Gaber, and F. Stahl, “A survey of data mining techniques for social media analysis,” 2013, arXiv:1312.4617. [Online]. Available: https://arxiv.org/abs/1312.4617

[104] J. Bian, Y. Liu, E. Agichtein, and H. Zha, “Finding the right facts in the crowd: Factoid question answering over social media,” in Proc. 17th Int. Conf. World Wide Web, Apr. 2008, pp. 467–476.

[105] E. Agichtein, C. Castillo, D. Donato, A. Gionis, and G. Mishne, “Finding high-quality content in social media,” in Proc. Int. Conf. Web Search Data Mining, Feb. 2008, pp. 183–194.

[106] K. Denecke and W. Nejdl, “How valuable is medical social media data? Content analysis of the medical Web,” Inf. Sci., vol. 179, no. 12, pp. 1870–1880, 2009.

[107] F. Chen, P. N. Tan, and A. K. Jain, “A co-classification framework for detecting Web spam and spammers in social media Web sites,” in Proc. 18th ACM Conf. Inf. Knowl. Manage., Nov. 2009, pp. 1807–1810.

[108] B. Markines, C. Cattuto, and F. Menczer, “Social spam detection,” in Proc. 5th Int. Workshop Adversarial Inf. Retr. Web, Apr. 2009, pp. 41–48.

[109] H. Sayyadi, M. Hurst, and A. Maykov, “Event detection and tracking in social streams,” in Proc. 3rd Int. AAAI Conf. Weblogs Social Media, Mar. 2009, pp. 1–4.

[110] H. Becker, M. Naaman, and L. Gravano, “Event identification in social media,” in Proc. WebDB, Jun. 2009, pp. 1–6.

[111] R. W. Lariscy, E. J. Avery, K. D. Sweetser, and P. Howes, “An examination of the role of online social media in journalists’ source mix,” Public Relations Rev., vol. 35, no. 3, pp. 314–316, 2009.

[112] C. M. De, Y. R. Lin, H. Sundaram, K. S. Candan, L. Xie, and A. Kelliher, “How does the data sampling strategy impact the discovery of information diffusion in social media?” in Proc. 4th Int. AAAI Conf. Weblogs Social Media, May 2010, pp. 1–8.

[113] H. Becker, M. Naaman, and L. Gravano, “Learning similarity metrics for event identification in social media,” in Proc. 3rd ACM Int. Conf. Web Search Data Mining, Feb. 2010, pp. 291–300.

[114] O. Oh, K. H. Kwon, and H. R. Rao, “An exploration of social media in extreme events: Rumor theory and Twitter during the haiti earthquake 2010,” in Proc. ICIS, vol. 231, Dec. 2010, pp. 7332–7336.

[115] M. Naaman, J. Boase, and C. H. Lai, “Is it really about me?: Message content in social awareness streams,” in Proc. ACM Conf. Comput. Supported Cooperat. Work, Feb. 2010, pp. 189–192.

[116] S. Asur and B. A. Huberman, “Predicting the future with social media,” in Proc. IEEE/WIC/ACM Int. Conf. Web Intell. Intell. Agent Technol., vol. 1, Aug. 2010, pp. 492–499.

[117] N. Diakopoulos, M. Naaman, and F. Kivran-Swaine, “Diamonds in the rough: Social media visual analytics for journalistic inquiry,” in Proc. IEEE Symp. Visual Anal. Sci. Technol., Oct. 2010, pp. 115–122.

[118] Z. Xiang and U. Gretzel, “Role of social media in online travel information search,” Tourism Manage., vol. 31, no. 2, pp. 179–188, 2010.

[119] W. C. Jacobsen and R. Forste, “The wired generation: Academic and social outcomes of electronic media use among University students,” Cyberpsychol., Behav., Social Netw., vol. 14, no. 5, pp. 275–280, 2011.

[120] K. Lee, J. Caverlee, Z. Cheng, and D. Z. Sui, “Content-driven detection of campaigns in social media,” in Proc. 20th ACM Int. Conf. Inf. Knowl. Manage., Oct. 2011, pp. 551–556.

[121] D. M. Romero, W. Galuba, S. Asur, and B. A. Huberman, “Influence and passivity in social media,” in Proc. Joint Eur. Conf. Mach. Learn. Knowl. Discovery Databases. Berlin, Germany: Springer, Sep. 2011, pp. 18–33.

[122] N. Michaelidou, N. T. Siamagka, and G. Christodoulides, “Usage, barriers and measurement of social media marketing: An exploratory investigation of small and medium B2B brands,” Ind. Marketing Manage., vol. 40, no. 7, pp. 1153–1159, 2011.

[123] K. Sedereviciute and C. Valentini, “Towards a more holistic stake- holder analysis approach. Mapping known and undiscovered stake- holders from social media,” Int. J. Strategic Commun., vol. 5, no. 4, pp. 221–239, 2011.

[124] F. Cheong and C. Cheong, “Social media data mining: A social network analysis of tweets during the 2010–2011 Australian floods,” PACIS, vol. 11, p. 46, Jul. 2011.

[125] E. Clark and K. Araki, “Text normalization in social media: Progress, problems and applications for a pre-processing system of casual English,” Procedia-Social Behav. Sci., vol. 27, pp. 2–11, Jan. 2011.

[126] C. C. Yang, L. Jiang, H. Yang, and X. Tang, “Detecting signals of adverse drug reactions from health consumer contributed content in social media,” in Proc. ACM SIGKDD Workshop Health Inf., Aug. 2012, pp. 1–8.

[127] J. Tang and H. Liu, “Feature selection with linked data in social media,” in Proc. SIAM Int. Conf. Data Mining, Apr. 2012, pp. 118–128.

[128] L. M. Aiello, A. Barrat, R. Schifanella, C. Cattuto, B. Markines, and F. Menczer, “Friendship prediction and homophily in social media,” ACM Trans. Web, vol. 6, no. 2, p. 9, 2012.

[129] B. Han, P. Cook, and T. Baldwin, “Geolocation prediction in social media data by finding location indicative words,” in Proc. COLING, Dec. 2012, pp. 1045–1062.

[130] G. Ver Steeg and A. Galstyan, “Information transfer in social media,” in Proc. 21st Int. Conf. World Wide Web, Apr. 2012, pp. 509–518.

[131] C. Sengstock and M. Gertz, “Latent geographic feature extraction from social media,” in Proc. 20th Int. Conf. Adv. Geograph. Inf. Syst., Nov. 2012, pp. 149–158.

[132] J. Cranshaw, R. Schwartz, J. Hong, and N. Sadeh, “The livehoods project: Utilizing social media to understand the dynamics of a city,” in Proc. 6th Int. AAAI Conf. Weblogs Social Media, May 2012, pp. 1–8.

[133] S. A. Moorhead, D. E. Hazlett, L. Harrison, J. K. Carroll, A. Irwin, and C. Hoving, “A new dimension of health care: Systematic review of the uses, benefits, and limitations of social media for health communication,” J. Med. Internet Res., vol. 15, no. 4, p. e85, 2013.

[134] H. A. Schwartz et al., “Personality, gender, and age in the language of social media: The open-vocabulary approach,” PLoS ONE, vol. 8, no. 9, p. e73791, 2013.

[135] M. Kandias, V. Stavrou, N. Bozovic, and D. Gritzalis, “Proactive insider threat detection through social media: The YouTube case,” in Proc. 12th ACM Workshop Privacy Electron. Soc., Nov. 2013, pp. 261–266.

[136] M. Thelwall, “The Heart and soul of the Web? Sentiment strength detection in the social Web with SentiStrength,” in Cyberemotions. Cham, Switzerland: Springer, 2013, pp. 119–134.

[137] F. Morstatter, J. Pfeffer, H. Liu, and K. M. Carley, “Is the sample good enough? Comparing data from twitter’s streaming api with Twitter’s firehose,” in Proc. 7th Int. AAAI Conf. Weblogs Social Media, Jun. 2013, pp. 1–9.

[138] Y. Hu, S. D. Farnham, and A. Monroy-Hernández, “Whoo. ly: Facil- itating information seeking for hyperlocal communities using social media,” in Proc. SIGCHI Conf. Hum. Factors Comput. Syst., Apr. 2013, pp. 3481–3490.

[139] F. Mitzlaff, M. Atzmueller, G. Stumme, and A. Hotho, “Semantics of user interaction in social media,” in Complex Networks IV. Berlin, Germany: Springer, 2013, pp. 13–25.

[140] M. De Choudhury, M. Gamon, S. Counts, and E. Horvitz, “Predicting depression via social media,” in Proc. 7th Int. AAAI Conf. Weblogs Social Media, Jun. 2013, pp. 1–10.

[141] W. He, S. Zha, and L. Li, “Social media competitive analysis and text mining: A case study in the pizza industry,” Int. J. Inf. Manage., vol. 33, no. 3, pp. 464–472, 2013.

[142] K. Casler, L. Bickel, and E. Hackett, “Separate but equal? A compar- ison of participants and data gathered via Amazon’s MTurk, social media, and face-to-face behavioral testing,” Comput. Hum. Behav., vol. 29, pp. 2156–2160, Nov. 2013.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

462 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

[143] T. Baldwin, P. Cook, M. Lui, A. MacKinlay, and L. Wang, “How noisy social media text, how diffrnt social media sources?” in Proc. 6th Int. Joint Conf. Natural Lang. Process., Oct. 2013, pp. 356–364.

[144] A. Majid, L. Chen, G. Chen, H. T. Mirza, and I. W. J. Hussain, “A context-aware personalized travel recommendation system based on geotagged social media data mining,” Int. J. Geograph. Inf. Sci., vol. 27, no. 4, pp. 662–684, 2013.

[145] M. Khan, M. Dickinson, and S. Kübler, “Does size matter? Text and grammar revision for parsing social media data,” in Proc. Workshop Lang. Anal. Social Media, 2013, pp. 1–10.

[146] L. Derczynski, A. Ritter, S. Clark, and K. Bontcheva, “Twitter part- of-speech tagging for all: Overcoming sparse and noisy data,” in Proc. Int. Conf. Recent Adv. Natural Lang. Process. (RANLP), 2013, pp. 198–206.

[147] Y. Zhang and M. Pennacchiotti, “Recommending branded products from social media,” in Proc. 7th ACM Conf. Recommender Syst., Oct. 2013, pp. 77–84.

[148] X. Ma, H. Wang, H. Li, J. Liu, and H. Jiang, “Exploring sharing patterns for video recommendation on YouTube-like social media,” Multimedia Syst., vol. 20, no. 6, pp. 675–691, 2014.

[149] S. Pei, L. Muchnik, J. S. Andrade, Jr., Z. Zheng, and H. A. Makse, “Searching for superspreaders of information in real-world social media,” Sci. Rep., vol. 4, Jul. 2014, Art. no. 5547.

[150] M. Abdul-Mageed, M. Diab, and S. Kübler, “SAMAR: Subjectivity and sentiment analysis for Arabic social media,” Comput. Speech Lang., vol. 28, no. 1, pp. 20–37, 2014.

[151] S. Mei, H. Li, J. Fan, X. Zhu, and C. R. Dyer, “Inferring air pollution by sniffing social media,” in Proc. IEEE/ACM Int. Conf. Adv. Social Netw. Anal. Mining, Aug. 2014, pp. 534–539.

[152] J. Choi et al., “The placing task: A large-scale geo-estimation challenge for social-media videos and images,” in Proc. 3rd ACM Multimedia Workshop Geotagging Appl. Multimedia, Nov. 2014, pp. 27–31.

[153] M. Sap et al., “Developing age and gender predictive lexica over social media,” in Proc. Conf. Empirical Methods Natural Lang. Process. (EMNLP), Oct. 2014, pp. 1146–1151.

[154] I. C. L. Memon, A. Majid, M. Lv, and I. C. G. Hussain, “Travel recommendation using geo-tagged photos in social media for tourist,” Wireless Pers. Commun., vol. 80, no. 4, pp. 1347–1362, 2015.

[155] D. Bamman, J. Eisenstein, and T. Schnoebelen, “Gender identity and lexical variation in social media,” J. Sociolinguistics, vol. 18, no. 2, pp. 135–160, 2014.

[156] H. Lin et al., “User-level psychological stress detection from social media using deep neural network,” in Proc. 22nd ACM Int. Conf. Multimedia, Nov. 2014, pp. 507–516.

[157] R. Nivedha and N. Sairam, “A machine learning based classification for social media messages,” Indian J. Sci. Technol., vol. 8, no. 16, p. 1, 2015.

[158] A. P. López-Monroy, M. Montes-y-Gómez, H. J. Escalante, L. Villaseñor-Pineda, and E. Stamatatos, “Discriminative subprofile- specific representations for author profiling in social media,” Knowl.- Based Syst., vol. 89, pp. 134–147, Nov. 2015.

[159] P. Barberá, “Birds of the same feather tweet together: Bayesian ideal point estimation using Twitter data,” Political Anal., vol. 23, no. 1, pp. 76–91, 2015.

[160] C. Cao and J. Caverlee, “Detecting spam URLs in social media via behavioral analysis,” in Proc. Eur. Conf. Inf. Retr. Cham, Switzerland: Springer, Mar. 2015, pp. 703–714.

[161] T. Ma et al., “Social network and tag sources based augmenting collaborative recommender system,” IEICE Trans. Inf. Syst., vol. 98, no. 4, pp. 902–910, 2015.

[162] M. Santillana, A. T. Nguyen, M. Dredze, M. J. Paul, E. O. Nsoesie, and J. S. Brownstein, “Combining search, social media, and traditional data sources to improve influenza surveillance,” PLoS Comput. Biol., vol. 11, no. 10, 2015, Art. no. e1004513.

[163] F. Gelli, T. Uricchio, M. Bertini, A. Del Bimbo, and S.-F. Chang, “Image popularity prediction in social media using sentiment and context features,” in Proc. 23rd ACM Int. Conf. Multimedia, Oct. 2015, pp. 907–910.

[164] W. Chen, C. K. Yeo, C. T. Lau, and B. S. Lee, “Real-time Twitter content polluter detection based on direct features,” in Proc. 2nd Int. Conf. Inf. Sci. Security (ICISS), Dec. 2015, pp. 1–4.

[165] J. Ma et al., “Detecting rumors from microblogs with recurrent neural networks,” in Proc. IJCAI, Jul. 2016, pp. 3818–3824.

[166] V. Kulkarni, B. Perozzi, and S. Skiena, “Freshman or fresher? Quanti- fying the geographic variation of language in online social media,” in Proc. 10th Int. AAAI Conf. Web Social Media, Mar. 2016, pp. 1–6.

[167] B. Bischke, D. Borth, C. Schulze, and A. Dengel, “Contextual enrich- ment of remote-sensed events with social media streams,” in Proc. 24th ACM Int. Conf. Multimedia, Oct. 2016, pp. 1077–1081.

[168] L. Liu, D. Preotiuc-Pietro, Z. R. Samani, M. E. Moghaddam, and L. Ungar, “Analyzing personality through social media profile picture choice,” in Proc. 10th Int. AAAI Conf. Web Social Media, Mar. 2016, pp. 1–10.

[169] J. P. Mazer et al., “Communication in the face of a school crisis: Examining the volume and content of social media mentions during active shooter incidents,” Comput. Hum. Behav., vol. 53, pp. 238–248, Dec. 2015.

[170] B. Dhingra, Z. Zhou, D. Fitzpatrick, M. Muehl, and W. W. Cohen, “Tweet2vec: Character-based distributed representations for social media,” 2016, arXiv:1605.03481. [Online]. Available: https://arxiv. org/abs/1605.03481

[171] L. Song, R. Y. K. Lau, R. C. W. Kwok, K. Mirkovski, and W. Dou, “Who are the spoilers in social media marketing? Incremental learning of latent semantics for social spam detection,” Electron. Commerce Res., vol. 17, no. 1, pp. 51–81, 2017.

[172] F. Krebs, B. Lubascher, T. Moers, P. Schaap, and G. Spanakis, “Social emotion mining techniques for Facebook posts reaction pre- diction,” 2017, arXiv:1712.03249. [Online]. Available: https://arxiv. org/abs/1712.03249

[173] J. Su, A. Shukla, S. Goel, and A. Narayanan, “De-anonymizing Web browsing data with social networks,” in Proc. 26th Int. Conf. World Wide Web, Apr. 2017, pp. 1261–1269.

[174] H. Aleid et al., “Framework to classify and analyze social media content,” Social Netw., vol. 7, no. 2, p. 79, 2018.

[175] F. Aisopos, G. Papadakis, and T. Varvarigou, “Sentiment analysis of social media content using N-Gram graphs,” in Proc. 3rd ACM SIGMM Int. Workshop Social Media, Nov. 2011, pp. 9–14.

[176] M. Taboada, J. Brooke, M. Tofiloski, K. Voll, and M. Stede, “Lexicon- based methods for sentiment analysis,” Comput. Linguistics, vol. 37, no. 2, pp. 267–307, 2011.

[177] I. Habernal, T. Ptácek, and J. Steinberger, “Reprint of ‘supervised sentiment analysis in Czech social media,” Inf. Process. Manage., vol. 51, no. 4, pp. 532–546, 2015.

[178] K. P. Murphy, “Naive Bayes classifiers,” Univ. Brit. Columbia, Vancouver, BC, Canada, Tech. Rep., 2006, p. 60, vol. 18.

[179] M. Salathé and S. Khandelwal, “Assessing vaccination sentiments with online social media: Implications for infectious disease dynamics and control,” PLoS Comput. Biol., vol. 7, no. 10, 2011, Art. no. e1002199.

[180] Y. Lu, X. Hu, F. Wang, S. Kumar, H. Liu, and R. Maciejewski, “Visualizing social media sentiment in disaster scenarios,” in Proc. 24th Int. Conf. World Wide Web, May 2015, pp. 1211–1215.

[181] R. Baracho, M. Bax, L. G. F. Ferreira, and G. Ca’ires Silva, Sentiment Analysis in Social Networks. Stanford, CA, USA: Association for the Advancement of Artificial Intelligence, 2012.

[182] B. Liu and L. Zhang, “A survey of opinion mining and sentiment analysis,” in Mining Text Data. Boston, MA, USA: Springer, 2012, pp. 415–463.

[183] A. Candelieri and F. Archetti, “Detecting events and sentiment on Twitter for improving urban mobility,” in Proc. ESSEM AAMAS, May 2015, pp. 106–115.

[184] C. J. Hutto and E. Gilbert, “VADER: A parsimonious rule-based model for sentiment analysis of social media text,” in Proc. 8th Int. AAAI Conf. Weblogs Social Media, May 2014, pp. 1–10.

[185] T. Chen, D. Borth, T. Darrell, and S. F. Chang, “Deepsen- tiBank: Visual sentiment concept classification with deep convolu- tional neural networks,” 2014, arXiv:1410.8586. [Online]. Available: https://arxiv.org/abs/1410.8586

[186] Q. You, J. Luo, H. Jin, and J. Yang, “Robust image sentiment analysis using progressively trained and domain transferred deep networks,” in Proc. 29th AAAI Conf. Artif. Intell., Feb. 2015, pp. 1–8.

[187] J. Yuan, Q. You, and J. Luo, “Sentiment analysis using social multi- media,” in Multimedia Data Mining and Analytics. Cham, Switzerland: Springer, 2015, pp. 31–59.

[188] A. Hanjalic, C. Kofler, and M. Larson, “Intent and its discontents: The user at the wheel of the online video search engine,” in Proc. 20th ACM Int. Conf. Multimedia, Oct. 2012, pp. 1239–1248.

[189] D. Cao, R. Ji, D. Lin, and S. Li, “A cross-media public sentiment analysis system for microblog,” Multimedia Syst., vol. 22, no. 4, pp. 479–486, 2016.

[190] W. He, H. Wu, G. Yan, V. Akula, and J. Shen, “A novel social media competitive analytics framework with sentiment benchmarks,” Inf. Manage., vol. 52, no. 7, pp. 801–812, 2015.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

CHAKRABORTY et al.: SURVEY OF SENTIMENT ANALYSIS FROM SOCIAL MEDIA DATA 463

[191] D. Zimbra, M. Ghiassi, and S. Lee, “Brand-related Twitter sentiment analysis using feature engineering and the dynamic architecture for artificial neural networks,” in Proc. 49th Hawaii Int. Conf. Syst. Sci. (HICSS), Jan. 2016, pp. 1930–1938.

[192] B. Jou, T. Chen, N. Pappas, M. Redi, M. Topkara, and S. F. Chang, “Visual affect around the world: A large-scale multilingual visual sen- timent ontology,” in Proc. 23rd ACM Int. Conf. Multimedia, Oct. 2015, pp. 159–168.

[193] T. Niu, S. Zhu, L. Pang, and A. E. Saddik, “Sentiment analysis on multi-view social data,” in Proc. Int. Conf. Multimedia Modeling. Cham, Switzerland: Springer, Jan. 2016, pp. 15–27.

[194] J. S. Jang, B. I. C. H. Lee Choi, J. H. Kim, D. M. Seo, and W. S. Cho, “Understanding pending issue of society and sentiment analysis using social media,” in Proc. 8th Int. Conf. Ubiquitous Future Netw. (ICUFN), Jul. 2016, pp. 981–986.

[195] F. H. Khan, U. Qamar, and S. Bashir, “SentiMI: Introducing point-wise mutual information with SentiWordNet to improve sentiment polarity detection,” Appl. Soft Comput., vol. 39, pp. 140–153, Feb. 2016.

[196] O. F. Villarroel, S. Ludwig, R. K. De, D. Grewal, and M. Wetzels, “Unveiling what is written in the stars: Analyzing explicit, implicit, and discourse patterns of sentiment in social media,” J. Consum. Res., vol. 43, no. 6, pp. 875–894, 2017.

[197] S. Poria, E. Cambria, D. Hazarika, and P. Vij, “A deeper look into sarcastic tweets using deep convolutional neural networks,” 2016, arXiv:1610.08815. [Online]. Available: https://arxiv.org/abs/ arXiv:1610.08815

[198] L. Wang and Z. Zhang, Support Vector Machines: Theory and Appli- cations. Berlin, Germany: Springer-Verlag, 2005.

[199] B. Liu, Sentiment Analysis and Opinion Mining (Synthesis Lectures on Human Language Technologies), vol. 5, no. 1. Ras al Khaimah, United Arab Emirates: Science Publishing Corporation, RAK Free Trade Zone, 2012, pp. 1–167.

[200] D. M. E.-D. M. Hussein, “A survey on sentiment analysis challenges,” J. King Saud Univ.-Eng. Sci., vol. 34, no. 4, pp. 330–338, 2016.

[201] S. Behdenna, F. Barigou, and G. Belalem, “Document level sentiment analysis: A survey,” EAI Endorsed Trans. Context-Aware Syst. Appl., vol. 4, Mar. 2018, Art. no. 154339, doi: 10.4108/eai.14-3-2018.154339.

[202] V. S. Jagtap and K. Pawar, “Analysis of different approaches to sentence-level sentiment classification,” Int. J. Sci. Eng. Technol., vol. 2, no. 3, pp. 164–170, 2013.

[203] K. Schouten and F. Frasincar, “Survey on aspect-level sentiment analysis,” IEEE Trans. Knowl. Data Eng., vol. 28, no. 3, pp. 813–830, Oct. 2015.

[204] M. K. Dalal and M. A. Zaveri, “Opinion mining from online user reviews using fuzzy linguistic hedges,” Appl. Comput. Intell. Soft Comput., vol. 2014, p. 2, Jan. 2014.

[205] P. Gonçalves, M. Araújo, F. Benevenuto, and M. Cha, “Comparing and combining sentiment analysis methods,” in Proc. 1st ACM Conf. Online Social Netw., Oct. 2013, pp. 27–38.

[206] F. Å. Nielsen, “A new ANEW: Evaluation of a word list for sentiment analysis in microblogs,” 2011, arXiv:1103.2903. [Online]. Available: https://arxiv.org/abs/1103.2903

[207] A. Hogenboom, B. Heerschop, F. Frasincar, U. Kaymak, and F. de Jong, “Multi-lingual support for lexicon-based sentiment analysis guided by semantics,” Decis. Support Syst., vol. 62, pp. 43–53, Jun. 2014.

[208] P. Ray and A. Chakrabarti, “Twitter sentiment analysis for product review using lexicon method,” in Proc. Int. Conf. Data Manage., Anal. Innov. (ICDMAI), Feb. 2017, pp. 211–216.

[209] R. Moraes, J. F. Valiati, and W. P. G. Neto, “Document-level sentiment classification: An empirical comparison between SVM and ANN,” Expert Syst. Appl., vol. 40, no. 2, pp. 621–633, 2013.

[210] B. Liu, Web Data Mining: Exploring Hyperlinks, Contents, and Usage Data. New York, NY, USA: Elsevier, 2007.

[211] J. Smailović, M. Grčar, N. Lavrač, and M. Žnidaršič, “Stream-based active learning for sentiment analysis in the financial domain,” Inf. Sci., vol. 285, pp. 181–203, Nov. 2014.

[212] A. Hasan, S. Moin, A. Karim, and S. Shamshirband, “Machine learning-based sentiment analysis for Twitter accounts,” Math. Comput. Appl., vol. 23, no. 1, p. 11, 2018.

[213] P. D. Turney, “Thumbs up or thumbs down?: Semantic orientation applied to unsupervised classification of reviews,” in Proc. 40th Annu. Meeting Assoc. Comput. Linguistics, Jul. 2002, pp. 417–424.

[214] F. Bravo-Marquez, M. Mendoza, and B. Poblete, “Meta-level sentiment models for big social data analysis,” Knowl.-Based Syst., vol. 69, pp. 86–99, Oct. 2014.

[215] M. Ahmad, S. Aftab, I. Ali, and N. Hameed, “Hybrid tools and techniques for sentiment analysis: A review,” Int. J. Multidiscip. Sci. Eng., vol. 8, no. 3, pp. 1–6, 2017.

[216] K. Elshakankery and M. F. Ahmed, “HILATSA: A hybrid incremental learning approach for Arabic tweets sentiment analysis,” Egyptian Inform. J., to be published.

[217] Q. T. Ain et al., “Sentiment analysis using deep learning techniques: A review,” Int. J. Adv. Comput. Sci. Appl., vol. 8, no. 6, p. 424, 2017.

[218] A. Kalaivani and D. Thenmozhi, “Sentiment analysis using deep learning techniques,” Int. J. Recent Technol. Eng., vol. 7, no. 6S5, pp. 1–7, 2019.

[219] A. M. El-Gazzar, T. M. Mohamed, and R. A. Sadek, “A hybrid SVD- HSV visual sentiment analysis system,” in Proc. 8th Int. Conf. Intell. Comput. Inf. Syst. (ICICIS), Dec. 2017, pp. 360–365.

[220] H. Zhang, A. C. Berg, M. Maire, and J. Malik, “SVM-KNN: Discrim- inative nearest neighbor classification for visual category recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., vol. 2, Jun. 2006, pp. 2126–2136.

[221] M. Lyons, S. Akamatsu, M. Kamachi, and J. Gyoba, “Coding facial expressions with gabor wavelets,” in Proc. 3rd IEEE Int. Conf. Autom. Face Gesture Recognit., Apr. 1998, pp. 200–205.

[222] I. S. P. James, “Face image retrieval with HSV color space using clustering techniques,” SIJ Trans. Comput. Sci. Eng. Appl., vol. 1, no. 1, pp. 1–4, 2013.

[223] K. Mahboob and F. Ali, “Sentiment analysis of pharmaceutical products evaluation based on customer review mining,” J. Comput. Sci. Syst. Biol., vol. 11, pp. 190–194, Mar. 2018, doi: 10.4172/jcsb.1000271.

[224] G. Liang, W. He, C. Xu, L. Chen, and J. Zeng, “Rumor identification in microblogging systems based on users’ behavior,” IEEE Trans. Computat. Social Syst., vol. 2, no. 3, pp. 99–108, Sep. 2015.

[225] R. Basak, S. Sural, N. Ganguly, and S. K. Ghosh, “Online public shaming on Twitter: Detection, analysis, and mitigation,” IEEE Trans. Comput. Social Syst., vol. 6, no. 2, pp. 208–220, Apr. 2019.

[226] S. Madisetty and M. S. Desarkar, “A neural network-based ensemble approach for spam detection in Twitter,” IEEE Trans. Comput. Social Syst., vol. 5, no. 4, pp. 973–984, Dec. 2018.

[227] H. Tajalizadeh and R. Boostani, “A novel stream clustering framework for spam detection in Twitter,” IEEE Trans. Comput. Social Syst., vol. 6, no. 3, pp. 525–534, Jun. 2019, doi: 10.1109/ TCSS.2019.2910818.

[228] Y. Tyshchuk and W. A. Wallace, “Modeling human behavior on social media in response to significant events,” IEEE Trans. Comput. Social Syst., vol. 5, no. 2, pp. 444–457, Jun. 2018.

[229] R. C. Chen, “User rating classification via deep belief network learning and sentiment analysis,” IEEE Trans. Comput. Social Syst., vol. 6, no. 3, pp. 535–546, Jun. 2019.

[230] T. Chowdhury, S. Muhuri, S. Chakraborty, and S. N. Chakraborty, “Analysis of adapted films and stories based on social network,” IEEE Trans. Comput. Social Syst., vol. 6, no. 5, pp. 858–869, Oct. 2019.

[231] K. Chakraborty, S. Bhattacharyya, R. Bag, and A. E. Hassanien, “Comparative sentiment analysis on a set of movie reviews using deep learning approach,” in Proc. Int. Conf. Adv. Mach. Learn. Technol. Appl. (AMLTA), Egypt, Cairo, Feb. 2018, pp. 311–318.

[232] K. Chakraborty, S. Bhattacharyya, R. Bag, and A. E. Hassanien, “Sentiment analysis on a set of movie reviews using deep learning techniques,” in Social Network Analytics—Computational Research Methods and Techniques. Amsterdam, The Netherlands: Elsevier, 2018.

[233] K. Chakraborty, R. Bag, and S. Bhattacharyya, “Relook into sentiment analysis performed on Indian languages using deep learning,” in Proc. 4th Int. Conf. Res. Comput. Intell. Commun. Netw. (ICRCICN), Nov. 2018, pp. 208–213.

Koyel Chakraborty was born in 1988. She received the B.Sc. degree in computer science from the University of Burdwan, Bardhaman, India, in 2009, the M.Sc. degree in computer science from West Bengal State University, Kolkata, India, in 2011, and the M.Tech. degree in computer science from the Maulana Abul Kalam Azad University of Technol- ogy, Kolkata, in 2016.

She is currently an Assistant Professor with the Department of Computer Science and Engineering, Supreme Knowledge Foundation Group of Institu-

tions, Chandannagar, India, under the Maulana Abul Kalam Azad University of Technology. Her research interest is in the fields of deep learning and sentiment analysis.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

464 IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, VOL. 7, NO. 2, APRIL 2020

Siddhartha Bhattacharyya (M’10–SM’13) received the bachelor’s degree in physics, and the bachelor’s and master’s degrees in optics and optoelectronics from the University of Calcutta, Kolkata, India, in 1995, 1998, and 2000, respectively, and the Ph.D. degree in computer science and engineering from Jadavpur University, Kolkata, in 2008.

He is currently serving as a Professor with the Department of Computer Science and Engineering, Christ University, Bengaluru, India. Prior to this,

he served as the Principal of the RCC Institute of Information Technology, Kolkata. He also served as a Senior Research Scientist with the Faculty of Electrical Engineering and Computer Science, VSB Technical University of Ostrava, Ostrava, Czech Republic. He has coauthored five books, coedited 40 books, and has authored or coauthored more than 250 research publications in international journals and conference proceedings. He holds three patents. His research interests include soft computing, pattern recognition, multimedia data processing, hybrid intelligence, social networks, and quantum computing.

Rajib Bag was born in 1969. He received the B.Sc. degree (Hons.) in physics from Calcutta Uni- versity, Kolkata, India, in 1991, the M.Sc. degree in physics from Vinoba Bhave University, Hazarib- agh, India, in 1996, and the M.Tech. degree and Ph.D. degree in control systems in engineering from Jadavpur University, Kolkata, in 2007 and 2012, respectively.

He is currently a Professor and the Head of the Department of Computer Science and Engineering, Supreme Knowledge Foundation Group of Institu-

tions, Chandannagar, India, under the Maulana Abul Kalam Azad University of Technology, Kolkata. Currently, five research scholars are doing their research work in different areas under his supervision. He has authored or coauthored more than 40 publications in reputed refereed journals and confer- ence proceedings. His research interest includes image and signal processing, education technology, machine learning, deep learning, and Internet-of-Things security besides control systems.

Authorized licensed use limited to: University of the Cumberlands. Downloaded on September 25,2021 at 01:55:11 UTC from IEEE Xplore. Restrictions apply.

<< /ASCII85EncodePages false /AllowTransparency false /AutoPositionEPSFiles true /AutoRotatePages /None /Binding /Left /CalGrayProfile (Black & White) /CalRGBProfile (sRGB IEC61966-2.1) /CalCMYKProfile (U.S. Web Coated \050SWOP\051 v2) /sRGBProfile (sRGB IEC61966-2.1) /CannotEmbedFontPolicy /Warning /CompatibilityLevel 1.4 /CompressObjects /Tags /CompressPages true /ConvertImagesToIndexed true /PassThroughJPEGImages true /CreateJobTicket false /DefaultRenderingIntent /Default /DetectBlends true /DetectCurves 0.0000 /ColorConversionStrategy /LeaveColorUnchanged /DoThumbnails true /EmbedAllFonts true /EmbedOpenType false /ParseICCProfilesInComments true /EmbedJobOptions true /DSCReportingLevel 0 /EmitDSCWarnings false /EndPage -1 /ImageMemory 524288 /LockDistillerParams true /MaxSubsetPct 100 /Optimize true /OPM 0 /ParseDSCComments false /ParseDSCCommentsForDocInfo true /PreserveCopyPage true /PreserveDICMYKValues true /PreserveEPSInfo true /PreserveFlatness true /PreserveHalftoneInfo true /PreserveOPIComments true /PreserveOverprintSettings true /StartPage 1 /SubsetFonts true /TransferFunctionInfo /Remove /UCRandBGInfo /Preserve /UsePrologue false /ColorSettingsFile () /AlwaysEmbed [ true /AdobeArabic-Bold /AdobeArabic-BoldItalic /AdobeArabic-Italic /AdobeArabic-Regular /AdobeHebrew-Bold /AdobeHebrew-BoldItalic /AdobeHebrew-Italic /AdobeHebrew-Regular /AdobeHeitiStd-Regular /AdobeMingStd-Light /AdobeMyungjoStd-Medium /AdobePiStd /AdobeSansMM /AdobeSerifMM /AdobeSongStd-Light /AdobeThai-Bold /AdobeThai-BoldItalic /AdobeThai-Italic /AdobeThai-Regular /ArborText /Arial-Black /Arial-BoldItalicMT /Arial-BoldMT /Arial-ItalicMT /ArialMT /BellGothicStd-Black /BellGothicStd-Bold /BellGothicStd-Light /ComicSansMS /ComicSansMS-Bold /Courier /Courier-Bold /Courier-BoldOblique /CourierNewPS-BoldItalicMT /CourierNewPS-BoldMT /CourierNewPS-ItalicMT /CourierNewPSMT /Courier-Oblique /CourierStd /CourierStd-Bold /CourierStd-BoldOblique /CourierStd-Oblique /EstrangeloEdessa /EuroSig /FranklinGothic-Medium /FranklinGothic-MediumItalic /Gautami /Georgia /Georgia-Bold /Georgia-BoldItalic /Georgia-Italic /Helvetica /Helvetica-Bold /Helvetica-BoldOblique /Helvetica-Oblique /Impact /KozGoPr6N-Medium /KozGoProVI-Medium /KozMinPr6N-Regular /KozMinProVI-Regular /Latha /LetterGothicStd /LetterGothicStd-Bold /LetterGothicStd-BoldSlanted /LetterGothicStd-Slanted /LucidaConsole /LucidaSans-Typewriter /LucidaSans-TypewriterBold /LucidaSansUnicode /Mangal-Regular /MicrosoftSansSerif /MinionPro-Bold /MinionPro-BoldIt /MinionPro-It /MinionPro-Regular /MinionPro-Semibold /MinionPro-SemiboldIt /MVBoli /MyriadPro-Black /MyriadPro-BlackIt /MyriadPro-Bold /MyriadPro-BoldIt /MyriadPro-It /MyriadPro-Light /MyriadPro-LightIt /MyriadPro-Regular /MyriadPro-Semibold /MyriadPro-SemiboldIt /PalatinoLinotype-Bold /PalatinoLinotype-BoldItalic /PalatinoLinotype-Italic /PalatinoLinotype-Roman /Raavi /Shruti /Sylfaen /Symbol /SymbolMT /Tahoma /Tahoma-Bold /Times-Bold /Times-BoldItalic /Times-Italic /TimesNewRomanPS-BoldItalicMT /TimesNewRomanPS-BoldMT /TimesNewRomanPS-ItalicMT /TimesNewRomanPSMT /Times-Roman /Trebuchet-BoldItalic /TrebuchetMS /TrebuchetMS-Bold /TrebuchetMS-Italic /Tunga-Regular /Verdana /Verdana-Bold /Verdana-BoldItalic /Verdana-Italic /Webdings /Wingdings-Regular /ZapfDingbats /ZWAdobeF ] /NeverEmbed [ true ] /AntiAliasColorImages false /CropColorImages true /ColorImageMinResolution 150 /ColorImageMinResolutionPolicy /OK /DownsampleColorImages true /ColorImageDownsampleType /Bicubic /ColorImageResolution 600 /ColorImageDepth -1 /ColorImageMinDownsampleDepth 1 /ColorImageDownsampleThreshold 1.50000 /EncodeColorImages true /ColorImageFilter /DCTEncode /AutoFilterColorImages true /ColorImageAutoFilterStrategy /JPEG /ColorACSImageDict << /QFactor 0.76 /HSamples [2 1 1 2] /VSamples [2 1 1 2] >> /ColorImageDict << /QFactor 0.15 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /JPEG2000ColorACSImageDict << /TileWidth 256 /TileHeight 256 /Quality 15 >> /JPEG2000ColorImageDict << /TileWidth 256 /TileHeight 256 /Quality 15 >> /AntiAliasGrayImages false /CropGrayImages true /GrayImageMinResolution 150 /GrayImageMinResolutionPolicy /OK /DownsampleGrayImages true /GrayImageDownsampleType /Bicubic /GrayImageResolution 600 /GrayImageDepth -1 /GrayImageMinDownsampleDepth 2 /GrayImageDownsampleThreshold 1.50000 /EncodeGrayImages true /GrayImageFilter /DCTEncode /AutoFilterGrayImages true /GrayImageAutoFilterStrategy /JPEG /GrayACSImageDict << /QFactor 0.76 /HSamples [2 1 1 2] /VSamples [2 1 1 2] >> /GrayImageDict << /QFactor 0.15 /HSamples [1 1 1 1] /VSamples [1 1 1 1] >> /JPEG2000GrayACSImageDict << /TileWidth 256 /TileHeight 256 /Quality 15 >> /JPEG2000GrayImageDict << /TileWidth 256 /TileHeight 256 /Quality 15 >> /AntiAliasMonoImages false /CropMonoImages true /MonoImageMinResolution 300 /MonoImageMinResolutionPolicy /OK /DownsampleMonoImages true /MonoImageDownsampleType /Bicubic /MonoImageResolution 900 /MonoImageDepth -1 /MonoImageDownsampleThreshold 1.33333 /EncodeMonoImages true /MonoImageFilter /CCITTFaxEncode /MonoImageDict << /K -1 >> /AllowPSXObjects false /CheckCompliance [ /None ] /PDFX1aCheck false /PDFX3Check false /PDFXCompliantPDFOnly false /PDFXNoTrimBoxError true /PDFXTrimBoxToMediaBoxOffset [ 0.00000 0.00000 0.00000 0.00000 ] /PDFXSetBleedBoxToMediaBox true /PDFXBleedBoxToTrimBoxOffset [ 0.00000 0.00000 0.00000 0.00000 ] /PDFXOutputIntentProfile (None) /PDFXOutputConditionIdentifier () /PDFXOutputCondition () /PDFXRegistryName () /PDFXTrapped /Unknown /CreateJDFFile false /Description << /ENU () >> >> setdistillerparams << /HWResolution [600 600] /PageSize [612.000 792.000] >> setpagedevice