Title :
Data intensive review mining for sentiment classification across heterogeneous domains
Author :
Bisio, Federica ; Gastaldo, Paolo ; Peretti, Chiara ; Zunino, Rodolfo ; Cambria, Erik
Author_Institution :
DITEN, Genoa Univ., Genoa, Italy
Abstract :
The automatic detection of orientation and emotions in texts is becoming increasingly important in the Web 2.0 scenario. There is a considerable need for innovative techniques and tools capable of identifying and detecting the attitude of unstructured text. The paper tackles two crucial aspects of the sentiment classification problem: first, the computational complexity of the deployed framework; second, the ability of the framework itself to operate effectively in heterogeneous commercial domains. The proposed approach adopts empirical learning to implement the sentiment-classification technology, and uses a distance-based predictive model to combine computational efficiency and modularity. A suitably designed semantic-based metric is the cognitive core that measures the distance between two user reviews, according to the sentiment they communicate. The framework ultimately nullifies the training process; at the same time, it takes advantage of a classification procedure whose computational cost increases linearly when the training corpus increases. To attain an objective measurement of the actual accuracy of the sentiment classification method, a campaign of tests involved a pair of complex, real-world scoring domains; the goal was to compare the predicted sentiment scores with actual scores provided by human assessors. Experimental results confirmed that the overall approach attained satisfactory performances in terms of both cross-domain classification accuracy and computational efficiency.
Keywords :
Internet; classification; computational complexity; data mining; learning (artificial intelligence); text analysis; Web 2.0 scenario; automatic detection; classification procedure; computational complexity; computational cost; computational efficiency; cross-domain classification accuracy; data intensive review mining; deployed framework; distance-based predictive model; heterogeneous commercial domain; heterogeneous domains; human assessors; innovative techniques; real-world scoring domain; semantic-based metric; sentiment classification; sentiment scores; sentiment-classification technology; unstructured text; Classification algorithms; Computational efficiency; Conferences; Measurement; Semantics; Social network services; Training; cross-domain sentiment classification; sentiment analysis;
Conference_Titel :
Advances in Social Networks Analysis and Mining (ASONAM), 2013 IEEE/ACM International Conference on
Conference_Location :
Niagara Falls, ON