ABSTRACT
Evaluative texts on the Web have become a valuable source of opinions on products, services, events, individuals, etc. Recently, many researchers have studied such opinion sources as product reviews, forum posts, and blogs. However, existing research has been focused on classification and summarization of opinions using natural language processing and data mining techniques. An important issue that has been neglected so far is opinion spam or trustworthiness of online opinions. In this paper, we study this issue in the context of product reviews, which are opinion rich and are widely used by consumers and product manufacturers. In the past two years, several startup companies also appeared which aggregate opinions from product reviews. It is thus high time to study spam in reviews. To the best of our knowledge, there is still no published study on this topic, although Web spam and email spam have been investigated extensively. We will see that opinion spam is quite different from Web spam and email spam, and thus requires different detection techniques. Based on the analysis of 5.8 million reviews and 2.14 million reviewers from amazon.com, we show that opinion spam in reviews is widespread. This paper analyzes such spam activities and presents some novel techniques to detect them
- E. Amitay, D. Carmel, A. Darlow, R. Lempel & A. Soffer. The connectivity sonar: detecting site functionality by structural patterns. Hypertext'03, 2003. Google ScholarDigital Library
- M. Andreolini, A. Bulgarelli, M. Colajanni & F. Mazzoni. Honeyspam: Honeypots fighting spam at the source. In Proc. USENIX SRUTI 2005, Cambridge, MA, July 2005. Google ScholarDigital Library
- R. Baeza-Yates, C. Castillo & V. Lopez. PageRank increase under different collusion topologies. AIRWeb'05, 2005.Google Scholar
- A. Z. Broder. On the resemblance and containment of documents. In Proceedings of Compression and Complexity of Sequences 1997, IEEE Computer Society, 1997. Google ScholarDigital Library
- C. Castillo, D. Donato, L. Becchetti, P. Boldi, S. Leonardi, M. Santini, S. Vigna. A reference collection for web spam, SIGIR Forum'06, 2006. Google ScholarDigital Library
- S. Chakrabarti. Mining the Web: discovering knowledge from hypertext data. Morgan Kaufmann, 2003. Google ScholarDigital Library
- K. Dave, S. Lawrence & D. Pennock. Mining the peanut gallery: opinion extraction and semantic classification of product reviews. WWW'2003. Google ScholarDigital Library
- I. Fette, N. Sadeh-Koniecpol, A. Tomasic. Learning to Detect Phishing Emails. WWW2007. Google ScholarDigital Library
- D. Fetterly, M. Manasse & M. Najork. Detecting phrase-level duplication on the World Wide Web. SIGIR'2005. Google ScholarDigital Library
- Z. Gyongyi & H. Garcia-Molina. Web Spam Taxonomy. Technical Report, Stanford University, 2004.Google Scholar
- M. R. Henzinger: Finding near-duplicate web pages: a large-scale evaluation of algorithms. SIGIR'06, 2006. Google ScholarDigital Library
- M. Hu & B. Liu. Mining and summarizing customer reviews. KDD'2004. Google ScholarDigital Library
- N. Jindal and B. Liu. Product Review Analysis. Technical Report, UIC, 2007.Google Scholar
- N. Jindal and B. Liu. Analyzing and Detecting Review Spam. ICDM2007. Google ScholarDigital Library
- W. Li, N. Zhong, C. Liu. Combining Multiple Email Filters Based on Multivariate Statistical Analysis. ISMIS 2006. Google ScholarDigital Library
- B. Liu. Web Data Mining: Exploring hyperlinks, contents and usage data. Springer, 2007. Google ScholarDigital Library
- A. Metwally, D. Agrawal, A. Abbadi. DETECTIVES: DETEcting Coalition hiT Inflation attacks in adVertising nEtworks Streams. WWW2007. Google ScholarDigital Library
- B. Mobasher, R. Burke & J. J Sandvig. Model-based collaborative filtering as a defense against profile injection attacks. AAAI'2006. Google ScholarDigital Library
- A. Ntoulas, M. Najork, M. Manasse & D. Fetterly. Detecting Spam Web Pages through Content Analysis. WWW'2006. Google ScholarDigital Library
- B. Pang, L. Lee & S. Vaithyanathan. Thumbs up? Sentiment classification using machine learning techniques. EMNLP'2002. Google ScholarDigital Library
- A-M. Popescu and O. Etzioni. Extracting Product Features and Opinions from Reviews. EMNLP'2005. Google ScholarDigital Library
- M. Sahami and S. Dumais and D. Heckerman and E. Horvitz. A Bayesian Approach to Filtering Junk {E}-Mail. AAAI Technical Report WS-98-05, 1998.Google Scholar
- P. Turney. Thumbs up or thumbs down? semantic orientation applied to unsupervised classification of reviews. ACL'2002. Google ScholarDigital Library
- Y. Wang, M. Ma, Y. Niu, H. Chen. Spam Double-Funnel: Connecting Web Spammers with Advertisers. WWW2007. Google ScholarDigital Library
- B. Wu and B. D. Davison. Identifying link farm spam pages. WWW'06, 2006. Google ScholarDigital Library
- B. Wu, V. Goel & B. D. Davison. Topical TrustRank: using topicality to combat Web spam. WWW'2006. Google ScholarDigital Library
- S. Ye, R. Song, J.-R. Wen, W.-Y. Ma. A Query-dependent duplicate detection approach for large scale search engines. APWeb'04, 2004.Google ScholarCross Ref
- Z. Zhang & B. Varadarajan, Utility scoring of product reviews, CIKM'2006. Google ScholarDigital Library
Index Terms
- Opinion spam and analysis
Recommendations
Detection of review spam
We have extracted all types of data that can be used in spam detection techniques.We have reviewed state of the art literature in the area of detection of spam reviews.In this research, we have categorized and classified spam detection methods and ...
Opinion spam detection framework using hybrid classification scheme
AbstractWith the advent of social networking sites, opinion-mining applications have attracted the interest of the online community on review sites to know about products for their purchase decisions. However, due to increasing trend of posting spam (fake)...
Exploring groups of opinion spam using sentiment analysis guided by nominated topics
Graphical abstractDisplay Omitted
Highlights- This is the first study using platform-offered aspects for spam detection.
- This ...
AbstractCurrently, it is common to see untruthful opinions (also known as review spam, fraud or shilling attack) that resemble each other explicitly or implicitly across multiple business-to-customer websites or opinion sharing communities. ...
Comments