Toward a Language Modeling Approach for Consumer Review Spam Detection

Numerous reports have indicated the severity of fake reviews (i.e., spam) posted to various e-Commerce or opinion sharing Web sites. Nevertheless, very few studies have been conducted to examine the trustworthiness of online consumer reviews because of the lack of an effective computational methodol...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Hauptverfasser:	Lai, C L, Xu, K Q, Lau, R Y K, Li, Y, Jing, L
Format:	Tagungsbericht
Sprache:	eng
Schlagworte:	Computational modeling Electronic Commerce Information services Internet Kullback-Leibler Divergence Language Models Probabilistic logic Review Spam Spam Detection Support vector machines Unsolicited electronic mail Web sites
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Numerous reports have indicated the severity of fake reviews (i.e., spam) posted to various e-Commerce or opinion sharing Web sites. Nevertheless, very few studies have been conducted to examine the trustworthiness of online consumer reviews because of the lack of an effective computational methodology. Unlike other kinds of Web spam, untruthful reviews could just look like other legitimate reviews (i.e., ham), and so it is difficult to apply any features to distinguish the two classes. One main contribution of our research work is the development of a novel computational methodology to combat online review spam. Our experimental results confirm that the KL divergence and the probabilistic language modeling based computational model is effective for the detection of untruthful reviews. Empowered by the proposed computational methods, our empirical study found that around 2% of the consumer reviews posted to a large e-Commerce site is spam.
DOI:	10.1109/ICEBE.2010.47