Securing Behavior-based Opinion Spam Detection

Reviews spams are prevalent in e-commerce to manipulate product ranking and customers decisions maliciously. While spams generated based on simple spamming strategy can be detected effectively, hardened spammers can evade regular detectors via more advanced spamming strategies. Previous work gave mo...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Hauptverfasser:	Ge, Shuaijun, Ma, Guixiang, Xie, Sihong, Yu, Philip S
Format:	Artikel
Sprache:	eng
Schlagworte:	Computer Science - Cryptography and Security Computer Science - Learning Statistics - Machine Learning
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page
container_issue
container_start_page
container_title
container_volume
creator	Ge, Shuaijun Ma, Guixiang Xie, Sihong Yu, Philip S
description	Reviews spams are prevalent in e-commerce to manipulate product ranking and customers decisions maliciously. While spams generated based on simple spamming strategy can be detected effectively, hardened spammers can evade regular detectors via more advanced spamming strategies. Previous work gave more attention to evasion against text and graph-based detectors, but evasions against behavior-based detectors are largely ignored, leading to vulnerabilities in spam detection systems. Since real evasion data are scarce, we first propose EMERAL (Evasion via Maximum Entropy and Rating sAmpLing) to generate evasive spams to certain existing detectors. EMERAL can simulate spammers with different goals and levels of knowledge about the detectors, targeting at different stages of the life cycle of target products. We show that in the evasion-defense dynamic, only a few evasion types are meaningful to the spammers, and any spammer will not be able to evade too many detection signals at the same time. We reveal that some evasions are quite insidious and can fail all detection signals. We then propose DETER (Defense via Evasion generaTion using EmeRal), based on model re-training on diverse evasive samples generated by EMERAL. Experiments confirm that DETER is more accurate in detecting both suspicious time window and individual spamming reviews. In terms of security, DETER is versatile enough to be vaccinated against diverse and unexpected evasions, is agnostic about evasion strategy and can be released without privacy concern.
doi_str_mv	10.48550/arxiv.1811.03739
format	Article
fullrecord	<record><control><sourceid>arxiv_GOX</sourceid><recordid>TN_cdi_arxiv_primary_1811_03739</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>1811_03739</sourcerecordid><originalsourceid>FETCH-LOGICAL-a679-61ddfae25dbee42680acbb81aa9b8f016995429bad6b407d54a246ff7e3f29e93</originalsourceid><addsrcrecordid>eNotzr1uwjAYhWEvDChwAUzkBpL6P_YIFGglJAbYo8_xZ7BEQmR-BHffQjsdvcvRQ8iE0VIapegHpEe8l8wwVlJRCTsk5Q6bW4rdIZ_jEe7xnAoHF_T5to9dPHf5roc2_8QrNtffHJFBgNMFx_-bkf1quV98FZvt-nsx2xSgK1to5n0A5Mo7RMm1odA4ZxiAdSZQpq1VklsHXjtJK68kcKlDqFAEbtGKjEz_bt_guk-xhfSsX_D6DRc_vUo9rw</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype></control><display><type>article</type><title>Securing Behavior-based Opinion Spam Detection</title><source>arXiv.org</source><creator>Ge, Shuaijun ; Ma, Guixiang ; Xie, Sihong ; Yu, Philip S</creator><creatorcontrib>Ge, Shuaijun ; Ma, Guixiang ; Xie, Sihong ; Yu, Philip S</creatorcontrib><description>Reviews spams are prevalent in e-commerce to manipulate product ranking and customers decisions maliciously. While spams generated based on simple spamming strategy can be detected effectively, hardened spammers can evade regular detectors via more advanced spamming strategies. Previous work gave more attention to evasion against text and graph-based detectors, but evasions against behavior-based detectors are largely ignored, leading to vulnerabilities in spam detection systems. Since real evasion data are scarce, we first propose EMERAL (Evasion via Maximum Entropy and Rating sAmpLing) to generate evasive spams to certain existing detectors. EMERAL can simulate spammers with different goals and levels of knowledge about the detectors, targeting at different stages of the life cycle of target products. We show that in the evasion-defense dynamic, only a few evasion types are meaningful to the spammers, and any spammer will not be able to evade too many detection signals at the same time. We reveal that some evasions are quite insidious and can fail all detection signals. We then propose DETER (Defense via Evasion generaTion using EmeRal), based on model re-training on diverse evasive samples generated by EMERAL. Experiments confirm that DETER is more accurate in detecting both suspicious time window and individual spamming reviews. In terms of security, DETER is versatile enough to be vaccinated against diverse and unexpected evasions, is agnostic about evasion strategy and can be released without privacy concern.</description><identifier>DOI: 10.48550/arxiv.1811.03739</identifier><language>eng</language><subject>Computer Science - Cryptography and Security ; Computer Science - Learning ; Statistics - Machine Learning</subject><creationdate>2018-11</creationdate><rights>http://arxiv.org/licenses/nonexclusive-distrib/1.0</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>228,230,776,881</link.rule.ids><linktorsrc>$$Uhttps://arxiv.org/abs/1811.03739$$EView_record_in_Cornell_University$$FView_record_in_$$GCornell_University$$Hfree_for_read</linktorsrc><backlink>$$Uhttps://doi.org/10.48550/arXiv.1811.03739$$DView paper in arXiv$$Hfree_for_read</backlink></links><search><creatorcontrib>Ge, Shuaijun</creatorcontrib><creatorcontrib>Ma, Guixiang</creatorcontrib><creatorcontrib>Xie, Sihong</creatorcontrib><creatorcontrib>Yu, Philip S</creatorcontrib><title>Securing Behavior-based Opinion Spam Detection</title><description>Reviews spams are prevalent in e-commerce to manipulate product ranking and customers decisions maliciously. While spams generated based on simple spamming strategy can be detected effectively, hardened spammers can evade regular detectors via more advanced spamming strategies. Previous work gave more attention to evasion against text and graph-based detectors, but evasions against behavior-based detectors are largely ignored, leading to vulnerabilities in spam detection systems. Since real evasion data are scarce, we first propose EMERAL (Evasion via Maximum Entropy and Rating sAmpLing) to generate evasive spams to certain existing detectors. EMERAL can simulate spammers with different goals and levels of knowledge about the detectors, targeting at different stages of the life cycle of target products. We show that in the evasion-defense dynamic, only a few evasion types are meaningful to the spammers, and any spammer will not be able to evade too many detection signals at the same time. We reveal that some evasions are quite insidious and can fail all detection signals. We then propose DETER (Defense via Evasion generaTion using EmeRal), based on model re-training on diverse evasive samples generated by EMERAL. Experiments confirm that DETER is more accurate in detecting both suspicious time window and individual spamming reviews. In terms of security, DETER is versatile enough to be vaccinated against diverse and unexpected evasions, is agnostic about evasion strategy and can be released without privacy concern.</description><subject>Computer Science - Cryptography and Security</subject><subject>Computer Science - Learning</subject><subject>Statistics - Machine Learning</subject><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2018</creationdate><recordtype>article</recordtype><sourceid>GOX</sourceid><recordid>eNotzr1uwjAYhWEvDChwAUzkBpL6P_YIFGglJAbYo8_xZ7BEQmR-BHffQjsdvcvRQ8iE0VIapegHpEe8l8wwVlJRCTsk5Q6bW4rdIZ_jEe7xnAoHF_T5to9dPHf5roc2_8QrNtffHJFBgNMFx_-bkf1quV98FZvt-nsx2xSgK1to5n0A5Mo7RMm1odA4ZxiAdSZQpq1VklsHXjtJK68kcKlDqFAEbtGKjEz_bt_guk-xhfSsX_D6DRc_vUo9rw</recordid><startdate>20181108</startdate><enddate>20181108</enddate><creator>Ge, Shuaijun</creator><creator>Ma, Guixiang</creator><creator>Xie, Sihong</creator><creator>Yu, Philip S</creator><scope>AKY</scope><scope>EPD</scope><scope>GOX</scope></search><sort><creationdate>20181108</creationdate><title>Securing Behavior-based Opinion Spam Detection</title><author>Ge, Shuaijun ; Ma, Guixiang ; Xie, Sihong ; Yu, Philip S</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-LOGICAL-a679-61ddfae25dbee42680acbb81aa9b8f016995429bad6b407d54a246ff7e3f29e93</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2018</creationdate><topic>Computer Science - Cryptography and Security</topic><topic>Computer Science - Learning</topic><topic>Statistics - Machine Learning</topic><toplevel>online_resources</toplevel><creatorcontrib>Ge, Shuaijun</creatorcontrib><creatorcontrib>Ma, Guixiang</creatorcontrib><creatorcontrib>Xie, Sihong</creatorcontrib><creatorcontrib>Yu, Philip S</creatorcontrib><collection>arXiv Computer Science</collection><collection>arXiv Statistics</collection><collection>arXiv.org</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext_linktorsrc</fulltext></delivery><addata><au>Ge, Shuaijun</au><au>Ma, Guixiang</au><au>Xie, Sihong</au><au>Yu, Philip S</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>Securing Behavior-based Opinion Spam Detection</atitle><date>2018-11-08</date><risdate>2018</risdate><abstract>Reviews spams are prevalent in e-commerce to manipulate product ranking and customers decisions maliciously. While spams generated based on simple spamming strategy can be detected effectively, hardened spammers can evade regular detectors via more advanced spamming strategies. Previous work gave more attention to evasion against text and graph-based detectors, but evasions against behavior-based detectors are largely ignored, leading to vulnerabilities in spam detection systems. Since real evasion data are scarce, we first propose EMERAL (Evasion via Maximum Entropy and Rating sAmpLing) to generate evasive spams to certain existing detectors. EMERAL can simulate spammers with different goals and levels of knowledge about the detectors, targeting at different stages of the life cycle of target products. We show that in the evasion-defense dynamic, only a few evasion types are meaningful to the spammers, and any spammer will not be able to evade too many detection signals at the same time. We reveal that some evasions are quite insidious and can fail all detection signals. We then propose DETER (Defense via Evasion generaTion using EmeRal), based on model re-training on diverse evasive samples generated by EMERAL. Experiments confirm that DETER is more accurate in detecting both suspicious time window and individual spamming reviews. In terms of security, DETER is versatile enough to be vaccinated against diverse and unexpected evasions, is agnostic about evasion strategy and can be released without privacy concern.</abstract><doi>10.48550/arxiv.1811.03739</doi><oa>free_for_read</oa></addata></record>
fulltext	fulltext_linktorsrc
identifier	DOI: 10.48550/arxiv.1811.03739
ispartof
issn
language	eng
recordid	cdi_arxiv_primary_1811_03739
source	arXiv.org
subjects	Computer Science - Cryptography and Security Computer Science - Learning Statistics - Machine Learning
title	Securing Behavior-based Opinion Spam Detection
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-02-03T14%3A18%3A14IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-arxiv_GOX&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=Securing%20Behavior-based%20Opinion%20Spam%20Detection&rft.au=Ge,%20Shuaijun&rft.date=2018-11-08&rft_id=info:doi/10.48550/arxiv.1811.03739&rft_dat=%3Carxiv_GOX%3E1811_03739%3C/arxiv_GOX%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_id=info:pmid/&rfr_iscdi=true