BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla

The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and orig...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Hauptverfasser:	Haider, Fabiha, Shifat, Fariha Tanjim, Ishmam, Md Farhan, Barua, Deeparghya Dutta, Sourove, Md Sakib Ul Rahman, Fahim, Md, Alam, Md Farhad
Format:	Artikel
Sprache:	eng
Schlagworte:	Computer Science - Computation and Language
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page
container_issue
container_start_page
container_title
container_volume
creator	Haider, Fabiha Shifat, Fariha Tanjim Ishmam, Md Farhan Barua, Deeparghya Dutta Sourove, Md Sakib Ul Rahman Fahim, Md Alam, Md Farhad
description	The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.
doi_str_mv	10.48550/arxiv.2410.13281
format	Article
fullrecord	<record><control><sourceid>arxiv_GOX</sourceid><recordid>TN_cdi_arxiv_primary_2410_13281</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2410_13281</sourcerecordid><originalsourceid>FETCH-arxiv_primary_2410_132813</originalsourceid><addsrcrecordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMgEKGBobWRhyMvg4JeaFeFgpOCr4luaUZOrmJCal5ih4JJakKgQXpKYmZyi4pJakJpdk5ucpuCSWJBanliik5RcphBQl5hXnZJakFgGVpigATUnPSeRhYE1LzClO5YXS3Azybq4hzh66YHvjC4oycxOLKuNB9seD7TcmrAIAMbE50w</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype></control><display><type>article</type><title>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</title><source>arXiv.org</source><creator>Haider, Fabiha ; Shifat, Fariha Tanjim ; Ishmam, Md Farhan ; Barua, Deeparghya Dutta ; Sourove, Md Sakib Ul Rahman ; Fahim, Md ; Alam, Md Farhad</creator><creatorcontrib>Haider, Fabiha ; Shifat, Fariha Tanjim ; Ishmam, Md Farhan ; Barua, Deeparghya Dutta ; Sourove, Md Sakib Ul Rahman ; Fahim, Md ; Alam, Md Farhad</creatorcontrib><description>The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.</description><identifier>DOI: 10.48550/arxiv.2410.13281</identifier><language>eng</language><subject>Computer Science - Computation and Language</subject><creationdate>2024-10</creationdate><rights>http://creativecommons.org/licenses/by/4.0</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>228,230,776,881</link.rule.ids><linktorsrc>$$Uhttps://arxiv.org/abs/2410.13281$$EView_record_in_Cornell_University$$FView_record_in_$$GCornell_University$$Hfree_for_read</linktorsrc><backlink>$$Uhttps://doi.org/10.48550/arXiv.2410.13281$$DView paper in arXiv$$Hfree_for_read</backlink></links><search><creatorcontrib>Haider, Fabiha</creatorcontrib><creatorcontrib>Shifat, Fariha Tanjim</creatorcontrib><creatorcontrib>Ishmam, Md Farhan</creatorcontrib><creatorcontrib>Barua, Deeparghya Dutta</creatorcontrib><creatorcontrib>Sourove, Md Sakib Ul Rahman</creatorcontrib><creatorcontrib>Fahim, Md</creatorcontrib><creatorcontrib>Alam, Md Farhad</creatorcontrib><title>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</title><description>The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.</description><subject>Computer Science - Computation and Language</subject><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><sourceid>GOX</sourceid><recordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMgEKGBobWRhyMvg4JeaFeFgpOCr4luaUZOrmJCal5ih4JJakKgQXpKYmZyi4pJakJpdk5ucpuCSWJBanliik5RcphBQl5hXnZJakFgGVpigATUnPSeRhYE1LzClO5YXS3Azybq4hzh66YHvjC4oycxOLKuNB9seD7TcmrAIAMbE50w</recordid><startdate>20241017</startdate><enddate>20241017</enddate><creator>Haider, Fabiha</creator><creator>Shifat, Fariha Tanjim</creator><creator>Ishmam, Md Farhan</creator><creator>Barua, Deeparghya Dutta</creator><creator>Sourove, Md Sakib Ul Rahman</creator><creator>Fahim, Md</creator><creator>Alam, Md Farhad</creator><scope>AKY</scope><scope>GOX</scope></search><sort><creationdate>20241017</creationdate><title>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</title><author>Haider, Fabiha ; Shifat, Fariha Tanjim ; Ishmam, Md Farhan ; Barua, Deeparghya Dutta ; Sourove, Md Sakib Ul Rahman ; Fahim, Md ; Alam, Md Farhad</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-arxiv_primary_2410_132813</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Computer Science - Computation and Language</topic><toplevel>online_resources</toplevel><creatorcontrib>Haider, Fabiha</creatorcontrib><creatorcontrib>Shifat, Fariha Tanjim</creatorcontrib><creatorcontrib>Ishmam, Md Farhan</creatorcontrib><creatorcontrib>Barua, Deeparghya Dutta</creatorcontrib><creatorcontrib>Sourove, Md Sakib Ul Rahman</creatorcontrib><creatorcontrib>Fahim, Md</creatorcontrib><creatorcontrib>Alam, Md Farhad</creatorcontrib><collection>arXiv Computer Science</collection><collection>arXiv.org</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext_linktorsrc</fulltext></delivery><addata><au>Haider, Fabiha</au><au>Shifat, Fariha Tanjim</au><au>Ishmam, Md Farhan</au><au>Barua, Deeparghya Dutta</au><au>Sourove, Md Sakib Ul Rahman</au><au>Fahim, Md</au><au>Alam, Md Farhad</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</atitle><date>2024-10-17</date><risdate>2024</risdate><abstract>The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.</abstract><doi>10.48550/arxiv.2410.13281</doi><oa>free_for_read</oa></addata></record>
fulltext	fulltext_linktorsrc
identifier	DOI: 10.48550/arxiv.2410.13281
ispartof
issn
language	eng
recordid	cdi_arxiv_primary_2410_13281
source	arXiv.org
subjects	Computer Science - Computation and Language
title	BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-02-03T22%3A25%3A10IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-arxiv_GOX&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=BanTH:%20A%20Multi-label%20Hate%20Speech%20Detection%20Dataset%20for%20Transliterated%20Bangla&rft.au=Haider,%20Fabiha&rft.date=2024-10-17&rft_id=info:doi/10.48550/arxiv.2410.13281&rft_dat=%3Carxiv_GOX%3E2410_13281%3C/arxiv_GOX%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_id=info:pmid/&rfr_iscdi=true