BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla

The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and orig...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Haider, Fabiha, Shifat, Fariha Tanjim, Ishmam, Md Farhan, Barua, Deeparghya Dutta, Sourove, Md Sakib Ul Rahman, Fahim, Md, Alam, Md Farhad
Format: Artikel
Sprache:eng
Schlagworte:
Online-Zugang:Volltext bestellen
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
container_end_page
container_issue
container_start_page
container_title
container_volume
creator Haider, Fabiha
Shifat, Fariha Tanjim
Ishmam, Md Farhan
Barua, Deeparghya Dutta
Sourove, Md Sakib Ul Rahman
Fahim, Md
Alam, Md Farhad
description The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.
doi_str_mv 10.48550/arxiv.2410.13281
format Article
fullrecord <record><control><sourceid>arxiv_GOX</sourceid><recordid>TN_cdi_arxiv_primary_2410_13281</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2410_13281</sourcerecordid><originalsourceid>FETCH-arxiv_primary_2410_132813</originalsourceid><addsrcrecordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMgEKGBobWRhyMvg4JeaFeFgpOCr4luaUZOrmJCal5ih4JJakKgQXpKYmZyi4pJakJpdk5ucpuCSWJBanliik5RcphBQl5hXnZJakFgGVpigATUnPSeRhYE1LzClO5YXS3Azybq4hzh66YHvjC4oycxOLKuNB9seD7TcmrAIAMbE50w</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype></control><display><type>article</type><title>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</title><source>arXiv.org</source><creator>Haider, Fabiha ; Shifat, Fariha Tanjim ; Ishmam, Md Farhan ; Barua, Deeparghya Dutta ; Sourove, Md Sakib Ul Rahman ; Fahim, Md ; Alam, Md Farhad</creator><creatorcontrib>Haider, Fabiha ; Shifat, Fariha Tanjim ; Ishmam, Md Farhan ; Barua, Deeparghya Dutta ; Sourove, Md Sakib Ul Rahman ; Fahim, Md ; Alam, Md Farhad</creatorcontrib><description>The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.</description><identifier>DOI: 10.48550/arxiv.2410.13281</identifier><language>eng</language><subject>Computer Science - Computation and Language</subject><creationdate>2024-10</creationdate><rights>http://creativecommons.org/licenses/by/4.0</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>228,230,776,881</link.rule.ids><linktorsrc>$$Uhttps://arxiv.org/abs/2410.13281$$EView_record_in_Cornell_University$$FView_record_in_$$GCornell_University$$Hfree_for_read</linktorsrc><backlink>$$Uhttps://doi.org/10.48550/arXiv.2410.13281$$DView paper in arXiv$$Hfree_for_read</backlink></links><search><creatorcontrib>Haider, Fabiha</creatorcontrib><creatorcontrib>Shifat, Fariha Tanjim</creatorcontrib><creatorcontrib>Ishmam, Md Farhan</creatorcontrib><creatorcontrib>Barua, Deeparghya Dutta</creatorcontrib><creatorcontrib>Sourove, Md Sakib Ul Rahman</creatorcontrib><creatorcontrib>Fahim, Md</creatorcontrib><creatorcontrib>Alam, Md Farhad</creatorcontrib><title>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</title><description>The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.</description><subject>Computer Science - Computation and Language</subject><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2024</creationdate><recordtype>article</recordtype><sourceid>GOX</sourceid><recordid>eNpjYJA0NNAzsTA1NdBPLKrILNMzMgEKGBobWRhyMvg4JeaFeFgpOCr4luaUZOrmJCal5ih4JJakKgQXpKYmZyi4pJakJpdk5ucpuCSWJBanliik5RcphBQl5hXnZJakFgGVpigATUnPSeRhYE1LzClO5YXS3Azybq4hzh66YHvjC4oycxOLKuNB9seD7TcmrAIAMbE50w</recordid><startdate>20241017</startdate><enddate>20241017</enddate><creator>Haider, Fabiha</creator><creator>Shifat, Fariha Tanjim</creator><creator>Ishmam, Md Farhan</creator><creator>Barua, Deeparghya Dutta</creator><creator>Sourove, Md Sakib Ul Rahman</creator><creator>Fahim, Md</creator><creator>Alam, Md Farhad</creator><scope>AKY</scope><scope>GOX</scope></search><sort><creationdate>20241017</creationdate><title>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</title><author>Haider, Fabiha ; Shifat, Fariha Tanjim ; Ishmam, Md Farhan ; Barua, Deeparghya Dutta ; Sourove, Md Sakib Ul Rahman ; Fahim, Md ; Alam, Md Farhad</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-arxiv_primary_2410_132813</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2024</creationdate><topic>Computer Science - Computation and Language</topic><toplevel>online_resources</toplevel><creatorcontrib>Haider, Fabiha</creatorcontrib><creatorcontrib>Shifat, Fariha Tanjim</creatorcontrib><creatorcontrib>Ishmam, Md Farhan</creatorcontrib><creatorcontrib>Barua, Deeparghya Dutta</creatorcontrib><creatorcontrib>Sourove, Md Sakib Ul Rahman</creatorcontrib><creatorcontrib>Fahim, Md</creatorcontrib><creatorcontrib>Alam, Md Farhad</creatorcontrib><collection>arXiv Computer Science</collection><collection>arXiv.org</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext_linktorsrc</fulltext></delivery><addata><au>Haider, Fabiha</au><au>Shifat, Fariha Tanjim</au><au>Ishmam, Md Farhan</au><au>Barua, Deeparghya Dutta</au><au>Sourove, Md Sakib Ul Rahman</au><au>Fahim, Md</au><au>Alam, Md Farhad</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla</atitle><date>2024-10-17</date><risdate>2024</risdate><abstract>The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and origin, multi-label classification of hateful content can help in comprehending hate motivation and enhance content moderation. While previous efforts have focused on monolingual or binary hate classification tasks, no work has yet addressed the challenge of multi-label hate speech classification in transliterated Bangla. We introduce BanTH, the first multi-label transliterated Bangla hate speech dataset comprising 37.3k samples. The samples are sourced from YouTube comments, where each instance is labeled with one or more target groups, reflecting the regional demographic. We establish novel transformer encoder-based baselines by further pre-training on transliterated Bangla corpus. We also propose a novel translation-based LLM prompting strategy for transliterated text. Experiments reveal that our further pre-trained encoders are achieving state-of-the-art performance on the BanTH dataset, while our translation-based prompting outperforms other strategies in the zero-shot setting. The introduction of BanTH not only fills a critical gap in hate speech research for Bangla but also sets the stage for future exploration into code-mixed and multi-label classification challenges in underrepresented languages.</abstract><doi>10.48550/arxiv.2410.13281</doi><oa>free_for_read</oa></addata></record>
fulltext fulltext_linktorsrc
identifier DOI: 10.48550/arxiv.2410.13281
ispartof
issn
language eng
recordid cdi_arxiv_primary_2410_13281
source arXiv.org
subjects Computer Science - Computation and Language
title BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla
url https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-02-03T22%3A25%3A10IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-arxiv_GOX&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=BanTH:%20A%20Multi-label%20Hate%20Speech%20Detection%20Dataset%20for%20Transliterated%20Bangla&rft.au=Haider,%20Fabiha&rft.date=2024-10-17&rft_id=info:doi/10.48550/arxiv.2410.13281&rft_dat=%3Carxiv_GOX%3E2410_13281%3C/arxiv_GOX%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_id=info:pmid/&rfr_iscdi=true