Improved POS Tagging Model for Malay Twitter Data based on Machine Learning Algorithm
Twitter is a popular social media platform in Malaysia that allows for 280-character microblogging. Almost everything that happens in a single day is tweeted by users. Because of the popularity of Twitter, most Malaysians use it daily, providing researchers and developers with a wealth of data on Ma...
Gespeichert in:
Veröffentlicht in: | International journal of advanced computer science & applications 2022-01, Vol.13 (7) |
---|---|
Hauptverfasser: | , |
Format: | Artikel |
Sprache: | eng |
Schlagworte: | |
Online-Zugang: | Volltext |
Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
container_end_page | |
---|---|
container_issue | 7 |
container_start_page | |
container_title | International journal of advanced computer science & applications |
container_volume | 13 |
creator | Ariffin, Siti Noor Allia Noor Tiun, Sabrina |
description | Twitter is a popular social media platform in Malaysia that allows for 280-character microblogging. Almost everything that happens in a single day is tweeted by users. Because of the popularity of Twitter, most Malaysians use it daily, providing researchers and developers with a wealth of data on Malaysian users. This paper explains why and how this study chose to create a new Malay Twitter corpus, Malay Part-of-Speech (POS) tags, and a Malay POS tagger model. The goal of this paper is to improve existing Malay POS tags so that they are more compatible with the newly created Malay Twitter corpus, as well as to build a POS tagging model specifically tailored for Malay Twitter data using various machine learning algorithms. For instance, Support Vector Machine (SVM), Naïve Bayes (NB), Decision Tree (DT), and K-Nearest Neighbor (KNN) classifiers. This study’s data was gathered by using Twitter's Advanced Search function and relevant and related keywords associated with informal Malay. The data was fed into machine learning algorithms after several stages of processing to serve as the training and testing corpus. The evaluation and analysis of the developed Malay POS tagger model show that the SVM classifier, as well as the newly proposed Malay POS tags, is the best machine learning algorithm for Malay Twitter data. Furthermore, the prediction accuracy and POS tagging results show that this research outperformed a comparable previous study, indicating that the Malay POS tagger model and its POS were successfully improved. |
doi_str_mv | 10.14569/IJACSA.2022.0130730 |
format | Article |
fullrecord | <record><control><sourceid>proquest_cross</sourceid><recordid>TN_cdi_proquest_journals_2707473345</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2707473345</sourcerecordid><originalsourceid>FETCH-LOGICAL-c204t-c5f9fa79c42a1ba17e76f52b01dff9100e26389a7298fbaa41d3d167972aa1613</originalsourceid><addsrcrecordid>eNotkEtrwzAMgM3YYKXrP9jBsHM6P2I7Pobu1dHSQVvYzSiJnaakceekG_33Sx-6SKBPEvoQeqRkTGMh9fP0M50s0zEjjI0J5URxcoMGjAoZCaHI7blOIkrU9z0ate2W9ME1kwkfoPV0tw_-1xb4a7HEKyjLqinx3Be2xs4HPIcajnj1V3WdDfgFOsAZtD3um76Xb6rG4pmF0JzG0rr0oeo2uwd056Bu7eiah2j99rqafESzxft0ks6inJG4i3LhtAOl85gBzYAqq6QTLCO0cE5TQiyTPNGgmE5cBhDTghdUKq0YAJWUD9HTZW__w8_Btp3Z-kNo-pOGKaJixXkseiq-UHnwbRusM_tQ7SAcDSXm7NBcHJqTQ3N1yP8BbhhjWQ</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2707473345</pqid></control><display><type>article</type><title>Improved POS Tagging Model for Malay Twitter Data based on Machine Learning Algorithm</title><source>EZB-FREE-00999 freely available EZB journals</source><creator>Ariffin, Siti Noor Allia Noor ; Tiun, Sabrina</creator><creatorcontrib>Ariffin, Siti Noor Allia Noor ; Tiun, Sabrina</creatorcontrib><description>Twitter is a popular social media platform in Malaysia that allows for 280-character microblogging. Almost everything that happens in a single day is tweeted by users. Because of the popularity of Twitter, most Malaysians use it daily, providing researchers and developers with a wealth of data on Malaysian users. This paper explains why and how this study chose to create a new Malay Twitter corpus, Malay Part-of-Speech (POS) tags, and a Malay POS tagger model. The goal of this paper is to improve existing Malay POS tags so that they are more compatible with the newly created Malay Twitter corpus, as well as to build a POS tagging model specifically tailored for Malay Twitter data using various machine learning algorithms. For instance, Support Vector Machine (SVM), Naïve Bayes (NB), Decision Tree (DT), and K-Nearest Neighbor (KNN) classifiers. This study’s data was gathered by using Twitter's Advanced Search function and relevant and related keywords associated with informal Malay. The data was fed into machine learning algorithms after several stages of processing to serve as the training and testing corpus. The evaluation and analysis of the developed Malay POS tagger model show that the SVM classifier, as well as the newly proposed Malay POS tags, is the best machine learning algorithm for Malay Twitter data. Furthermore, the prediction accuracy and POS tagging results show that this research outperformed a comparable previous study, indicating that the Malay POS tagger model and its POS were successfully improved.</description><identifier>ISSN: 2158-107X</identifier><identifier>EISSN: 2156-5570</identifier><identifier>DOI: 10.14569/IJACSA.2022.0130730</identifier><language>eng</language><publisher>West Yorkshire: Science and Information (SAI) Organization Limited</publisher><subject>Algorithms ; Classifiers ; Decision trees ; Machine learning ; Marking ; Social networks ; Support vector machines ; Tags</subject><ispartof>International journal of advanced computer science & applications, 2022-01, Vol.13 (7)</ispartof><rights>2022. This work is licensed under https://creativecommons.org/licenses/by/4.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>314,776,780,27901,27902</link.rule.ids></links><search><creatorcontrib>Ariffin, Siti Noor Allia Noor</creatorcontrib><creatorcontrib>Tiun, Sabrina</creatorcontrib><title>Improved POS Tagging Model for Malay Twitter Data based on Machine Learning Algorithm</title><title>International journal of advanced computer science & applications</title><description>Twitter is a popular social media platform in Malaysia that allows for 280-character microblogging. Almost everything that happens in a single day is tweeted by users. Because of the popularity of Twitter, most Malaysians use it daily, providing researchers and developers with a wealth of data on Malaysian users. This paper explains why and how this study chose to create a new Malay Twitter corpus, Malay Part-of-Speech (POS) tags, and a Malay POS tagger model. The goal of this paper is to improve existing Malay POS tags so that they are more compatible with the newly created Malay Twitter corpus, as well as to build a POS tagging model specifically tailored for Malay Twitter data using various machine learning algorithms. For instance, Support Vector Machine (SVM), Naïve Bayes (NB), Decision Tree (DT), and K-Nearest Neighbor (KNN) classifiers. This study’s data was gathered by using Twitter's Advanced Search function and relevant and related keywords associated with informal Malay. The data was fed into machine learning algorithms after several stages of processing to serve as the training and testing corpus. The evaluation and analysis of the developed Malay POS tagger model show that the SVM classifier, as well as the newly proposed Malay POS tags, is the best machine learning algorithm for Malay Twitter data. Furthermore, the prediction accuracy and POS tagging results show that this research outperformed a comparable previous study, indicating that the Malay POS tagger model and its POS were successfully improved.</description><subject>Algorithms</subject><subject>Classifiers</subject><subject>Decision trees</subject><subject>Machine learning</subject><subject>Marking</subject><subject>Social networks</subject><subject>Support vector machines</subject><subject>Tags</subject><issn>2158-107X</issn><issn>2156-5570</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2022</creationdate><recordtype>article</recordtype><sourceid>8G5</sourceid><sourceid>BENPR</sourceid><sourceid>GUQSH</sourceid><sourceid>M2O</sourceid><recordid>eNotkEtrwzAMgM3YYKXrP9jBsHM6P2I7Pobu1dHSQVvYzSiJnaakceekG_33Sx-6SKBPEvoQeqRkTGMh9fP0M50s0zEjjI0J5URxcoMGjAoZCaHI7blOIkrU9z0ate2W9ME1kwkfoPV0tw_-1xb4a7HEKyjLqinx3Be2xs4HPIcajnj1V3WdDfgFOsAZtD3um76Xb6rG4pmF0JzG0rr0oeo2uwd056Bu7eiah2j99rqafESzxft0ks6inJG4i3LhtAOl85gBzYAqq6QTLCO0cE5TQiyTPNGgmE5cBhDTghdUKq0YAJWUD9HTZW__w8_Btp3Z-kNo-pOGKaJixXkseiq-UHnwbRusM_tQ7SAcDSXm7NBcHJqTQ3N1yP8BbhhjWQ</recordid><startdate>20220101</startdate><enddate>20220101</enddate><creator>Ariffin, Siti Noor Allia Noor</creator><creator>Tiun, Sabrina</creator><general>Science and Information (SAI) Organization Limited</general><scope>AAYXX</scope><scope>CITATION</scope><scope>3V.</scope><scope>7XB</scope><scope>8FE</scope><scope>8FG</scope><scope>8FK</scope><scope>8G5</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>ARAPS</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>GNUQQ</scope><scope>GUQSH</scope><scope>HCIFZ</scope><scope>JQ2</scope><scope>K7-</scope><scope>M2O</scope><scope>MBDVC</scope><scope>P5Z</scope><scope>P62</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>Q9U</scope></search><sort><creationdate>20220101</creationdate><title>Improved POS Tagging Model for Malay Twitter Data based on Machine Learning Algorithm</title><author>Ariffin, Siti Noor Allia Noor ; Tiun, Sabrina</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-LOGICAL-c204t-c5f9fa79c42a1ba17e76f52b01dff9100e26389a7298fbaa41d3d167972aa1613</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2022</creationdate><topic>Algorithms</topic><topic>Classifiers</topic><topic>Decision trees</topic><topic>Machine learning</topic><topic>Marking</topic><topic>Social networks</topic><topic>Support vector machines</topic><topic>Tags</topic><toplevel>online_resources</toplevel><creatorcontrib>Ariffin, Siti Noor Allia Noor</creatorcontrib><creatorcontrib>Tiun, Sabrina</creatorcontrib><collection>CrossRef</collection><collection>ProQuest Central (Corporate)</collection><collection>ProQuest Central (purchase pre-March 2016)</collection><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>ProQuest Central (Alumni) (purchase pre-March 2016)</collection><collection>Research Library (Alumni Edition)</collection><collection>ProQuest Central (Alumni Edition)</collection><collection>ProQuest Central UK/Ireland</collection><collection>Advanced Technologies & Aerospace Collection</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central Korea</collection><collection>ProQuest Central Student</collection><collection>Research Library Prep</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Computer Science Collection</collection><collection>Computer Science Database</collection><collection>Research Library</collection><collection>Research Library (Corporate)</collection><collection>Advanced Technologies & Aerospace Database</collection><collection>ProQuest Advanced Technologies & Aerospace Collection</collection><collection>Publicly Available Content Database</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>ProQuest Central Basic</collection><jtitle>International journal of advanced computer science & applications</jtitle></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Ariffin, Siti Noor Allia Noor</au><au>Tiun, Sabrina</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>Improved POS Tagging Model for Malay Twitter Data based on Machine Learning Algorithm</atitle><jtitle>International journal of advanced computer science & applications</jtitle><date>2022-01-01</date><risdate>2022</risdate><volume>13</volume><issue>7</issue><issn>2158-107X</issn><eissn>2156-5570</eissn><abstract>Twitter is a popular social media platform in Malaysia that allows for 280-character microblogging. Almost everything that happens in a single day is tweeted by users. Because of the popularity of Twitter, most Malaysians use it daily, providing researchers and developers with a wealth of data on Malaysian users. This paper explains why and how this study chose to create a new Malay Twitter corpus, Malay Part-of-Speech (POS) tags, and a Malay POS tagger model. The goal of this paper is to improve existing Malay POS tags so that they are more compatible with the newly created Malay Twitter corpus, as well as to build a POS tagging model specifically tailored for Malay Twitter data using various machine learning algorithms. For instance, Support Vector Machine (SVM), Naïve Bayes (NB), Decision Tree (DT), and K-Nearest Neighbor (KNN) classifiers. This study’s data was gathered by using Twitter's Advanced Search function and relevant and related keywords associated with informal Malay. The data was fed into machine learning algorithms after several stages of processing to serve as the training and testing corpus. The evaluation and analysis of the developed Malay POS tagger model show that the SVM classifier, as well as the newly proposed Malay POS tags, is the best machine learning algorithm for Malay Twitter data. Furthermore, the prediction accuracy and POS tagging results show that this research outperformed a comparable previous study, indicating that the Malay POS tagger model and its POS were successfully improved.</abstract><cop>West Yorkshire</cop><pub>Science and Information (SAI) Organization Limited</pub><doi>10.14569/IJACSA.2022.0130730</doi><oa>free_for_read</oa></addata></record> |
fulltext | fulltext |
identifier | ISSN: 2158-107X |
ispartof | International journal of advanced computer science & applications, 2022-01, Vol.13 (7) |
issn | 2158-107X 2156-5570 |
language | eng |
recordid | cdi_proquest_journals_2707473345 |
source | EZB-FREE-00999 freely available EZB journals |
subjects | Algorithms Classifiers Decision trees Machine learning Marking Social networks Support vector machines Tags |
title | Improved POS Tagging Model for Malay Twitter Data based on Machine Learning Algorithm |
url | https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-02-05T05%3A45%3A10IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest_cross&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=Improved%20POS%20Tagging%20Model%20for%20Malay%20Twitter%20Data%20based%20on%20Machine%20Learning%20Algorithm&rft.jtitle=International%20journal%20of%20advanced%20computer%20science%20&%20applications&rft.au=Ariffin,%20Siti%20Noor%20Allia%20Noor&rft.date=2022-01-01&rft.volume=13&rft.issue=7&rft.issn=2158-107X&rft.eissn=2156-5570&rft_id=info:doi/10.14569/IJACSA.2022.0130730&rft_dat=%3Cproquest_cross%3E2707473345%3C/proquest_cross%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=2707473345&rft_id=info:pmid/&rfr_iscdi=true |