On Language Clustering: A Non-parametric Statistical Approach

Any approach aimed at pasteurizing and quantifying a particular phenomenon must include the use of robust statistical methodologies for data analysis. With this in mind, the purpose of this study is to present statistical approaches that may be employed in nonparametric nonhomogeneous data framework...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	arXiv.org 2022-09
Hauptverfasser:	Chattopadhyay, Anagh, Ghosh, Soumya Sankar, Karmakar, Samir
Format:	Artikel
Sprache:	eng
Schlagworte:	Cartesian coordinates Classification Clustering Computer Science - Computation and Language Data analysis Data mining Language Multivariate statistical analysis Natural language processing Outliers (statistics) Pasteurizing Robustness Statistical analysis Statistics - Applications
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page
container_issue
container_start_page
container_title	arXiv.org
container_volume
creator	Chattopadhyay, Anagh Ghosh, Soumya Sankar Karmakar, Samir
description	Any approach aimed at pasteurizing and quantifying a particular phenomenon must include the use of robust statistical methodologies for data analysis. With this in mind, the purpose of this study is to present statistical approaches that may be employed in nonparametric nonhomogeneous data frameworks, as well as to examine their application in the field of natural language processing and language clustering. Furthermore, this paper discusses the many uses of nonparametric approaches in linguistic data mining and processing. The data depth idea allows for the centre-outward ordering of points in any dimension, resulting in a new nonparametric multivariate statistical analysis that does not require any distributional assumptions. The concept of hierarchy is used in historical language categorisation and structuring, and it aims to organise and cluster languages into subfamilies using the same premise. In this regard, the current study presents a novel approach to language family structuring based on non-parametric approaches produced from a typological structure of words in various languages, which is then converted into a Cartesian framework using MDS. This statistical-depth-based architecture allows for the use of data-depth-based methodologies for robust outlier detection, which is extremely useful in understanding the categorization of diverse borderline languages and allows for the re-evaluation of existing classification systems. Other depth-based approaches are also applied to processes such as unsupervised and supervised clustering. This paper therefore provides an overview of procedures that can be applied to nonhomogeneous language classification systems in a nonparametric framework.
doi_str_mv	10.48550/arxiv.2209.06720
format	Article
fullrecord	<record><control><sourceid>proquest_arxiv</sourceid><recordid>TN_cdi_arxiv_primary_2209_06720</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2714787606</sourcerecordid><originalsourceid>FETCH-LOGICAL-a950-28aec17ab2551fbe9f246918344fb3de827d27de8d9bca3cd18ad134a32425123</originalsourceid><addsrcrecordid>eNotj0FLxDAUhIMguKz7AzxZ8NyavKRNKngoRVehuAf3Xl7TtGbptjVNRf-9dVd4MIc3zMxHyA2jkVBxTO_RfduvCICmEU0k0AuyAs5ZqATAFdlM04FSCssnjvmKPO76oMC-nbE1Qd7NkzfO9u1DkAVvQx-O6PBovLM6ePfo7eStxi7IxtENqD-uyWWD3WQ2_7om--enff4SFrvta54VIaYxDUGh0UxitVSypjJpAyJJmeJCNBWvjQJZL2dUnVYaua6ZwppxgRwExAz4mtyeY09s5ejsEd1P-cdYnhgXx93Zsez6nM3ky8Mwu37ZVIJkQiqZ0IT_Amc6Uzs</addsrcrecordid><sourcetype>Open Access Repository</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2714787606</pqid></control><display><type>article</type><title>On Language Clustering: A Non-parametric Statistical Approach</title><source>arXiv.org</source><source>Free E- Journals</source><creator>Chattopadhyay, Anagh ; Ghosh, Soumya Sankar ; Karmakar, Samir</creator><creatorcontrib>Chattopadhyay, Anagh ; Ghosh, Soumya Sankar ; Karmakar, Samir</creatorcontrib><description>Any approach aimed at pasteurizing and quantifying a particular phenomenon must include the use of robust statistical methodologies for data analysis. With this in mind, the purpose of this study is to present statistical approaches that may be employed in nonparametric nonhomogeneous data frameworks, as well as to examine their application in the field of natural language processing and language clustering. Furthermore, this paper discusses the many uses of nonparametric approaches in linguistic data mining and processing. The data depth idea allows for the centre-outward ordering of points in any dimension, resulting in a new nonparametric multivariate statistical analysis that does not require any distributional assumptions. The concept of hierarchy is used in historical language categorisation and structuring, and it aims to organise and cluster languages into subfamilies using the same premise. In this regard, the current study presents a novel approach to language family structuring based on non-parametric approaches produced from a typological structure of words in various languages, which is then converted into a Cartesian framework using MDS. This statistical-depth-based architecture allows for the use of data-depth-based methodologies for robust outlier detection, which is extremely useful in understanding the categorization of diverse borderline languages and allows for the re-evaluation of existing classification systems. Other depth-based approaches are also applied to processes such as unsupervised and supervised clustering. This paper therefore provides an overview of procedures that can be applied to nonhomogeneous language classification systems in a nonparametric framework.</description><identifier>EISSN: 2331-8422</identifier><identifier>DOI: 10.48550/arxiv.2209.06720</identifier><language>eng</language><publisher>Ithaca: Cornell University Library, arXiv.org</publisher><subject>Cartesian coordinates ; Classification ; Clustering ; Computer Science - Computation and Language ; Data analysis ; Data mining ; Language ; Multivariate statistical analysis ; Natural language processing ; Outliers (statistics) ; Pasteurizing ; Robustness ; Statistical analysis ; Statistics - Applications</subject><ispartof>arXiv.org, 2022-09</ispartof><rights>2022. This work is published under http://creativecommons.org/licenses/by-nc-nd/4.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><rights>http://creativecommons.org/licenses/by-nc-nd/4.0</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>228,230,776,780,881,27904</link.rule.ids><backlink>$$Uhttps://doi.org/10.1007/978-3-031-27609-5_4$$DView published paper (Access to full text may be restricted)$$Hfree_for_read</backlink><backlink>$$Uhttps://doi.org/10.48550/arXiv.2209.06720$$DView paper in arXiv$$Hfree_for_read</backlink></links><search><creatorcontrib>Chattopadhyay, Anagh</creatorcontrib><creatorcontrib>Ghosh, Soumya Sankar</creatorcontrib><creatorcontrib>Karmakar, Samir</creatorcontrib><title>On Language Clustering: A Non-parametric Statistical Approach</title><title>arXiv.org</title><description>Any approach aimed at pasteurizing and quantifying a particular phenomenon must include the use of robust statistical methodologies for data analysis. With this in mind, the purpose of this study is to present statistical approaches that may be employed in nonparametric nonhomogeneous data frameworks, as well as to examine their application in the field of natural language processing and language clustering. Furthermore, this paper discusses the many uses of nonparametric approaches in linguistic data mining and processing. The data depth idea allows for the centre-outward ordering of points in any dimension, resulting in a new nonparametric multivariate statistical analysis that does not require any distributional assumptions. The concept of hierarchy is used in historical language categorisation and structuring, and it aims to organise and cluster languages into subfamilies using the same premise. In this regard, the current study presents a novel approach to language family structuring based on non-parametric approaches produced from a typological structure of words in various languages, which is then converted into a Cartesian framework using MDS. This statistical-depth-based architecture allows for the use of data-depth-based methodologies for robust outlier detection, which is extremely useful in understanding the categorization of diverse borderline languages and allows for the re-evaluation of existing classification systems. Other depth-based approaches are also applied to processes such as unsupervised and supervised clustering. This paper therefore provides an overview of procedures that can be applied to nonhomogeneous language classification systems in a nonparametric framework.</description><subject>Cartesian coordinates</subject><subject>Classification</subject><subject>Clustering</subject><subject>Computer Science - Computation and Language</subject><subject>Data analysis</subject><subject>Data mining</subject><subject>Language</subject><subject>Multivariate statistical analysis</subject><subject>Natural language processing</subject><subject>Outliers (statistics)</subject><subject>Pasteurizing</subject><subject>Robustness</subject><subject>Statistical analysis</subject><subject>Statistics - Applications</subject><issn>2331-8422</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2022</creationdate><recordtype>article</recordtype><sourceid>ABUWG</sourceid><sourceid>AFKRA</sourceid><sourceid>AZQEC</sourceid><sourceid>BENPR</sourceid><sourceid>CCPQU</sourceid><sourceid>DWQXO</sourceid><sourceid>GOX</sourceid><recordid>eNotj0FLxDAUhIMguKz7AzxZ8NyavKRNKngoRVehuAf3Xl7TtGbptjVNRf-9dVd4MIc3zMxHyA2jkVBxTO_RfduvCICmEU0k0AuyAs5ZqATAFdlM04FSCssnjvmKPO76oMC-nbE1Qd7NkzfO9u1DkAVvQx-O6PBovLM6ePfo7eStxi7IxtENqD-uyWWD3WQ2_7om--enff4SFrvta54VIaYxDUGh0UxitVSypjJpAyJJmeJCNBWvjQJZL2dUnVYaua6ZwppxgRwExAz4mtyeY09s5ejsEd1P-cdYnhgXx93Zsez6nM3ky8Mwu37ZVIJkQiqZ0IT_Amc6Uzs</recordid><startdate>20220914</startdate><enddate>20220914</enddate><creator>Chattopadhyay, Anagh</creator><creator>Ghosh, Soumya Sankar</creator><creator>Karmakar, Samir</creator><general>Cornell University Library, arXiv.org</general><scope>8FE</scope><scope>8FG</scope><scope>ABJCF</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>HCIFZ</scope><scope>L6V</scope><scope>M7S</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>PTHSS</scope><scope>AKY</scope><scope>EPD</scope><scope>GOX</scope></search><sort><creationdate>20220914</creationdate><title>On Language Clustering: A Non-parametric Statistical Approach</title><author>Chattopadhyay, Anagh ; Ghosh, Soumya Sankar ; Karmakar, Samir</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-LOGICAL-a950-28aec17ab2551fbe9f246918344fb3de827d27de8d9bca3cd18ad134a32425123</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2022</creationdate><topic>Cartesian coordinates</topic><topic>Classification</topic><topic>Clustering</topic><topic>Computer Science - Computation and Language</topic><topic>Data analysis</topic><topic>Data mining</topic><topic>Language</topic><topic>Multivariate statistical analysis</topic><topic>Natural language processing</topic><topic>Outliers (statistics)</topic><topic>Pasteurizing</topic><topic>Robustness</topic><topic>Statistical analysis</topic><topic>Statistics - Applications</topic><toplevel>online_resources</toplevel><creatorcontrib>Chattopadhyay, Anagh</creatorcontrib><creatorcontrib>Ghosh, Soumya Sankar</creatorcontrib><creatorcontrib>Karmakar, Samir</creatorcontrib><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>Materials Science & Engineering Collection</collection><collection>ProQuest Central (Alumni Edition)</collection><collection>ProQuest Central UK/Ireland</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central Korea</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Engineering Collection</collection><collection>Engineering Database</collection><collection>Publicly Available Content Database</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>Engineering Collection</collection><collection>arXiv Computer Science</collection><collection>arXiv Statistics</collection><collection>arXiv.org</collection><jtitle>arXiv.org</jtitle></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Chattopadhyay, Anagh</au><au>Ghosh, Soumya Sankar</au><au>Karmakar, Samir</au><format>journal</format><genre>article</genre><ristype>JOUR</ristype><atitle>On Language Clustering: A Non-parametric Statistical Approach</atitle><jtitle>arXiv.org</jtitle><date>2022-09-14</date><risdate>2022</risdate><eissn>2331-8422</eissn><abstract>Any approach aimed at pasteurizing and quantifying a particular phenomenon must include the use of robust statistical methodologies for data analysis. With this in mind, the purpose of this study is to present statistical approaches that may be employed in nonparametric nonhomogeneous data frameworks, as well as to examine their application in the field of natural language processing and language clustering. Furthermore, this paper discusses the many uses of nonparametric approaches in linguistic data mining and processing. The data depth idea allows for the centre-outward ordering of points in any dimension, resulting in a new nonparametric multivariate statistical analysis that does not require any distributional assumptions. The concept of hierarchy is used in historical language categorisation and structuring, and it aims to organise and cluster languages into subfamilies using the same premise. In this regard, the current study presents a novel approach to language family structuring based on non-parametric approaches produced from a typological structure of words in various languages, which is then converted into a Cartesian framework using MDS. This statistical-depth-based architecture allows for the use of data-depth-based methodologies for robust outlier detection, which is extremely useful in understanding the categorization of diverse borderline languages and allows for the re-evaluation of existing classification systems. Other depth-based approaches are also applied to processes such as unsupervised and supervised clustering. This paper therefore provides an overview of procedures that can be applied to nonhomogeneous language classification systems in a nonparametric framework.</abstract><cop>Ithaca</cop><pub>Cornell University Library, arXiv.org</pub><doi>10.48550/arxiv.2209.06720</doi><oa>free_for_read</oa></addata></record>
fulltext	fulltext
identifier	EISSN: 2331-8422
ispartof	arXiv.org, 2022-09
issn	2331-8422
language	eng
recordid	cdi_arxiv_primary_2209_06720
source	arXiv.org; Free E- Journals
subjects	Cartesian coordinates Classification Clustering Computer Science - Computation and Language Data analysis Data mining Language Multivariate statistical analysis Natural language processing Outliers (statistics) Pasteurizing Robustness Statistical analysis Statistics - Applications
title	On Language Clustering: A Non-parametric Statistical Approach
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-25T14%3A43%3A10IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest_arxiv&rft_val_fmt=info:ofi/fmt:kev:mtx:journal&rft.genre=article&rft.atitle=On%20Language%20Clustering:%20A%20Non-parametric%20Statistical%20Approach&rft.jtitle=arXiv.org&rft.au=Chattopadhyay,%20Anagh&rft.date=2022-09-14&rft.eissn=2331-8422&rft_id=info:doi/10.48550/arxiv.2209.06720&rft_dat=%3Cproquest_arxiv%3E2714787606%3C/proquest_arxiv%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=2714787606&rft_id=info:pmid/&rfr_iscdi=true