CorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction

This study propose a fully automated system for speech correction and accent reduction. Consider the application scenario that a recorded speech audio contains certain errors, e.g., inappropriate words, mispronunciations, that need to be corrected. The proposed system, named CorrectSpeech, performs...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	arXiv.org 2022-10
Hauptverfasser:	Tan, Daxin, Deng, Liqun, Zheng, Nianzu, Yeung, Yu Ting, Jiang, Xin, Chen, Xiao, Tan, Lee
Format:	Artikel
Sprache:	eng
Schlagworte:	Automation Editing Reduction Speech Speech recognition
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

container_end_page
container_issue
container_start_page
container_title	arXiv.org
container_volume
creator	Tan, Daxin Deng, Liqun Zheng, Nianzu Yeung, Yu Ting Jiang, Xin Chen, Xiao Tan, Lee
description	This study propose a fully automated system for speech correction and accent reduction. Consider the application scenario that a recorded speech audio contains certain errors, e.g., inappropriate words, mispronunciations, that need to be corrected. The proposed system, named CorrectSpeech, performs the correction in three steps: recognizing the recorded speech and converting it into time-stamped symbol sequence, aligning recognized symbol sequence with target text to determine locations and types of required edit operations, and generating the corrected speech. Experiments show that the quality and naturalness of corrected speech depend on the performance of speech recognition and alignment modules, as well as the granularity level of editing operations. The proposed system is evaluated on two corpora: a manually perturbed version of VCTK and L2-ARCTIC. The results demonstrate that our system is able to correct mispronunciation and reduce accent in speech recordings. Audio samples are available online for demonstration https://daxintan-cuhk.github.io/CorrectSpeech/ .
format	Article
fullrecord	<record><control><sourceid>proquest</sourceid><recordid>TN_cdi_proquest_journals_2649832662</recordid><sourceformat>XML</sourceformat><sourcesystem>PC</sourcesystem><sourcerecordid>2649832662</sourcerecordid><originalsourceid>FETCH-proquest_journals_26498326623</originalsourceid><addsrcrecordid>eNqNi80KgkAURocgSMp3uNBasDs6WTuRpLW2DxmvlOiMzc_Cty_KB2j1wTnnW7EAOT9EWYK4YaG1fRzHKI6YpjxgVaGNIenqiUg-zpBD6Ydhhtw7PTaOWqhn62iEThv4RbBcnlpBo1rIpSTloKLWf-GOrbtmsBQuu2X78nIrrtFk9MuTdfdee6M-6o4iOWUchUD-X_UGn8k_Ng</addsrcrecordid><sourcetype>Aggregation Database</sourcetype><iscdi>true</iscdi><recordtype>article</recordtype><pqid>2649832662</pqid></control><display><type>article</type><title>CorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction</title><source>Freely Accessible Journals</source><creator>Tan, Daxin ; Deng, Liqun ; Zheng, Nianzu ; Yeung, Yu Ting ; Jiang, Xin ; Chen, Xiao ; Tan, Lee</creator><creatorcontrib>Tan, Daxin ; Deng, Liqun ; Zheng, Nianzu ; Yeung, Yu Ting ; Jiang, Xin ; Chen, Xiao ; Tan, Lee</creatorcontrib><description>This study propose a fully automated system for speech correction and accent reduction. Consider the application scenario that a recorded speech audio contains certain errors, e.g., inappropriate words, mispronunciations, that need to be corrected. The proposed system, named CorrectSpeech, performs the correction in three steps: recognizing the recorded speech and converting it into time-stamped symbol sequence, aligning recognized symbol sequence with target text to determine locations and types of required edit operations, and generating the corrected speech. Experiments show that the quality and naturalness of corrected speech depend on the performance of speech recognition and alignment modules, as well as the granularity level of editing operations. The proposed system is evaluated on two corpora: a manually perturbed version of VCTK and L2-ARCTIC. The results demonstrate that our system is able to correct mispronunciation and reduce accent in speech recordings. Audio samples are available online for demonstration https://daxintan-cuhk.github.io/CorrectSpeech/ .</description><identifier>EISSN: 2331-8422</identifier><language>eng</language><publisher>Ithaca: Cornell University Library, arXiv.org</publisher><subject>Automation ; Editing ; Reduction ; Speech ; Speech recognition</subject><ispartof>arXiv.org, 2022-10</ispartof><rights>2022. This work is published under http://arxiv.org/licenses/nonexclusive-distrib/1.0/ (the “License”). Notwithstanding the ProQuest Terms and Conditions, you may use this content in accordance with the terms of the License.</rights><oa>free_for_read</oa><woscitedreferencessubscribed>false</woscitedreferencessubscribed></display><links><openurl>$$Topenurl_article</openurl><openurlfulltext>$$Topenurlfull_article</openurlfulltext><thumbnail>$$Tsyndetics_thumb_exl</thumbnail><link.rule.ids>776,780</link.rule.ids></links><search><creatorcontrib>Tan, Daxin</creatorcontrib><creatorcontrib>Deng, Liqun</creatorcontrib><creatorcontrib>Zheng, Nianzu</creatorcontrib><creatorcontrib>Yeung, Yu Ting</creatorcontrib><creatorcontrib>Jiang, Xin</creatorcontrib><creatorcontrib>Chen, Xiao</creatorcontrib><creatorcontrib>Tan, Lee</creatorcontrib><title>CorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction</title><title>arXiv.org</title><description>This study propose a fully automated system for speech correction and accent reduction. Consider the application scenario that a recorded speech audio contains certain errors, e.g., inappropriate words, mispronunciations, that need to be corrected. The proposed system, named CorrectSpeech, performs the correction in three steps: recognizing the recorded speech and converting it into time-stamped symbol sequence, aligning recognized symbol sequence with target text to determine locations and types of required edit operations, and generating the corrected speech. Experiments show that the quality and naturalness of corrected speech depend on the performance of speech recognition and alignment modules, as well as the granularity level of editing operations. The proposed system is evaluated on two corpora: a manually perturbed version of VCTK and L2-ARCTIC. The results demonstrate that our system is able to correct mispronunciation and reduce accent in speech recordings. Audio samples are available online for demonstration https://daxintan-cuhk.github.io/CorrectSpeech/ .</description><subject>Automation</subject><subject>Editing</subject><subject>Reduction</subject><subject>Speech</subject><subject>Speech recognition</subject><issn>2331-8422</issn><fulltext>true</fulltext><rsrctype>article</rsrctype><creationdate>2022</creationdate><recordtype>article</recordtype><sourceid>ABUWG</sourceid><sourceid>AFKRA</sourceid><sourceid>AZQEC</sourceid><sourceid>BENPR</sourceid><sourceid>CCPQU</sourceid><sourceid>DWQXO</sourceid><recordid>eNqNi80KgkAURocgSMp3uNBasDs6WTuRpLW2DxmvlOiMzc_Cty_KB2j1wTnnW7EAOT9EWYK4YaG1fRzHKI6YpjxgVaGNIenqiUg-zpBD6Ydhhtw7PTaOWqhn62iEThv4RbBcnlpBo1rIpSTloKLWf-GOrbtmsBQuu2X78nIrrtFk9MuTdfdee6M-6o4iOWUchUD-X_UGn8k_Ng</recordid><startdate>20221014</startdate><enddate>20221014</enddate><creator>Tan, Daxin</creator><creator>Deng, Liqun</creator><creator>Zheng, Nianzu</creator><creator>Yeung, Yu Ting</creator><creator>Jiang, Xin</creator><creator>Chen, Xiao</creator><creator>Tan, Lee</creator><general>Cornell University Library, arXiv.org</general><scope>8FE</scope><scope>8FG</scope><scope>ABJCF</scope><scope>ABUWG</scope><scope>AFKRA</scope><scope>AZQEC</scope><scope>BENPR</scope><scope>BGLVJ</scope><scope>CCPQU</scope><scope>DWQXO</scope><scope>HCIFZ</scope><scope>L6V</scope><scope>M7S</scope><scope>PIMPY</scope><scope>PQEST</scope><scope>PQQKQ</scope><scope>PQUKI</scope><scope>PRINS</scope><scope>PTHSS</scope></search><sort><creationdate>20221014</creationdate><title>CorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction</title><author>Tan, Daxin ; Deng, Liqun ; Zheng, Nianzu ; Yeung, Yu Ting ; Jiang, Xin ; Chen, Xiao ; Tan, Lee</author></sort><facets><frbrtype>5</frbrtype><frbrgroupid>cdi_FETCH-proquest_journals_26498326623</frbrgroupid><rsrctype>articles</rsrctype><prefilter>articles</prefilter><language>eng</language><creationdate>2022</creationdate><topic>Automation</topic><topic>Editing</topic><topic>Reduction</topic><topic>Speech</topic><topic>Speech recognition</topic><toplevel>online_resources</toplevel><creatorcontrib>Tan, Daxin</creatorcontrib><creatorcontrib>Deng, Liqun</creatorcontrib><creatorcontrib>Zheng, Nianzu</creatorcontrib><creatorcontrib>Yeung, Yu Ting</creatorcontrib><creatorcontrib>Jiang, Xin</creatorcontrib><creatorcontrib>Chen, Xiao</creatorcontrib><creatorcontrib>Tan, Lee</creatorcontrib><collection>ProQuest SciTech Collection</collection><collection>ProQuest Technology Collection</collection><collection>Materials Science & Engineering Collection</collection><collection>ProQuest Central (Alumni Edition)</collection><collection>ProQuest Central UK/Ireland</collection><collection>ProQuest Central Essentials</collection><collection>ProQuest Central</collection><collection>Technology Collection</collection><collection>ProQuest One Community College</collection><collection>ProQuest Central Korea</collection><collection>SciTech Premium Collection</collection><collection>ProQuest Engineering Collection</collection><collection>Engineering Database</collection><collection>Publicly Available Content Database</collection><collection>ProQuest One Academic Eastern Edition (DO NOT USE)</collection><collection>ProQuest One Academic</collection><collection>ProQuest One Academic UKI Edition</collection><collection>ProQuest Central China</collection><collection>Engineering Collection</collection></facets><delivery><delcategory>Remote Search Resource</delcategory><fulltext>fulltext</fulltext></delivery><addata><au>Tan, Daxin</au><au>Deng, Liqun</au><au>Zheng, Nianzu</au><au>Yeung, Yu Ting</au><au>Jiang, Xin</au><au>Chen, Xiao</au><au>Tan, Lee</au><format>book</format><genre>document</genre><ristype>GEN</ristype><atitle>CorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction</atitle><jtitle>arXiv.org</jtitle><date>2022-10-14</date><risdate>2022</risdate><eissn>2331-8422</eissn><abstract>This study propose a fully automated system for speech correction and accent reduction. Consider the application scenario that a recorded speech audio contains certain errors, e.g., inappropriate words, mispronunciations, that need to be corrected. The proposed system, named CorrectSpeech, performs the correction in three steps: recognizing the recorded speech and converting it into time-stamped symbol sequence, aligning recognized symbol sequence with target text to determine locations and types of required edit operations, and generating the corrected speech. Experiments show that the quality and naturalness of corrected speech depend on the performance of speech recognition and alignment modules, as well as the granularity level of editing operations. The proposed system is evaluated on two corpora: a manually perturbed version of VCTK and L2-ARCTIC. The results demonstrate that our system is able to correct mispronunciation and reduce accent in speech recordings. Audio samples are available online for demonstration https://daxintan-cuhk.github.io/CorrectSpeech/ .</abstract><cop>Ithaca</cop><pub>Cornell University Library, arXiv.org</pub><oa>free_for_read</oa></addata></record>
fulltext	fulltext
identifier	EISSN: 2331-8422
ispartof	arXiv.org, 2022-10
issn	2331-8422
language	eng
recordid	cdi_proquest_journals_2649832662
source	Freely Accessible Journals
subjects	Automation Editing Reduction Speech Speech recognition
title	CorrectSpeech: A Fully Automated System for Speech Correction and Accent Reduction
url	https://sfx.bib-bvb.de/sfx_tum?ctx_ver=Z39.88-2004&ctx_enc=info:ofi/enc:UTF-8&ctx_tim=2025-01-23T17%3A19%3A04IST&url_ver=Z39.88-2004&url_ctx_fmt=infofi/fmt:kev:mtx:ctx&rfr_id=info:sid/primo.exlibrisgroup.com:primo3-Article-proquest&rft_val_fmt=info:ofi/fmt:kev:mtx:book&rft.genre=document&rft.atitle=CorrectSpeech:%20A%20Fully%20Automated%20System%20for%20Speech%20Correction%20and%20Accent%20Reduction&rft.jtitle=arXiv.org&rft.au=Tan,%20Daxin&rft.date=2022-10-14&rft.eissn=2331-8422&rft_id=info:doi/&rft_dat=%3Cproquest%3E2649832662%3C/proquest%3E%3Curl%3E%3C/url%3E&disable_directlink=true&sfx.directlink=off&sfx.report_link=0&rft_id=info:oai/&rft_pqid=2649832662&rft_id=info:pmid/&rfr_iscdi=true