Speech Recognition of Accented Mandarin Based on Improved Conformer

The convolution module in Conformer is capable of providing translationally invariant convolution in time and space. This is often used in Mandarin recognition tasks to address the diversity of speech signals by treating the time-frequency maps of speech signals as images. However, convolutional net...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Sensors (Basel, Switzerland) Switzerland), 2023-04, Vol.23 (8), p.4025
Hauptverfasser:	Yang, Xing-Yao, Zhang, Shao-Dong, Xiao, Rui, Yu, Jiong, Li, Zi-Yang
Format:	Artikel
Sprache:	eng
Schlagworte:	Accentuation Algorithms Channels Conformer Convolution Datasets Dialects Dictionaries Error analysis Error reduction Feature maps Language Mandarin mandarin accent Neural networks Recognition, Psychology Regression analysis Sequences spectrogram Speech Speech Perception Speech recognition temporal convolutional network Time Voice recognition
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	The convolution module in Conformer is capable of providing translationally invariant convolution in time and space. This is often used in Mandarin recognition tasks to address the diversity of speech signals by treating the time-frequency maps of speech signals as images. However, convolutional networks are more effective in local feature modeling, while dialect recognition tasks require the extraction of a long sequence of contextual information features; therefore, the SE-Conformer-TCN is proposed in this paper. By embedding the squeeze-excitation block into the Conformer, the interdependence between the features of channels can be explicitly modeled to enhance the model's ability to select interrelated channels, thus increasing the weight of effective speech spectrogram features and decreasing the weight of ineffective or less effective feature maps. The multi-head self-attention and temporal convolutional network is built in parallel, in which the dilated causal convolutions module can cover the input time series by increasing the expansion factor and convolutional kernel to capture the location information implied between the sequences and enhance the model's access to location information. Experiments on four public datasets demonstrate that the proposed model has a higher performance for the recognition of Mandarin with an accent, and the sentence error rate is reduced by 2.1% compared to the Conformer, with only 4.9% character error rate.
ISSN:	1424-8220 1424-8220
DOI:	10.3390/s23084025