Speech Models Training Technologies Comparison Using Word Error Rate

The main purpose of this work is to analyze and compare several technologies used for training speech models, including traditional approaches as Hidden Markov Models (HMMs) and more recent methods as Deep Neural Networks (DNNs). The technologies have been explained and compared using word error rat...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Advances in Cyber-Physical Systems 2023-05, Vol.8 (1), p.74-80
Hauptverfasser:	Yakubovskyi, Roman, Morozov, Yuriy
Format:	Artikel
Sprache:	eng
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	The main purpose of this work is to analyze and compare several technologies used for training speech models, including traditional approaches as Hidden Markov Models (HMMs) and more recent methods as Deep Neural Networks (DNNs). The technologies have been explained and compared using word error rate metric based on the input of 1000 words by a user with 15 decibel background noise. Word error rate metric has been ex- plained and calculated. Potential replacements for com- pared technologies have been provided, including: Atten- tion-based, Generative, Sparse and Quantum-inspired models. Pros and cons of those techniques as a potential replacement have been analyzed and listed. Data analyzing tools and methods have been explained and most common datasets used for HMM and DNN technologies have been described. Real life usage examples of both methods have been provided and systems based on them have been ana- lyzed.
ISSN:	2524-0382 2707-0069
DOI:	10.23939/acps2023.01.074