Improved Voice Activity Detection Using Contextual Multiple Hypothesis Testing for Robust Speech Recognition

This paper shows an improved statistical test for voice activity detection in noise adverse environments. The method is based on a revised contextual likelihood ratio test (LRT) defined over a multiple observation window. The motivations for revising the original multiple observation LRT (MO-LRT) ar...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on audio, speech, and language processing speech, and language processing, 2007-11, Vol.15 (8), p.2177-2189
Hauptverfasser:	Ramirez, J., Segura, J.C., Gorriz, J.M., Garcia, L.
Format:	Artikel
Sprache:	eng
Schlagworte:	Applied sciences Design methodology Detectors Exact sciences and technology Hypothesis testing Information, signal and communications theory Light rail systems Likelihood ratio Molybdenum Multiple hypothesis testing Noise robustness robust speech recognition Signal processing Signal to noise ratio Speech Speech processing Speech recognition Statistical tests Studies Tasks Technological innovation Telecommunication standards Telecommunications and information theory Testing Voice voice activity detection (VAD) Voice activity detectors Working environment noise
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	This paper shows an improved statistical test for voice activity detection in noise adverse environments. The method is based on a revised contextual likelihood ratio test (LRT) defined over a multiple observation window. The motivations for revising the original multiple observation LRT (MO-LRT) are found in its artificially added hangover mechanism that exhibits an incorrect behavior under different signal-to-noise ratio (SNR) conditions. The new approach defines a maximum a posteriori (MAP) statistical test in which all the global hypotheses on the multiple observation window containing up to one speech-to-nonspeech or nonspeech-to-speech transitions are considered. Thus, the implicit hangover mechanism artificially added by the original method was not found in the revised method so its design can be further improved. With these and other innovations, the proposed method showed a higher speech/nonspeech discrimination accuracy over a wide range of SNR conditions when compared to the original MO-LRT voice activity detector (VAD). Experiments conducted on the AURORA databases and tasks showed that the revised method yields significant improvements in speech recognition performance over standardized VADs such as ITU T G.729 and ETSI AMR for discontinuous voice transmission and the ETSI AFE for distributed speech recognition (DSR), as well as over recently reported methods.
ISSN:	1558-7916 2329-9290 1558-7924 2329-9304
DOI:	10.1109/TASL.2007.903937