An Automatic Lipreading System for Spoken Digits With Limited Training Data

It is well known that visual cues of lip movement contain important speech relevant information. This paper presents an automatic lipreading system for small vocabulary speech recognition tasks. Using the lip segmentation and modeling techniques we developed earlier, we obtain a visual feature vecto...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on circuits and systems for video technology 2008-12, Vol.18 (12), p.1760-1765
Hauptverfasser:	Wang, S.L., Liew, A.W.C., Lau, W.H., Leung, S.H.
Format:	Artikel
Sprache:	eng
Schlagworte:	Algorithms Applied sciences Discrete transforms Exact sciences and technology Hidden Markov models Image recognition Image segmentation Image sequences Information, signal and communications theory Lipreading Miscellaneous Mouth Pattern recognition Signal processing Speech processing Speech recognition Spline Studies Telecommunications and information theory Training data visual feature extraction visual speech recognition Vocabulary
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	It is well known that visual cues of lip movement contain important speech relevant information. This paper presents an automatic lipreading system for small vocabulary speech recognition tasks. Using the lip segmentation and modeling techniques we developed earlier, we obtain a visual feature vector composed of outer and inner mouth features from the lip image sequence for recognition. A spline representation is employed to transform the discrete-time sampled features from the video frames into the continuous domain. The spline coefficients in the same word class are constrained to have similar expression and are estimated from the training data by the EM algorithm. For the multiple-speaker/speaker-independent recognition task, an adaptive multimodel approach is proposed to handle the variations caused by various talking styles. After building the appropriate word models from the spline coefficients, a maximum likelihood classification approach is taken for the recognition. Lip image sequences of English digits from 0 to 9 have been collected for the recognition test. Two widely used classification methods, HMM and RDA, have been adopted for comparison and the results demonstrate that the proposed algorithm deliver the best performance among these methods.
ISSN:	1051-8215 1558-2205
DOI:	10.1109/TCSVT.2008.2004924