Comparing performance of spectral distance measures and neural network methods for vowel recognition

Neural networks were trained to classify single 20 ms frames of vowels using either perceptually-based spectral representations or LPC spectra as input. Classification performance was compared with performance of several distance measures using nearest-neighbor and mean-distance decision criteria. T...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:Computer speech & language 1989, Vol.3 (1), p.21-34
Hauptverfasser: Kamm, Candace A., Streeter, Lynn A., Kane-Esrig, Yana, Burr, David J.
Format: Artikel
Sprache:eng
Online-Zugang:Volltext
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Beschreibung
Zusammenfassung:Neural networks were trained to classify single 20 ms frames of vowels using either perceptually-based spectral representations or LPC spectra as input. Classification performance was compared with performance of several distance measures using nearest-neighbor and mean-distance decision criteria. The non-network distance measures included LPC-residual and cepstral distance measures used in conventional automatic speech recognition systems, as well as a formant-based measure and a new elastic distance measure that explicitly corrects for the effects of spectral tilt. Using an optimal error rate criterion, vowels were discriminated best using the elastic distance measure with the perceptually-based spectrum. Neural networks with LPC spectra as input performed comparably to the better conventional distance measures. While the performance of networks trained with perceptually-based spectral inputs was poorer than that of networks trained with LPC spectra, the features represented by the hidden nodes of this network were more consistent with factors related to human vowel perception.
ISSN:0885-2308
1095-8363
DOI:10.1016/0885-2308(89)90012-0