Block-online multi-channel speech enhancement using deep neural network-supported relative transfer function estimates

This work addresses the problem of block-online processing for multi-channel speech enhancement. Such processing is vital in scenarios with moving speakers and/or when short utterances are processed, e.g. in voice assistant applications. We consider several variants of a system that performs beamfor...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IET signal processing 2020-05, Vol.14 (3), p.124-133
Hauptverfasser:	Malek, Jiri, Koldovský, Zbynĕk, Bohac, Marek
Format:	Artikel
Sprache:	eng
Schlagworte:	array signal processing baseline automatic speech recognition system batch processing block length block‐online multichannel speech enhancement block‐online processing deep neural network‐based voice activity detection deep neural network‐supported relative transfer function estimates enhancement method highly dynamic environments neural nets perceptual evaluation processed block processing regime relative transfer functions Research Article short utterances speech enhancement speech quality speech recognition time 250.0 ms transfer functions voice assistant scenarios
Online-Zugang:	Volltext bestellen
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	This work addresses the problem of block-online processing for multi-channel speech enhancement. Such processing is vital in scenarios with moving speakers and/or when short utterances are processed, e.g. in voice assistant applications. We consider several variants of a system that performs beamforming supported by deep neural network-based voice activity detection followed by post-filtering. The speaker is targeted through estimating relative transfer functions between microphones. Each block of the input signals is processed independently to make the method applicable in highly dynamic environments. Due to short processed blocks, the statistics required by the beamformer are estimated less precisely. The influence of this inaccuracy is studied and compared to batch processing regime, when recordings are treated as one block. The experimental evaluation is performed on large datasets of CHiME-4 and another dataset featuring moving target speaker. The experiments are evaluated in terms of objective and perceptual criteria. Moreover, word error rate (WER) of a speech recognition system is evaluated, for which the method serves as a front-end. The results indicate that the proposed method is robust for short length of the processed block. Significant improvements in terms of the criteria and WER are observed even for the block length of 250 ms.
ISSN:	1751-9675 1751-9683 1751-9683
DOI:	10.1049/iet-spr.2019.0304