Front-end for Far-field Speech Recognition based on Frequency Domain Linear Prediction

We use cookies

This website uses cookies and other tracking technologies to improve your browsing experience for the following purposes: to enable basic functionality of the website, to provide a better experience on the website, to measure your interest in our products and services and to personalize marketing interactions, to deliver ads that are more relevant to you.

[BibTeX] [Marc21]

Type of publication:	Conference paper
Citation:	tsamuel:interspeech-1:2008
Booktitle:	Interspeech 2008
Year:	2008
Note:	IDIAP-RR 08-17
Crossref:	tsamuel:rr08-17: Front-end for Far-field Speech Recognition based on Frequency Domain Linear Prediction, Ganapathy, Sriram, Thomas, Samuel and Hermansky, Hynek, Idiap-RR-17-2008
Abstract:	Automatic Speech Recognition (ASR) systems usually fail when they encounter speech from far-field microphone in reverberant environments. This is due to the application of short-term feature extraction techniques which do not compensate for the artifacts introduced by long room impulse responses. In this paper, we propose a front-end, based on Frequency Domain Linear Prediction (FDLP,',','), that tries to remove reverberation artifacts present in far-field speech. Long temporal segments of far-field speech are analyzed in narrow frequency sub-bands to extract FDLP envelopes and residual signals. Filtering the residual signals with gain normalized inverse FDLP filters result in a set of sub-band signals which are synthesized to reconstruct the signal back. ASR experiments on far-field speech data processed by the proposed front-end show significant improvements (relative reduction of $30 \%$ in word error rate) compared to other robust feature extraction techniques.
Userfields:	ipdmembership={speech},
Keywords:
Projects	Idiap
Authors	Ganapathy, Sriram Thomas, Samuel Hermansky, Hynek
Added by:	[UNK]
Total mark:	0
Attachments
tsamuel-interspeech-1-2008.pdf tsamuel-interspeech-1-2008.ps.gz
Notes

processing time: 0.0011 seconds.