Recognition Of Reverberant Speech Using Frequency Domain Linear Prediction

Type of publication:	Idiap-RR
Citation:	tsamuel:rr08-41
Number:	Idiap-RR-41-2008
Year:	2008
Institution:	IDIAP
Note:	To appear in IEEE Signal Processing Letters 2008
Abstract:	Performance of a typical automatic speech recognition (ASR) system severely degrades when it encounters speech from reverberant environments. Part of the reason for this degradation is the feature extraction techniques that use analysis windows which are much shorter than typical room impulse responses. We present a feature extraction technique based on modeling temporal envelopes of the speech signal in narrow sub-bands using Frequency Domain Linear Prediction (FDLP). FDLP provides an all-pole approximation of the Hilbert envelope of the signal obtained by linear prediction on cosine transform of the signal. ASR experiments on speech data degraded with a number of room impulse responses (with varying degrees of distortion) show significant performance improvements for the proposed FDLP features when compared to other robust feature extraction techniques (average relative reduction of $24 \%$ in word error rate). Similar improvements are also obtained for far-field data which contain natural reverberation in background noise. These results are achieved without any noticeable degradation in performance for clean speech.
Userfields:	ipdmembership={speech},
Keywords:
Projects:	Idiap
Authors:	Thomas, Samuel Ganapathy, Sriram Hermansky, Hynek
Crossref by	tsamuel:ieee-letters:2008
Added by:	[UNK]
Total mark:	0
Attachments
tsamuel-idiap-rr-08-41.pdf tsamuel-idiap-rr-08-41.ps.gz
Notes

processing time: 0.0356 seconds.