DEEP NEURAL NETWORK BASED POSTERIORS FOR TEXT-DEPENDENT SPEAKER VERIFICATION

Type of publication:	Conference paper
Citation:	Dey_ICASSP_2016
Publication status:	Published
Booktitle:	Proceedings of 2016 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2016)
Year:	2016
Month:	March
Pages:	5050-5054
Publisher:	IEEE
Location:	Shanghai
Crossref:	Dey_Idiap-RR-08-2016: DEEP NEURAL NETWORK BASED POSTERIORS FOR TEXT-DEPENDENT SPEAKER VERIFICATION, Dey, Subhadeep, Madikeri, Srikanth, Ferras, Marc and Motlicek, Petr, Idiap-RR-08-2016
Abstract:	The i-vector and Joint Factor Analysis (JFA) systems for text- dependent speaker verification use sufficient statistics computed from a speech utterance to estimate speaker models. These statis- tics average the acoustic information over the utterance thereby losing all the sequence information. In this paper, we study ex- plicit content matching using Dynamic Time Warping (DTW) and present the best achievable error rates for speaker-dependent and speaker-independent content matching. For this purpose, a Deep Neural Network/Hidden Markov Model Automatic Speech Recog- nition (DNN/HMM ASR) system is used to extract content-related posterior probabilities. This approach outperforms systems using Gaussian mixture model posteriors by at least 50% Equal Error Rate (EER) on the RSR2015 in content mismatch trials. DNN posteriors are also used in i-vector and JFA systems, obtaining EERs as low as 0.02%.
Keywords:
Projects:	Idiap SIIP
Authors:	Dey, Subhadeep Madikeri, Srikanth Ferras, Marc Motlicek, Petr
Added by:	[UNK]
Total mark:	0
Attachments
Dey_ICASSP_2016.pdf
Notes

processing time: 0.0003 seconds.