Joint Phoneme Segmentation Inference and Classification using CRFs

We use cookies

This website uses cookies and other tracking technologies to improve your browsing experience for the following purposes: to enable basic functionality of the website, to provide a better experience on the website, to measure your interest in our products and services and to personalize marketing interactions, to deliver ads that are more relevant to you.

[BibTeX] [Marc21]

Type of publication:	Conference paper
Citation:	Palaz_GLOBALSIP_2014
Publication status:	Published
Booktitle:	Global Conference on Signal and Information Processing
Year:	2014
Month:	December
Pages:	587 - 591
Publisher:	IEEE
Location:	Atlanta, GA
DOI:	10.1109/GlobalSIP.2014.7032185
Abstract:	State-of-the-art phoneme sequence recognition systems are based on hybrid hidden Markov model/artificial neural networks (HMM/ANN) framework. In this framework, the local classifier, ANN, is typically trained using Viterbi expectation-maximization algorithm, which involves two separate steps: phoneme sequence segmentation and training of ANN. In this paper, we propose a CRF based phoneme sequence recognition approach that simultaneously infers the phoneme segmentation and classifies the phoneme sequence. More specifically, the phoneme sequence recognition system consists of a local classifier ANN followed by a conditional random field (CRF) whose parameters are trained jointly, using a cost function that discriminates the true phoneme sequence against all competing sequences. In order to efficiently train such a system we introduce a novel CRF based segmentation using acyclic graph. We study the viability of the proposed approach on TIMIT phoneme recognition task. Our studies show that the proposed approach is capable of achieving performance similar to standard hybrid HMM/ANN and ANN/CRF systems where the ANN is trained with manual segmentation.
Keywords:
Projects	Idiap
Authors	Palaz, Dimitri Magimai-Doss, Mathew Collobert, Ronan
Added by:	[UNK]
Total mark:	0
Attachments
Palaz_GLOBALSIP_2014.pdf
Notes

processing time: 0.0003 seconds.