CONF aradilla:simpe:2007/IDIAP Detection and Recognition of Number Sequences in Spoken Utterances Aradilla, Guillermo Ajmera, Jitendra EXTERNAL http://publications.idiap.ch/attachments/papers/2007/aradilla-simpe-2007.pdf PUBLIC http://publications.idiap.ch/index.php/publications/showcite/aradilla:rr07-42 Related documents 2nd Workshop on Speech in Mobile and Pervasive Environments (SiMPE) 2007 IDIAP-RR 07-42 In this paper we investigate the detection and recognition of sequences of numbers in spoken utterances. This is done in two steps: first, the entire utterance is decoded assuming that only numbers were spoken. In the second step, non-number segments (garbage) are detected based on word confidence measures. We compare this approach to conventional garbage models. Also, a comparison of several phone posterior based confidence measures is presented in this paper. The work is evaluated in terms of detection task (hit rate and false alarms) and recognition task (word accuracy) within detected number sequences. The proposed method is tested on German continuous spoken utterances where target content (numbers) is only 20\%. REPORT aradilla:rr07-42/IDIAP Detection and Recognition of Number Sequences in Spoken Utterances Aradilla, Guillermo Ajmera, Jitendra EXTERNAL http://publications.idiap.ch/attachments/reports/2007/aradilla-idiap-rr-07-42.pdf PUBLIC Idiap-RR-42-2007 2007 IDIAP In this paper we investigate the detection and recognition of sequences of numbers in spoken utterances. This is done in two steps: first, the entire utterance is decoded assuming that only numbers were spoken. In the second step, non-number segments (garbage) are detected based on word confidence measures. We compare this approach to conventional garbage models. Also, a comparison of several phone posterior based confidence measures is presented in this paper. The work is evaluated in terms of detection task (hit rate and false alarms) and recognition task (word accuracy) within detected number sequences. The proposed method is tested on German continuous spoken utterances where target content (numbers) is only 20\%.