logo Idiap Research Institute        
 [BibTeX] [Marc21]
A Sector-Based, Frequency-Domain Approach to Detection and Localization of Multiple Speakers
Type of publication: Conference paper
Citation: lathoud05a
Booktitle: Proceedings of ICASSP 2005
Year: 2005
Month: 3
Address: Philadelphia, USA
Note: IDIAP-RR 04-54
Crossref: lathoud-rr-04-54:
Abstract: Detection and localization of speakers with microphone arrays is a difficult task due to the wideband nature of speech signals, the large amount of overlaps between speakers in spontaneous conversations, and the presence of noise sources. Many existing audio multi-source localization methods rely on prior knowledge of the sectors containing active sources and/or the number of active sources. This paper proposes sector-based, frequency-domain approaches that address both detection and localization problems by measuring relative phases between microphones. The first approach is similar to delay-sum beamforming. The second approach is novel: it relies on systematic optimization of a centroid in phase space, for each sector. It provides major, systematic improvement over the first approach as well as over previous work. Very good results are obtained on more than one hour of recordings in real meeting room conditions, including cases with up to 3 concurrent speakers.
Userfields: ipdinar={2004}, ipdmembership={speech},
Keywords:
Projects Idiap
Authors Lathoud, Guillaume
Magimai.-Doss, Mathew
Added by: [UNK]
Total mark: 0
Attachments
  • lathoud05a.pdf
  • lathoud05a.ps.gz
Notes