logo Idiap Research Institute        
 [BibTeX] [Marc21]
Fast latent semantic indexing of spoken documents by using self-organizing maps
Type of publication: Conference paper
Citation: kurimo-icassp00b
Booktitle: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP'2000
Year: 2000
Month: 6
Address: Istanbul, Turkey
Note: IDIAP-RR 99-20
Crossref: kurimo-icassp00:
Abstract: This paper describes a new latent semantic indexing (LSI) method for spoken audio documents. The framework is indexing broadcast news from radio and TV as a combination of large vocabulary continuous speech recognition (LVCSR,',','), natural language processing (NLP) and information retrieval (IR). For indexing, the documents are presented as vectors of word counts, whose dimensionality is rapidly reduced by random mapping (RM). The obtained vectors are projected into the latent semantic subspace determined by SVD, where the vectors are then smoothed by a self-organizing map (SOM). The smoothing by the closest document clusters is important here, because the documents are often short and have a high word error rate (WER). As the clusters in the semantic subspace reflect the news topics, the SOMs provide an easy way to visualize the index and query results and to explore the database. Test results are reported for TREC's spoken document retrieval databases.
Userfields: ipdmembership={speech},
Keywords:
Projects Idiap
Authors Kurimo, Mikko
Added by: [UNK]
Total mark: 0
Attachments
  • icassp00.pdf
  • kurimo-icassp00.ps.gz
Notes