No Audio, No Transcripts, No Problem: LLM-based ASR Adaptation Using Only Domain Documents
| Type of publication: | Conference paper |
| Citation: | Burdisso_SLT_2026 |
| Publication status: | Accepted |
| Booktitle: | IEEE SLT |
| Year: | 2026 |
| Abstract: | Can a speech recognizer be adapted to a new domain with no target audio and no transcripts, using only its written documentation? This is the realistic onboarding setting: a new customer hands over raw documentation, not the transcripts that text-only adaptation normally assumes. We turn a domain's documents into adaptation text in two ways, directly as written sentences or rendered by a large language model (LLM) into synthetic spoken dialogs, and feed them to a recently proposed denoising-based text-only adaptation method. We show that, across in-domain, out-of-domain, and cross-domain settings on two corpora, document-only adaptation can match or closely approach real transcriptions, and whether the written or conversational rendering is best depends on speaking style and the amount of generated text. To our knowledge, this is the first study of document-only adaptation of LLM-based ASR; as no corpus pairs audio, transcripts, and documents, we build and release a document-grounded benchmark. |
| Additional Research Programs: |
AI for Everyone |
| Keywords: | |
| Projects: |
UNIPHORE ELOQUENCE |
| Authors: | |
| Added by: | [UNK] |
| Total mark: | 0 |
|
Attachments
|
|
|
Notes
|
|
|
|
|