Improving Domain-Specific ASR for Soccer Broadcast Commentary
| Type of publication: | Idiap-RR |
| Citation: | Awada_Idiap-RR-03-2026 |
| Number: | Idiap-RR-03-2026 |
| Year: | 2026 |
| Month: | 8 |
| Institution: | Idiap |
| Address: | Rue Marconi 19, Martigny, 1920, Switzerland |
| Note: | Semester project as part of Bachelor programme at EPFL. |
| Abstract: | Pre-trained automatic speech recognition (ASR) models such as OpenAI’s Whisper achieve excellent performance on clean, read speech but degrade significantly on domain-specific audio. This project investigates this degradation on soccer broadcast commentary, where crowd noise, fast-paced delivery, and domain-specific vocabulary, particularly player and team names, pose challenges. Using the GOAL benchmark and SoccerNet broadcast audio, we fine-tune Whisper Medium on soccer commentary data and explore three inference-time techniques for improving entity recognition without retraining: decoder prompting with match rosters, shallow fusion via sequence biasing, and their combination. The combined approach achieves 88.3% entity detection accuracy, within 2.0 percentage points of the oracle upper bound, and raises unseen entity detection from 27.9% to 67.0%. We find that fine-tuning is a prerequisite for these inference-time techniques: without domain adaptation, prompting and shallow fusion cause severe hallucination. This contrasts with prior results on air traffic control speech, suggesting that domain difficulty determines whether prompting can work without fine-tuning. A companion reproducibility guide details how to replicate all experiments on the EPFL Izar cluster. |
| Main Research Program: | Human-AI Teaming |
| Additional Research Programs: |
AI for Everyone |
| Keywords: | Automatic Speech Recognition, entity recognition, model fine-tuning, shallow fusion, speech LLMs |
| Projects: |
Idiap ELOQUENCE |
| Authors: | |
| Added by: | [ADM] |
| Total mark: | 0 |
|
Attachments
|
|
|
Notes
|
|
|
|
|