Szczegóły publikacji
Opis bibliograficzny
HALAS: a human-annotated dataset of hallucinations of modern ASR systems / Mateusz BARAŃSKI, Jan JASIŃSKI, Julitta BARTOLEWSKA, Marcin WITKOWSKI, Konrad KOWALCZYK // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 2958-1796 ). — S. 6698–6702. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/baranski26_inte... [2026-09-22]. — Bibliogr. s. 6702, Abstr.
Autorzy (5)
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 170181 |
|---|---|
| Data dodania do BaDAP | 2026-09-23 |
| DOI | 10.21437/Interspeech.2026-337 |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2026 |
| Czasopismo/seria | Interspeech |
Abstract
End-to-end Automatic Speech Recognition (ASR) systems hallucinate on natural speech, yet existing mitigation methods are typically evaluated on non-speech or artificially corrupted audio. We introduce HALAS, the first human-annotated dataset of naturally occurring hallucinations from seven state-of-the-art ASR models on real unprocessed earnings call recordings. HALAS provides span-level labels, enabling analysis of hallucination patterns and their severity. Our analysis reveals strong cross-model vocabulary overlap and confirms that hallucinations also occur for almost correctly transcribed speech (characterized by a low Word Error Rate). The proposed benchmark with HALAS shows that the character and semantic-level metrics used as a proxy for hallucination detection reach 81% ROC-AUC, while state-of-the-art detection methods achieve an F1 score of only 53.1%. As such, HALAS establishes the first rigorous non-artificial benchmark for the detection and mitigation of ASR hallucinations.