Szczegóły publikacji
Opis bibliograficzny
From text metrics to model internals: a study of whisper ASR hallucination detection / Jan JASIŃSKI, Mateusz BARAŃSKI, Julitta BARTOLEWSKA, Marcin WITKOWSKI, Konrad KOWALCZYK // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 2958-1796 ). — S. 6098–6102. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/jasinski26_inte... [2026-09-22]. — Bibliogr. s. 6102, Abstr.
Autorzy (5)
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 170180 |
|---|---|
| Data dodania do BaDAP | 2026-09-23 |
| DOI | 10.21437/Interspeech.2026-338 |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2026 |
| Czasopismo/seria | Interspeech |
Abstract
Hallucinations of ASR models — fluent transcriptions with no basis in audio — degrade system performance and pose risks in downstream applications. Robust detection of such errors remains a challenge. This paper studies Whisper large v3 hallucination detection on real-speech human-annotated data across three paradigms: text-based, LLM-based, and internal decoder state probing. Text classifiers utilizing metrics for text evaluation achieve high recall but degrade without reference transcripts. LLM-based detection improves precision with domain-specific prompt conditioning, yet remains less competitive than the lightweight text-based methods. Probing Whisper's decoder representations, without a ground-truth reference, yields the strongest performance, revealing that hallucination traits are encoded across intermediate decoding layers. A late-fusion meta-classifier combining text and internal-state outputs achieves the best overall detection performance.