Szczegóły publikacji

Opis bibliograficzny

Do what I say: a spoken prompt dataset for instruction-following / Maike Züfle, Sara Papi, Fabian Retkowski, Szymon MAZUREK, Marek KASZTELNIK, Alexander Waibel, Luisa Bentivogli, Jan Niehues // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN  2958-1796 ). — S. 3820–3824. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/zufle26_intersp... [2026-09-22]. — Bibliogr. s. 3824, Abstr.

Autorzy (8)

Słowa kluczowe

spoken promptpromptinstruction followingdatasetbenchmarkcrosslingualmultilingual

Dane bibliometryczne

ID BaDAP170027
Data dodania do BaDAP2026-09-16
DOI10.21437/Interspeech.2026-685
Rok publikacji2026
Typ publikacjimateriały konferencyjne (aut.)
Otwarty dostęptak
KonferencjaInterspeech 2026
Czasopismo/seriaInterspeech

Abstract

Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompts, which may not reflect real-world scenarios where users interact with speech. To address this gap, we introduce DoWhatISay (DOWIS), a multilingual dataset of human-recorded spoken and written prompts designed to pair with any existing benchmark for realistic evaluation of SLLMs under spoken instruction conditions. Spanning 9 tasks and 11 languages, it provides 10 prompt variants per task-language pair, across five styles. Using DOWIS, we benchmark state-of-the-art SLLMs, analyzing the interplay between prompt modality, style, language, and task type. Results show that text prompts consistently outperform spoken prompts, particularly for low-resource and cross-lingual settings. Only for tasks with speech output, spoken prompts do close the gap, highlighting the need for speech-based prompting in SLLM evaluation.

Publikacje, które mogą Cię zainteresować

fragment książki
#170181Data dodania: 23.9.2026
HALAS: a human-annotated dataset of hallucinations of modern ASR systems / Mateusz BARAŃSKI, Jan JASIŃSKI, Julitta BARTOLEWSKA, Marcin WITKOWSKI, Konrad KOWALCZYK // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN  2958-1796 ). — S. 6698–6702. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/baranski26_inte... [2026-09-22]. — Bibliogr. s. 6702, Abstr.
fragment książki
#170180Data dodania: 23.9.2026
From text metrics to model internals: a study of whisper ASR hallucination detection / Jan JASIŃSKI, Mateusz BARAŃSKI, Julitta BARTOLEWSKA, Marcin WITKOWSKI, Konrad KOWALCZYK // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN  2958-1796 ). — S. 6098–6102. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/jasinski26_inte... [2026-09-22]. — Bibliogr. s. 6102, Abstr.