Szczegóły publikacji
Opis bibliograficzny
Do what I say: a spoken prompt dataset for instruction-following / Maike Züfle, Sara Papi, Fabian Retkowski, Szymon MAZUREK, Marek KASZTELNIK, Alexander Waibel, Luisa Bentivogli, Jan Niehues // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 2958-1796 ). — S. 3820–3824. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/zufle26_intersp... [2026-09-22]. — Bibliogr. s. 3824, Abstr.
Autorzy (8)
- Züfle Maike
- Papi Sara
- Retkowski Fabian
- AGHMazurek Szymon
- AGHKasztelnik Marek
- Waibel Alexander
- Bentivogli Luisa
- Niehues Jan
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 170027 |
|---|---|
| Data dodania do BaDAP | 2026-09-16 |
| DOI | 10.21437/Interspeech.2026-685 |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2026 |
| Czasopismo/seria | Interspeech |
Abstract
Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompts, which may not reflect real-world scenarios where users interact with speech. To address this gap, we introduce DoWhatISay (DOWIS), a multilingual dataset of human-recorded spoken and written prompts designed to pair with any existing benchmark for realistic evaluation of SLLMs under spoken instruction conditions. Spanning 9 tasks and 11 languages, it provides 10 prompt variants per task-language pair, across five styles. Using DOWIS, we benchmark state-of-the-art SLLMs, analyzing the interplay between prompt modality, style, language, and task type. Results show that text prompts consistently outperform spoken prompts, particularly for low-resource and cross-lingual settings. Only for tasks with speech output, spoken prompts do close the gap, highlighting the need for speech-based prompting in SLLM evaluation.