Szczegóły publikacji

Opis bibliograficzny

Task-conditioned audio-text-image fusion for cognitive score estimation from speech-based assessments / Justyna KRZYWDZIAK, Władysław Średniawa, Agnieszka Pruszek, Wojciech Szecówka, Michał K. Grzeszczyk, Teresa Brzozka, Bartłomiej Eljasiak, Zofia Marciniak, Łukasz Łazarski, Mateusz Matuszewski, Daria HEMMERLING // W: Interspeech 2026 [Dokument elektroniczny] : 27 September – 1 October 2026, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ISCA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN  2958-1796 ). — S. 2358–2362. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/krzywdziak26_in... [2026-09-25]. — Bibliogr. s. 2362, Abstr. — J. Krzywdziak, W. Szecówka, D. Hemmerling - dod. afiliacja: Samsung R&D Institute, Poland

Autorzy (11)

Słowa kluczowe

mild cognitive impairmentmulti-modal fusionspeech processing

Dane bibliometryczne

ID BaDAP170245
Data dodania do BaDAP2026-09-25
DOI10.21437/Interspeech.2026-1424
Rok publikacji2026
Typ publikacjimateriały konferencyjne (aut.)
Otwarty dostęptak
KonferencjaInterspeech 2026
Czasopismo/seriaInterspeech

Abstract

People with mild cognitive impariment tend to perform speech-based tasks differently than the healthy population. In picture description tasks, they often describe a narrower portion of the scene, missing details and spatial relations. We propose a multimodal framework for estimating cognitive assessment scores (MoCA and MMSE) from speech-based cognitive tasks. The model uses pretrained encoders for audio, text, and optional vision, followed by modality-specific projections and task-conditioned fusion. We compare staged variants from acoustic features to foundation-model embeddings, late fusion, mid-level cross-attention fusion, and a final mixture-of-experts model. We also introduce a label-conditioned image-text alignment loss for picture description samples to model differences in transcript-image consistency between healthy controls and patients with mild cognitive impairment. The best model achieved 2.11(+-0.03) MoCA RMSE and 1.85(+-0.08) MMSE RMSE in picture description task.

Publikacje, które mogą Cię zainteresować

fragment książki
#170180Data dodania: 23.9.2026
From text metrics to model internals: a study of whisper ASR hallucination detection / Jan JASIŃSKI, Mateusz BARAŃSKI, Julitta BARTOLEWSKA, Marcin WITKOWSKI, Konrad KOWALCZYK // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN  2958-1796 ). — S. 6098–6102. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/jasinski26_inte... [2026-09-22]. — Bibliogr. s. 6102, Abstr.
fragment książki
#170181Data dodania: 23.9.2026
HALAS: a human-annotated dataset of hallucinations of modern ASR systems / Mateusz BARAŃSKI, Jan JASIŃSKI, Julitta BARTOLEWSKA, Marcin WITKOWSKI, Konrad KOWALCZYK // W: Interspeech 2026 [Dokument elektroniczny] : speaking together : 27 September–1 October, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ICMSA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN  2958-1796 ). — S. 6698–6702. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/baranski26_inte... [2026-09-22]. — Bibliogr. s. 6702, Abstr.