Szczegóły publikacji
Opis bibliograficzny
Task-conditioned audio-text-image fusion for cognitive score estimation from speech-based assessments / Justyna KRZYWDZIAK, Władysław Średniawa, Agnieszka Pruszek, Wojciech Szecówka, Michał K. Grzeszczyk, Teresa Brzozka, Bartłomiej Eljasiak, Zofia Marciniak, Łukasz Łazarski, Mateusz Matuszewski, Daria HEMMERLING // W: Interspeech 2026 [Dokument elektroniczny] : 27 September – 1 October 2026, Sydney, Australia. — Wersja do Windows. — Dane tekstowe. — [Australia : ISCA], [2026]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 2958-1796 ). — S. 2358–2362. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2026/krzywdziak26_in... [2026-09-25]. — Bibliogr. s. 2362, Abstr. — J. Krzywdziak, W. Szecówka, D. Hemmerling - dod. afiliacja: Samsung R&D Institute, Poland
Autorzy (11)
- AGHKrzywdziak Justyna
- Średniawa Władysław
- Pruszek Agnieszka
- AGHSzecówka Wojciech
- Grzeszczyk Michal K.
- Brzozka Teresa
- Eljasiak Bartłomiej
- Marciniak Zofia
- Łazarski Łukasz
- Matuszewski Mateusz
- AGHHemmerling Daria
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 170245 |
|---|---|
| Data dodania do BaDAP | 2026-09-25 |
| DOI | 10.21437/Interspeech.2026-1424 |
| Rok publikacji | 2026 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2026 |
| Czasopismo/seria | Interspeech |
Abstract
People with mild cognitive impariment tend to perform speech-based tasks differently than the healthy population. In picture description tasks, they often describe a narrower portion of the scene, missing details and spatial relations. We propose a multimodal framework for estimating cognitive assessment scores (MoCA and MMSE) from speech-based cognitive tasks. The model uses pretrained encoders for audio, text, and optional vision, followed by modality-specific projections and task-conditioned fusion. We compare staged variants from acoustic features to foundation-model embeddings, late fusion, mid-level cross-attention fusion, and a final mixture-of-experts model. We also introduce a label-conditioned image-text alignment loss for picture description samples to model differences in transcript-image consistency between healthy controls and patients with mild cognitive impairment. The best model achieved 2.11(+-0.03) MoCA RMSE and 1.85(+-0.08) MMSE RMSE in picture description task.