Szczegóły publikacji
Opis bibliograficzny
Multi-task learning for speech emotion recognition in naturalistic conditions / Bartłomiej Zgórzyński, Juliusz Wójtowicz-Kruk, Piotr MASZTALSKI, Władysław Średniawa // W: Interspeech 2025 [Dokument elektroniczny] : 17–21 August 2025, Rotterdam, The Netherlands. — Wersja do Windows. — Dane tekstowe. — [France : ISCA], [2025]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 2958-1796 ). — S. 4678–4682. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 4682, Abstr. — P. Masztalski - dod. afiliacja: Samsung R&D Institute Poland
Autorzy (4)
- Zgórzyński Bartłomiej
- Wójtowicz-Kruk Juliusz
- AGHMasztalski Piotr
- Średniawa Władysław
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 162078 |
|---|---|
| Data dodania do BaDAP | 2025-09-08 |
| Tekst źródłowy | URL |
| DOI | 10.21437/Interspeech.2025-2033 |
| Rok publikacji | 2025 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2025 |
| Czasopismo/seria | Interspeech |
Abstract
This work introduces a multi-encoder joint classification and regression training framework for speech emotion recognition. We present our solution for the Interspeech 2025 Speech Emotion Recognition in Naturalistic Conditions Challenge, leveraging a multi-modal, multi-encoder architecture with a fusion module. Our results demonstrate the effectiveness of the multi-task approach for both classification and regression tasks, achieving a top 10 spot in categorical emotion classification and 2nd place in emotional attribute prediction among competing teams. Furthermore, an ablation study shows that employing multi-task learning outperforms separate task-specific training. These findings highlight the potential of multi-task, multi-encoder systems for speech emotion recognition.