Szczegóły publikacji
Opis bibliograficzny
Joint diarization and separation using sepformer with non-autoregressive attractors / Magdalena RYBICKA, Konrad KOWALCZYK, Thomas Thebaud, Najim Dehak, Jesús Villalba // IEEE Signal Processing Letters ; ISSN 1070-9908 . — 2025 — vol. 32, s. 2913–2917. — Bibliogr. s. 2917, Abstr. — Publikacja dostępna online od: 2025-07-18. — M. Rybicka - dod. afiliacja: Johns Hopkins University, Baltimore, MD, USA
Autorzy (5)
- AGHRybicka Magdalena
- AGHKowalczyk Konrad
- Thebaud Thomas
- Dehak Najim
- Villalba Jesús
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 162376 |
|---|---|
| Data dodania do BaDAP | 2025-09-15 |
| Tekst źródłowy | URL |
| DOI | 10.1109/LSP.2025.3590325 |
| Rok publikacji | 2025 |
| Typ publikacji | artykuł w czasopiśmie |
| Otwarty dostęp | |
| Creative Commons | |
| Czasopismo/seria | IEEE Signal Processing Letters |
Abstract
Speaker diarization and speech separation both aim to track speaker activity in multi-speaker recordings, but they differ in their granularity. Diarization provides a binary indication of whether a speaker is active within a given time frame, whereas speech separation produces individual audio signals, each containing the isolated speech of a specific speaker. Recently, there has been growing interest in approaches that unify diarization and speech separation, particularly those leveraging neural models trained jointly to enhance performance in both tasks. In this letter, we propose a single neural model for joint speaker diarization and speech separation. Our model estimates speaker representations using a non-autoregressive attractor generation mechanism integrated into a modified SepFormer model. We present two variants of the model, designed for scenarios with sparse or highly overlapping speech, which achieve relative improvements of 51% for both separation and diarization over state-of-the-art methods, as evaluated on the LibriMix, LibriheavyMix and CALLHOME datasets.