Szczegóły publikacji

Opis bibliograficzny

Joint diarization and separation using sepformer with non-autoregressive attractors / Magdalena RYBICKA, Konrad KOWALCZYK, Thomas Thebaud, Najim Dehak, Jesús Villalba // IEEE Signal Processing Letters ; ISSN  1070-9908 . — 2025 — vol. 32, s. 2913–2917. — Bibliogr. s. 2917, Abstr. — Publikacja dostępna online od: 2025-07-18. — M. Rybicka - dod. afiliacja: Johns Hopkins University, Baltimore, MD, USA

Autorzy (5)

Słowa kluczowe

clusteringnon autoregressive modelspeaker diarizationspeech separationattractor mechanismend-to-end

Dane bibliometryczne

ID BaDAP162376
Data dodania do BaDAP2025-09-15
Tekst źródłowyURL
DOI10.1109/LSP.2025.3590325
Rok publikacji2025
Typ publikacjiartykuł w czasopiśmie
Otwarty dostęptak
Creative Commons
Czasopismo/seriaIEEE Signal Processing Letters

Abstract

Speaker diarization and speech separation both aim to track speaker activity in multi-speaker recordings, but they differ in their granularity. Diarization provides a binary indication of whether a speaker is active within a given time frame, whereas speech separation produces individual audio signals, each containing the isolated speech of a specific speaker. Recently, there has been growing interest in approaches that unify diarization and speech separation, particularly those leveraging neural models trained jointly to enhance performance in both tasks. In this letter, we propose a single neural model for joint speaker diarization and speech separation. Our model estimates speaker representations using a non-autoregressive attractor generation mechanism integrated into a modified SepFormer model. We present two variants of the model, designed for scenarios with sparse or highly overlapping speech, which achieve relative improvements of 51% for both separation and diarization over state-of-the-art methods, as evaluated on the LibriMix, LibriheavyMix and CALLHOME datasets.

Publikacje, które mogą Cię zainteresować

artykuł
#155335Data dodania: 20.9.2024
End-to-end neural speaker diarization with non-autoregressive attractors / Magdalena RYBICKA, Jesús Villalba, Thomas Thebaud, Najim Dehak, Konrad KOWALCZYK // IEEE/ACM Transactions on Audio, Speech and Language Processing ; ISSN  2329-9290 . — Tytuł poprz.: IEEE Transactions on Audio, Speech, and Language Processing ; ISSN:  1558-7916. — 2024 — vol. 32, s. 3960-3973. — Bibliogr. s. 3972, Abstr. — Publikacja dostępna online od: 2024-08-07. — M. Rybicka - dod. afiliacja: Johns Hopkins University, Baltimor, USA
fragment książki
#141902Data dodania: 6.9.2022
End-to-end neural speaker diarization with an iterative refinement of non-autoregressive attention-based attractors / Magdalena RYBICKA, Jesús Villalba, Najim Dehak, Konrad KOWALCZYK // W: INTERSPEECH 2022 [Dokument elektroniczny] : September 18–22, Incheon, Korea. — Wersja do Windows. — Dane tekstowe. — [Seoul : The Acoustical Society of Korea], [2022]. — S. 5090–5094. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://isca-speech.org/archive/pdfs/interspeech_2022/rybicka... [2022-09-03]. — Bibliogr. s. 5094, Abstr.