Szczegóły publikacji

Opis bibliograficzny

End-to-end neural speaker diarization with non-autoregressive attractors / Magdalena RYBICKA, Jesús Villalba, Thomas Thebaud, Najim Dehak, Konrad KOWALCZYK // IEEE/ACM Transactions on Audio, Speech and Language Processing ; ISSN  2329-9290 . — Tytuł poprz.: IEEE Transactions on Audio, Speech, and Language Processing ; ISSN:  1558-7916. — 2024 — vol. 32, s. 3960-3973. — Bibliogr. s. 3972, Abstr. — Publikacja dostępna online od: 2024-08-07. — M. Rybicka - dod. afiliacja: Johns Hopkins University, Baltimor, USA

Autorzy (5)

Słowa kluczowe

speaker diarizationnon autoregressive modelattractor mechanismself attentionclusteringend-to-enditerative refinement

Dane bibliometryczne

ID BaDAP155335
Data dodania do BaDAP2024-09-20
Tekst źródłowyURL
DOI10.1109/TASLP.2024.3439993
Rok publikacji2024
Typ publikacjiartykuł w czasopiśmie
Otwarty dostęptak
Creative Commons
Czasopismo/seriaIEEE/ACM Transactions on Audio, Speech and Language Processing

Abstract

Despite many recent developments in speaker diarization, it remains a challenge and an active area of research to make diarization robust and effective in real-life scenarios. Well-established clustering-based methods are showing good performance and qualities. However, such systems are built of several independent, separately optimized modules, which may cause non-optimum performance. End-to-end neural speaker diarization (EEND) systems are considered the next stepping stone in pursuing high-performance diarization. Nevertheless, this approach also suffers limitations, such as dealing with long recordings and scenarios with a large (more than four) or unknown number of speakers in the recording. The appearance of EEND with encoder-decoder-based attractors (EEND-EDA) enabled us to deal with recordings that contain a flexible number of speakers thanks to an LSTM-based EDA module. A competitive alternative over the referenced EEND-EDA baseline is the EEND with non-autoregressive attractor (EEND-NAA) estimation, proposed recently by the authors of this article. NAA back-end incorporates k-means clustering as part of the attractor estimation and an attractor refinement module based on a Transformer decoder. However, in our previous work on EEND-NAA, we assumed a known number of speakers, and the experimental evaluation was limited to 2-speaker recordings only. In this article, we describe in detail our recent EEND-NAA approach and propose further improvements to the EEND-NAA architecture, introducing three novel variants of the NAA back-end, which can handle recordings containing speech of a variable and unknown number of speakers. Conducted experiments include simulated mixtures generated using the Switchboard and NIST SRE datasets and real-life recordings from the CALLHOME and DIHARD II datasets. In experimental evaluation, the proposed systems achieve up to 51% relative improvement for the simulated scenario and up to 15% for real recordings over the baseline EEND-EDA.

Publikacje, które mogą Cię zainteresować

fragment książki
#141902Data dodania: 6.9.2022
End-to-end neural speaker diarization with an iterative refinement of non-autoregressive attention-based attractors / Magdalena RYBICKA, Jesús Villalba, Najim Dehak, Konrad KOWALCZYK // W: INTERSPEECH 2022 [Dokument elektroniczny] : September 18–22, Incheon, Korea. — Wersja do Windows. — Dane tekstowe. — [Seoul : The Acoustical Society of Korea], [2022]. — S. 5090–5094. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://isca-speech.org/archive/pdfs/interspeech_2022/rybicka... [2022-09-03]. — Bibliogr. s. 5094, Abstr.
artykuł
#162376Data dodania: 15.9.2025
Joint diarization and separation using sepformer with non-autoregressive attractors / Magdalena RYBICKA, Konrad KOWALCZYK, Thomas Thebaud, Najim Dehak, Jesús Villalba // IEEE Signal Processing Letters ; ISSN  1070-9908 . — 2025 — vol. 32, s. 2913–2917. — Bibliogr. s. 2917, Abstr. — Publikacja dostępna online od: 2025-07-18. — M. Rybicka - dod. afiliacja: Johns Hopkins University, Baltimore, MD, USA