Szczegóły publikacji

Opis bibliograficzny

Spine2Net: SpineNet with Res2Net and time-squeeze-and-excitation blocks for speaker recognition / Magdalena RYBICKA, Jesús Villalba, Piotr Żelasko, Najim Dehak, Konrad KOWALCZYK // W: INTERSPEECH 2021 [Dokument elektroniczny] : proceedings of the 22nd annual conference of the International Speech Communication Association : 30 August–3 September 2021, Brno, Czechia. — Brno : Brno University of Technology, 2021. — (Interspeech : Proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 1990-9772). — e-ISBN: 978-171383690-2. — S. 496–500. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-speech.org/archive/pdfs/interspeech_2021/ryb... [2021-09-21]. — Bibliogr. s. 500, Abstr. — W bazie Scopus numery stron: 491-495. — P. Żelasko - afiliacja: Johns Hopkins University, Baltimore, USA

Autorzy (5)

Słowa kluczowe

SpineNet modelspeaker recognitionResNet modeldeep neural networksscale-permuted network

Dane bibliometryczne

ID BaDAP136305
Data dodania do BaDAP2021-09-22
DOI10.21437/Interspeech.2021-1163
Rok publikacji2021
Typ publikacjimateriały konferencyjne (aut.)
Otwarty dostęptak
KonferencjaInterspeech 2021
Czasopismo/seriaInterspeech

Abstract

Modeling speaker embeddings using deep neural networks is currently state-of-the-art in speaker recognition. Recently, ResNet-based structures have gained a broader interest, slowly becoming the baseline along with the deep-rooted Time Delay Neural Network based models. However, the scale-decreased design of the ResNet models may not preserve all of the speaker information. In this paper, we investigate the SpineNet structure with scale-permuted design to tackle this problem, in which feature size either increases or decreases depending on the processing stage in the network. Apart from the presented adjustments of the SpineNet model for the speaker recognition task, we also incorporate popular modules dedicated to the residual-like structures, namely the Res2Net and Squeeze-and-Excitation blocks, and modify them to work effectively in the presented neural network architectures. The final proposed model, i.e., the SpineNet architecture with Res2Net and Time-Squeeze-and-Excitation blocks, achieves remarkable Equal Error Rates of 0.99 and 0.92 for the Extended and Original trial lists of the well-known VoxCeleb1 dataset.

Publikacje, które mogą Cię zainteresować

fragment książki
#130943Data dodania: 6.11.2020
On parameter adaptation in softmax-based cross-entropy loss for improved convergence speed and accuracy in DNN-based speaker recognition / Magdalena RYBICKA, Konrad KOWALCZYK // W: INTERSPEECH 2020 [Dokument elektroniczny] : October 25–29, Shanghai, China. — Wersja do Windows. — Dane tekstowe. — [China] : ISCA, cop. 2020. — (Interspeech : Proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 1990-9772). — S. 3805–3809. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://isca-speech.org/archive/Interspeech_2020/pdfs/2264.pdf [2020-11-05]. — Bibliogr. s. 3809, Abstr.
fragment książki
#136293Data dodania: 27.9.2021
Combating reverberation in NTF-based speech separation using a sub-source weighted multichannel Wiener filter and linear prediction / Mieszko FRAŚ, Marcin WITKOWSKI, Konrad KOWALCZYK // W: INTERSPEECH 2021 [Dokument elektroniczny] : proceedings of the 22nd annual conference of the International Speech Communication Association : 30 August–3 September 2021, Brno, Czechia. — Brno : Brno University of Technology, 2021. — (Interspeech : Proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 1990-9772). — e-ISBN: 978-171383690-2. — S. 3895–3899. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-speech.org/archive/pdfs/interspeech_2021/fra... [2021-09-21]. — Bibliogr. s. 3899. Abstr. — W bazie Scopus zakres stron: 2403-2407