Szczegóły publikacji

Opis bibliograficzny

On parameter adaptation in softmax-based cross-entropy loss for improved convergence speed and accuracy in DNN-based speaker recognition / Magdalena RYBICKA, Konrad KOWALCZYK // W: INTERSPEECH 2020 [Dokument elektroniczny] : October 25–29, Shanghai, China. — Wersja do Windows. — Dane tekstowe. — [China] : ISCA, cop. 2020. — (Interspeech : Proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 1990-9772). — S. 3805–3809. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://isca-speech.org/archive/Interspeech_2020/pdfs/2264.pdf [2020-11-05]. — Bibliogr. s. 3809, Abstr.

Autorzy (2)

Słowa kluczowe

ResNetspeaker recognitionspeaker embeddingsoftmax activation functionsdeep neural networks

Dane bibliometryczne

ID BaDAP130943
Data dodania do BaDAP2020-11-06
DOI10.21437/Interspeech.2020-2264
Rok publikacji2020
Typ publikacjimateriały konferencyjne (aut.)
Otwarty dostęptak
KonferencjaInterspeech 2020
Czasopismo/seriaInterspeech

Abstract

In various classification tasks the major challenge is in generating discriminative representation of classes. By proper selection of deep neural network (DNN) loss function we can encourage it to produce embeddings with increased inter-class separation and smaller intra-class distances. In this paper, we develop softmax-based cross-entropy loss function which adapts its parameters to the current training phase. The proposed solution improves accuracy up to 24% in terms of Equal Error Rate (EER) and minimum Detection Cost Function (minDCF). In addition, our proposal also accelerates network convergence compared with other state-of-the-art softmax-based losses. As an additional contribution of this paper, we adopt and subsequently modify the ResNet DNN structure for the speaker recognition task. The proposed ResNet network achieves relative gains of up to 32% and 15% in terms of EER and minDCF respectively, compared with the well-established Time Delay Neural Network (TDNN) architecture for x-vector extraction.

Publikacje, które mogą Cię zainteresować

fragment książki
#136305Data dodania: 22.9.2021
Spine2Net: SpineNet with Res2Net and time-squeeze-and-excitation blocks for speaker recognition / Magdalena RYBICKA, Jesús Villalba, Piotr Żelasko, Najim Dehak, Konrad KOWALCZYK // W: INTERSPEECH 2021 [Dokument elektroniczny] : proceedings of the 22nd annual conference of the International Speech Communication Association : 30 August–3 September 2021, Brno, Czechia. — Brno : Brno University of Technology, 2021. — (Interspeech : Proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 1990-9772). — e-ISBN: 978-171383690-2. — S. 496–500. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-speech.org/archive/pdfs/interspeech_2021/ryb... [2021-09-21]. — Bibliogr. s. 500, Abstr. — W bazie Scopus numery stron: 491-495. — P. Żelasko - afiliacja: Johns Hopkins University, Baltimore, USA
fragment książki
#162187Data dodania: 10.9.2025
Efficient low-latency speech enhancement with mobile audio streaming networks / Michal Romaniuk, Piotr Masztalski, Karol Piaskowski, Mateusz Matuszewski // W: INTERSPEECH 2020 [Dokument elektroniczny] : October 25–29, Shanghai, China. — Wersja do Windows. — Dane tekstowe. — [China] : ISCA, cop. 2020. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN  2958-1796 ). — S. 3296-3300. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.isca-archive.org/interspeech_2020/romaniuk20_inte... [2025-09-09]. — Bibliogr. s. 3299-3300, Abstr.