Szczegóły publikacji
Opis bibliograficzny
On parameter adaptation in softmax-based cross-entropy loss for improved convergence speed and accuracy in DNN-based speaker recognition / Magdalena RYBICKA, Konrad KOWALCZYK // W: INTERSPEECH 2020 [Dokument elektroniczny] : October 25–29, Shanghai, China. — Wersja do Windows. — Dane tekstowe. — [China] : ISCA, cop. 2020. — (Interspeech : Proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 1990-9772). — S. 3805–3809. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://isca-speech.org/archive/Interspeech_2020/pdfs/2264.pdf [2020-11-05]. — Bibliogr. s. 3809, Abstr.
Autorzy (2)
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 130943 |
|---|---|
| Data dodania do BaDAP | 2020-11-06 |
| DOI | 10.21437/Interspeech.2020-2264 |
| Rok publikacji | 2020 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2020 |
| Czasopismo/seria | Interspeech |
Abstract
In various classification tasks the major challenge is in generating discriminative representation of classes. By proper selection of deep neural network (DNN) loss function we can encourage it to produce embeddings with increased inter-class separation and smaller intra-class distances. In this paper, we develop softmax-based cross-entropy loss function which adapts its parameters to the current training phase. The proposed solution improves accuracy up to 24% in terms of Equal Error Rate (EER) and minimum Detection Cost Function (minDCF). In addition, our proposal also accelerates network convergence compared with other state-of-the-art softmax-based losses. As an additional contribution of this paper, we adopt and subsequently modify the ResNet DNN structure for the speaker recognition task. The proposed ResNet network achieves relative gains of up to 32% and 15% in terms of EER and minDCF respectively, compared with the well-established Time Delay Neural Network (TDNN) architecture for x-vector extraction.