Szczegóły publikacji
Opis bibliograficzny
How to merge your embeddings: statistical vs attention-based speaker embedding aggregation for speaker verification with multiple enrollments / Justyna Krzywdziak, Piotr MASZTALSKI, Michał Romaniuk, Miłosz DUDEK, Joanna Stępień, Mateusz Matuszewski, Daria HEMMERLING // W: Eusipco2025 [Dokument elektroniczny] : 33rd European Signal Processing Conference : signal processing in the land of art, culture and beauty : Palermo, [Italy], 8–12 September 2025. — Wersja do Windows. — Dane tekstowe. — [Belgium] : European Association for Signal Processing (EURASIP), [2025]. — ( EUSIPCO... : European Signal Proceedings Conference ... ; ISSN 2076-1465 ). — e-ISBN: 978-9-46-459362-4. — S. 26–30. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://eusipco2025.org/wp-content/uploads/pdfs/0000026.pdf [2025-09-24]. — Bibliogr. s. 30, Abstr. — P. Masztalski, M. Dudek, J. Stępień, D. Hemmerling, J. Krzywdziak - dod. afiliacja: Samsung R&D Institute, Poland
Autorzy (7)
- Krzywdziak Justyna
- AGHMasztalski Piotr
- Romaniuk Michał
- AGHDudek Miłosz
- AGHStępień Joanna
- Matuszewski Mateusz
- AGHHemmerling Daria
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 162911 |
|---|---|
| Data dodania do BaDAP | 2025-09-25 |
| DOI | 10.23919/EUSIPCO63237.2025.11226451 |
| Rok publikacji | 2025 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | European Signal Processing Conference 2025 |
| Czasopismo/seria | EUSIPCO... |
Abstract
In this paper, we evaluate existing state-of-the-art approaches to enrollment utterance aggregation, while proposing alternative methods to improve the effectiveness of speaker verification (SV) systems, when dealing with multiple utterances of the target speaker during the enrollment phase. We investigate multiple facets of the problem, including text-dependent and text-independent scenarios, as well as near-field and far-field speech. Additionally, we assess the impact of several enrollment utterance augmentation methods on aggregation quality. Our research evaluates embedding aggregation approaches, ranging from straightforward techniques such as calculating the average, max and median, to more advanced attention-based models. We propose a modified attention-based architecture that outperforms other techniques by 5 percentage points of Equal Error Rate (EER) in the performance of the verification system. Moreover we suggest a data augmentation method that can improve presented aggregation methods by almost 4 percentage points EER.