Szczegóły publikacji

Opis bibliograficzny

How to merge your embeddings: statistical vs attention-based speaker embedding aggregation for speaker verification with multiple enrollments / Justyna Krzywdziak, Piotr MASZTALSKI, Michał Romaniuk, Miłosz DUDEK, Joanna Stępień, Mateusz Matuszewski, Daria HEMMERLING // W: Eusipco2025 [Dokument elektroniczny] : 33rd European Signal Processing Conference : signal processing in the land of art, culture and beauty : Palermo, [Italy], 8–12 September 2025. — Wersja do Windows. — Dane tekstowe. — [Belgium] : European Association for Signal Processing (EURASIP), [2025]. — ( EUSIPCO... : European Signal Proceedings Conference ... ; ISSN  2076-1465 ). — e-ISBN: 978-9-46-459362-4. — S. 26–30. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://eusipco2025.org/wp-content/uploads/pdfs/0000026.pdf [2025-09-24]. — Bibliogr. s. 30, Abstr. — P. Masztalski, M. Dudek, J. Stępień, D. Hemmerling, J. Krzywdziak - dod. afiliacja: Samsung R&D Institute, Poland

Autorzy (7)

Słowa kluczowe

augmentationspeaker verificationspeaker embedding aggregationenrollmentattention

Dane bibliometryczne

ID BaDAP162911
Data dodania do BaDAP2025-09-25
DOI10.23919/EUSIPCO63237.2025.11226451
Rok publikacji2025
Typ publikacjimateriały konferencyjne (aut.)
Otwarty dostęptak
KonferencjaEuropean Signal Processing Conference 2025
Czasopismo/seriaEUSIPCO...

Abstract

In this paper, we evaluate existing state-of-the-art approaches to enrollment utterance aggregation, while proposing alternative methods to improve the effectiveness of speaker verification (SV) systems, when dealing with multiple utterances of the target speaker during the enrollment phase. We investigate multiple facets of the problem, including text-dependent and text-independent scenarios, as well as near-field and far-field speech. Additionally, we assess the impact of several enrollment utterance augmentation methods on aggregation quality. Our research evaluates embedding aggregation approaches, ranging from straightforward techniques such as calculating the average, max and median, to more advanced attention-based models. We propose a modified attention-based architecture that outperforms other techniques by 5 percentage points of Equal Error Rate (EER) in the performance of the verification system. Moreover we suggest a data augmentation method that can improve presented aggregation methods by almost 4 percentage points EER.