Szczegóły publikacji
Opis bibliograficzny
Clustering-based hard negative sampling for supervised contrastive speaker verification / Piotr MASZTALSKI, Michał Romaniuk, Jakub Żak, Mateusz Matuszewski, Konrad KOWALCZYK // W: Interspeech 2025 [Dokument elektroniczny] : 17–21 August 2025, Rotterdam, The Netherlands. — Wersja do Windows. — Dane tekstowe. — [France : ISCA], [2025]. — ( Interspeech : proceedings of the ... Annual Conference of the International Speech Communication Association ; ISSN 2958-1796 ). — S. 3698–3702. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 3702, Abstr. — P. Masztalski - dod. afiliacja: Samsung R&D Institute Poland
Autorzy (5)
- AGHMasztalski Piotr
- Romaniuk Michał
- Żak Jakub
- Matuszewski Mateusz
- AGHKowalczyk Konrad
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 162076 |
|---|---|
| Data dodania do BaDAP | 2025-09-08 |
| Tekst źródłowy | URL |
| DOI | 10.21437/Interspeech.2025-442 |
| Rok publikacji | 2025 |
| Typ publikacji | materiały konferencyjne (aut.) |
| Otwarty dostęp | |
| Konferencja | Interspeech 2025 |
| Czasopismo/seria | Interspeech |
Abstract
In speaker verification, contrastive learning is gaining popularity as an alternative to the traditionally used classificationbased approaches. Contrastive methods can benefit from an effective use of hard negative pairs, which are different-class samples particularly challenging for a verification model due to their similarity. In this paper, we propose CHNS - a clusteringbased hard negative sampling method, dedicated for supervised contrastive speaker representation learning. Our approach clusters embeddings of similar speakers, and adjusts batch composition to obtain an optimal ratio of hard and easy negatives during contrastive loss calculation. Experimental evaluation shows that CHNS outperforms a baseline supervised contrastive approach with and without loss-based hard negative sampling, as well as a state-of-the-art classification-based approach to speaker verification by as much as 18 % relative EER and minDCF on the VoxCeleb dataset using two lightweight model architectures.