Szczegóły publikacji

Opis bibliograficzny

Improving noisy label learning via transition matrix estimation from feature embeddings / Mateusz WOJTULEWICZ, Piotr DUDA, Dacheng Tao, Leszek RUTKOWSKI // Knowledge-Based Systems / Butterworths ; ISSN  0950-7051 . — 2026 — vol. 351 pt. A art. no. 116676, s. 1–20. — Bibliogr. s. 19–20, Abstr. — Publikacja dostępna online od: 2026-07-20. — M. Wojtulewicz, P. Duda. L. Rutkowski - dod. afiliacja: CDSI AGH. — P. Duda - dod. afiliacja: Faculty of Computer Science and Artificial Intelligence, Częstochowa University of Technology, Częstochowa, Poland. — L. Rutkowski - dod. afiliacja: Systems Research Institute, Polish Academy of Sciences, Warsaw, Poland

Autorzy (4)

Słowa kluczowe

noisy label learningdata centric artificial intelligencetransition matrix estimationfeature embedding geometryloss correction

Dane bibliometryczne

ID BaDAP169597
Data dodania do BaDAP2026-09-23
Tekst źródłowyURL
DOI10.1016/j.knosys.2026.116676
Rok publikacji2026
Typ publikacjiartykuł w czasopiśmie
Otwarty dostęptak
Creative Commons
Czasopismo/seriaKnowledge-Based Systems

Abstract

Learning in the presence of noisy labels remains a major challenge for modern machine learning systems. A central concept in this setting is the noise transition matrix, which encodes how clean labels are corrupted into noisy observations and enables principled loss correction. However, estimating this matrix is difficult when no clean labels, anchor points, or strong structural priors are available. In this paper, we propose an unsupervised framework that estimates the transition matrix directly from the geometry of feature embeddings obtained either from noisy-label training or from related pretrained feature extractors. Our approach exploits neighborhood consistency in the embedding space and introduces two complementary variants: Hard Neighbor Voting (HNV) and Soft Neighbor Voting (SNV). Both methods construct empirical confusion matrices without access to clean labels or anchor samples and estimate both the transition matrix and the global noise rate. Experiments on synthetic noise (MNIST, EMNIST, Fashion-MNIST, and Kuzushiji-MNIST) and real-world human label noise (CIFAR-10N) show that the proposed estimators improve direct transition-matrix recovery over the Anchor Points (AP) baseline in the studied settings. When integrated into standard training pipelines through loss correction, the resulting estimates can also improve downstream classification performance and remain competitive with representative state-of-the-art robust learning methods. © 2026 The Authors