Szczegóły publikacji
Opis bibliograficzny
RB-CCR: Radial-Based Combined Cleaning and Resampling algorithm for imbalanced data classification / Michał KOZIARSKI, Colin Bellinger, Michał Woźniak // Machine Learning ; ISSN 0885-6125. — 2021 — vol. 110 iss. 11-12, s. 3059–3093. — Bibliogr. s. 3091–3093, Abstr. — Publikacja dostępna online od: 2021-10-14
Autorzy (3)
- AGHKoziarski Michał
- Bellinger Colin
- Woźniak Michał
Słowa kluczowe
Dane bibliometryczne
| ID BaDAP | 138846 |
|---|---|
| Data dodania do BaDAP | 2022-01-19 |
| Tekst źródłowy | URL |
| DOI | 10.1007/s10994-021-06012-8 |
| Rok publikacji | 2021 |
| Typ publikacji | artykuł w czasopiśmie |
| Otwarty dostęp | |
| Creative Commons | |
| Czasopismo/seria | Machine Learning |
Abstract
Real-world classification domains, such as medicine, health and safety, and finance, often exhibit imbalanced class priors and have asynchronous misclassification costs. In such cases, the classification model must achieve a high recall without significantly impacting precision. Resampling the training data is the standard approach to improving classification performance on imbalanced binary data. However, the state-of-the-art methods ignore the local joint distribution of the data or correct it as a post-processing step. This can causes sub-optimal shifts in the training distribution, particularly when the target data distribution is complex. In this paper, we propose Radial-Based Combined Cleaning and Resampling (RB-CCR). RB-CCR utilizes the concept of class potential to refine the energy-based resampling approach of CCR. In particular, RB-CCR exploits the class potential to accurately locate sub-regions of the data-space for synthetic oversampling. The category sub-region for oversampling can be specified as an input parameter to meet domain-specific needs or be automatically selected via cross-validation. Our 5 x 2 cross-validated results on 57 benchmark binary datasets with 9 classifiers show that RB-CCR achieves a better precision-recall trade-off than CCR and generally out-performs the state-of-the-art resampling methods in terms of AUC and G-mean.