Szczegóły publikacji

Opis bibliograficzny

A hybrid feature-enhanced IndoBERT framework with controlled semi-supervised learning for low-resource Indonesian hate speech detection / Shoffan SAIFULLAH, Rafał DREŻEWSKI // Applied Sciences (Basel) [Dokument elektroniczny]. — Czasopismo elektroniczne ; ISSN  2076-3417 . — 2026 — vol. 16 iss. 13 art. no. 6478, s. 1–51. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 47–51, Abstr. — Publikacja dostępna online od: 2026-06-29. — S. Saifullah – dod. afiliacja: Department of Informatics, Universitas Pembangunan Nasional Veteran Yogyakarta, Indonesia

Autorzy (2)

Słowa kluczowe

IndoBERThate speech detectionpseudo-labelingTF-IDFIndonesian language processingsemi-supervised learninglow-resource NLPhybrid feature fusion

Dane bibliometryczne

ID BaDAP168986
Data dodania do BaDAP2026-07-30
Tekst źródłowyURL
DOI10.3390/app16136478
Rok publikacji2026
Typ publikacjiartykuł w czasopiśmie
Otwarty dostęptak
Creative Commons
Czasopismo/seriaApplied Sciences (Basel)

Abstract

Low-resource hate speech detection remains a challenging task for Indonesian social media due to limited labeled annotations, highly informal linguistic expressions, and substantial lexical variability. Under such conditions, purely supervised transformer models often suffer from unstable semantic generalization, while conventional pseudo-labeling methods are vulnerable to noisy unlabeled sample propagation. To address these limitations, this study proposes a hybrid feature-enhanced IndoBERT framework integrated with a controlled semi-supervised learning strategy. The proposed model combines contextual IndoBERT embeddings with abusive lexicon cues, handcrafted linguistic indicators, and TF-IDF–SVD statistical representations through a lightweight concatenation–projection feature fusion mechanism, while unlabeled data are incorporated via adaptive confidence thresholding and class-balanced pseudo-label selection to improve pseudo-label reliability. Extensive experiments were conducted under realistic low-resource supervision settings using only 5%, 10%, and 20% labeled data, and the proposed framework was systematically compared against representative baselines, including sparse lexical machine learning models, shallow neural architectures, multilingual transformers, IndoBERTweet, naive pseudo-labeling, and LLM-based prompting. The results show that model effectiveness is strongly supervision-dependent. Under the most extreme low-resource setting, compact statistical augmentation provides the most stable complementary signal, whereas under moderate low-resource supervision, the full hybrid representation combined with controlled semi-supervised learning yields the strongest and most consistent gains. The proposed Hybrid IndoBERT + controlled SSL framework outperforms all baselines at the 20% labeled setting, reaching an accuracy of 0.8654, Macro-F1 of 0.8633, and ROC-AUC of 0.9334. Additional analyses of pseudo-label reliability, calibration behavior, computational efficiency, and qualitative error patterns further show that the proposed framework improves low-resource robustness while maintaining comparable inference-time efficiency. These findings demonstrate that low-resource hate speech detection benefits most from the staged integration of contextual semantic modeling, interpretable linguistic cues, global lexical–statistical structure, and carefully regulated unlabeled data exploitation. Additional experiments using GPT-4o-mini and Llama-3.1-8B further demonstrate that the proposed framework remains competitive against general-purpose large language model prompting approaches under low-resource Indonesian hate speech detection scenarios. The proposed framework provides a practical and reproducible direction for hate speech detection in annotation-constrained social media environments.

Publikacje, które mogą Cię zainteresować

artykuł
#151683Data dodania: 30.1.2024
Automated text annotation using a semi-supervised approach with meta vectorizer and machine learning algorithms for hate speech detection / Shoffan SAIFULLAH, Rafał DREŻEWSKI, Felix Andika DWIYANTO, Agus Sasmito Aribowo, Yuli Fauziah, Nur Heri Cahyana // Applied Sciences (Basel) [Dokument elektroniczny]. — Czasopismo elektroniczne ; ISSN 2076-3417. — 2024 — vol. 14 iss. 3 art. no. 1078, s. 1–19. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 17–19, Abstr. — Publikacja dostępna online od: 2024-01-26. — S. Saifullah - dod. afiliacja: Department of Informatics, Universitas Pembangunan Nasional Veteran Yogyakarta, Indonesia. — R. Dreżewski - dod. afiliacja: Artificial Intelligence Research Group (AIRG), Informatics Department, Faculty of Industrial Technology, Universitas Ahmad Dahlan, Indonesia. — F. A. Dwiyanto - dod. afiliacja: Department of Electrical Engineering, Universitas Negeri Malang, Malang, Indonesia