Szczegóły publikacji

Opis bibliograficzny

Investigating code similarity patterns in LLM-generated and human-written programming solutions / Paulina GACEK, Bartosz Gdowski, Konrad Szymański, Wojciech Żmuda // W: Proceedings of the 18th International Conference on Computer Supported Education [Dokument elektroniczny] : May 18–20, 2026, Benidorm, Spain , Vol. 1 / eds. Edmundo Tovar, Tania Di Mascio, Christoph Meinel. — Wersja do Windows. — Dane tekstowe. — [Spain] : ScitePress, [2026]. — ( CSEDU ; ISSN  2184-5026 ). — e-ISBN: 978-989-758-833-4. — S. 478–485. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.scitepress.org/PublicationsDetail.aspx?ID=pYkGcjH... [2026-09-01]. — Bibliogr. s. 485, Abstr. — Dostęp po zalogowaniu

Autorzy (4)

Słowa kluczowe

computer science educationplagiarism detectioncode similaritycode generationsource code plagiarismlarge language models

Dane bibliometryczne

ID BaDAP169730
Data dodania do BaDAP2026-09-02
DOI10.5220/0014949400004021
Rok publikacji2026
Typ publikacjimateriały konferencyjne (aut.)
Otwarty dostęptak
Creative Commons
KonferencjaInternational Conference on Computer Supported Education 2026

Abstract

As Large Language Models (LLMs) reshape programming education, detecting AI-generated submissions using traditional similarity-based plagiarism tools presents a novel challenge. This paper investigates whether LLM-generated code exhibits structural and lexical patterns that systematically differ from human-written solutions. Analyzing a dataset of 9,000 solutions to constrained algorithmic tasks, we compare pre-2023 human submissions against synthetic outputs from GPT-5.2 and Gemini 3.0 Flash. Our findings reveal that AI solutions occupy a highly concentrated space, exhibiting extreme structural convergence (¿75% similarity compared to 28% for humans) and highly constrained lexical overlap. While commenting behaviors vary significantly by model, these overarching convergences suggest similarity-based detection is theoretically viable. However, due to severe risks of model-specific statistical masking and false positives, strict pedagogical precautions are required before deployment.

Publikacje, które mogą Cię zainteresować

fragment książki
#169727Data dodania: 2.9.2026
Comparative analysis of non-commercial plagiarism detectors for computer science education / Paulina GACEK, Bartosz Gdowski, Konrad Szymański, Wojciech Żmuda // W: Proceedings of the 18th International Conference on Computer Supported Education [Dokument elektroniczny] : May 18–20, 2026, Benidorm, Spain , Vol. 3 / eds. Edmundo Tovar, Tania Di Mascio, Christoph Meinel. — Wersja do Windows. — Dane tekstowe. — [Spain] : ScitePress, [2026]. — ( CSEDU ; ISSN  2184-5026 ). — e-ISBN: 978-989-758-833-4. — S. 1972–1982. — Wymagania systemowe: Adobe Reader. — Tryb dostępu: https://www.scitepress.org/Link.aspx?doi=10.5220/001483650000... [2026-09-01]. — Bibliogr. s. 1981–1982, Abstr. — Dostęp po zalogowaniu
fragment książki
#161865Data dodania: 3.9.2025
Regarding context size in LLM-based metaheuristic design / Adam Viktorin, Michal PLUHÁČEK, Jozef Kovac, Tomas Kadavy, Roman Senkerik // W: GECCO'25 Companion [Dokument elektroniczny] : proceedings of the 2025 Genetic and Evolutionary Computation Conference Companion : July 14–18, 2025, Málaga, Spain. — Wersja do Windows. — Dane tekstowe. — New York : Association for Computing Machinery, 2025. — e-ISBN: 979-8-4007-1464-1. — S. 2345–2353. — Wymagania systemowe: Adobe Reader. — Bibliogr. s. 2352–2353, Abstr. — Publikacja dostępna online od: 2025-08-11