Discriminating Human-authored from ChatGPT-Generated Code Via Discernable Feature Analysis

June 26, 2023 · Declared Dead · 🏛 2023 IEEE 34th International Symposium on Software Reliability Engineering Workshops (ISSREW)

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Li Ke, Hong Sheng, Fu Cai, Zhang Yunhe, Liu Ming arXiv ID 2306.14397 Category cs.SE: Software Engineering Cross-listed cs.CY Citations 14 Venue 2023 IEEE 34th International Symposium on Software Reliability Engineering Workshops (ISSREW) Last Checked 4 months ago

Abstract

The ubiquitous adoption of Large Language Generation Models (LLMs) in programming has underscored the importance of differentiating between human-written code and code generated by intelligent models. This paper specifically aims to distinguish code generated by ChatGPT from that authored by humans. Our investigation reveals disparities in programming style, technical level, and readability between these two sources. Consequently, we develop a discriminative feature set for differentiation and evaluate its efficacy through ablation experiments. Additionally, we devise a dataset cleansing technique, which employs temporal and spatial segmentation, to mitigate the dearth of datasets and to secure high-caliber, uncontaminated datasets. To further enrich data resources, we employ "code transformation," "feature transformation," and "feature customization" techniques, generating an extensive dataset comprising 10,000 lines of ChatGPT-generated code. The salient contributions of our research include: proposing a discriminative feature set yielding high accuracy in differentiating ChatGPT-generated code from human-authored code in binary classification tasks; devising methods for generating extensive ChatGPT-generated codes; and introducing a dataset cleansing strategy that extracts immaculate, high-grade code datasets from open-source repositories, thus achieving exceptional accuracy in code authorship attribution tasks.