lnu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Code Correctness and Quality in the Era of AI Code Generation: Examining ChatGPT and GitHub Copilot
Linnaeus University, Faculty of Technology, Department of computer science and media technology (CM).
Linnaeus University, Faculty of Technology, Department of computer science and media technology (CM).
2023 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

The use of AI tools for code generation is increasing in popularity, and two of these tools are ChatGPT and GitHub Copilot. These tools could potentially reduce development time and costs for developers and companies, however, ensuring the correctness and quality of AI-generated code is crucial for its adoption. This study conducted a quantitative controlled experiment to evaluate the code generation capabilities of Copilot and ChatGPT in terms of code correctness and quality. The experiment aimed to address research questions regarding the performance of these AI tools. The results indicate that both ChatGPT and Copilot can generate correct code from given instructions, though there is room for improvement. ChatGPT achieved a correctness rate of 87.33%, while Copilot performed slightly better at 89%. Statistical analysis revealed no significant difference in code correctness between the two tools. Regarding code quality, ChatGPT demonstrated impressive performance, with 98.52% of generated lines free from quality rule violations. Furthermore, 80.7% of ChatGPT-generated algorithms had no quality rule violations. Copilot generated correct lines for 94.07% of total lines but only achieved 64.7% of algorithms with no quality rule violations. The statistical analysis showed a statistically significant difference in code quality between ChatGPT and Copilot, indicating that ChatGPT generally produces higher quality code. This research contributes to understanding the capabilities of AI code generation tools and highlights their potential to produce correct and high-quality code. 

Place, publisher, year, edition, pages
2023. , p. 69
Keywords [en]
AI Code Generation, ChatGPT, Copilot, Code Correctness, Code Quality
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:lnu:diva-121545OAI: oai:DiVA.org:lnu-121545DiVA, id: diva2:1764568
Subject / course
Computer Science
Educational program
Web Development Programme, 180 credits
Supervisors
Examiners
Available from: 2023-07-03 Created: 2023-06-08 Last updated: 2023-07-03Bibliographically approved

Open Access in DiVA

fulltext(2381 kB)2130 downloads
File information
File name FULLTEXT01.pdfFile size 2381 kBChecksum SHA-512
b80650c8efa568cfaa8e8c2f19aff57dbee9a61d008c6b9b799a2536d4b56bca81803b626e27263aefc6f81d16006c9d8678f45f47bcb2b9b820c7ee14956265
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Hansson, EmiliaEllréus, Oliwer
By organisation
Department of computer science and media technology (CM)
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 2130 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 8150 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf