lnu.sePublications
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
The Future of Grading: Can LLMs Accurately Score Student Work?
Norwegian University of Science and Technology, Norway.
Norwegian University of Science and Technology, Norway.
Linnaeus University, Faculty of Technology, Department of Informatics.ORCID iD: 0000-0002-0199-2377
Kristiania University of Applied Sciences, Norway.
Show others and affiliations
2025 (English)In: Proceedings of the Future Technologies Conference (FTC) 2025, Volume 4, Springer Nature, 2025, p. 268-289Conference paper, Published paper (Refereed)
Abstract [en]

Exploring the transformative potential of large language models (LLMs) in revolutionizing the assessment process in education is a pressing need. LLMs have the capability to automatically evaluate student submissions, significantly enhancing the educational landscape and providing a second opinion and additional support to human evaluators. Therefore, our evaluation of various LLMs aims to assess their ability to provide automated, detailed, consistent, unbiased, and efficient feedback and scoring in educational settings, instilling a sense of optimism about the future of assessment. To achieve our objective, we comprehensively evaluated various state-of-the-art LLMs, such as GPT-4, GPT-4o, and Mixtral 8x22B on diverse datasets, including ASAP SAS, ASAP AES, and a real-world BWD dataset specifically collected for this study. The experimental results on these datasets employing sound prompt engineering techniques demonstrate that LLMs possess the potential not only to automate the scoring of student submissions but also to accurately match the scores of human assessors for actual courses taught in universities. Notably, GPT-4o exhibited promising capabilities in scoring short-answer submissions. These models were particularly good in STEM-related domains for tasks with clear structure, well-defined rubrics, and minimal subjective interpretation. However, the study identifies specific challenges for certain tasks, particularly creative tasks, underscoring the need for further research in this area.

Place, publisher, year, edition, pages
Springer Nature, 2025. p. 268-289
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:lnu:diva-142224DOI: 10.1007/978-3-032-07992-3_18Scopus ID: 2-s2.0-105021830683ISBN: 9783032079923 (electronic)OAI: oai:DiVA.org:lnu-142224DiVA, id: diva2:2009803
Conference
Future Technologies Conference
Available from: 2025-10-28 Created: 2025-10-28 Last updated: 2026-01-21Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Kastrati, Zenun

Search in DiVA

By author/editor
Kastrati, Zenun
By organisation
Department of Informatics
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 41 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf