lnu.sePublikationer
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
The Future of Grading: Can LLMs Accurately Score Student Work?
Norwegian University of Science and Technology, Norway.
Norwegian University of Science and Technology, Norway.
Linnéuniversitetet, Fakulteten för teknik (FTK), Institutionen för informatik (IK).ORCID-id: 0000-0002-0199-2377
Kristiania University of Applied Sciences, Norway.
Visa övriga samt affilieringar
2025 (Engelska)Ingår i: Proceedings of the Future Technologies Conference (FTC) 2025, Volume 4, Springer Nature, 2025, s. 268-289Konferensbidrag, Publicerat paper (Refereegranskat)
Abstract [en]

Exploring the transformative potential of large language models (LLMs) in revolutionizing the assessment process in education is a pressing need. LLMs have the capability to automatically evaluate student submissions, significantly enhancing the educational landscape and providing a second opinion and additional support to human evaluators. Therefore, our evaluation of various LLMs aims to assess their ability to provide automated, detailed, consistent, unbiased, and efficient feedback and scoring in educational settings, instilling a sense of optimism about the future of assessment. To achieve our objective, we comprehensively evaluated various state-of-the-art LLMs, such as GPT-4, GPT-4o, and Mixtral 8x22B on diverse datasets, including ASAP SAS, ASAP AES, and a real-world BWD dataset specifically collected for this study. The experimental results on these datasets employing sound prompt engineering techniques demonstrate that LLMs possess the potential not only to automate the scoring of student submissions but also to accurately match the scores of human assessors for actual courses taught in universities. Notably, GPT-4o exhibited promising capabilities in scoring short-answer submissions. These models were particularly good in STEM-related domains for tasks with clear structure, well-defined rubrics, and minimal subjective interpretation. However, the study identifies specific challenges for certain tasks, particularly creative tasks, underscoring the need for further research in this area.

Ort, förlag, år, upplaga, sidor
Springer Nature, 2025. s. 268-289
Nationell ämneskategori
Språkbehandling och datorlingvistik
Identifikatorer
URN: urn:nbn:se:lnu:diva-142224DOI: 10.1007/978-3-032-07992-3_18Scopus ID: 2-s2.0-105021830683ISBN: 9783032079923 (digital)OAI: oai:DiVA.org:lnu-142224DiVA, id: diva2:2009803
Konferens
Future Technologies Conference
Tillgänglig från: 2025-10-28 Skapad: 2025-10-28 Senast uppdaterad: 2026-01-21Bibliografiskt granskad

Open Access i DiVA

Fulltext saknas i DiVA

Övriga länkar

Förlagets fulltextScopus

Person

Kastrati, Zenun

Sök vidare i DiVA

Av författaren/redaktören
Kastrati, Zenun
Av organisationen
Institutionen för informatik (IK)
Språkbehandling och datorlingvistik

Sök vidare utanför DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetricpoäng

doi
isbn
urn-nbn
Totalt: 53 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf