Human versus machine

Comparing human and large-language-model assessments of students’ writing through benchmark ratings

Autor/innen

  • Afra Sturm
  • Valentin Unger
  • Fabian Grünig

DOI:

https://doi.org/10.21248/dideu.963

Schlagworte:

text rating, large language models, narrative texts, instructive texts, elementary school

Abstract

For teachers to effectively use large-language-model(LLM)-based ratings in the formative or summative assessment of texts, it is essential to ensure that such ratings can assess student writing in a valid and reliable manner. This study investigates whether a validated human text-rating procedure (benchmark rating) can be replicated by an LLM-based rating procedure. We tested the replication with two genres of elementary school students’ text—narrative and instructive—using nine LLMs from three providers (OpenAI, Anthropic, Mistral). Each LLM generated three independent scores per text via structured, benchmark-aligned prompts that were then aggregated into a consensus score. Results showed that intrarater reliability was high to excellent, ICC(3, k) ≈ .68–.97, and alignment with human ratings ranged from moderate to strong, ICC(3, 1) ≈ .47–.85, with larger models consistently outperforming smaller ones. Systematic bias patterns emerged, varying by model and genre, indicating a need for calibration. Increasing output token windows and reducing temperature parameters mitigated truncation and schema-related failures. Although LLM-based benchmark ratings can approximate expert judgments and reduce the need for labor-intensive human triple coding, limitations remain regarding cost (for larger models), genre- and task specificity, and sensitivity to text presentation and student grade level—factors that constrain immediate classroom use, particularly for formative feedback.

Downloads

Veröffentlicht

2026-09-03

Zitationsvorschlag

Sturm, A., Unger, V., & Grünig, F. (2026). Human versus machine: Comparing human and large-language-model assessments of students’ writing through benchmark ratings. Didaktik Deutsch, 42–64. https://doi.org/10.21248/dideu.963

Ausgabe

Rubrik

Empirical Research Articles