Measuring the quality of AI-generated feedback
From theoretical modelling to empirical evidence
DOI:
https://doi.org/10.21248/dideu.966Keywords:
Writing, AI, Automated Writing Evaluation, Large Language Models, FeedbackAbstract
This study takes a theoretical and empirical approach to exploring the quality of AI-generated feedback on texts. First, it reviews the current state of research and demonstrates challenges involved in operationalizing feedback quality. Furthermore, it develops a theoretical three-level model that establishes the necessary terminology. The study then compares how 75 experienced teachers and three AI systems (Mistral Large, Llama 4 Maverick, GPT-4.1) provide feedback on three pupils’ texts. The feedback takes the form of criterion-based scores, overall grades, and short feedback texts. Eight trained raters assessed the quality of the feedback texts based on set criteria.
The teachers and AI systems (across 10 iterations) showed high interrater reliability (ICC = 0.7–0.9). The AI models consistently gave higher grades, but in the same ranking order as the teachers. The criterion-based scores differed significantly. Teachers weighted the individual criteria independently when predicting the overall grade, whereas AI systems did not. The strongest predictors of the quality of the feedback texts were concreteness and explanation. However, the source (AI vs. teacher) had no significant influence on the perceived usefulness.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Maurice Fürstenberg

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
