Measuring the quality of AI-generated feedback

From theoretical modelling to empirical evidence

Authors

  • Maurice Fürstenberg

DOI:

https://doi.org/10.21248/dideu.966

Keywords:

Writing, AI, Automated Writing Evaluation, Large Language Models, Feedback

Abstract

This study takes a theoretical and empirical approach to exploring the quality of AI-generated feedback on texts. First, it reviews the current state of research and demonstrates challenges involved in operationalizing feedback quality. Furthermore, it develops a theoretical three-level model that establishes the necessary terminology. The study then compares how 75 experienced teachers and three AI systems (Mistral Large, Llama 4 Maverick, GPT-4.1) provide feedback on three pupils’ texts. The feedback takes the form of criterion-based scores, overall grades, and short feedback texts. Eight trained raters assessed the quality of the feedback texts based on set criteria.

The teachers and AI systems (across 10 iterations) showed high interrater reliability (ICC = 0.7–0.9). The AI models consistently gave higher grades, but in the same ranking order as the teachers. The criterion-based scores differed significantly. Teachers weighted the individual criteria independently when predicting the overall grade, whereas AI systems did not. The strongest predictors of the quality of the feedback texts were concreteness and explanation. However, the source (AI vs. teacher) had no significant influence on the perceived usefulness.

Downloads

Published

2026-09-03

How to Cite

Fürstenberg, M. (2026). Measuring the quality of AI-generated feedback: From theoretical modelling to empirical evidence. Didaktik Deutsch, 65–96. https://doi.org/10.21248/dideu.966

Issue

Section

Empirical Research Articles