Sturm, A., Unger, V., & Grünig, F. (2026). Human versus machine: Comparing human and large-language-model assessments of students’ writing through benchmark ratings. Didaktik Deutsch, 42–64. https://doi.org/10.21248/dideu.963