AI Feedback in Higher Education: Why Design Beats Tools

The most useful thing about this paper isn’t only the framework it builds. It’s the failure that comes first. Zapata, Tzirides, Cope, Kalantzis, and Searsmith (2026) spent several instructional cycles putting generative AI feedback in front of graduate students, and the students rejected it. That rejection, and what the team did next, is why this study earns your attention.

I advocate for AI in education. I also think most AI-feedback pilots skip the part that actually decides success. This paper doesn’t skip it.

The Study That Watched AI Feedback Fail, Then Flip

Here’s the setup. The researchers built a formative feedback ecology in graduate education courses, where students received comments from peers, instructors, and an AI reviewer running on GPT-3. Early on, the AI lost badly. Students described peer feedback as more specific, more context-aware, and warmer. Some compared the AI reviewer to the machines in Blade Runner and 2001: A Space Odyssey. The anxiety was real, and the authors report it plainly.

Then they changed one thing. In 2024 the team rebuilt the AI reviewer with retrieval-augmented generation, feeding it a corpus of 35 million tokens: five years of graduate drafts plus instructor publications. After that, perceptions flipped. Students and instructors began describing the calibrated AI as a genuine contributor, and in some cases said its feedback rivaled peer comments.

That flip is the whole argument in miniature. The tool didn’t get smarter on its own. The team made it fit the context.

AI Feedback in Higher Education

Why Calibration Changed the AI Feedback Story

I want to be careful here, and so are the authors. When students called the calibrated AI a “collaborative partner,” the researchers didn’t take that as proof the machine had become a colleague. They note that “we treat these characterisations as empirically meaningful descriptors in the feedback activity system and not as claims about AI agency or social personhood” (p. 6). That restraint is the kind I wish more AI research showed.

What changed wasn’t the AI’s soul. It was the alignment. A generic model produced generic comments. A model grounded in the real rubric, the real discipline, and years of actual student writing produced feedback students could use. The calibrated AI gave them something worth weighing. This also tracks with the case Bearman and colleagues make about evaluative judgement, where students learn to weigh a source against their own reasoning before accepting it.

The framework itself, GenAI-by-Design, marries two approaches many of my readers will recognize: Cope and Kalantzis’s Learning by Design and UNESCO’s AI Competency Framework for Students. Learning by Design supplies eight knowledge processes that move students between experience and reflection. UNESCO contributes four competency areas across three levels, Understand, Apply, and Create. The paper maps one onto the other so AI literacy grows as a spiral across a course, not a single lecture.

Grade the Process, Not the Product

The assessment argument is where this paper does its best work. Zapata et al. say the old model breaks. They argue that “traditional models that assess only final products are inadequate in GenAI-enhanced environments” (p. 16). If a student can generate a polished artifact in seconds, grading the artifact tells you almost nothing.

Their answer is to assess the process. They want transparent prompt-and-revision logs, and evidence of how students questioned, refined, and justified AI output. Give credit for the reasoning, not the polish. This is the same lesson Fan et al. (2025) taught with metacognitive laziness, where AI improved the product without improving the thinking behind it. Grade the thinking, and the incentive changes.

I’d push this into everyday practice. Ask students to disagree with the AI at least once and defend the disagreement. Make the prompt log part of the grade. Treat the revision history as the real submission.

References

  • Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. _Assessment & Evaluation in Higher Education, 49_(6), 893-905. https://doi.org/10.1080/02602938.2024.2335321
  • Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. _British Journal of Educational Technology, 56 (2), 489–530. https://doi.org/10.1111/bjet.13544
  • Kalantzis, M., & Cope, B. (2025). Literacy in the time of artificial intelligence. Reading Research Quarterly, 60_, e591. https://doi.org/10.1002/rrq.591
  • UNESCO. (2024). AI competency framework for students. United Nations Educational, Scientific and Cultural Organization. https://doi.org/10.54675/JKJB9835
  • Zapata, G. C., Tzirides, A. O., Cope, B., Kalantzis, M., & Searsmith, D. (2026). GenAI-by-Design: A theoretically grounded, research-informed pedagogical framework for classroom practice and assessment in higher education. Pedagogy, Culture & Society. Advance online publication. https://doi.org/10.1080/14681366.2026.2667388

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top