AI Assessment Guidelines: What Top Universities Got Right and Wrong

When ChatGPT dropped in late 2022, universities scrambled. Some banned it outright. Others released cautious memos. A few tried to get ahead of the disruption with actual guidelines. Moorhouse, Yeo, and Wan (2023) decided to find out what the world’s top-ranking institutions actually told their instructors to do, and the results tell a story that’s equal parts progress and missed opportunity.

The study reviewed publicly available assessment guidelines from the top 50 universities according to the Times Higher Education 2023 World University Rankings. The search was conducted on June 15, 2023, which makes it an early snapshot, not a comprehensive audit.

What the University Guidelines Covered

Only 23 of the 50 universities had publicly available, university-level assessment guidelines written for instructors. Fifteen of those institutions were in the United States, four were in the United Kingdom, two were in Canada, and one each was in Australia and Japan. Most of the guidance came from university teaching and learning centres.

This is an important limitation of the study. The analysis describes guidance available on one day in June 2023, during the first institutional response to ChatGPT. It also reflects a small group of highly ranked and well-resourced universities, with little representation from institutions in the Global South.

The terminology was already becoming a problem. Fifty-seven percent of the guidelines referred only to ChatGPT. Treating one product as a synonym for generative AI could leave instructors unprepared for image generators, coding assistants, writing tools, and other systems that affect assessment in different ways.

Across the 23 universities, the guidance concentrated on three areas: academic integrity, assessment design, and communication with students. Sixty percent discussed plagiarism, while approximately 61 percent mentioned AI-detection tools. Most of the institutions that addressed detection warned instructors about its limitations, including inaccuracy, privacy concerns, and the possibility of damaging trust with students.

AI Assessment Guidelines

On academic integrity, Moorhouse et al. found that 60% of the guidelines addressed plagiarism directly. They identified three forms of GAI-related plagiarism that institutions were worried about: copying and pasting AI-generated responses, running material through multiple AI paraphrasers to dodge detection, and failing to document GAI use.

About 61% mentioned detection tools like GPTZero and Turnitin, but most actually discouraged relying on them. The reasons? Inaccuracy, privacy concerns, and the risk of eroding trust between faculty and students. That’s a significant finding, because it shows that even at an early stage, many institutions recognized that detection is a dead end. I’ve written about the postplagiarism argument before, where the focus shifts from catching students to rethinking what we ask them to do. The fact that top universities were already backing away from detection tools in mid-2023 supports that trajectory.

Five Recommendations for Redesigning Assessment

Assessment design received considerable attention in the reviewed guidance. Seventy-four percent of the universities offered instructors specific recommendations for modifying assessment in response to generative AI.

Five approaches appeared across the guidelines:

  • test existing assignments with generative AI tools;
  • design tasks that require critical thinking, creativity, and personal reflection;
  • assess the development process through drafts and scaffolded stages;
  • incorporate generative AI into selected assessment activities;
  • use supervised or in-class assessment when it matches the learning outcome.

Testing an assignment with AI can help an instructor see how easily the task can be completed and which parts of the prompt need reconsideration. This should be treated as an exploratory step. An assignment redesigned around a current limitation may become vulnerable as the technology improves.

Process evidence offers a more durable response. Drafts, planning notes, conferences, reflections, and oral follow-ups allow instructors to examine how students developed their work. These elements can make student thinking visible even when AI contributes to the finished product.

Ten universities recommended incorporating generative AI directly into some assessments. Students might generate an AI response and evaluate its accuracy, compare several outputs, identify missing perspectives, or revise weak reasoning. These activities treat AI output as material for analysis and keep the student responsible for intellectual judgment.

Communication was the most common recommendation. Eighty-seven percent of the guidelines advised instructors to discuss expectations clearly with students. Suggested approaches included syllabus statements, classroom conversations, and collaboration with librarians. Effective guidance should explain which uses are permitted, when disclosure is required, what students remain responsible for, and how the rules connect to learning.

The tension running through the whole study is this: most guidelines still lean defensive. Moorhouse et al. put it this way: “The emphasis in the guidelines still seem to focus on limiting or preventing GAI use in assessment tasks” (p. 9). And then they argue for a different direction: “Allowing or even requiring students to use GAI at various stages of the assessment process would, in fact, enhance the authenticity of assessments” (p. 9). I think they’re right about where the field needs to go. If professionals in every industry are using generative AI, then asking students to prove they can work without it doesn’t prepare them for anything real. The authenticity argument is strong.

Moorhouse et al. also propose a new competency they call “generative artificial intelligence assessment literacy.” It has three components: recognizing GAI’s implications for academic integrity, designing assessments that make room for both student learning and GAI use, and communicating with students about productive and ethical GAI use.

The authors argue that “the specific nature of assessments means that instructors need to develop GAI assessment literacy” (p. 9). I’d complicate that slightly. It’s not that instructors need a whole new literacy. It’s that the assessment literacy they’ve always needed now has a new dimension. The principles of good assessment design, alignment, authenticity, transparency, don’t change. What changes is the environment those principles have to work in.

This paper is a useful baseline. It documented what the world’s most resourced institutions were thinking in the first half of 2023. But a baseline is only useful if we measure the distance traveled since. The question now is whether these guidelines evolved or whether they calcified.

References

  • Eaton, S. E. (2023). Postplagiarism: Transdisciplinary ethics and integrity in the age of artificial intelligence and neurotechnology. International Journal for Educational Integrity, 19(23). https://doi.org/10.1007/s40979-023-00144-1
  • Moorhouse, B. L., Yeo, M. A., & Wan, Y. (2023). Generative AI tools and assessment: Guidelines of the world’s top-ranking universities. Computers and Education Open, 5, 100151. https://doi.org/10.1016/j.caeo.2023.100151
  • Perkins, M., Roe, J., & Furze, L. (2024). The AI Assessment Scale revisited: A framework for educational assessment (Preprint). December 2024. https://arxiv.org/abs/2412.09029

The AI in Education Research Digest

Keep exploring the research with me.

Get my reviews of new studies on AI, teaching, and learning, with thoughtful takeaways for your work as an educator.

Free to join. Unsubscribe anytime. Privacy policy.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top