Assessment Validity in the Age of AI

Every few months a fresh panic moves through education circles. Are our assessments still valid now that students can prompt their way to a polished essay? I’ve been in enough of these conversations to notice the question itself rests on a confusion. Michael Kane’s 2013 paper in the Journal of Educational Measurement clears it up, and it does so years before anyone outside a research lab had touched a generative mode

Kane’s central move is deceptively simple. We don’t validate a test, and we don’t even validate the scores. We validate a proposed interpretation or use of those scores. The same exam can support a modest claim, a wild one, or a dozen in between, and each one has to earn its own evidence. Once you internalize that, the whole “is AI breaking assessment” debate starts to look misframed.

Why Assessment Validity Was Never a Property of the Test

The old instinct treats validity like a stamp on the instrument. A test is valid, or it isn’t. Kane dismantles that. He argues a test score becomes meaningful only when we attach a claim to it, and the claim is what gets evaluated. An essay score might support the claim that a student can organize an argument on a familiar topic under timed conditions. Stretch it into a claim about their independent reasoning across any context, and you’ve taken on a far heavier evidential burden.

Kane builds this into two linked arguments. The first, the interpretation/use argument, lays out every inference needed to travel from what a student actually did to whatever conclusion we want to draw. The second, the validity argument, evaluates how plausible that chain turns out to be.

He gives the structure a memorable image: the inferences “can be envisioned as the spans of a bridge leading from the test performances to the conclusions and decisions included in the proposed interpretation and use; if one span falls, the bridge is out, even if the other spans are strongly supported” (p. 13).

That bridge metaphor is the whole game. A pristine rubric and flawless scoring mean nothing if the extrapolation to real-world skill collapses. And this is exactly where AI lands its punch. When a student uses a model to draft, the scoring span can stay perfectly intact while the extrapolation span silently buckles. The score still measures the artifact. It no longer tells you what you assumed it told you about the person.

The Ambition Tax on Every Claim

Kane’s second principle is one I wish more assessment designers carried around. The bigger your claim, the more evidence you owe. He puts it with unusual bite: “If scores are to do more work, we have to pay them more by supplying the evidence needed to support the additional claims being made” (p. 34).

I read that as a discipline, not a limitation. Most assessment trouble in the AI era comes from claims we never paid for. A teacher gives a take-home essay and treats the score as evidence of independent writing ability, a construct the task was never designed to isolate. Kane would call that begging the question. The fix isn’t a better plagiarism detector. The fix is matching the claim to what the task can actually support, which lands squarely on his thesis and mine: pedagogy decides whether the assessment holds.

This is where Kane’s framework connects to the current wave of work on assessment redesign. Corbin et al. (2025a/2025b/2025c/2025d; Dawson et al., 2024, and Bearman et al. (2023; 2024), and more recently Roe and colleagues with their notion of assessment twins, are circling the same question from the GenAI side. Kane gives that conversation its older spine. He hands us the vocabulary to say precisely which inference AI threatens, and the bridge tells us why fixing one span doesn’t save the others.

assessment validity

Score Uses Need Their Own Justification

Even a perfectly valid interpretation does not license any particular use of the score. Kane is firm that the two require separate cases. A valid measure of math achievement does not, by itself, justify using that score to track a child, gate a diploma, or rank a teacher. Each use has to be evaluated on its own consequences.

For anyone watching AI accountability schemes bloom, this should land hard. Schools are rushing to build decisions on top of AI-flagged work, AI-scored writing, and detection probabilities. Kane’s point is that validating the underlying signal, even if you could, says nothing about whether the decision built on it is defensible. The consequences carry their own weight, and negative ones can sink a use no matter how clean the measurement looks.

He frames the appraisal stage as adversarial by design, and quotes Cronbach to sharpen it. As Cronbach (1980) put it, cited in Kane’s paper, “the job of validation is not to support an interpretation, but to find out what might be wrong with it. A proposition deserves some degree of trust only when it has survived serious attempts to falsify it” (p. 103). That’s the attitude missing from most AI assessment policy I read. We defend our tools. We rarely try to break our own claims about them.

How I Read Kane in 2026

A 2013 psychometrics paper covering licensing exams and standardized tests can feel a generation removed from a classroom where students argue with chatbots. The temporal gap is real, and Kane never imagined the mess we’re in. Some of his examples now read like artifacts from a quieter era.

The framework, though, has aged better than almost anything written about AI assessment since. Kane gives us a way to stop asking the unanswerable question, “is this still valid,” and start asking the useful ones. What claim am I making about this score? Which inference does AI actually threaten? Have I paid for the claim, or am I begging it? Is the use justified on its own terms?

References

  • Bearman, M., Nieminen, J. H., & Ajjawi, R. (2023). Designing assessment in a digital world: An organising framework. _Assessment & Evaluation in Higher Education_, 48(3), 291-304. https://doi.org/10.1080/02602938.2022.2069674 .
  • Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. _Assessment & Evaluation in Higher Education, 49_(6), 893-905. https://doi.org/10.1080/02602938.2024.2335321
  • Corbin, T., Bearman, M., Boud, D., & Dawson, P. (2025a). The wicked problem of AI and assessment. Assessment & Evaluation in Higher Education. 1–17. https://doi.org/10.1080/02602938.2025.2553340
  • Corbin, T., Dawson, P., Nicola-Richmond, K., & Partridge, H. (2025b). ‘Where’s the line? It’s an absurd line’: Towards a framework for acceptable uses of AI in assessment. Assessment & Evaluation in Higher Education, 50(5), 705-717. https://doi.org/10.1080/02602938.2025.2456207
  • Corbin, T., Dawson, P., & Liu, D. (2025c). Talk is cheap: Why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 50(7), 1087–1097. https://doi.org/10.1080/02602938.2025.2503964
  • Corbin, T., Tai, J., & Flenady, G. (2025d). Understanding the place and value of GenAI feedback: A recognition-based framework. Assessment & Evaluation in Higher Education, 50(5), 718–731. https://doi.org/10.1080/02602938.2025.2459641
  • Dawson, P., Bearman, M., Dollinger, M., & Boud, D. (2024). Validity matters more than cheating. _Assessment & Evaluation in Higher Education_, 49(7), 1005–1016. https://doi.org/10.1080/02602938.2024.2386662 
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
  • Roe, J., Perkins, M., & Giray, L. (2025). Assessment twins: A protocol for AI-vulnerable summative assessment. arXiv preprint arXiv:2510.02929.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top