Picture a college sorting freshmen into calculus or remedial algebra on the strength of one placement test. It looks like a clean, almost mechanical decision. Michael Kane took that exact scenario apart in 1992 and found seven separate assumptions buried inside it, each one carrying the weight of a real student’s path. I came back to this paper after writing about his 2013 expansion of the same ideas, and the older version lands harder in a classroom now crowded with AI.
We make these stacked decisions constantly, and we almost never lay the stack out. A take-home essay scored as evidence of writing ability assumes the student wrote it, assumes the task taps the skill we care about, assumes the score travels to performance in settings the task never sampled. An AI-detection flag driving a misconduct decision carries its own tower of assumptions about what the tool measures and what the number means.
Kane’s whole method is the antidote, and it’s deceptively plain. He argues validity belongs to the interpretation we put on scores, not to the test and not to the scores. That interpretation is an argument, running from the score to the statement or decision we want to make, and it stands or falls on how plausible that argument turns out to be. You validate by spelling out the inferences, naming the rival readings, then gathering evidence for your chain and against the competitors.
The discipline he asks for is the discipline most AI assessment skips. State the claim clearly enough that you can see what you’re assuming. Kane is firm that the buried assumption is the dangerous one. Nobody supports a premise they never noticed they were leaning on, and that’s exactly where score-based decisions fall apart.

Why the Weakest Link Decides Everything
Kane treats these as practical arguments and that difference does real work. Practical arguments live with incomplete, sometimes shaky evidence, so they aim for convincing, not certain. One consequence runs against our instinct to pile up reassuring data. Kane notes that “in practical arguments, redundancy can be a virtue” (p. 528). Several independent lines of weak evidence, combined, can make a single assumption sturdy.
The flip side cuts deeper. A validity argument is only as strong as its shakiest assumption, which means heaping evidence on the parts you already trust buys you almost nothing. He points to a problem here that AI assessment keeps reproducing. We measure what’s easy to measure, an AI score, a similarity percentage, a fluency rating, and treat that clean number as the whole case. Kane would call that supporting the link that was never in doubt while the real weak span goes unexamined.
He leans on Cronbach to make the point that a single observed behavior almost never maps to a single skill. As Cronbach (1971) put it, cited in Kane’s paper, “an item qua item cannot be matched with a single behavioral process” (p. 453). Any clean story about what a task measures rests on collateral assumptions, and those assumptions are precisely what an honest validity argument drags into the open.
What Survives From 1992
A 1992 paper predates the web most of us learned to teach on, never mind generative models. Kane’s examples are placement tests and reading comprehension items, the quiet machinery of a slower era. Some of the surrounding detail reads like a museum piece now.
The framework, though, has barely aged. Kane closes with four properties of these arguments that explain why. They’re built, not found. As he writes, “interpretive arguments are artifacts. They are made, not discovered” (p. 533). They shift as evidence and social priorities move. They bend to fit particular students and odd circumstances. And they’re judged by degree of plausibility, never by a tidy valid-or-invalid stamp.
Read that list against the AI moment and it almost feels written for us. We keep treating assessment validity like a fixed property a tool either has or doesn’t. Kane says we made the interpretation, which means we can remake it, and we’re on the hook to defend whatever we claim. That’s a pedagogical stance before it’s a technical one, and it lands squarely on the thesis I keep returning to across this blog: pedagogy decides whether AI helps or hurts.
His 2013 paper, the one I covered recently, adds the consequences of score use and sharpens the separation between interpreting a score and justifying a decision. The bones, though, were already here in 1992. Specify the claim. Find the weak assumption. Build the case where the case is actually thin.
References
- Kane, M. T. (1992). An argument-based approach to validity. Psychological Bulletin, 112(3), 527–535. https://doi.org/10.1037/0033-2909.112.3.527
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. https://doi.org/10.1111/jedm.12000
