The Homework–Exam Inversion Registers Substitution
The tool amplifies capability only at the moment of use; whether raw capability then improves is the relation to the tool, not the tool.
Homework scores once predicted exam performance. A large panel of more than 26,000 Chinese secondary-school students found that generative-AI adoption raised average homework scores by about 18 percent and cut completion time by roughly 30 percent, while closed-book monthly exam scores fell by around 20 percent within six months and high-stakes entrance-exam scores declined by 18 to 24 percent over a longer period. The losses concentrated among approximately 80 percent of AI users whose homework behavior — high scores paired with unusually short completion times — pointed to heavy outsourcing of effort. The inversion, and a with-AI exam curve sitting strictly below a without-AI curve, is presented as the tool eroding ability. It is an artifact of the measurement: one score is raw capability plus AI, the other is raw capability only.
The circulating graph locates the cause in the tool
The empirical pattern is real. Homework scores rose. Completion time fell. Unaided exam scores fell, and the fall was largest among those whose homework times revealed substitution. The graph of that inversion is a useful instrument. It densifies comparison between two registers that used to move together.
The freeze is treating the inversion as an effect the tool performs on the student. Mistaking the expression for the intelligence is that freeze when a score is held as the activity. Here the homework score becomes an expression of student-plus-tool while the exam remains an expression of the student alone. The two instruments no longer share a subject. The gap is an artifact of that mismatch. Reading it as what AI did to capability relocates initiation into the medium.
Tools do nothing automatically
Tools, including generative AI, do nothing automatically to any person. Their effect on performance is determined entirely by how the individual relates to them. The decisive variable is not the technology’s presence but the user’s stance toward effort, output, and personal agency.
As with all tools, the instrument amplifies the difference already present in how individuals relate to effort. The majority use it to do less themselves. Those who keep agency at the edge use it to do more. The tool does not make people less capable. The less one’s capability, the more motivation to use AI to get a better score when that is allowed. Those already disposed to do less are more likely to use the tool to do less. Causality stays at the edge that steers is that placement under discovery costume: the model recombines residue; every initiating distinction remains the user’s own act.
Substitution attributes results to the tool
When a person attributes results primarily to the tool, they treat it as a substitute. In that mode the tool produces surface results with reduced personal investment, and independent capability declines. Skills that are no longer exercised thin. Most users follow this path because the majority prefer to accomplish tasks with the least possible effort. The study’s data match the pattern exactly. The largest learning losses appear among those whose homework behavior reveals substitution.
The allocation of causal power in validation tracks which side is treated as supplying the next step. Once the answer is treated as arriving from the tool, the student’s own following is no longer the source of continuation. Homework still completes. The update that would have been the work does not.
Amplification keeps results at the edge
The tool raises performance at the moment of use. Whether that increment becomes a reason to improve raw capability or a reason not to is the relation the individual holds. When a person continues to attribute outcomes to their own judgment and effort, and treats the tool as an amplifier of that effort, independent capability can be retained and even strengthened. The individual uses the technology to reach further, verify more carefully, or explore problems that would otherwise remain out of reach. Amplification is real, yet it is neither automatic nor equal. It depends on the user’s preexisting ability and, more importantly, on the attitude they bring to the interaction.
The amplification paradox is that geometry as tools densify: the same instrument in different hands produces radically different outcomes, because selection and sensing stay with the user. The homework–exam split is the school face of the same widening.
Generating the solution is not helping to figure it out
A still deeper and largely overlooked dimension of this relationship appears in the concrete way students use AI for homework. There is a fundamental difference between asking the AI to generate the finished solution and asking it to help the student figure out the solution. In the first case the student receives an answer and typically copies or lightly edits it; cognitive work is largely bypassed and little lasting capability is built. In the second case the AI functions as a guide — offering hints, explanations, counter-examples, or feedback on the student’s own attempts — while the student remains responsible for constructing the solution. The two practices embody entirely different relationships with the tool. One replaces the student’s thinking; the other supports and extends it. Neither the circulating graph nor most public discussion of the study distinguishes between these modes of use, yet the distinction is central to whether independent capability erodes or grows.
Token efficiency, emulation, and the unclosable gap is that cut under compression costume: the shortcut that arrives after the path is walked is efficiency; the shortcut that skips the path is a different act under the same name. Asking for the finished solution is the skip. Asking for help in figuring it out keeps the inefficient path in the sequence from which later compression can be the effect.
Closed-book exams register substitution as decline
This mechanism also explains why current educational measurements produce the striking inversion. School assessments, especially high-stakes closed-book examinations, continue to evaluate unaided performance. In a system that still rewards independent problem-solving, the substitution pattern registers as clear decline. That decline is real under existing rules. It does not demonstrate that AI universally diminishes the capacity to solve problems when the tool is available. The difficulty is that many students are not developing skilled collaborative use; they are developing dependence by allowing the AI to generate solutions rather than helping them arrive at solutions themselves.
Closed reality in benchmark maxing is the sealed-field face: a scoreboard measures what it was drawn to measure. The closed-book exam is such a hold. Under that hold, substitution is decline. Treating the hold as the whole of capability — including capability exercised with the tool — is a further freeze. The exam remains a valid instrument for unaided performance. Exhaustiveness does not travel with it.
The inversion is an artifact of the measurement
If one score includes raw capability plus AI and the other is raw capability only, one will always sit above the other. That is not a further fact about the tool. It is the composition of the two instruments. Homework for an AI user is the raw term plus the tool’s contribution. The exam is the raw term alone. Compare those two numbers and the homework number sits above.
Using the tool when the homework score is allowed to include it is rational. There is no reason to use a tool that does not raise the allowed score above raw capability. If it did not, the AI would be useless, and the use would not make sense. Take the homework number — AI allowed — as the reference, then test without AI, and the result is lower than they would score with the tool. That is what a useful instrument looks like under that rule. The graph reads the usefulness as decline.
The graph is more misleading still once the homework score is held fixed. For the student who did not use AI, that number is already raw capability. For the student who did, that number is raw plus the tool. The exam then measures raw only. Whoever did not use AI will test better than whoever did, at the same homework score. The comparison is guaranteed by the measurement.
Another factor stacks the same comparison. The lower the raw capability, the larger the gap between that remainder and the homework score the tool can supply, and the stronger the motivation to use AI when it is allowed. Selection into the AI-user group is already toward lower raw. The graph then reads that selection as the tool’s effect.
The with-AI exam curve that sits strictly below the without-AI curve is the raw term, averaged. Individuals who use the tool to amplify their own capability exist. They are a minority among AI users. Their raw capability is averaged out by the rest of the group, whose homework behavior shows substitution.
People who never had a chance to take modern transportation would on average be more capable of walking than those who do. The walking difference is real. It is not a measure of mobility. Thank God we do not actually only walk around.
Whatever is one prompt away is that freeze when the old surface is scored as the whole of the field: walking remains; its share of reach shrinks.
Locating the cause in the tool trains the same substitution
Presenting the data primarily as evidence that AI harms capability allocates causality to the tool. Studying the effect of AI relocates causality is that allocation already at the experimental premise: access treated as the independent variable, the average of opposed postures then named as a property of the technology. That allocation is the substitution pattern speaking as explanation. By locating the problem in the technology, such presentations encourage the belief that the tool is the decisive agent. In doing so they obscure the more consequential question of how individuals choose to relate to the tool — whether they ask it to replace their effort or to support the exercise of that effort. A reaction that relocates the cause into the instrument cannot address the stance that produced the pattern.
Individual choices as the only causal levers is that restore under social costume: externalizing explanations are further discrete acts that relocate the registration of the choice. Here the relocated name is AI.
The tool amplifies capability only at the moment of use. Whether that increment is used as a reason to improve raw capability, or as a reason not to, depends on the individual’s choice — the relation to the tool, not the tool. Generating the finished solution and asking for help in figuring it out are those two relations. The empirical pattern revealed by the homework–exam inversion is real. What it reveals is the diversification of individual choice in relating to the tool, and of raw capability — not diversification by the tool. The interpretation that treats the instrument as the causal agent relocates the choice. The unobservable driver of learning is that relocation generalized: the look holds only the scores; the driver is where causal power is successively located.