Studying the Effect of AI Relocates Causality
The experimental premise that treats the tool as the causal agent is the same substitution the data then register.
A growing body of randomized trials and field experiments has examined what happens when students are given access to generative AI during learning. In a high-school mathematics experiment of nearly a thousand students, those using an unrestricted GPT-4 interface scored about 48 percent higher on practice problems while the tool was available, then 17 percent lower on an unassisted exam than peers who had never used it. A guarded tutor that withheld finished answers raised practice scores still further and left the unaided exam statistically indistinguishable from the control. Similar patterns appear in the field: the homework–exam inversion registers homework gains around 18 percent against closed-book exam losses around 20 percent. Meta-analyses sometimes report moderate positive average effects on short-term achievement.
The data register short-term assistance and unaided decline
The data themselves are revealing. They make visible a consistent pattern of short-term assistance followed by thinner independent performance once the tool is withdrawn. Because the studies were conducted under a shared premise, they generate the conclusion that premise already contained. Both the experimental designs and the popular interpretation rest on the same allocation: the AI itself is the causal agent that either helps or harms learning. Access is treated as the independent variable and test performance as the dependent variable. The circulating test of the claim is: if AI helped students learn, a randomly assigned group using it would score better on unaided tests. They do not, in the unrestricted condition. Under that shared assumption the findings become almost self-explanatory. Once individuals believe the tool is what does the work, they reduce their own effort. Cognitive load is off-loaded, struggle is avoided, and the capacities required for independent performance thin. When the tool is later removed, scores fall. The data do not reveal a surprising property of the technology. They reveal the predictable consequence of treating a tool as an agent.
The studies isolate a tool interacting with a posture already adopted
That premise is the deeper error. The studies do not isolate the effect of a tool. They isolate the effect of a tool interacting with whatever habits, intentions, and postures the individual already brings — and those postures are themselves successive choices, not a second nature the instrument then meets. When students treat the AI as a substitute for cognitive effort — asking it for finished answers, letting it carry the load of reasoning, using it to avoid struggle — the tool amplifies that choice. Performance rises while the crutch is present and collapses when it is removed. The apparent failure of the tool is substitution succeeding. Conversely, when an individual uses the same tool as a lever — requesting explanations, testing understanding, deliberately confronting difficulty before consulting it — the identical technology amplifies improvement. The difference is not located in the model weights. It is located in the relationship the person chooses to form with the instrument. The amplification paradox is that geometry as tools densify: the same instrument in different hands produces radically different outcomes because selection stays with the user. Causality stays at the edge that steers is that placement under discovery costume: the model recombines residue; every initiating distinction remains the user’s own act.
The guarded tutor in the same experiment is still this geometry, one step later. Interaction logs showed unrestricted users asking for and copying solutions, and tutor users asking for help or attempting answers themselves. The paper reads the exam difference as a property of guardrails. Guardrails pattern which alignments are available. They do not author the next act. Locating the remedy in a better-designed tool is the original premise retraced as policy.
Substitute and lever are both self-imposed
Both substitution and amplification are therefore self-imposed. The tool never initiates either condition. It receives and magnifies the posture already adopted. The allocation of causal power in validation tracks which side is treated as supplying the next step. Once the answer is treated as arriving from the tool, the student’s own following is no longer the source of continuation. Randomized trials that average across users necessarily average across these opposing postures. The null or negative average result does not prove the tool is inert or toxic. It proves that, in the aggregate, self-imposed substitution outweighed self-imposed amplification. The rational choice postulate is the original sin is that averaging under economic costume: the mean redistributes residue already produced and re-labels it as if it authored the pattern. Here the mean of opposed relations is re-labeled as the effect of AI.
Studying the effect of the tool retraces the same substitution
Focusing research and public discussion on “the effect of AI” further entrenches the original premise. It invites people to locate causality outside themselves, to wait for a better tool rather than examine their own engagement, and thereby reproduces the very pattern the data reveal. Individual choices as the only causal levers is that restore under social costume: externalizing explanations are further discrete acts that relocate the registration of the choice. Here the relocated name is access.
The studies remain valuable precisely because the data are revealing. What they measure is not the power of the technology but the distribution of agency among the people using it. Once the hidden premise is set aside, the practical implication inverts: the decisive variable is never the tool. It is the individual choice, repeatedly made, to treat the tool as amplifier rather than replacement. The unobservable driver of learning is that choice generalized past the instrument: instruction, curriculum, and study design receive the same allocation; observation holds only the downstream scores.