The Real AI Risk in Education Is Cognitive Outsourcing, Not Cheating
Input
Modified
AI raises output while weakening the junior talent pipeline Routine work builds judgment as well as products Firms must preserve human verification and unaided practice

Early-career workers in the jobs most touched by generative AI have already lost ground. Economists at Stanford's Digital Economy Lab tracked millions of payroll records and found a 16 percent relative drop in employment for workers aged 22 to 25 in the most AI-exposed occupations, even as older workers in those same jobs kept their footing. That number did not come from a survey about feelings toward AI. It came from paychecks. It says something concrete is already happening to the bottom rung of the career ladder and it is happening faster than most workplace policy has caught up with. The usual conversation about AI at work stops at productivity. This one has to start further down, at the question of who is still learning anything at all.
The Productivity Story Is Only Half True
Generative AI does raise output and the data on this point is not in dispute. A study from MIT and Harvard, run inside a real customer support operation with more than five thousand agents, found that access to an AI assistant raised novice productivity by roughly one-third. The gain was not spread evenly. Novice and lower-skilled agents improved by roughly a third. The most experienced agents barely moved at all. The tool was doing something specific: it was handing newer workers without necessarily building the same expertise.
That sounds like good news and in a narrow sense it is. But it also reveals what the tool is actually good at. It is good at retrieving and repackaging things that are already known and already written down somewhere. It is much less reliable the moment a task drifts outside that territory. A field study of consultants working with AI found that when a task fell just past the edge of what the tool could actually handle, performance dropped by about 19 percentage points compared with consultants working unaided. The tool did not signal that it had run out of depth. It kept answering in the same confident tone whether it was right or wrong. That gap between fluent and correct is where the real cost hides.

Where the Bill Comes Due
The uncomfortable part of this story is not about today's output. It is about who is being trained to replace the people currently good enough to catch an AI system when it is wrong. Judgment like that is not downloaded. It is built slowly, through years of doing unglamorous work without a shortcut, getting some of it wrong and being corrected. That is exactly the work a capable AI assistant now removes first because it is the most tedious and the easiest to automate. Controlled evidence from learning settings shows the same gap: AI can improve assisted performance without building the independent capability needed once the tool is removed.

The employment numbers among young workers suggest firms are already making this trade, mostly without meaning to. When entry-level hiring quietly shrinks in the occupations AI touches most, while senior roles in the same fields hold steady, the pattern is not really about entry-level workers being unnecessary. It is about firms consuming the routine work that used to double as training, without replacing the training itself. Researchers studying resume and job posting data across tens of thousands of firms have described this same shift as seniority-biased, meaning the disappearing jobs cluster overwhelmingly at the junior end. No single company chooses this outcome on purpose. Each one is simply responding to what looks like a good deal this quarter. The bill for that deal comes due in roughly fifteen years, when the pipeline of people capable of directing AI, rather than just accepting its output turns out to be thinner than anyone planned for.
Why This Is Not Just the Calculator Argument Again
The obvious objection is that people said the same thing about calculators, then search engines, then Wikipedia and none of those tools destroyed anyone's ability to think. There is real truth in that comparison and it deserves a fair hearing rather than a dismissal. But it breaks down in three specific ways once it meets an actual workplace.
A calculator is reliably correct inside its domain. A language model can be confidently and fluently wrong and only a person's own background knowledge will catch it. A calculator also only executes the step a person has already decided on. A language model frequently chooses the framing itself, quietly performing the part of the task that used to require expertise: deciding what question is actually being asked. And the scale of the shortcut is not comparable either. A calculator saves someone a few seconds of arithmetic. A language model can produce an entire memo or analysis in the time it takes to read the prompt, which means the temptation to skip the developmental work is proportionally much larger.
The paper’s behavioral evidence supports this in a more direct way. A study out of MIT's Media Lab had people write essays using either an AI assistant, a search engine or no outside help at all, while monitoring their brain activity. The group writing with AI assistance showed the weakest recall of their own work and many participants struggled afterward to recall or quote from work they had just submitted. This is not a new story about memory the way search engines were. Researchers who study cognitive offloading have long distinguished between handing off busywork to free up mental space and handing off the actual thinking that was the point of the task in the first place. What the MIT findings suggest is that generative AI, used carelessly, sits closer to the second category than most people assume.
What Firms Can Actually Do About It
None of this argues for banning the technology and it should not be read as an argument against using it well. The fix is not a ban and it is not a lecture about personal discipline either, since neither one addresses where the failure actually sits. The failure is structural and organizational. Most firms have adopted a powerful capability without deciding, on purpose, who checks the output before it counts as finished and which categories of work still require someone to do the job unaided as part of their own development.
That distinction has to be made at the level of the task, inside the firm, not by government regulation of the technology itself. Where public policy does have a legitimate and fairly narrow role is in fixing the coordination problem no single company can solve alone. A firm that keeps training junior staff the hard way looks less efficient, in the short run, than a competitor that has cut that cost. That is precisely the kind of shared, slow-building loss that professional licensing standards, accreditation rules and public procurement conditions already exist to correct in other contexts. The same logic applies here.
The jobs data already shows where an unmanaged version of this shift leads. A 16 percent drop in youth employment in the most exposed fields is not an abstract worry about the future. It is a measurement of something happening right now inside a labor market that has not yet built the systems to tell the difference between a worker who checked the machine's work and one who simply signed off on it. That difference is going to matter more, not less, as these tools keep improving. The firms and institutions that decide now who is responsible for verifying AI output and who is still required to build real judgment the slow way are the ones that will have someone capable of directing, verifying and taking responsibility for AI output. The ones that do not are spending down a form of expertise they are not replacing.
This article is based on an original research article published by The SIAI Research. For the original version, please refer to Cognitive Outsourcing in Education: Why AI’s Real Classroom Crisis Is Verification, Not Cheating.
The views expressed in this article are those of the author(s) and do not necessarily reflect the official position of The SIAI or its affiliates.
References
Brynjolfsson, E., Chandar, B., & Chen, R. (2025). Canaries in the coal mine? Six facts about the recent employment effects of artificial intelligence. Stanford Digital Economy Lab.
Brynjolfsson, E., Li, D., & Raymond, L. R. (2023, revised 2023). Generative AI at work (NBER Working Paper No. 31161). National Bureau of Economic Research.
Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X. H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. MIT Media Lab.
Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688.
Sparrow, B., Liu, J., & Wegner, D. M. (2011). Google effects on memory: Cognitive consequences of having information at our fingertips. Science, 333(6043), 776–778.
Yaraghi, N. (2026). Repaying the inheritance: How education and research policy can address AI's borrowed expertise. Brookings Institution.