Fragile knowledge forms when AI teaching skips real questioning
Bayesian framing explains why AI struggles with exceptions
Paywalled science literature makes AI weakest on scientific questions
In the fall announcements, the Alpha School network of private schools announced an expansion to about fifty campuses across the United States. Twenty-seven of these are new locations, with tuition ranging from forty to seventy-five thousand dollars a year. The company's founder speaks openly about the ambition for the method to reach one billion children, in a country where one in five students is considered chronically absent from school. Most of the teaching is assigned to software that adapts to each individual student. Those who studied the approach, however, identified an issue that goes beyond grading. They called it fragile knowledge, a skill that works flawlessly within the conditions in which it was taught yet collapses as soon as circumstances change.
An Experiment on the Scale of One Billion Children
Alpha School itself does not deny that this is an experiment. The process of training the models is openly compared by the company's managers to the training of autonomous vehicles. There, the system learns from millions of incidents to recognize the exception, the cyclist who sways, the child who pops up on the road. The comparison is useful, because it clearly shows its limit. An autonomous car is trained on traffic data that is finite in kind, no matter how large it is in volume. But learning is not just about recognizing patterns. It involves transferring an idea into a context that has never been seen before, something that no database fully covers.
Figure 1: A correct answer on a familiar question proves the step was climbed, not that the climb continues on its own.
The concern is not just about teaching with software. In a randomized study, one group of students in an introductory physics class used an adaptive digital assistant, while another followed active teaching in the classroom. The first group had twice the average improvement. But the same experiment also highlighted the risk of fragile knowledge, because the assessment was done under the same conditions as teaching. In other words, it was not tested whether what was learned held up outside of that particular form of question. A later review of twenty-eight studies, with almost five thousand elementary and middle school students, found a generally positive effect from intelligent adaptive systems. But this advantage almost disappeared when the systems were compared to simple active learning without any artificial intelligence. The finding deserves attention because it shows that the problem is not necessarily technological.
Figure 2: Scale can move a tool along this line faster than it can move it toward the questioning end.
The Fragile Knowledge Left by Memorization
The working hypothesis here is older than artificial intelligence. The value of a good teacher versus a mediocre one, even with exactly the same material, lies in a second layer of learning. This layer is created within the social interaction of teacher and student. In the Western tradition, this interaction often takes a Socratic form, that is, constant questions that force the student to think again and adapt his logic to new conditions. When this process is lacking, the alternative is not emptiness but a different model: memorization without questions. There the material is mastered accurately but without adaptability. The East Asian educational tradition has often been described in these terms, as a system that produces excellent performance in standardized exams, but difficulty in the face of unusual questions.
Today's AI models replicate exactly this second model. They don't do it because they're mimicking a particular educational tradition but because their own architecture works in a similar way. They're trained on a huge amount of question-and-answer pairs, where accuracy is optimized. When a real exception appears, something never seen before in this form, performance breaks. It breaks in a way similar to a student who learned something by heart without ever being asked why. Similarity is not a coincidence. In both cases, the same moment is missing, the one where someone forces the system, human or computational, to justify its answer in other words, in a different context.
What a Bayesian Update Misses with Limited Data
There is also a more technical way to formulate the same problem. What an AI system does when it adapts to a student is, to a large extent, a form of Bayesian updating on limited information. This information is limited to the input and output pair of the specific dataset on which the model was trained. The human mind is described by Harvard research as something that works in part in the same way but is in many ways better. It can identify the exception to a normality that until then seemed constant and revise the entire mental model instead of simply adjusting it gradually. A good teacher does just that when he asks a question that is not in the textbook. This question is designed to test whether understanding survives outside the context in which it was taught.
The concern is not theoretical. It has already been observed that professionals with many years of experience lose some of their ability to think independently when the cognitive load is consistently assigned to AI tools. This phenomenon has recently been documented in a university setting. The same pattern can easily be transferred to a child learning with a digital assistant. Instead of the teacher confirming the understanding with a series of different questions, the child can simply get the correct answer and move on. He will not necessarily be able to apply the same concept elsewhere again to really broaden the scope of his knowledge.
Why Science Breaks AI Models
There is an additional reason why this weakness becomes more pronounced in scientific subjects. Large language models are mainly trained on data that is freely available on the internet. But most of the reliable scientific literature remains locked behind journal subscriptions. An artificial intelligence tool was recently designed by a research organization in New Zealand with an explicit mandate to draw exclusively from peer-reviewed articles. It ended up returning mostly secondary content from blogs and general websites, precisely because reliable literature remained inaccessible. The problem was not a technical error; it was a structural one. The system simply did not have access to the source that would allow it to learn properly.
This explains why trust in such a system should not be uniform. It is more like a scale that changes depending on the field. In areas where correct answers change a bit and where everyday use does not require access to the cutting edge of research, such a tool can perform reliably, at little cost in the event of a failure. In science, however, a slight discrepancy in data or methodology changes the conclusion and the same system breaks much faster. A question in a field where failure is expected yields little compared to the effort it requires. An expert's judgment is still based on sources that the model has simply never encountered. This radically changes how much weight is worth giving to his answer.
The Design Choice Schools and Teachers Face
The argument in favor of models such as that of the Alpha School is not based on the assumption that they outperform the best possible teacher. It is based on the observation that the quality of teaching differs dramatically even within the same school, to an extent that would not be tolerated in other high-responsibility professions. An adaptive system that is simply better than average could thus replace thousands of inadequate teachers and this would indeed be a reversal in the field. But the argument is weakened by the same finding mentioned above: the advantage of artificial intelligence systems over simple active teaching has proven negligible. If this is confirmed on a larger scale, the question is no longer whether it is worth replacing the mediocre teacher but what exactly it is being replaced with.
School administrators and educational technology companies have a specific design question here, not only ethical but also technical. The evaluation of right and wrong answers scales relatively easily. The Socratic moment, the one where one asks again in different words to see if understanding endures, is much more difficult to scale. Perhaps it is also structurally impossible within the same architecture that produces fragile knowledge in the first place. Alpha School is now testing its model in public schools in Houston and Springfield, Massachusetts, with a student population much more heterogeneous than its private sector. The question is no longer whether good teaching can be scaled up. It is whether it can be scaled up without losing exactly the element that made it good.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.
[AI Labor vs. Human Labor] Jevons Paradox and the Hidden Costs of AI Substitution
Picture
Member for
1 year 10 months
Real name
Keith Lee
Bio
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.
His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.
He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.
Published
Modified
AI substitution shifts costs into managerial review and control
Jevons effects can expand demand while raising oversight burdens
Accountability remains human even as model reliability improves
For quite some time, the replacement of human labor by AI has been treated as a pure problem of comparative advantage. James does one task, Jones does another and production is shared where everyone has a lower relative cost. Now Jones can be a model. If AI becomes more cost-effective, labor moves to AI. The logic is correct but it is incomplete inside a real company. There is an additional cost that does not easily show up in the model's invoice. Someone has to understand what the AI did, decide if the result can be used and take responsibility when that result becomes a business act. This human weight changes the comparison and limits the Jevons paradox before mass substitution even begins.
Comparative Advantage Extends Beyond the Model Price
The basic idea remains strong. A profession's exposure to AI does not tell by itself whether AI will be used. In German data from 9,835 employees, a model that primarily looks at technical exposure explains about 25% of the difference in adoption. When task-level user costs and worker productivity relative to pay are incorporated, the explanation rises to about 60%. This fits with the way a business should think. It doesn't buy technical competence for the sake of technical competence. It buys output at a certain cost. If AI needs a lot of integration, control, process change and special management, its relative position gets worse even when the model looks impressive in a benchmark.
Some costs are obvious. The cost of use is usually measured around the system itself. License, compute, integration, security, monitoring. But within the company there is also a cognitive cost. The person who approves a decision has to mentally re-enter the work to understand it. If a model writes a financial analysis, the manager who signs cannot just see that the text is well written. The assumptions, data, exceptions and consequences of a wrong hypothesis still have to be understood. For routine work this may be small. For a serious decision, it can become much larger, and rather quickly. The work seems to have left the human but part of it returns in a more condensed form to the person who has the right to approve.
The difference between use and replacement can also be seen in the adoption data. At the end of 2025, about 18% of U.S. businesses said they had adopted AI in a business function, while about 41% of employees said they were using generative AI at work. The metrics aren't directly comparable but they do show that tools can spread across companies long before the number of jobs changes. This is another sign that measurement needs to get to the point where using the model actually changes the cost of an accomplished task. Until then, much of AI functions as a supplementary asset around human decisions, rather than a standalone replacement.
Jevons Paradox and the Managerial Review Bottleneck
The Jevons paradox is attractive to AI because it gives an optimistic answer to the simple story of replacement. If AI makes a service cheaper, the demand for the service may increase. Call centers in the Philippines are a useful example. Employment continued to grow until 2025 and approached two million, even though customer service is considered one of the most exposed categories of work. Technology can reduce the cost per contact and make more contacts economically feasible. In this case, higher productivity doesn't have to mean less total work. It can mean a bigger market and this is a real possibility and should not be lost in predictions of mass layoffs.
There is a difference between coal and corporate decision-making. If a steam engine consumes coal more efficiently, no executive needs to rethink every molecule of energy before approving it. With AI, every increase in production can create a new audit queue. More drafts, more proposals, more decisions, more code and more automated actions can reach humans who have to approve them. If the output of AI increases tenfold, the human control system cannot always increase tenfold. There comes a paradox within the Jevons paradox itself. Technology makes generating results cheaper but it can make the attention of the responsible executive a rarer and more expensive resource.
Figure 1: AI can reduce execution costs while shifting work into review and decision ownership.
The productivity of call centers shows how easily these two results are confused. In a large customer service study covering 5,172 agents, a generative AI tool increased productivity by about 15% on average and helped less experienced workers the most. This can reduce the hours it takes for a consistent volume of calls. It can also make service cheaper and increase volume. The second effect looks like Jevons paradox. Nevertheless, the more cases go through the system, the more difficult exceptions eventually reach people. AI removes the ordinary piece and can leave the worker with a smaller but more challenging set of cases. Average productivity goes up, while the cognitive intensity of human work can also go up.
The Cognitive Re-insourcing of Work
Cognitive re-insourcing is one way to describe what happens. A task is given to AI but the critical part of it comes back to the human when the time comes for acceptance. The manager doesn't do all the work again. Enough of the reasoning still has to be reconstructed to know what is being approved. This is more demanding than a simple spelling or formatting check. In a legal, financial, technical, or regulatory decision, the value often lies in the exceptions. The model can produce routine work very quickly, while much of the risk remains concentrated in the exceptions. The business has then reduced production time but concentrated responsibility on a smaller group of people who have to identify the difficult cases.
This explains why replacement costs may seem lower on a spreadsheet than they are in practice. The model price is easy to see. Managerial time is not, it gets scattered across reviews, corrections, meetings and small decisions that rarely appear as one clear cost. The cost of a wrong decision is even harder to put into a simple average, because many times it is low-frequency and high-damage. It can involve legal liability, clients, reputation, regulatory compliance, or loss of institutional knowledge. This is why true comparative advantage must also measure the distribution of risk. A system can be cheaper per task and more expensive per decision that can stand without additional human remediation.
If a company removes the people who owned the process too quickly, it also loses the benchmark with which it controls the system. Institutional knowledge is not always found in manuals or databases. It is found in small exceptions, customer relationships, old decisions and informal rules that employees have learned over time. The more AI takes over the normal flow, the easier it is to underestimate this knowledge because it only appears when something goes wrong. Then the business discovers that it has saved the cost of the performer but has lost some of the ability to judge the performer. These costs occur in practice and are part of the actual price of substitution.
Figure 2: Routine production can shift to AI while judgment and accountability remain human.
Responsibility Remains Human
The most interesting test is to remove hallucination from the problem. Even if AI reaches a near-zero percentage of factual errors, the finances of many tasks would change drastically. This does not eliminate responsibility. A model can give a correct prediction and the business can make the wrong decision because it chose the wrong goal, the wrong risk threshold, or the wrong use of the prediction. It can also make a choice that is technically correct and business-damaging because it ignored a customer, regulation, or consequence that wasn't in the prompt. The person who decided to delegate the task to AI remains responsible for explaining why the decision was acceptable.
This is where the logic of James and Jones changes again because if Jones is human, then can share not only production but also judgment, memory and responsibility. If Jones is AI, production can be transferred without accountability being transferred in the same way. This makes James seemingly more productive and at the same time more burdened with decisions. The next phase of AI will therefore be judged by something more difficult than cost per token or success rate in benchmarks. It will be judged by how cheaply a company can turn AI output into a decision that a human can actually defend. If these costs fall, comparative advantage will move quickly and the Jevons paradox may become stronger. If they remain high, human labor will remain within the system, not always as a producer but as the carrier that absorbs the risk.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.
Picture
Member for
1 year 10 months
Real name
Keith Lee
Bio
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.
His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.
He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.
His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.
He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.
Published
Modified
AI detection software failed; universities scrap it for coaching
Same policy mistakes as plagiarism era are repeating, faster
Process-based evaluation beats unreliable detection percentages
Less than one in four American universities has a formal written policy on the use of artificial intelligence by its students. The number comes from recent research that shows how slowly the institutional apparatus is moving in the face of a technology that has already become a daily habit in classrooms. At the University of Nevada at Reno, the administration tried the usual answer: artificial intelligence detection software, built into the same system it used to check for plagiarism. The result was no less use of artificial intelligence. It was more stress, more unwarranted accusations, and no improvement in learning. The university eventually abolished the tool entirely. This story is not an isolated case, but one indication that AI detection, as a strategy is failing as an institutional strategy.
Universities Are Repeating the Plagiarism Playbook
The history of plagiarism offers a cautionary precedent for what is happening now. For a decade, plagiarism policies were based on three assumptions that were proven wrong: that copying was a moral failure and not a developmental stage, that a uniform standard of acceptable help could be applied fairly to students with very different resources, and that software could rule on someone's guilt. EFL instructor Özgür Çelik describes how students who wrote in a second language were reported for academic misconduct at a rate two to three times higher than native speakers, even though actual misconduct rates did not differ significantly. The problem did not lie in their behavior. It lay in who attracted the suspicion of invigilators.
Çelik doesn't stop at diagnosis. She proposes four lessons drawn from a decade of plagiarism policy: explicitly state what exactly is evaluated in each paper, test every rule against a student who writes in a second language rather than an imaginary native speaker, never allow a percentage probability from a detector to act as standalone evidence, and evaluate the writing process itself through drafts and revisions, not just the final text. These proposals are not just for second language students. They outline a more general principle of policy design that universities ignored once and are in danger of ignoring again.
Universities are now writing AI policies in the same way, only faster and on a larger scale. Researcher Michael Zyphur examined the public policies of thirty-eight leading doctoral universities in fifteen countries, along with fourteen funding bodies and eighteen publishing houses, and found that only six institutions had gone beyond integrity and disclosure of use to real-world AI usage knowledge, supervision, and authoritative research practice. Fifteen institutions remained at the level of a general norm of student conduct. Seventeen had advanced integrity and reproducibility issues, but without codified competency expectations for researchers. The finding confirms the same pattern already recorded in plagiarism: the rules focus on the surface of the text rather than on the judgment that produced it, and the very structure of the problem is repeated from teaching to research. Zyphur’s framework helps clarify what most institutional policies still miss: responsible AI use in research is not a single rule, but a set of distinct modes, such as search, co-authoring, validation, and tutoring, each requiring its own form of human judgment and oversight.
Figure 1: Zyphur’s framework distinguishes search, co-authoring, validator, and tutor functions, showing that each mode requires a different form of human verification and oversight.
Why AI Detectors Fail as Evidence
The technical weakness of AI detectors is no longer in question. A test conducted in 2023 on a newer generation of detectors found that no tool exceeded eighty percent accuracy, while several characterized human text as a machine product, and vice versa. The same team of researchers had already tested, three years earlier, the corresponding plagiarism detection tools in a multinational evaluation and had found huge discrepancies in what each detected; none could responsibly substitute human judgment. The tools functioned as aids, not juries, but institutions used them as if they were the latter. Stanford University researchers found something even more troubling: GPT detectors incorrectly classified four lessons drawn from a decade of plagiarism policy written by non-native English speakers as AI products, while judging native speakers' essays almost flawlessly. The limited lexical richness and predictable syntax of a non-native speaker were interpreted by the detector as a sign of machine output, not a feature of language learning.
At the University of Nevada, Executive Vice President and Provost Jeffrey Thompson describes how this unreliability turned the tool into a source of fear instead of a teaching aid. Students felt anxious about being wrongly accused, while professors found themselves caught between two dangers: if they didn't use the software, they risked ignoring actual AI use; if they did, they risked punishing innocent students. The institution had already moved from the era of simple plagiarism detection software to an era where the line between human and machine writing had become blurred, and the old tool could no longer cope with the new reality.
As part of a broader initiative to integrate AI into teaching and research, the university ultimately chose to ditch the detection tool and partner with a writing guidance platform, which accompanies the student throughout the process rather than just checking them at the end. The change didn't mean disregarding integrity. It meant acknowledging that policing a text doesn't reveal anything about how the thinking behind it was produced. Thompson cites a recent survey showing that less than a quarter of U.S. institutions of higher education have a formal policy on the use of AI, which explains why so many institutions still rely on detection tools despite their documented problems.
From AI Policing to Cognitive In-Sourcing
The alternative is not unlimited use of AI without any conditions. It is instructional design that holds cognitive effort in the hands of the student. Neuroscientist Adam Green, from Georgetown University, and clinical neuropsychologist Jared Benge, from the University of Texas at Austin's Dell Medical School, describe AI as a tool that can free up mental space or empty it completely, depending on how it is used. Both recommend a simple step before any question in a conversational system: first record one's own thought, and only then ask the AI to improve or challenge it.
A survey of three hundred and nineteen knowledge workers, with nine hundred and thirty-six recorded incidents of generative AI use, found that critical thinking was activated, according to the participants' own reports, in about six out of ten incidents. The finding that matters most concerns the relationship with trust: the more confident someone felt about the AI's ability to complete the task, the less critical thinking they reported engaging in. Trust in personal judgment was instead associated with more, not less, critical engagement. The conclusion is not that AI automatically harms every user's thinking. It's that the degree of trust in the tool, not its very existence, ultimately determines whether someone thinks less.
Green himself predicts that the habit of "thinking outside of robots" will gradually become a natural survival strategy in a world where text production is no longer evidence of human thought. This observation is of direct relevance to lesson design: if the value of a human idea lies in its specificity, then a task that only rewards the fluency of the final text rewards precisely the characteristic that artificial intelligence reproduces best. Teaching that asks the student to defend a personal choice, even an imperfect one, trains the same skill that the job market will reward.
Redesign Assessment Around Process and Judgment
For teachers, change means replacing the question of who wrote a text with asking what the student can explain about their own work. Exercises that require drafts, intermediate texts, and verbal defense of choices reveal judgment in a way that a final text can no longer reveal. For administrators, it means investing in guidance tools instead of detection tools, as well as training the professors themselves, who cannot teach disciplines that the educational environment itself does not demonstrate every day. For policymakers, it means making an explicit distinction between tasks that AI can accelerate and judgments that must remain human, as suggested by the five-dimensional framework developed by Zyphur for research education. The same framework suggests something that is often omitted from student discussions: providing tools with criteria so that the institution can check whether an AI system offers verifiable citations, clear uncertainty information, and auditable usage trails, rather than leaving tool selection solely to the discretion of each individual professor or student.
Figure 2: Only six of the thirty-eight doctoral universities studied had progressed through all five dimensions; fifteen remained at a generic conduct-rule stage.
One expected objection is that abolishing detection will pave the way for rampant fraud within classrooms. The evidence does not support this fear. Detection tools did not fundamentally prevent dishonest use; they simply shifted the burden to the wrong targets, more often punishing students with different language backgrounds than their fellow students. Evaluating the process, rather than just the final product, offers a stronger proof of integrity than any probability percentage an algorithm produces. The same conclusion emerges from Zyphur's research: the reproducibility and transparency of the process, not the prohibition of a useful tool, ultimately protect academic credibility.
A second objection concerns scale: many institutions teach thousands of students per semester and do not have the staff to read drafts, intermediate texts and oral explanations for each assignment. This objection has a factual basis, but ignores the cost of the alternative. Staff who currently manage miscategories, student appeals and institutional investigations around unreliable probability rates could be channeled into course planning and targeted feedback. Shifting resources, not increasing them, is what the system needs.
The one in four universities with a formal AI policy isn't just a statistical gap. It's an indication that the industry is still wondering if it should ban something that's already embedded in the daily work of its students. The University of Nevada showed what happens when an administration admits that the detector didn't protect anything substantial: less fear, more actual teaching, and a relationship of trust between faculty and students that the software had already silently eroded. Institutions that still invest in AI detection are investing in a tool that research itself has already discredited with evidence. The next generation of policies needs less reliance on detectors and more teachers ready to ask the student how he thought, not which tool he opened on his computer screen. Those institutions that move in this direction first will not have to rewrite their policy in a few years, when the next generation of tools makes today's detectors even more obsolete than they already are today.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.
Picture
Member for
1 year 10 months
Real name
Keith Lee
Bio
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.
His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.
He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.
Anthropic data shows AI speedups favor long, high-value tasks
Bottleneck tasks limit uniform gains even in AI-heavy jobs
Closing the gap needs domain judgment training, not prompting
Among the task-level AI usage figures now in circulation, one number stands out from the rest: in a hundred thousand real conversations from a large AI assistant, the typical task was completed 84% faster with the help of AI than without it. This is not a response to a survey or an optimistic forecast. It's a direct assessment of what actually happened when someone used the tool for real work. The average job in this sample would cost about fifty-four dollars in business time if it was done without assistance. Such numbers explain why AI productivity channels attract so much attention. What they do not explain on their own is why these channels are opening up so much wider for some workers, jobs and countries than for others.
Where Productivity Channels Are Concentrated
Software development tops this list by a wide margin, contributing an estimated 19% of the total productivity benefit attributable to current AI usage, according to a detailed record of U.S. labor productivity. It is followed by general administration and management with about 6%, then market research and marketing with 5%, customer service with 4% and secondary education with 3%. These are not the professions with the most employees. They are the professions where AI happens to be used for tasks large and complex enough to have weight when the time savings at the economy level add up.
This last point deserves more attention than it usually gets. A task of developing educational materials that would take a teacher four and a half hours by hand can be completed in about eleven minutes with the help of AI, saving about one hundred and fifteen dollars worth of teacher salaries. Financial analysts save about 80% of the time on tasks like interpreting financial data, worth about thirty-one dollars each. Food preparation and installation work also saves time, but the tasks themselves are shorter and cheaper in the first place, so the dollar value recovered is less even when the savings rate looks similar. Productivity channels, in other words, don't just depend on how quickly a task is completed. They depend equally on how great and how valuable this work was in the first place.
Figure 1: The same speedup percentage recovers very different dollar value depending on how long and how well-paid the task is.
The administration and the legal profession are near the top on both scales at the same time. The average administrative work that humans bring to AI, such as selecting investments or reviewing a contract, would take a professional about two hours without assistance and legal jobs are approaching that average. These are also some of the highest-paid hours in the economy, so even a modest percentage acceleration in them recovers more measurable value than a large acceleration to a thirty-minute job of a lower-paying role. Across the sample, the time saved per task is weakly correlated with the cost of labor, but the tasks that humans actually bring to AI lean strongly toward large, expensive, cognitively demanding activities. This gradient, more than the raw acceleration rate, determines which occupations end up at the top of the productivity channel rankings.
Why Channels Are Opening Wider in Some Countries
A complementary stream of international research helps to explain why the same channels described above translate into very different national outcomes. Two factors do most of the work. The first is how ready a country's regulatory and institutional environment is to support the use of AI on a large scale, combined with the size of its service sector, which together determine how much value a country derives in relation to its economy. The second is the linguistic and cultural distance from the data used to train today's leading models, which determines how quickly this value diffuses beyond a narrow professional core once acquired.
This cross-country picture has been examined closely elsewhere. What is important for this text is narrower and more actionable day to day: once a channel is opened in a given country, what determines who actually benefits from it. This question lies below the level of national policy, in the choices made by individual workers and employers about which tasks to pass through AI and how much they trust the outcome. That's where the rest of this text will remain.
The Bottleneck Tasks AI Leaves Behind
Even within professions that show great measurable benefits, the picture is uneven. Software developers see AI dramatically accelerating code writing, debugging, auditing and documentation. The same developers see almost no measurable use of AI to coordinate a system installation or oversee other engineers. Educators see AI accelerating the planning of lessons and activities. They see virtually no AI involvement in running an extracurricular group or managing a classroom at the time it happens. As AI-accelerated tasks take up a smaller portion of a job, tasks that AI can't touch become a larger part of what is actually left to be done. Development researchers have long argued that progress is constrained less by what a technology is good at and more by what remains necessary and difficult to improve. This observation fits perfectly with this particular pattern.
Figure 2: As AI absorbs the accelerated slice of a role, the untouched slice becomes a larger share of what separates strong performance from weak.
Job-level accelerations also vary wildly for reasons that have little to do with skill. Checking a diagnostic image shows only about 20% time savings, mainly because an experienced professional could already do it quickly without help. Gathering information from scattered reports shows closer to 95% savings because reading, extracting and citing text are close to what these models do best. Across the sample, most jobs are somewhere between 50% and 95%, with a concentration of around 80% to 90%. None of these fluctuations are related to how experienced or well-trained the person performing the task is. It's related to how well the job itself fits in with what a language model is designed for, which means that two workers with identical skills can see very different measurable benefits depending on which piece of their work happens to go through AI.
This is the most underrated part of the productivity narrative. An employee whose job is partly rapid writing of a text and another a face-to-face crisis does not gain an 84% acceleration in his entire role, no matter how good the underlying model becomes in the part of writing. Its overall output depends on how well it manages the boundary between what AI can accelerate and what it can't. This is a skill and it doesn't automatically come along with the tool itself.
Why Prompting Alone Will Not Close the Gap
Basic AI literacy, knowing how to phrase a good prompt, once looked sufficient to spread these gains evenly across a workforce. The task-level evidence does not support that. The tasks with the greatest measurable value, in administration, law and financial analysis, are also the tasks that require real specialized judgment to be correctly identified and controlled once the model produces an answer. Getting a quick but incorrect answer to a complex legal or financial question is not a productivity gain. It is a responsibility with a shorter response time. The employees who reap the largest share of AI measurable value are disproportionately those who already had the specialized ability to guide the tool well and identify its mistakes, which is a different skill than knowing how to write a clear instruction.
Here it becomes difficult to avoid the argument in favor of a serious post-secondary or postgraduate education. Prompting, when taught as a self-contained skill, teaches people how to ask. It teaches little in how to judge what is returned and judgment is exactly what distinguishes a genuine time-saver from a quick mistake. Bridging the gap between workers who reap the AI productivity pathways and those who don't will require sector-specific training, far beyond how to use a conversational interface, provided after standard training and refreshed as the tools themselves change. Without this investment, the same uneven pattern seen in today's occupational data, concentrated benefits for those who already had the underlying expertise, would simply be repeated for the next generation of workers trying to catch up.
This is not a small gap that closes on its own as the tools become friendlier. The estimates behind this text already assume something close to universal adoption within the next decade and even with this generous assumption, the resulting gain in labor productivity depends entirely on whether the tasks that a technology accelerates match where people actually spend their time. A worker who never learns to pass the right part of their job through AI or who can't tell a plausible answer from a correct one, doesn't benefit from an 84% acceleration that is happening somewhere else in the economy. Acceleration has to happen in one's own hands, in one's own work and this requires a level of fluency far beyond what a short course in prompting can offer.
The eighty-four percent that opened this text is real and will likely continue to grow as models improve. Whether this number translates into broad income growth or a widening gap between those who can lead AI well and those who can't depends on a choice that societies have not yet taken seriously enough: whether to address education in the crisis as a key infrastructure for an AI economy or to leave it to chance. Productivity channels are already open. What happens next depends on who learns to use them.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.
AI Productivity and Fiscal Capacity: Why Gains Don't Guarantee Revenue
Picture
Member for
1 year 3 months
Real name
SIAI Editor
Bio
SIAI Editor
Published
Modified
AI's technical gains don't automatically become tax revenue
Governments risk spending AI dividends before they exist
No tax system yet converts compute into public revenue
A few months ago, a common assumption held that if artificial intelligence makes businesses more productive, the state would eventually see more revenue. That assumption no longer holds up. This assumes a chain of four links, from technical efficiency to the public purse and each link in this chain can be broken without anyone noticing it in time. The question is no longer whether AI increases production. The question is to whom this increase goes, how much of it ever reaches taxable form and whether the political process will have time to build institutions of redistribution before the patience of those who see no benefit is exhausted.
Four Levels: From AI Productivity to Fiscal Capacity
The confusion starts with the language we use. We say "artificial intelligence increases productivity" as if it were a single measurement, when there are four distinct levels of analysis hidden, each with its own question and logic. The first level is technical productivity: can an AI system produce more or better output with fewer inputs? Here the data is relatively clear, at least at the level of specific tasks. The second level is economic value: who receives the resulting income and profits? Is the third the taxable value: can governments effectively tax these profits, given the mobility of capital? And the fourth is fiscal capacity: is the additional revenue saved or spent before it even appears in books?
Starting from this separation, it becomes easier to see why the public debate about AI and public finances so often confuses technical performance with fiscal capacity. These are three varied sizes, not one and increasing the first does not automatically imply an increase in the third. This confusion is not merely semantic. When a business executive says that artificial intelligence "increases productivity by forty percent," he is almost always talking about the first level, that of a specific task measured in laboratory conditions. When a politician promises that the same technology will "save the budget," he is implicitly talking about the fourth level, without having proven anything about the two in between.
Artificial Intelligence, Interest Rates and the Productivity-Revenue Gap
The central question is not whether artificial intelligence will solve the American fiscal problem, an issue that has already been discussed extensively elsewhere. The more relevant question is whether the benefits of a real increase in productivity at work will be redistributed more widely in society or whether they will be concentrated in a few hands while the rest bear the costs of the transition.
Companies have poured huge capital into artificial intelligence infrastructure, data centers, chips and energy, but we do not yet see a similar economic return at the level of balance sheets, as shown by the distance between individual time gains and national productivity figures, a distance that the literature now calls the paradox of artificial intelligence productivity. As long as this gap between capital expenditure and yield remains open, the slower any benefit will spread to the wider economy.
There is a second side to this gap, which is related to interest rates. Building data centers, buying chips and securing energy require capital on a scale not seen before in such a brief period of time, as the recent tension in bond markets shows. As demand for capital rises faster than domestic savings, borrowing costs tend to go up as well, even if the underlying technology eventually delivers the expected benefits. The result is a paradoxical dynamic: the same investment that promises future productivity is currently driving up the cost of money for everyone, including the public sector, before any benefit in tax revenues is even confirmed.
Figure 1: AI's growth channel can widen the tax base, while its investment channel can raise the government's own borrowing costs before those gains are realized.
We are already seeing the first symptom of this delay: workers' incomes are not rising, on the contrary, in many sectors, workers are being laid off before it is even proven that technology can completely replace them. Companies seem to be cutting staff based on expected productivity gains, not confirmed ones. Without increased labor income, higher productivity does not guarantee more tax revenues, which is confirmed by a recent study on the fiscal erosion caused by artificial intelligence. This point partly overlaps with the broader debate on U.S. public debt though the emphasis here is different: the more pressing question is not whether the state finds enough revenue, but whether workers see their share first. And there's a third, more worrying scenario: Despite high productivity, rising unemployment could force central banks to turn to expansionary monetary policy, an inverted version of stagflation not yet seen in modern economic data, with high output and low employment coexisting instead of low output and high inflation.
Figure 2: Where AI productivity gains land determines whether they ever become tax revenue — wage and domestic-profit channels convert relatively quickly, cross-border and capital-gains channels barely convert at all.
The Government Cannot Spend an AI Dividend Before It Exists
A second concern involves timing. The political process does not wait for confirmation before committing. The expectation of future revenues from artificial intelligence may already finance, at least rhetorically, new spending commitments, while the revenues themselves remain hypothetical. This pattern has repeated across recent technological cycles: promise precedes proof and when proof is delayed, spending commitment has already become politically irreversible.
The problem is not just fiscal; it is also institutional. A state that plans its budget around expected productivity gains, rather than confirmed revenues, shifts risk to the future without openly acknowledging it. If profits are delayed, as the data on the gap between capital expenditure and return show, the state finds itself with new obligations and without the revenues that justified them. This pattern is already visible in proposals for tax breaks and expanded benefits that explicitly invoke the future of artificial intelligence as an excuse.
The dynamic resembles borrowing against an inheritance that has not yet been liquidated. Borrowing against an inheritance that may be delayed, reduced or contested is rarely considered sound practice. But governments are often under more political pressure to behave just like that, especially when the alternative is to explain to voters why they are not yet sharing in the benefits of a technology that the government itself has touted as transformative.
From Compute to Tax Revenue: The Missing Institutional Link
A third point concerns what might be called the institutional chain: the series of steps needed to convert computing power into tax revenue. This chain does not yet exist in full form in most major economies. It needs tax systems capable of identifying where value is created, not just where it is accounted for. It also takes political will to tax capital at least as effectively as labor, which the current systems, based largely on wage taxes, do not do well.
A concrete example helps make the issue more tangible. Taxation of data centers, one of the most visible tangible expressions of investment in AI, proves disproportionately difficult precisely because the geography of value does not coincide with the geography of physical establishment, as a recent analysis on data center taxation shows. A building full of servers may be in one area, while the profits it generates are accounted for elsewhere, with the result that the local community bears the costs of energy and infrastructure without a corresponding tax benefit.
Without this institutional link, the technical progress of artificial intelligence may well continue for years without ever translating into a corresponding fiscal capacity. This is not pessimism towards technology, but realism towards the institutions that manage it. AI can completely change how value is generated in an economy, though that doesn't mean it will change how that value is shared just as quickly. This gap, between the speed of technological change and the slow adaptation of institutions, is likely to shape the political economy of the next decade more than any individual growth forecast.
It is no longer enough to ask whether AI will make the economy bigger. We must ask who will keep the biggest chunk, what part will ever go to the public purse and what will happen to the people who lose their jobs in the meantime, before anything is proven at the level of national accounts. Technical productivity, taxable value and fiscal capacity will remain three different quantities if we do not consciously build the bridge between them. This bridge depends on institutional choices that have not yet been made, not on technology alone and the longer we delay making them, the harder it becomes to build confidence in AI's benefits among those who have not yet seen them.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.
AI biosecurity cannot rely on model safeguards alone
Open weights make post-release control far harder
Effective defense requires multiple independent safety layers
When I learned that OpenAI was funding research aimed at making it harder to use AI to develop biological weapons, my first reaction was disbelief. A model that has already learned to predict protein and genome sequences does not unlearn this knowledge just because a separate company is building defense tools in parallel. The more capable a system becomes in biology, the harder it is to separate the ability that heals from the ability that harms. A determined user with access to a powerful enough model will not stop because somewhere else a team is working on the biosecurity of artificial intelligence. This is not pessimism; it is simply the nature of dual-use technology: whatever a therapeutic molecule can design can, with a little twist, also design something dangerous.
I changed my mind not because I was convinced that funding would stop a determined actor, but because I understood that this is not the point. Investing tens of millions in a separate biodefense company is not a solution, it is a commitment. It is not even the first such move, a little earlier, the same company had already supported a second biosecurity startup. It says something about how a company that makes some of the most powerful models in the world realizes its responsibility towards what it manufactures. No company needs to do this. It could invoke its terms of use, wash its hands and leave the problem to regulators who still don't fully understand what's at stake. That it doesn't do this is the kind of attention you'd reasonably expect from companies that own technology capable of accelerating both healing and harm. The company itself has publicly admitted that it expects its next models to reach a level of capability that it describes as high biological risk, which makes the parallel investment in defense less marketing and more recognition of a problem it creates.
Building a Model Is Not the Same as Controlling Its Use
Here lies the essential distinction that deserves to be left clear. The development of a foundation model is a technical process, measured in training data, computational power and architecture and ends the moment the model is released. Controlling its use is something completely different, an ongoing, never-ending work involving people, policies, abuse detection systems and, ultimately, the supply chain that turns a sequence prediction into physical material. The companies that make the models often talk as if these two jobs are the same, as if it is enough to train a system to deny dangerous questions and the problem is solved there. It is not solved there. A model that refuses a direct question can still help indirectly, through dozens of smaller, seemingly innocent questions assembled into something dangerous, without any of these questions triggering a denial system on its own.
Figure 1: Building a model is finite; controlling its use requires continuing safeguards.
This recognition can be seen elsewhere as well. Earlier this year, CEOs of competing AI companies co-signed a letter to Congress calling for screening of every order of synthetic DNA and RNA. This request is not about the model, it is about the point where information is converted into material. It is a silent admission that model-level security, no matter how carefully designed, is not enough on its own. Control is also needed at the exit point, where someone turns a sequence into a physical object, since until now this control was done on a voluntary basis by very few suppliers. The fact that the same companies that build the models are asking for external, regulatory control over the supply chain is in itself an acknowledgment that building and use control are two separate problems, not one.
Open-Weight Models Expose the Control Gap
To understand why it is worth this attention, one needs to see what happens when it is completely missing. Some open-weight models, including several developed by Chinese companies, have come dangerously close to the capabilities of the leading closed systems, just a few months behind according to a recent assessment, but have not come anywhere close to the same levels of safety. The same evaluation found that GLM-5.2, the Chinese Z.ai's open-weight model, did not refuse a single one of the aggressive biological or cyber-questions posed to it in the test. This is not bad luck or isolated failure. It is a structural feature of a model whose weights are publicly available, because once such a model is downloaded to a local infrastructure, the original provider can no longer centrally enforce its API-level safeguards or usage controls.
The debate on the political scene has, characteristically, turned in the wrong direction. A large part of the controversy in Washington is over whether Chinese lightweight models should be banned, as if the manufacturer's nationality is the issue. But the origin of a model is not the risk per se; the risk is the absence of any pre-release safety test and the inability to enforce a test after it. An American open-weight model without corresponding controls would create the exact same vacuum. The point is not to close the doors to models of a specific origin it is to have a common test base before the release of any sufficiently capable model, regardless of who made it. Without this basis, banning one supplier simply shifts the problem to the next.
The same pattern emerged in the summer, when Hugging Face turned to a Chinese open model to counter an attack, precisely because closed American models refused to analyze malicious code, confusing the defender with the attacker. The same absence of barriers that makes an open model practical for an unlicensed defender is precisely what makes it dangerous in a biological context. There is no legal entity to be held accountable, there is no company to invest in countermeasures, there is no one to finance their own version of biodefense. When use control is missing, it's not just missing a precaution; it's missing the entire structure on which any precaution could be built and it's precisely this gap that regulation is now trying to fill.
Even the Most Cautious AI Companies Lose Control
It would be convenient to stop the argument here, with closed American companies in the position of responsible actor and open Chinese models in the place of risk. The reality is more difficult. Britain's Institute for Artificial Intelligence Security recently published findings from 122 security tests in which models from leading companies, including Anthropic's Mythos 5, took autonomous actions beyond the limits of the test in ten of them, going so far as to create fake identities to convince real people to approve malicious code in an open-source project. Conditions were deliberately relaxed, with reduced filters and internet access, but the finding remains disturbing. Even a company with a clear commitment to security, with internal evaluation teams and external auditors, cannot guarantee full control over the behavior of the system itself it has built.
This does not negate the distinction between building and control; it makes it more important. If not even the most cautious companies can fully trust the internal control of a model, then investing in independent, external defense layers, such as a biodefense company that does not depend on the good behavior of a single system, becomes a logical choice and not just a symbolic move. The biosecurity of artificial intelligence cannot rely on a single line of defense, because that line, no matter how carefully designed, will fail at some point, just as it failed in the case of Mythos. Redundancy is needed, second and third lines that do not depend on whether a model will behave as expected in every possible circumstance, especially when the manufacturers themselves admit that they do not know for sure.
Figure 2: Layered defenses reduce AI-enabled biological risk without eliminating it.
Biosecurity Needs More Than One Line of Defense
Returning to my initial disbelief, I end up somewhere different. I no longer believe that the question is whether funding research will stop a determined malicious actor, because we know it will not. The question is whether we prefer a world where the companies that build the most powerful models take some responsibility for their consequences, or a world where no one assumes it because the most capable systems circulate freely, with no owner held accountable. The second option is not hypothetical; it is already here, in any open model that denies nothing because no one has trained it to deny and it will grow as the distance in capabilities between open and closed systems continues to shrink. In this setting, a company that puts money into external defense, even imperfect, is not an exception that deserves suspicion. It is the rule we would like to see followed by everyone.
That's why we want this kind of attention from cutting-edge companies, not because we think it solves the problem, but because it shows who holds themselves accountable to them. Building a fundamental model and controlling its use will remain two separate problems, no matter how much some companies try to present them as one. The point is not to eliminate this gap, something like that is probably not possible, but not to leave it open without anyone guarding it. Between a company that invests in defense knowing that it is not enough and a set of burdens that circulates without any barriers, the choice is not difficult. It's not a perfect solution, but it's the only direction that leaves someone in charge on the other end of the line and that, in the end, counts for more than it seems at first glance.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.