Superhuman Labor: Hiring, Training and Tools for Causal Reasoning
Authored on
Modified
Causal training plus LLM access produced the strongest group Firms can hire, train and equip for this combination Whether the advantage lasts across careers remains untested

Superhuman labor has been described in this series of notes as the work of a human working with AI tools and producing more, faster and cheaper than he would produce on his own. The argument has been revisited several times with the same insistence on one point: the tool is not enough, because a loop fed only by the output of the models recycles what already exists and the value is judged by the quality of the human input. A randomized experiment with 1,053 first-year students of Bocconi in Milan, made public in early October 2026, adds another confirmation, this time with a control group. The next move of reasoning is more interesting than the confirmation itself. If the human input that makes the difference is causal reasoning and if it is taught cheaply, then superhuman labor ceases to be an accidental privilege of a few talented people and becomes something that a company can build with its recruitment, with internal training and with the tools it gives to its people.
What the Milan Experiment with 1,053 Students Showed
The thirteen sections of an introductory management course were randomly divided into four conditions: training in causal thinking, access to ChatGPT, a combination of the two, or no intervention. The training was a twelve-question game with feedback, teaching students to build explicit chains of cause and effect, state when a claim ceases to be valid and explain the mechanism linking cause to effect. This was followed by a 45-minute assignment: an 180-word recommendation on how more alumni would learn about and use the university's merchandise shop, a common marketing problem. The texts were graded by graduate students who were not familiar with the condition, compared with solutions from three experts, a marketing professor, the store manager and an experienced alumnus and analyzed for the coherence of logic, mechanisms, conditions of failure and variety of ideas, in total in eight measures.
ChatGPT raised the scores and brought the texts closer to the experts' solutions. Causal training did not improve these scores and adding it to the tool did not raise them further. In terms of depth, the picture is reversed, since only those who had been trained stated when their proposal would fail, explained the mechanisms behind it and submitted ideas different from those of their fellow students, while ChatGPT alone did not achieve any of these. The tool increased the volume of ideas and trained thinking changed their character. The team with the two interventions outperformed the control group in all eight measures and never lagged behind a team with one intervention. One point deserves special attention: the addition of ChatGPT did not take anything away from those who had been trained, which answers, at least for a 45-minute horizon, the fear that the tool is eroding a way of thinking that has already been taught.
Superhuman Labor Depends on the Quality of Human Input
Previous versions of the argument placed human input somewhat vaguely, in judgment, in framing and in questions worth asking. The experiment narrows the definition to a specific skill to be taught and at the same time corrects an implicit assumption. Superhuman labor was often read as a property of pairing the tool with any user, with the quality of the user being taken for granted. The evidence shows that the tool in the hands of anyone produces the standard answer faster, which has value but does not feed the loop with anything new, while the tool in the hands of someone who has learned to ask when and why an idea fails produces the non-standard answer that the loop needs to keep it from turning to the same point. The combination group is therefore the superhuman group in the strict sense. A theoretical analysis of the viability of careers in the age of AI reaches a similar threshold from another starting point, with young people crossing a knowledge threshold supplementing AI and those staying below it being replaced.

SIAI's recent research on the changing value of human labor explains why this matters for a business and not just for an auditorium. Checking the result of a model requires the same knowledge as producing it and often more, because the reviewer has to recognize a mistake that seems right without having gone through the steps that produced it himself. In a large evaluation by experienced workers, about four out of ten model responses to professional text tasks needed correction or repetition and someone with knowledge of the profession had to figure out which ones. How to use it also weighs heavily. In a randomized study with developers learning a new Python library, those who asked the model for conceptual explanations scored comprehension scores of 65 percent to 86 percent, while those who assigned it writing or debugging stayed between 24 percent and 39 percent. The tool was the same and the difference was made by the user's habit of asking why, i.e. exactly the habit that training cultivates in causal thinking.
Workforce Strategy: Recruitment, Training and Tailored LLMs
The first lever is the recruitment of people who have already been trained in causal reasoning, with a short written case study where the candidate must state when their proposal would fail. But the supply of such people is limited and external recruitment is expensive. In personnel data from an American investment bank, those hired from outside were initially paid about 18 percent more than colleagues who were promoted internally to similar positions, had lower ratings in the first two years and left more often. The companies themselves seem to know this. In the US Census Bureau survey, 20.8 percent of those who used AI had trained existing staff and only 2.0 percent had hired already-trained staff, while in the European Union, among companies that considered AI without adopting it, 70.3 percent cited a lack of relevant knowledge in 2025.
If the market is not enough, the second lever is internal training. The recent SIAI Workforce Strategy Working Paper compares seven alternatives over a 24-month horizon and, in the hypothetical example of a mid-size insurance service center with 20 employees and 6,000 cases per month, training domain professionals yields around €80,000 net contribution relative to maintaining the existing flow, while hiring an applied AI engineer ends up at minus €163,000 and a tool-only deployment at minus €101,000. Training reaches two-thirds of full capacity in about five months versus eight for recruitment. Its advantage comes mainly from the mistakes avoided and less from the hours saved and if the mistakes cost nothing, its contribution would drop to €16,000. Avoiding mistakes is precisely the skill taught by the question of when a claim ceases to be valid, so causal reasoning belongs in the content of the training.
The third lever is custom tools. The AI assistant that increased the productivity of customer service beginners by 34 percent had been trained on the same company's successful conversations and appeared within the workflow, with suggestions that the employee could accept, modify or ignore. For people trained in causal thinking, corresponding customization means tools that are fed with the data, customers and quality criteria of the business and that explicitly ask for a mechanism and conditions of failure before giving a final text, so that the user's judgment enters the flow instead of being bypassed. TO THE NEW, which on October 5, 2026 opened Coursera's catalog to all of its over 2,200 employees with role-specific paths curated by practice leads, is moving in the same direction on the training side.

What Remains Open for the Entry-Level Pipeline
The three-lever strategy has a point that's easily missed in a two-year budget. In data from resumes and job postings for approximately 62 million employees at 285,000 U.S. businesses, the employment of junior employees at companies that adopted generative AI fell by 7.7 percent in the six quarters following the first quarter of 2023, while senior employees were not affected. The SIAI working paper captures these costs as a forecast that only appears over the 36-month horizon, when today's beginners should have become tomorrow's experienced reviewers. A company that hires only ready-made people with judgment and assigns to the model the tasks on which young people were learning consumes a stock that does not renew and training in causal thinking, if done early and cheaply, is one of the few ways to renew that stock without the firm returning to the old, expensive model of apprenticeship next to an experienced colleague for years.
The Milan experiment lasted 45 minutes. No dataset has yet tracked for years a group of employees who were trained in causal thinking and worked with LLMs from the start, so whether the combination maintains its lead on a career horizon is an open question. The SIAI paper also shows that recruitment performs best when the gap is large and the deadline falls between search time and learning time, so training is not always the right answer. TO THE NEW's 2,200 engineers join the program in phases and the announcement doesn't say whether their routes include something like a twelve-question game or just fluency with tools, something that will show, if at all, in their evaluations sometime in 2027.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.