Skip to main content

Causal Reasoning and LLMs: Why Employers Now Want Both

Picture

Member for

1 year 4 months
Real name
SIAI Editor
Bio
SIAI Editor is the institutional editorial identity of the Swiss Institute of Artificial Intelligence (SIAI). It covers research and analysis across AI policy and governance, economics and finance, law and regulation, workforce and education, scientific applications, computational methods, infrastructure, and the strategic adoption of artificial intelligence.

Publications under SIAI Editor are prepared or coordinated by SIAI’s research and editorial team and include research synthesis, policy and industry analysis, technical interpretation, and interdisciplinary work connecting artificial intelligence with established fields of research and professional practice.

Modified

Causal training plus ChatGPT beat control on all eight measures
Employers now seek analytical thinking alongside AI fluency
Short MOOCs pay off only when they teach reasoning

In thirteen sections of an introductory management course at Bocconi University in Milan, 1,053 first-year students had 45 minutes to write a 180-word recommendation on a real-world problem, how more alumni would learn about the university's merchandise shop and how they would use it more often. Those who had ChatGPT at their disposal wrote texts that the evaluators rated higher. Only those who had gone through causal training, however, articulated when their proposal would fail, explained the mechanism behind it and submitted ideas that were unlike their fellow students. The team that had both did not lag by any measure. The finding on causal reasoning and LLMs describes in miniature a change that is already being recorded in hiring, training budgets and corporate agreements with platforms such as Coursera, where employers are now looking for people who know how to ask why and at the same time know how to put a model to work on the answer.

What ChatGPT Improves and What Causal Training Changes

The experiment was pre-registered and randomized at the class level, with four conditions divided into thirteen sections: training in causal thinking, access to ChatGPT through the university's institutional subscription in a separate browser tab, a combination of the two, or no intervention. The training was not a weeks-long seminar. It took the form of a twelve-question game with feedback, teaching students to build explicit chains of cause and effect, stating under what conditions their claim would collapse and explaining the mechanism by which a cause produces its effect, while the rest played a placebo game with the same questions without feedback. Eight outcomes were measured. Reasoning was graded by LLMs based on a rubric that had been declared in advance, performance by twenty trained graduate students, three per text and without knowing the condition and the similarity of each recommendation was measured against solutions independently written by a marketing professor, the store manager and an experienced alumnus.

The results are clearly divided. ChatGPT raised the performance, i.e. the evaluators' scores and the proximity to the experts' solutions, while the training in causal thinking did not improve it and its addition to the tool did not raise it further. Both interventions made the texts more coherent and richer in ideas, the tool to a greater extent,

and the combination is more than each one separately. In terms of depth, however, i.e. the reporting of conditions of failure, the explanation of mechanisms and the differentiation of ideas from those of fellow students, only the trained team moved. ChatGPT alone did not achieve any of this and adding it to the training did not take anything away. The group with the two interventions surpassed the control group in all eight outcomes and never lagged behind a group with a single intervention, while in coherent logic, mechanisms and number of ideas, the two interventions reinforced each other.

Figure 1: ChatGPT lifts scores; only the trained reasoning moves mechanisms and falsifiability.

There is a detail in this finding that directly concerns HR managers. The evaluators rewarded texts that resembled the experts' solutions and this is exactly what produces an LLM with ease, since it reproduces the standard answer to a typical marketing problem. Standardized criteria, useful when thousands of students or employees have to be rated in a way that seems fair, end up discouraging the variety of ideas and first-principles thinking and a company that judges its executives with rubrics designed before the advent of tools risks rewarding polished compliance at the very moment it became cheap, while the hard-to-find qualification, the ability to tell when a proposal won't work, goes unnoticed on the evaluation sheet.

Superhuman Labor: Causal Reasoning and LLMs Together Take the Lead

The concept of superhuman labor, as formulated by SIAI in August 2026, describes a human working with AI tools and producing work better, faster and cheaply enough to count, without implying general artificial intelligence. The human gives the judgment, the context and the questions worth answering, the tool gives coverage and speed. In the same analysis, it was pointed out that a feedback loop fed only by the output of the models recycles material that already exists and returns to the same point, unless someone introduces a truly new question or framing. The Milan experiment gives this distinction measurable content. The LLM increased the volume of ideas and raised the evaluations, the trained thinking changed their character and the combination alone produced texts that were well graded and at the same time original. The causal thinking and LLM group is, in this sense, the closest empirical picture of superhuman labor measured with a control group.

Data from the workplace shows why the combination has financial weight and is not an academic ornament. In an experiment with 758 consultants from the Boston Consulting Group, those who had GPT-4 completed 12.2 percent more tasks, 25.1 percent faster and with quality over 40 percent higher, as long as the tasks remained within the limits of the tool's capabilities, while in a task that was outside these limits the probability of a correct solution was 19 percentage points lower for those who used it. The limits are not marked anywhere on the screen. The ability to state when a claim ceases to be valid, which is exactly what the game of twelve questions taught, is what allows the user to understand that the model's convincing answer has gone out of the field. In customer service, for more than 5,000 employees of an American software company, an AI assistant increased resolved requests per hour by 14 percent on average and by 34 percent for beginners, but in tasks where the correct answer was already in the conversations of the best colleagues.

The reverse scenario was recorded in Kenya, where 640 entrepreneurs gained access via WhatsApp to a GPT-4-based consultant for five months. The average did not differ statistically significantly from the control group, because it covered two opposite paths: those who were already doing well gained over 15 percent, while those who were struggling recorded a performance almost 10 percent lower. The questions to the assistant and his answers were similar in the two groups and the distance appeared in the selection and application of the advice, where the problem did not have a predetermined correct answer and the tool multiplied the judgment in front of it, whatever it was.

Employers Now Want Analytical Thinking and AI Fluency in the Same Hire

The job market has begun to price this combination before it even makes a name for itself in the hiring departments. The World Economic Forum's Future of Jobs 2025 report, based on a survey of more than 1,000 employers with over 14 million workers in 55 economies, ranks analytical thinking as the first basic skill, with seven out of ten companies considering it essential, while AI and big data are at the top of the skills with the fastest rise and the share of employers considering them basic increased by 17 percentage points compared to the 2023 version. The two rankings do not compete. They describe the same worker from two sides. 63 percent of employers consider skills gaps to be the biggest obstacle to business transformation by 2030 and 85 percent plan to prioritise upskilling existing staff, significantly higher than the 70 percent who expect to hire people with new skills.

Figure 2: Employers expect to retrain nearly half their workforce in place rather than replace it.

How the gap is closed in practice is shown by data from the US Census Bureau. In the supplementary questionnaire of the Business Trends and Outlook survey, collected from December 2023 to February 2024, 20.8 percent of businesses using AI had trained existing staff and only 2.0 percent had hired employees already trained in AI. The ratio of ten to one says something about the talent supply, since ready-made employees who combine industry knowledge, judgment and fluency with tools are few and expensive and says something more about demand, as companies prefer to build capacity on people whose judgment they already know from years of evaluations. On the part of the trainees, the movement is accelerating. Coursera reported to its shareholders that enrollment in AI courses reached 20 per minute in the first quarter of 2026, up from 15 in 2025 and 8 in 2024.

The pressure is also reaching the entry-level positions. Data from resumes and job advertisements for approximately 62 million employees in 285,000 US companies show that the employment of new employees in companies that adopted generative AI fell by 7.7 percent in the six quarters after the first quarter of 2023 compared to the rest, while senior management was not affected. With fewer entry-level positions, the new hire is expected to contribute at a higher level from day one and the profile that meets this expectation is very similar to the Milan team that had both the training and the tool.

Short MOOCs Become the Default Route to Company-Wide AI Training

On October 5, 2026, TO THE NEW, a Singapore-based digital engineering firm, announced a partnership with Coursera that opens up the platform's full enterprise catalog to all of its over 2,200 employees in India, the US, the United Arab Emirates and Australia. The agreement supports an internal program to accelerate AI adoption in each technical specialty, with role-specific learning paths curated by practice leads and teams onboarded in phases, while the company states that it now runs all client engagements as AI-enabled. The move is not isolated. BSI announced a global training partnership with Coursera on May 5, 2026; Coursera on May 11 completed its merger with Udemy into an organization with 18,000 enterprise customers and the U.S. Department of Labor opened on April 29 a website to integrate AI skills into approved apprenticeship programs. Businesses are turning to short, digital courses that scale easily because the market doesn't give them ready-made people at the pace they need and because external candidates cost more and take a while to perform.

The strongest counterargument argues that training in the tools is sufficient, since AI benefits the less experienced more and will smooth out the differences on its own, so a separate thinking program would be an unnecessary expense. Data from customer service supports this reading for structured tasks. In the Milan experiment, however, ChatGPT alone did not produce failure conditions, mechanisms, or nuanced ideas and in Kenya the normalization was reversed for the weakest once the problem became open. A course that teaches instruction writing without causal thinking yields polished texts that look similar to each other. The Milan training, moreover, fit into a dozen questions with feedback, which fits easily into a MOOC learning path. For recruiters this translates into a short written case study where the candidate has to state when their proposal would fail, while for training departments it means a first-principles stage before the tools stage, with assessments that also measure the differentiation of ideas, not just proximity to the formal answer.

In Milan, the team that had both the training and the tool was the only one to outperform the control team in all eight measures, in a 45-minute test and after a game of twelve questions. The cost of the training was small, the cost of the tool was an institutional subscription. The companies that currently distribute MOOC catalogs to thousands of employees have already covered the second half. TO THE NEW's announcement does not indicate whether the role-specific paths curated by the practice leads include anything similar to that game, nor how the progress of the 2,200 engineers will be measured when the last onboarding phase is completed. The answer will show whether the company bought fluency with the tool or the combination that, in the Milan experiment, was nowhere behind.


This article reflects the analytical judgment of The SIAI Editorial Board and does not constitute policy advice or the official position of any affiliated institution.


References

Asirvatham, H. et al. (2026) Training Novices to Think, or Giving Them LLMs? Evidence from an RCT. CEPR Discussion Paper 21882.
Bonney, K. et al. (2024) Tracking Firm Use of AI in Real Time. CES Working Paper 24-16, U.S. Census Bureau.
Brynjolfsson, E., Li, D. and Raymond, L. (2025) 'Generative AI at work', The Quarterly Journal of Economics, 140(2), pp. 889-942.
BSI (2026) 'BSI and Coursera partner on global workforce skills development', press release, 5 May.
Coursera, Inc. (2026a) Q1 2026 Shareholder Letter.
Coursera, Inc. (2026b) 'Coursera completes combination with Udemy', press release, 11 May.
CXOtoday News Desk (2026) 'TO THE NEW partners with Coursera to bring AI learning to its entire workforce', CXOtoday, 5 October.
Dell'Acqua, F. et al. (2023) Navigating the Jagged Technological Frontier. Harvard Business School Working Paper 24-013.
Fumagalli, C. and Gambardella, A. (2026) 'Training novices to think in the age of large language models', VoxEU.org, 1 October.
Hosseini Maasoum, S.M. and Lichtinger, G. (2025) Generative AI as Seniority-Biased Technological Change. SSRN Working Paper 5425555.
Leopold, T.A. et al. (2025) The Future of Jobs Report 2025. World Economic Forum.
Otis, N.G. et al. (2026) 'The uneven impact of generative artificial intelligence on entrepreneurial performance', Management Science.
SIAI Editor (2026) 'Superhuman labor and the feedback loop no one is watching', SIAI AI Memo, 17 August.
U.S. Department of Labor (2026) 'US Department of Labor launches website to build artificial intelligence skills', news release, 29 April.

Picture

Member for

1 year 4 months
Real name
SIAI Editor
Bio
SIAI Editor is the institutional editorial identity of the Swiss Institute of Artificial Intelligence (SIAI). It covers research and analysis across AI policy and governance, economics and finance, law and regulation, workforce and education, scientific applications, computational methods, infrastructure, and the strategic adoption of artificial intelligence.

Publications under SIAI Editor are prepared or coordinated by SIAI’s research and editorial team and include research synthesis, policy and industry analysis, technical interpretation, and interdisciplinary work connecting artificial intelligence with established fields of research and professional practice.