Skip to main content

AI Labor and Human Labor: Why Replacement Is Slower Than the Hype, but More Serious Than Workers Think

AI is not yet a universal substitute for human labor. Enterprise deployment remains expensive, operationally fragile, and dependent on human review The immediate transition is labor compression: fewer AI-leveraged workers may be expected to produce more output before end-to-end automation becomes viable Falling inference costs, expanding infrastructure, and better workflow design will make more tasks economically contestable - so firms and workers should redesign work now

[AI and Tax] Data-Center Taxation and the Geography of AI Value

[AI and Tax] Data-Center Taxation and the Geography of AI Value

Keith Lee*

*Swiss Institute of Artificial Intelligence, Chaltenbodenstrasse 26, 8834 Schindellegi, Schwyz, Switzerland

Abstract

Data centres are increasingly presented as an attractive tax base, given their capital intensity, immobility and local demands for electricity, land, water and other infrastructure. The article argues, however, that the physical location of servers cannot automatically identify the wider economic rents generated by AI intrinsic to the wider AI value chain. These value streams are dispersed across geographies and layers in an international value chai- from semiconductors, hardware and energy, through cloud orchestration, foundation models and application software, to proprietary data and ultimately to adopting entities.  These locations may differ substantially from the physical location of the server and are linked through contractual, licensing, transfer-pricing and corporate ownership arrangements which, together, separate compute geography from the location of accrued economic rent and associated profit. The article contends that host jurisdictions do have a legitimate claim to recover for identifiable land, grid, water, environmental and other infrastructure costs, but do not automatically acquire taxing rights over all model, cloud, softwar and market-based rents.  It proceeds to differentiate the locational characteristics of training and inference workloads and to note how mixed workloads weaken the case for a broad facility levy. Concluding that Europe's key regulatory dilemma is to balance legitimate cost-reflective charges with an interest in growth and innovation, the article advocates taxing site-specific infrastructure and environmental burdens locally while addressing wider rents through separate profit-allocation and jurisdictional rules.


[AI and Tax] is an independent research series developed by Professor Keith Lee following his presentation at the conference Inequalities in Longevity, held at Fondazione Giorgio Cini in Venice on 3–4 July 2026, and a subsequent substantive discussion with Federico Fubini of Corriere della Sera concerning AI-driven productivity, labour income, and fiscal capacity.


Public-facing adaptation

A public-facing adaptation of this research, written by the author for a wider readership, is available from The Economy Review: "[AI and Tax] Tax the Load, Not the Algorithm: A Smarter Model for AI Data Centre Taxation."


1. Introduction - Why Data Centres Are an Attractive but Incomplete Tax Base for AI

The attraction of a city-based and physical tax base for data centers is straightforward. If AI diminishes aspects of the tax base that have traditionally fallen on labor income, its employment, its plant and equipment or its labor-intensive processes, then the tangible, capital-intensive, polluting, geographically fixed over the short to medium term and state-plannable nature of the data center infrastructure seems to offer a recognizable alternative handle-one that is easier to notice, measure and charge than less tangible digital profits. In a digital, and indeed AI, age, where value appears to be drifting partly across borders via codified lines of code in the cloud, the data center tower seems to have prima facie appeal as an analog solution. However, that intuition is incomplete. The tangible placement of data center infrastructure can justly underpin certain local taxation and fee-recovery rights, but it does not secure a rightful claim in law or economy as to the entire AI-generated economic rent.[1]

The trade-off is that a data center is just one node in a layered and globally distributed value chain.[2] OECD's recent benchmarking of AI supply shows chips, data centers, clouds, models, connectivity and end users form part of the same supply chain, which has a structure characterized by large fixed costs, scale economies, vertical integration and bottlenecks not exactly aligned with taxes across national frontiers.[3] OECD analysis shows that leading AI firms increasingly operate across several layers of the infrastructure stack, while countries differ markedly in their specialization across compute, cloud, models and applications. A host country can thus earn property taxes, construction-related activity, operator and contractor wages, electricity payments and utility revenues from hosting the data center, while much larger rents go elsewhere to chip design, semiconductor fabrication, hardware suppliers, cloud hyperscalers, model vendors, integration software, proprietary data owners, adopting firms and final investors.

A single charge against the coal in the mines, or those contracts, or those silicon wafers, involves other geographies of rent. For computing, geography has to determine where the servers actually run. For energy geography, the question is where electricity, capacity, storage, cooling and transmission lines are supplied from and reinforced. For labor geography, there is the question of where the engineers, technicians, managers, AI researchers, software engineers, facility technicians, clients and end-users are. For intangibles, it is where legally protected models, patents, the software program, or the data rights are owned or exercised. For corporate, it is where the subsidiary posts its contract, invoices and reports. For markets, it is where the paying users or consumers are. Each separate geography may imply a participant in value creation, a legal interface and a political policy base. Entry into the data center, as if it encapsulated all those six geographies simultaneously, erodes distinctions maintained for inextricable reasons by international tax law, by utilities regulation and by industrial economics.

This differentiation is clearly observable in tax doctrine.[4] According to the OECD Model Tax Convention, a server may qualify as a fixed place of business under certain conditions, but a website by itself has no taxable presence and ordinary hosting arrangements do not usually bring the server into use for the business tenant at the hosting.[5] Where a permanent establishment does exist, the location of the hosting service does not alone tell how much profit is attributable to it. Transfer-pricing rules draw another difference; the legal owner of an intangible property does not by itself receive the residual return if the development, enhancement, maintenance, protection and exploitation of that intangible property are carried out or controlled elsewhere.[6] Put otherwise, neither the server nor the nominal owner of the source code is sufficient to define the locus of AI rent.

The truth is therefore more subtle and the mainstream rhetoric is simply too coarse to capture what is realistic. Data-center taxation is feasible when it recovers identifiable local costs, when it accounts for local externalities and when it only taxes the incremental operating income of the data center that is genuinely attributable to it. It ceases to be analytically sound and may even prove distortionary, when it is believed to be a surrogate for the broader economic rent created along the entire AI value chain. The server has a place in the matrix, but indispensability should not be confused with platform dominance. Sound policy must distinguish he difference between taxing local environmental and physical infrastructure externalities and taxing facility profits and claiming taxing rights over rent extracted from cloud, model, software, data and IP ownership and design and capturing the tax domain of the jurisdiction in which paying users live. The rest of the paper explores this nuance in terms of the AI value stack, the multiple geographies of AI, the justifiable claims of host jurisdictions, the shortcomings of data facility taxation as a proxy for AI rent and an application to the particular strategic dilemma now facing Europe.

2. The AI Value Stack: Seven Layers of Value Creation and Rent Capture

The first is that an incomplete tax proxy for the data center exists because AI value is not generated on a lone asset but on an entire stack. Closest to the value chain are semiconductors. The need for powerful AI capabilities pushes for GPU and AI accelerators, high-bandwidth memory, foundry services such as TSMC’s and manufacturing capacity, all of which are characterized by high capital intensity and high concentration and differentiation.[7] Previous OECD research on AI infrastructure already demonstrated the fact that the GPU market is highly concentrated, that complementary software (e.g., CUDA) is integral and that specialized upstream suppliers, including TSMC, ASML, SK Hynix and Samsung, occupy critical positions without necessarily participating in consumer cloud application markets.. Whenever compute becomes limiting, or advanced-generation chips are constrained due to supply problems, then a large share of rent is specific to this layer, rather than the later data centre.[8]

Figure 1. Infrastructure pressure is spreading across electricity, chips and capital expenditure rather than concentrating solely in server facilities.

The second layer includes servers, networking, storage and the construction of the data center.[9] This is the layer most visible to local decision makers since it encompasses the buildings, racks, cables, switchgear, site preparation and complex engineering works. It is also the layer most often called upon in the eyes of many by public debates about the locus of AI. Yet even here, the economic footprint is quite equivocal. Some value is gained on-site in terms of civil engineering works, construction services and routine operations. Much of the value in the provision of advanced AI flows to globally active equipment vendors, engineering consulting firms and integrated infrastructure suppliers, whose earnings do not remain local. OECD studies emphasize that modern AI systems are reliant not just on an anonymous warehouse of servers but on advanced interconnects, resilient networks and digital backbone infrastructure -a good proportion of which is owned or coordinated by the world's most major technology players with a global presence. The physical plant thus remains a requirement, but is often not the primary owner of the space of strategic value, which renders the shell itself valuable.

The third layer includes electricity, the grid, access, cooling, land and water.[10] Here, the argument for local public benefits is strongest. Power is almost always the largest operational expense of running a data center; heating or cooling loads and power densities are rising dramatically for AI-centric data centers; and according to the International Energy Agency (IEA), it may not be long before a single AI server rack consumes maximum electrical power similar to that of dozens of households.[11] OECD analysis concludes that these data centers consume enormous amounts of electricity, often excessively consume resources like water and require large, costly cooling infrastructure. Electricity supply and water supplies and discharge systems are geographically sticky because the location of grid connection points, substations and water will be explicitly local. This lets hosts claim the value for the environmental costs and burdens of the local infrastructure and resources, even if the rent flows up the stack to many global owners.

The fourth layer is the cloud-computing and orchestration services.[12] Economics is now no longer about physical raw materials. Rather, they are about having control over the shared capacity, interfaces, middleware, scheduling, managed services and ecosystem access. Since the OECD shows that the largest hyperscalers are especially well positioned to claim large proportions of the AI cloud market, that combined hyperscaler market shares are consistently very large across national and regional studies and that a growing proportion of large AI firms go to market primarily via partnerships through which cloud providers provide both capital and compute, then it follows that much of the product is in fact the contract, not the server hall in Finland, Ireland, or Virginia; while large proportions of the quasi-rent can be claimed where the racks are in another country and the customer never actually sees its servers.

Figure 2. Cloud concentration shows how control over contracts and orchestration can capture value above the physical facility layer.

The fifth layer is the foundation models and model APIs.[13] Frontier-model developers can capture rents through scarce expertise, safety fine-tuning, brand, switching costs, first-mover benefits, proprietary evaluation pipelines and the ability to commercialize a pretrained model over numerous downstream applications. However, these companies are not necessarily completely vertically self-reliant. BIS evidence demonstrates that private frontier model companies like OpenAI and Anthropic have limited direct participation in compute and cloud markets and utilize close partner and infrastructure arrangements. This matters for tax analysis because the model rent may not sit with the data-center operator or with the customer-facing user, but with a third party who licenses or provides access to the model via cross-border contracts, which are themselves possibly routed by different connected firms.

The sixth layer comprises data, software, applications and distribution. This is where narrow model capacity can translate into industry-specific commercial value. Firms that own or maintain proprietary industrial data, customer workflows, embedded software, enterprise integration capacity and distribution channels may be able to claim economic returns that aren't readily reducible to the raw compute or frontier model. Evidence on national AI ecosystems indicates that economies without leading positions in compute or frontier models may nevertheless specialize in downstream applications, software and distribution. This is an important antidote to data-center-tax narratives. Even if the most intensive compute activities are thousands of kilometers away, a significant portion of the value that can be effectively appropriated may be claimed by those firms that translate the models to the relevant markets, own the relevant workflow, or maintain the route to the customer. The seventh layer is the adopting firm and its final customers.[14] Economically, not all of the value of AI is captured as taxable profit by upstream digital companies. Some of the value is transferred down as lower prices, better/cheaper products, better quality, less downtime, faster design cycles, higher yields, or newer output in the adopting company. Some accrue to end consumers. Some is earned by low-cost producers of an imported model, on an imported cloud, who build a factory, hospital, supply chain, or professional service that is more productive. So the largest social or private gain from AI adoption can potentially be realized outside the data center and outside the model vendor company. The AI value chain, thus, does not end when an individual facility powers up a GPU cluster; economically, it ends where the adoption shifts output, margins, quality and market power.

In principle, these seven levels can also see their rents shift over time.[15] During times of chip shortage, rents may accrue to upstream suppliers of design, foundries and memory. When power-rather than efficiency-becomes the critical bottleneck, tremendous rents could go to the few rare grid access providers or the handful of players capable of quick energization. When cloud monopolies tighten their grip in the coordination layer, tremendous rents might eventually go to those few providers who link platforms together and those firms that lock in entire ecosystems of users. Where model interface measures become commoditized and where open alternatives flourish, these rents may migrate to all but the earliest stages of industrial applications, to industrial data, or to the distribution layer. Where adoption diffuses to all firms, the distribution of rents may shift again toward those firms that best reorganize production itself. The available evidence therefore, suggests that policy should not assume that current rent-holders will necessarily hold power forever because their facility is eventually heavily taxed and therefore fixed in place.

3. Six Geographies of AI Value: Compute, Energy, Labor, Intellectual Property, Corporate Structure and Markets

Since the AI value stack is layered, the production of AI services and products also occurs across multiple geographies that are only partially overlapping.[16] Compute geography specifies site of server and accelerator installation and job execution; energy geography specifies site of electric power and grid capacity, backup, cooling and transmission reinforcement; labor geography identifies where model engineers, site of facility operators, where associated professionals work, site of people working with the output; intellectual-property geography identifies wher ownership of models, software, patents, datasets, or licenses; corporate geography specifies the location of the parent and subsidiary contracting, invoicing, funding and profit-reporting site; market geography specifies the location of consumers and end users of AI products and services. The location of the economic value stems not from any one of these, but from the interaction of all six. Compute geography and energy geography are not the same thing, even on the same site. The IEA finds that the siting and expansion of data centers is highly dependent upon local grid variables, queue times and capacity to bring large loads online.[17] The OECD finds that the availability of water and cooling can influence siting decisions. A Finnish or Swedish site may be selected because its temperature, electrical mix, water availability and permitting process will enable it to support inference or training loads efficiently. Such a choice says a lot about local infrastructure scarcity and local system costs, but it does not mean that Finland and Sweden become the sole source of the firm's model, customer base, or reported profits.

Labor geography and market geography often tend to differ even more from compute geography. The team working on the initial development of the model might be based in California, London, Paris, or Berlin. The operations team administering and updating the site might be based in Finland. The customer success team might be based in Dublin or Warsaw. The engineers actually using the system might come from Bavaria and the head of production might come from Poland. The organization paying for the service might be a German manufacturer, but its customers will be spread across the rest of the European Union. In terms of taxation, these differences matter because production, consumption and business activity can all stay within the jurisdiction withoutsubstantial local compute capacity. In terms of economics, they matter because the productivity effects of artificial intelligence are powered by organizing a firm once it adopts the technology – not by the servers in the data center being cooled down.

Another level of stratification was created between intellectual-property geography and corporate geography. In many cases where commercial AI services are sold into the EEA, the services may be contracted through Irish entities with respect to the service, even if the technology is created elsewhere. OpenAI's European terms state that when European EEA and Switzerland residents seek to contract with OpenAI, customers that reside within the EEA or Switzerland shall enter into the terms of the agreement with OpenAI Ireland Ltd; whereas the contractor-entity schedule for Google Cloud states that for much of EMEA, the contracting-entity name is Google Cloud EMEA Limited, stated in Dublin.[18] The mere presence of such foreign locations (for services, contracting entities, or even ultimate parent companies does not necessarily reflect where profits are ultimately taxed and potentially where profits will be taxed, even if no activity is directly linked to holding that specific location), but the location of an Irish contracting entity remains a possibility for a European customer where the model is created in the US, run on servers within another Member State, and at the same time, under an ultimately U.S. parent company. For representative factors such as who is legally invoiced, where the server is located, where the ultimate parent company is located, and who is the beneficial owner, the distinctions are difficult to discern.

The hypothetical commissioned by a German manufacturer, using an American model, via a contracting Irish entity, running inference in a Finnish data center, on chips designed in the United States, fabricated in Taiwan, applying German industrial data to improve the output of a Polish factory but still owned by a parent corporation in Germany, is economically and legally plausible. Each element is a described pattern elsewhere in the current AI economy. Google publicizes Hamina in Finland as one of its live data-center locations.[19] OpenAI and Google both contract some of their European work through Irish entities. Publications by the OECD demonstrate that chip design, fabrication, cloud infrastructure, models, and applications are themselves geographically distinct and frequently located in separate jurisdictions. The real legal point is not to claim the example as exhaustive or universal, but rather, to demonstrate that it is theoretically, legally, and economically feasible for the one unit of AI-enabled value to pass through no fewer than six jurisdictions before taxing rights must be allocated among jurisdictions.

Once that multi-jurisdictional chain is recognized, the boundaries of a facility tax are clear.[20] An Irish or German tax on the value generated in a data center in Ireland or Germany might be justified to the extent that the data centers are the means by which land-use authorizations, access to the grid, cooling infrastructure, ecological regulation, connected transport and communication links, and the public services and streets of a neighborhood are supplied. An Irish, German, or other national taxing the profits generated by servers within a particular permanent establishment or facility business is probably justified to the extent that they are collecting from the operation of the servers taxes on the resources used by the servers, be they the services of a Polish, U.S., Taiwanese, or other chip designer or the running costs of an Irish contract manufacturer. But the home country of the server is not in and of itself in a position to determine what part of the overall AI surplus should go to the Taiwanese foundry, the Irish contractor, the German consortium, the American chip home, and the European end user.

4. What Data-Centre Host Jurisdictions Can Legitimately Tax and Charge

By untangling the broader geography of value in AI, the fiscal claim of host jurisdictions is not any weaker but rather clearer. There is a compelling normative argument for a data-center host to charge for what it supplies in practice, and for adjusting for the additional burdens that the facility places on the local systems.[21] These norms are familiar: benefit taxation, user charging, recovering costs, and Pigovian correction. If a data center occupies scarce land, necessitates rezoning, consumes public infrastructure, and its operation strains constrained grid capacity, the host jurisdiction can legitimately charge to recoup the costs. This claim carries its greatest force when it is implementation-neutral, can be relatively quantified, and is defined in relation to the immediate, local facility rather than certain purported rights over the rents of the distant AI stack. That logic clearly leans towards natural land-use controls, property taxation, building- and civil engineering-related charges, and local authority permitting fees. A server hall is a building on land. It may necessitate road access, public safety supervision, flood or fire planning, and utility cooperation. The IEA points out that data centers can become operational in between 2 and 3 years, which on its own indicates a significant and concentrated development process, even if the broader energy system required to support them proceeds at a slower rate. There is therefore nothing conceptually unusual about considering the facility in its initial form to be no different than other large industrial or warehousing assets when it comes to land, building, and property taxation. These are not levies on AI rent in the abstract; they are levies on land-use and built environment infrastructure on a site-specific capital asset.

Grid-connection costs, transmission and distribution upgrades, and capacity-reservation charges are defensible if they correspond to system costs.[22] One must make a special case here because data centers are capable of creating system costs that extend well beyond their fenced perimeter. For instance, IEA analysis of Europe shows waiting times for grid connections averaging over two years, core European hubs experiencing seven-to ten-year queues, and the European Union needing some policy coordination simply to match project pipelines with existing electricity infrastructure.[23] ACER reports that in 2024 EU transmission system operators spent 4.3 billion on various remedial actions to address grid congestion.[24] If a facility needs a dedicated network line, a new substation, flexible connection arrangements, or network reinforcement, then the host system is not just hosting a digital service but securing limited grid capacity. Costing that scarcity with market forces is internally consistent and can help guard against other ratepayers being burdened by implicit cross-subsidy.

Figure 3. Rising electricity demand strengthens the case for local infrastructure charges, but not for assigning the entire AI rent to the host jurisdiction.

Water use, cooling, carbon emissions, externalities of backup generation, and environmental monitoring are other justifiable local compensations. OECD analysis emphasizes that many AI data centers are water-dependent in cooling requirements and that water access is a cause of location; it also concludes that cooling technologies are changing, so local policy should focus on actual resource use rather than a blanket anti-data-center approach. The EU's second-generation data-center energy-performance framework has explicitly included monitoring of water usage once the facility becomes operational. Carbon or local air-pollution taxes are equally justifiable if on-site generation or backup-generator testing, or local particulate matter emissions impinge on the inflowing community. Again, the relevant tax or charging base is not the global AI but the physically present load and its measurable externalities. Ireland is an example of how such a claim can be meaningful. Central Statistics Office Ireland has stated that data centers were responsible for 5 percent of metered electricity consumption in 2015, 22 percent in 2024 and 23 percent in 2025.[25] At these levels, the host-country validity of concerns over the adequacy of the grid, capacity planning, limited locations, and fair distribution should be clear. Such concerns are clearly not merely rhetorical; they are indicative of how a significant proportion of an actual national electricity system should be allocated. Slightly smaller-scale data-center concentrations could cause similar host country concerns in other parts of the world. This justifies the prediction that the charging of infrastructure and pollution costs is not merely a pretext for taxing intangible profit flows, but an actual attempt to deal with the substantial, geographically distributed consequences of digital loads that require substantial infrastructure.

Figure 4. Ireland’s rising data-centre share of metered electricity demonstrates the scale of a legitimate host-country infrastructure claim without locating the entire AI surplus there.

European law and policy now increasingly mirror this logic of host jurisdiction precisely. The recast Energy Efficiency Directive requires Member States to cover reporting by significant data centers, and the Commission's 2024 delegated regulation established the database and KPIs for that reporting.[26] The Commission's data-center performance page states that the database collects information on energy performance and water footprint for sites with significant energy consumption, and earlier Commission documents established the reporting threshold at built-in (IT) power consumption of at least 500 kW.[27] The very same framework makes clear that the legal regime in place, not yet evolved into an EU tax instrument, is essentially a transparency and monitoring one. It is not a mature EU tax instrument even now, and no attempt should be made to categorize it as such.

As of mid-2026, the Commission is preparing a Data Center Energy Efficiency Package that would include baseline data collection, a rating scheme, and an electronic label, followed by work on minimum performance standards.[28] The Commission's March 2026 call for evidence outlined the proposed rating regulation as a follow-up to the 2024 delegated regulation and the Commission's broader energy efficiency roadmap, which identified minimum standards in a consultation phase.[29] In other words, EU reporting and collecting obligations are in place, but the rating-label regime and any mandatory minimum performance standards are a work in progress, not a finalized, fully functioning tariff and regulatory scheme. This distinction is central to accurate analysis. Existing standards and labels could influence costs and investment signals, but they should not be conflated with implemented fiscal measures or an EU decision to extract rents from global AI by taxing server venues.

Figure 5. The EU framework has advanced from reporting toward rating and prospective standards, not toward an implemented tax on global AI rent.

The correct policy position, therefore, becomes the disciplined rather than anti-host one. Data-center hosts could legitimately charge for land, property, building, and connection, congestion, environment, monitoring, and facility-flexibility profits, impose locational conditions, and negotiate obligations of system-integration (such as use of waste-heat or flexible loads) where appropriate. However, these charges should be aimed at as site-specific public financing tools, not policy rhetoric used as an improper substitute for taxing the whole of AI. After the financial purpose is changed from recovering local costs to capturing the entire surplus of global AI, this data center is no longer an appropriate tax base, and a comparatively inappropriate substitute.

5. Why Physical Server Location Is an Incomplete Proxy for AI Rent

The most significant logical error in the data-center tax proposal is a category error. A data centre is a physical input into AI, but it cannot be the only residual claimant on the model, be the sole owner of a customer relationship in AI, or stand to gain from the returns to AI adoption. In many commercial arrangements, it is a cost center, an infrastructure asset, or an inception node. The rents that really do matter may not be sitting on a facility but on a cloud platform distributing finite compute time among users, on a model supplier turning a boundary API into revenue, on a software company leveraging AI in the enterprise's processes, or on an in-house company sharpening its margins and output.

This is further underscored by demand heterogeneity. Data centers are not solely for AI use, and even the abrupt current spate of activity in AI does not change the multipurpose character of the infrastructure it is built upon.[30] The IEA estimates that accelerators (of which the majority are AI accelerators) constitute nearly half of the net growth in global data-center electricity use by 2030; conventional data-center servers still account for just under a fifth, and the remaining part is accounted for by data-center infrastructure and other IT equipment.[31] All three main types of data centers: enterprise, colocation, and server provider, and hyperscale (including AI-driven workloads of any type) contribute to the overall demand increase. Any facility-level tax base attracted by AI rent will therefore almost certainly appear as some combination of AI and non-AI workload within a mixed-use facility (absent a transparent way of individualizing them in meters, which the public debate does not seem to assume); the potential end result could be a tax on non-labor digital infrastructure with only partial links to frontier-AI business phenomena.

Training and inference also differ in ways that weaken the argument for a simple facility proxy.[32] OECD analysis describes training as extremely expensive, requiring high-density clusters, specialized accelerators, high-bandwidth networking, power, and cooling. Inference inevitably also consumes hardware optimized for energy efficiency (or latency). But as services reach the point of commercial scale, they harness device-to-edge latency effects, and cloud providers also benefit from developing optimum local or regional infrastructure, for example, for security reasons. Carnegie's analysis of data-center competitiveness shows that for major compute investments, time to power, project speed, and the availability of infrastructure may dominate over even huge differentials in local taxes or electricity prices.[33] In fact, these two facts point in different local directions. Frontier-model training may often be best concentrated where power, chips, financing, and permitting are all readily available. Inference, by contrast, needs to be assembled along the geographies where latency, customer proximity, security, and institutional conditions are optimum. A single, uniform facility tax may therefore not be adequate to capture the value spread across all AI workloads.

Legal complications extend the physical/taxable location divide. The OECD Model Tax Convention supports three propositions that are relevant here: a website by itself is not a place of business, a hosting arrangement will generally not involve a server being made available to an enterprise, and even where a server or related facility does constitute a permanent establishment, that alone will not determine the division of profit attributable to it. Further, the business conducted on the server must be analyzed on a case-by-case basis; attribution is a separate determination. Modern cloud architectures heighten this importance because a cloud client accessing AI may never own, lease, or have managerial control over the specific server involved in generating a model response. The transfer-pricing rules introduce a second form of incompleteness. The 2022 Transfer Pricing Guidelines issued by the OECD are explicit that legal ownership of an intangible asset does not confer the right to a 100 percent share of its returns: what counts is who does, supervises, and bears the risks for development, enhancement, maintenance, protection, and exploitation. If the foreign-owned affiliate holding the legal title to a platform, portfolio of patents, or software stack doesn’t control or perform those functions, that affiliate shouldn’t be guaranteed to get all the income flowing from those activities. The other side of this coin is equally significant for the facility argument: a data-center company that supplies power, space and operations but doesn’t own or develop the intangible shouldn’t automatically be where the model rent is located. All manner of revenue flows-royalties, license fees, cloud charges, cost-sharing arrangements, services agreements-can be used to redistribute income from the physical location, subject to arm’s-length and treaty limitations.

A blunt facility charge can be distortionary because, rather than falling on the parties best able to bear or pass on it, it targets a plausible global bottleneck. OECD evidence demonstrates that the most substantial AI and cloud companies have scale, scope, vertical integration, and multi-layer presence. Firms active across compute, infrastructure, models and applications are generally those best able to push workloads to others, repackage contracts, secure high-powered deals or bear higher facility costs for long periods. These advantages, which are unequally distributed within and across communities of local and global research users, universities, start-ups, and incumbent providers, will come into play in the event of a facility tax and may prevent it from falling solely on the incumbent rent recipient. Private providers, domestic research organizations, universities, and local start-up providers will be less able to pass that fee through to their own constituents. So targeting the rent through a facility tax may be costly even if the tax is invoked rhetorically.

The difference between operational and announced capacity aggravates the problem.[34] Many public conversations about AI data-center booms often move from project announcements into assumptions about realized rent. However, the IEA Europe analysis explicitly distinguishes between pipeline capacity and installed capacity. It reports that the project pipeline in Europe implies a capacity of 130 percent of the current installed capacity, but that installed capacity by 2030 will increase by only about 70 percent in relation to 2024, due to delays and local restrictions (for example, on emissions or controls on end-user access due to grid congestion). In other words, levies or obligations related to announced capacity, memoranda of understanding, or speculative AI-hub announcements could be put in place long before the manufacturing activity is capital-intensive, operational, profitable, or yet to be built. For an already tenuous attempt to tax emerging AI rents, such a timing mismatch makes the proxy even less robust. On the whole, the takeaway is not that taxation of facilities is misguided-it is just a more limited tool than most supporters believe. It is effective at valuing local externalities and taxing the operations of facilities. It is ineffective at embodying the putative makeup of AI-generated rent when the chips, the models, the contracts, the data, the producers, the adopters, and the shareholders are frequently offsite. The farther away from local cost return a policy goes, and the closer it gets to claiming a share of AI's total economic surplus, the more likely it is to be invalid analytically and economically distortionary.

Figure 6. The broad projection range demonstrates why prospective capacity should not be treated as already operational or profitable.
6. Europe’s Policy Trade-Off: Cost Recovery, Compute Expansion and the Risk of Taxing the Wrong Base

The European policy risk results directly from the above analysis.[35] Europe is now pursuing two objectives. On one side, it is tightening up data centers' energy efficiency, use, and integration into systems. On the other hand, it is explicitly trying to grow its domestic cloud and compute capacity, as part of a new overall agenda for AI and technology-sovereignty. The Commission's AI Continent plan envisages tripling EU data-center capacity at least within five to seven years, and its proposal for a Cloud and AI Development Act aims at easing deployment and access to energy, land, water, and finance, towards building up strategic capacity for AI, cloud, and computing-intensive applications.[36] Such a policy mix does not sound inconsistent at face value. It can only sound inconsistent if Europe is going to regard the data center equally as a limited factor of production that must be grown and as a fictional tax base convenient for extracting the entire rent of AI.

The pain is genuine because the pain of Europe's energy constraints is genuine. The IEA finds EU grid-connection waits that can range from two to ten years, and concludes that Europe will not realize the full volume signaled by its project pipeline under current constraints. Data centers are also expected to contribute 10 percent to EU electricity-demand growth by 2030 in current policy settings.[37] In that setting, cost-reflective local levies are not a policy indulgence; they are a necessity both for economic management and public support. Host communities and network users cannot reasonably be expected to freely carry additional congestion, capacity shortages, water use, or system-upgrade costs. The mistake is elsewhere: in stacking on non-cost-reflective fiscal loads to imitate an AI rent tax, and simultaneously expecting a faster rollout of private digital infrastructure in a region already constrained in the squeeze between time to power.

Figure 7. Differences in deployment time can materially affect where new compute capacity is built.

These goals can, however, be partially reconciled. The first step is an honest assessment of the mission of each instrument. Cost-reflective connection fees, capacity-reservation charges, water charges, carbon tax, and environmental reporting covered local costs and externalities, whereas easier permitting, queue management, flexible interconnection design, and better planning addressed deployment bottlenecks. The IEA states that non-firm connections and queue management would provide a way to upgrade ready-to-build projects; and the Commission is developing so-called tripartite models between data center operators, other energy players, and public authorities precisely to make data centers an integrated part of the local energy system & heat networks. A Europe that makes data centers pay for the costs they impose and also reduces unnecessary permitting and grid frictions is not contradicting itself; a Europe that seeks to morph facility charges into a new, lower global AI profit surcharge does contradict itself.

The idea that certain frontier-model training might take place in the Middle East and other resource-plentiful regions, while inference occurs nearer to European endpoints, should thus be put to the test rather than simply accepted or dogmatically thrown out. Carnegie's cross-national competitiveness index places the UAE high in part because of project speed, and because prolonged delay dominates other cost factors; but it also illustrates how political, economic, and security shocks, or conflict-driven delay, can cause rapid oscillations in such rankings. The economics of locating some training in power-providing and fast-permitting economies and geographies is sound. Large training loads are capital-dependent, schedule-sensitive, and more easily relocated than many policymakers appreciate. But that alone is not necessarily enough of a cure to Europe's need to bring certain inference workloads loser to EU users, or to preserve some strategically necessary compute capacity, either within or beyond the EU's borders, for resilience, security, and public sector control; or to realize the EU digital-ecosystem objectives that cannot be achieved simply by externalizing certain capability-creating activities.

Figure 8. For the illustrative facility, deployment delay produces a larger modeled value loss than higher electricity prices or the removal of tax incentives.

Both of the strongest potential counterarguments to the paper's thesis, therefore, warrant serious consideration. Data centers have the potential to produce significant investments, local construction expenditure, utility requirements, local government receipts, resilience gains, and ecosystem externalities; none of these arguments necessitates the assumption that the data center captures all of the rent accruing to AI, and none of them would be diminished if the Commission moved to regulate, tax, or subsidize facilities at the least-cost jurisdiction. The danger would come if one slides from those sensible host-jurisdiction arguments into the further conclusion that a tax on the host facility is therefore the best way of taxing AI.

This is also where the conceptual difference between facility taxation and market-jurisdiction taxation becomes essential. The 2026 briefing document on cross-border services taxation produced by the European Parliament states that digitalization means that services can be supplied from afar at near-zero additional costs, and describes one policy lever as trying to reverse the balance of taxation taxing rights – nations where the services are supplied to global markets instead of where the services come from.[38] But whether or not one favors any digital-levy measure, this particular debate involves a different kind of tax entity, a different tax object entirely from a local charge on a data-centre facility. One involves global customers and globally sourced services. The other involves a locatable infrastructure asset. Mixing the two leads to bad policy in compelling a local utility-and-environment measure to do the job that a global profit allocation inquiry should.

The simplest and most justifiable European stance is therefore a tiered one. Taxes and charges on AI facility use should be narrowly correlated with local costs, congestion externalities, and environmental impact, and levied universally among all workloads, inclusive of AI, proportionately to their demands. Any policy to expand AI compute capacity should be targeted at shortening power-up times, integrating more additions to national grids, siting capacity where there is genuinely existing headroom, and protecting critical inference and public-interest AI computing within the Union and allied safe havens. Where broader AI profits and rents are taxed, the taxation of broader AI profits and rents should be directed at taxing those profits and rents rather than relying on the mistaken assumption that the full equation of taxable surpluses resides wherever a server rack is sited, as long as it happens to be getting the power from a server rack. Europe's mistake would not be if data centers were made to pay their way locally. It would be treating that local payment as a substitute for taxing AI value where it is actually earned.

7. Conclusion - Tax the Facility for Local Costs, Not the Entire AI Value Chain

The case for taxing data centers starts from a solid intuition and would end up in a narrower conclusion than many of its proponents seem to desire. Data centers are fixed, resource-use-heavy sites that impose tangible burdens on land, grids, cooling, water, and local public services; the home jurisdiction can thus fairly seek to recover local expenditures, impose local externality prices, and tax the profits that are truly generated there by data center operation. It does not contain an unearned claim to the full AI rent in the form of the full economic rent, simply on the grounds of having a few servers stacked within its geographical jurisdiction. AI's value is spread across chips, clouds, model-builders, software system-stacks, proprietary data, user firms, and final customers: each with different geographies, jurisdictions, and legal considerations. A server hall is a single essential point on an AI value chain, but by no means is it the natural geographical location of that value.

The strategic implication is unambiguous. By using facility taxation as an approximation of AI rent taxation, Europe (or another state) may be collecting too much tax on a relatively immobile taxable input, and too little on more mobile, legally complicated forms of ability-to-pay and value creation. The right way forward is institutional separation. Local facility-specific charges to reflect relative resource costs, taxed in a tightly calibrated manner at the facility level, and a separate global-policy debate to attribute other profits, intangible assets, and jurisdictional rights. The priority is not the abolition of data center taxation, but the restoration of its proper function.


This article was prepared as an independent research contribution following Professor Keith Lee’s presentation at the conference Inequalities in Longevity, held at Fondazione Giorgio Cini in Venice on 3–4 July 2026, and a subsequent substantive discussion with Federico Fubini of Corriere della Sera. The discussion helped sharpen the motivating questions concerning AI-driven productivity, the declining employment intensity of economic growth, and the resulting pressures on the public tax base.

The research, analysis, and conclusions were developed independently by the author. This publication is separate from the official conference proceedings and from the editorial coverage published by Corriere della Sera.

Unless expressly stated otherwise, this publication has not been commissioned, reviewed, or endorsed by Fondazione Giorgio Cini or Corriere della Sera. References to the conference and the subsequent discussion describe the intellectual context in which the research developed and do not imply co-authorship, institutional affiliation, formal collaboration, or partnership. The analysis, interpretations, and conclusions are those of the author and do not necessarily reflect the official positions of Fondazione Giorgio Cini, Corriere della Sera, the Swiss Institute of Artificial Intelligence (SIAI), or their respective affiliates.


References

[1, 2, 3, 7, 12, 15, 16, 32] OECD (2025) Competition in Artificial Intelligence Infrastructure. OECD Roundtables on Competition Policy Papers, No. 330. Paris: OECD Publishing.

[4, 5, 20] OECD (2017) Model Tax Convention on Income and on Capital: Condensed Version 2017. Paris: OECD Publishing.

[6] OECD (2022) OECD Transfer Pricing Guidelines for Multinational Enterprises and Tax Administrations 2022. Paris: OECD Publishing.

[8] Sastry, G., Heim, L., Belfield, H., Anderljung, M. et al. (2024) ‘Computing Power and the Governance of Artificial Intelligence’. arXiv, 2402.08797.

[9] Pilz, K. and Heim, L. (2023) ‘Compute at Scale: A Broad Investigation into the Data Center Industry’. arXiv, 2311.02651.

[10, 27, 28, 30] European Commission (2026) Energy Performance of Data Centres. Brussels: Directorate-General for Energy.

[11, 22, 23, 31, 34, 37] International Energy Agency (2026) Key Questions on Energy and AI. Paris: IEA.

[13] Stanford Institute for Human-Centered Artificial Intelligence (2026) Artificial Intelligence Index Report 2026. Stanford, CA: Stanford University.

[14] International Monetary Fund (2024) Broadening the Gains from Generative AI: The Role of Fiscal Policies. Washington, DC: International Monetary Fund.

[17] International Energy Agency (2026) ‘Data Centre Electricity Use Surged in 2025, Even with Tightening Bottlenecks Driving a Scramble for Solutions’. Press release, 16 April.

[18] OpenAI (2026) Europe Terms of Use. Updated 16 January 2026.

[18] Google Cloud (2026) Google Contracting Entity. Google Cloud.

[19] Google (n.d.) Hamina, Finland — Google Data Center Location. Google Data Centers.

[21] OECD (2019) Taxing Energy Use 2019: Using Taxes for Climate Action. Paris: OECD Publishing.

[24] European Union Agency for the Cooperation of Energy Regulators (2025) Market Monitoring Report: Electricity Wholesale Markets. Ljubljana: ACER.

[25] Central Statistics Office Ireland (2026) Data Centres Metered Electricity Consumption 2025. Cork: Central Statistics Office.

[26] European Parliament and Council (2023) ‘Directive (EU) 2023/1791 of 13 September 2023 on Energy Efficiency and Amending Regulation (EU) 2023/955’. Official Journal of the European Union, L 231, pp. 1–111.

[26] European Commission (2024) ‘Commission Delegated Regulation (EU) 2024/1364 of 14 March 2024 on the First Phase of the Establishment of a Common Union Rating Scheme for Data Centres’. Official Journal of the European Union, L 2024/1364.

[29] European Commission (2026) ‘Rating Scheme for Data Centres in the EU — Commission Launches Call for Feedback’. Directorate-General for Energy, 27 March.

[33] Phillips-Robins, A., Tawil, T. and Winter-Levy, S. (2026) The Compute Coalition: How to Build the Future of AI in the Free World. Washington, DC: Carnegie Endowment for International Peace.

[35, 36] European Commission (2025) The AI Continent Action Plan. Brussels: Directorate-General for Communications Networks, Content and Technology.

[36] European Commission (2026) Cloud and AI Development Act. Brussels: Directorate-General for Communications Networks, Content and Technology.

[38] Amaro, F. and Picciotto, S. (2026) Possible EU Own Resource Based on a Digital Levy: Cross-Border Services Trade, Digital Transformation and Tax Implications. Brussels: European Parliament.

[AI and Tax] Labor Income, AI Rents and Fiscal Erosion

[AI and Tax] Labor Income, AI Rents and Fiscal Erosion

Keith Lee*

*Swiss Institute of Artificial Intelligence, Chaltenbodenstrasse 26, 8834 Schindellegi, Schwyz, Switzerland

Abstract

Artificial intelligence does not mechanically erode the tax base; the fiscal impact depends upon how the productivity gains are shared and where the income generated is taxed. This paper studies AI as a productivity shock, an income-distribution shock and a fiscal-transmission shock. It separates task automation from job displacement, firm productivity from macro employment effects and ordinary business returns from AI rents. The evidence suggests that generative AI can create productivity gains in selected tasks, particularly for lower-skilled workers,but has not yet produced broad employment displacement. Nevertheless, it is uncertain whether output resulting from such gains will always offset labor substitution in sectors with weak demand growth, in concentrated markets or where complementary investments are lacking. Accordingly, the fiscal impact will rest on whether such gains are ultimately paid to workers, consumers, domestic firms, investors, or foreign AI suppliers. Countries that rely heavily on labor taxes and social contributions, import substantial AI services or tax capital income weakly may be most affected. The most pressing policy dilemma is not in taxing AI as a technology, but in safeguarding the domestic tax base as digitalization shifts income from broad, immediately taxed payroll remuneration to more concentrated, mobile, or deferred forms of income.
 


[AI and Tax] is an independent research series developed by Professor Keith Lee following his presentation at the conference Inequalities in Longevity, held at Fondazione Giorgio Cini in Venice on 3–4 July 2026, and a subsequent substantive discussion with Federico Fubini of Corriere della Sera concerning AI-driven productivity, labour income, and fiscal capacity.


Public-facing adaptation

A public-facing adaptation of this research, written by the author for a wider readership, is available from The Economy Review: “[AI and Tax] Why AI Fiscal Erosion Begins Before Jobs Disappear


1. Introduction - Does AI Replace Tasks or Complement Workers?

AI can achieve superhuman performance on selected tasks, but it cannot create a superhuman worker. For example, modern classification, summarising, translating, editing, compiling, searching, document comparison, pattern recognition and generation technologies can perform certain activities at a speed that is significantly faster than unassisted humans. Many processes that took days may now take hours. But the largest demonstrations of task acceleration can hardly be taken as evidence that the overall productivity of a professional, such as a lawyer, has been multiplied ten times. The agent’s job may again involve problem framing, fact checking, forecasting, communication, coordination and responsibility for error. AI may accelerate one step of an otherwise unchanged procedure or a new requirement for review or oversight may be introduced.

The strongest experimental and workplace evidence points to material but bounded productivity improvements. In a randomized trial of professional writing tasks, providing generative AI access cut average task duration by 40 percent and boosted average assessed quality by 18 percent.[1] A study of customer-support agents shows that an AI assistant resulted in about 14 percent more issues resolved per hour on average, with substantially larger improvements among relatively inexperienced and initially lower-performing workers.[2] In a randomized trial of management consultants, AI usage led to more task completions, hastened completion speed and better answers when assignments hovered within the model's capability frontier, but to poorer answers on some tasks designed to be at least partly outside of that frontier.[3] The empirical case thus tilts towards selective amplification rather than a universal productivity multiple across intellectual work.

This distinction is important because the work task is often conflated with the entire job. A model could generate a draft version that automates but does not displace the task of a worker who defines the goal, checks the output and assumes professional responsibility. A worker can complete a routine job analysis faster but needs more time to do unusual analyses or advise customers. A company can output less labor content to accomplish one internal procedure, while increasing sales, quality, or new services. An economy can eliminate some tasks and still create new jobs in implementation, verification, data stewardship and customer service.The International Labor Organization's refined global index of occupational exposure underscores this gap between predicted exposure and employment displacement. It estimates that approximately 25 percent of global jobs are exposed to some degree of generative AI, with only 3.3 percent experiencing the highest levels of exposure.[4] Exposure level is highest in high-income economies (34 percent) but still low in the lowest-income countries at 11 percent[5] due to the preponderance of clerical, professional and digitized service work in the global North. Clerical jobs are the most exposed, but the International Labor Organization finds this unlikely to lead to complete job replacement, as most occupations blend activities subject to augmentation or automation with other activities requiring human input nonetheless. Measures of exposure are thus indicators of the technical feasibility of automation should it be adopted, not the probability of unemployment, nor a timeline for displacement.[6]

Macroeconomic estimates are closer to the more moderate of the task-level results. Acemoglu's task-based study, for which he approximates the effects of presently available AI applications, reports them as being potentially substantial but, in aggregate terms, modest over ten years, with total-factor-productivity improvements below the more revolutionary estimates typically ascribed to generative AI.[7] Early labor-market evidence has also failed to justify a prediction of employment losses across the economy. Administrative evidence from Denmark, tracking exposed occupations and thousands of workplaces, found no statistically significant net effect on earnings or recorded hours in the first two years of widespread chatbot adoption.[8] That survey nonetheless documents several new tasks associated with implementation, oversight, integration, ethics and compliance. An alternative analysis of occupational task exposure indicates that increased mean AI exposure in affected occupations can decrease labor demand in the occupations themselves, whereas only mixing AI concentrations over a subset of a job’s tasks allows workers to reallocate effort, while enough productivity growth at adopting firms to counteract its direct substitution is somewhat, yet not fully, offset by scale effects, yielding much more modest total employment changes than one would expect based solely on simulated task exposure effects.[9]

AI should therefore be interpreted as comprising three shocks at once. It is a productivity shock, in the sense that it may lessen the human effort and time necessary to undertake some particular types of activity. It is an income-distribution shock, in that the generated surplus will flow to one or another of labor, consumers, adopting firms, technology suppliers, investors, or scarce complementary employees. It is a fiscal-transmission shock, in so far as such recipients will be taxed through various mechanisms, at various effective rates, at different times and in different jurisdictions. Personal income tax, social contributions and payroll taxes are collected relatively quickly, whereas retained profits and unrealized capital gains may find their way into the domestic fiscal system more distantly. Hence, the central case should be put to the test and not taken for granted. The increased fiscal erosion will be more likely if substitution effects outweigh augmentation and scale expansion, if output increases are too small, if wage pass-through is weak, if rents are highly concentrated and if a large share of the technology payment or the proceeds of technology ownership arise outside the country that is home to the workers and customers. Gross domestic product is not equivalent to domestic fiscal capture.

2. Do Firms Expand Output Enough to Offset Labour Displacement?

Given a level of output, successful automation will require less labor for a given output. This accounting relation is commonly cited as a prediction of job destruction, though it describes only the initial phase of adjustment. Cheaper unit costs can be used to lower prices, increase quality and delivery speed, raise margins, or free up resources for new products and if demand grows enough, a firm may still be able to produce more goods with the same number of jobs, or raise employment while still reducing the labor hours needed per unit. The ability of expanded scale to offset displacements will depend on the elasticity of demand, competition, availability of capital and organizational capacity and on how much of the production process AI is capable of improving.

The relevant level of analysis is the task. An automated task can be defined as a task for which an AI-enabled workflow can produce a sufficiently acceptable output with little or no additional human input. An advantaged task involves the use of AI in a way that extends the speed, scope, or quality of work performed by a worker. An affected task is one in which there is some amount of partial automation, but which produces additional work in areas such as verification, communication, documentation, or other transaction costs. Jobs are constituted with some mixture of such task categories.[10] Firms use compositions of job task bundles that vary substantially. An accounting clerk sending data to or from a standardized system, such as VAT returns, might be reduced to ordinary substitution; by contrast, an accountant using the system to analyze complex transactions might receive means of augmentation.

Figure 1. The highest-exposure gender gap more than doubles in high-income economies.

Firm-level European evidence supports the possibility of expansion. Akin to a finding reported in the European Investment Bank’s 2026 assessment of about 13,000 firms in the European Union and the United States, where AI adoption in European firms is associated with stronger productivity performance without establishing a uniform employment effect[11] and the strongest benefits seem to accrue to medium-sized and large firms that can combine AI with investments in software, information and skills and organizational capacity.[12] While not being able to establish that all adopting firms expand and that all occupations are protected in all adopting firms, these results can hardly be interpreted as contradicting the idea that the direct labor effect of AI can be offset at the level of the firm when the investments in complementary factor inputs and the output growth are sufficiently large.

A firm may grow, even if industry employment falls. It can happen if more productive early adopters take market share from other poor performers, raising their employment but reducing sectoral labor demand; or if the cost savings give rise to firm-specific new demand by boosting output, thus raising employment in the whole sector. Hampole et al. assist in distinguishing these processes. Their task-based proxies indicate that occupations in which the average exposure to AI/machine-learning abilities is higher suffer lower relative labor demand.[13] However, this effect is mitigated when only some of an occupation's task package is exposed, as workers can compensate by intensifying other, less exposed activities. At the firm level, demand for labor related to increased productivity tends to compensate for much of the reduction in demand caused by the exposure of existing tasks.[14] The overall impact thus remains relatively weak vis-à-vis the occupational substitution effect. This does not imply that all former employees will be retained, since expansion might draw more workers into other occupations, plants, or skills. It demonstrates that automation in tasks does not equate with total employment.

Figure 2. The refined framework places fewer jobs in intermediate exposure categories while slightly expanding the highest-exposure group.

The scale response is sector-specific. In software, consulting, marketing and some professional services, lower production costs can enable doing more projects because the incremental cost of serving one more customer is low and there might be a significant latent demand. In the private and public healthcare, education and administration sectors, it won’t necessarily generate commercial sales but will accelerate throughput or improve quality. The fiscal benefit may appear by decreasing waiting time or allowing the same staff to serve more users. On the other hand, relatively fixed demand seems to apply to long-established back-office functions. Consequently, double the number of reconciliations, reports, or activities would not be realized just because AI enables them to be cheaper and faster. How the increased saving translates into expansion depends on the level of contestability. The firm, in a competitive market, is willing to cut prices or increase quality to the extent that it keeps its market share, which would increase demand and share some of this gain with consumers. If the market power is high, the savings are reflected in a bigger margin, thus curbing the scale effect. Thus, employment effects may be stronger in contestable markets and weaker in concentrated ones.

The market structure influences the share of static surplus, which goes to the innovator, customers, or the upstream technology supplier. The pattern of adjustment makes interpretation less straightforward. In the short run, firms frequently implement AI within current procedures. Employment contracts, uncertainty regarding precision, legal standards and organizational inertia constrain direct substitution. Employees may utilize the time savings to clear workloads, generate more exhaustive outputs, or participate in more roles. In the long run, firms can develop procedures, alter occupational structures and trim hiring in exposed positions. The lack of sizable employment impacts in the initial years of chatbot installation argues against a swift implosion, but does not prove it to be settled in the long term.

Entry-level jobs are especially critical. Young employees do many codifiable tasks from which they may build experience and tacit knowledge. AI can be a complement to inexperienced workers, guiding their work immediately and therefore at a lower cost to firms than on-the-job training, enabling its use to a greater range of labor market entrants.[15] It can also cut back on the number of apprentices needed; experienced workers get their routine work done faster. The fiscal implications are even more far-reaching: weaker entry routes diminish lifetime income, revenue taxes and the future stock of skilled workers.

Creation of new tasks is the key to the long-term saturation of automation as a force. Adoption of artificial intelligence (AI) gives rise to work in the areas of workflow design, model interpretation and evaluation, data governance, security, compliance, customer interaction and specialized applications.[16] It may also render feasible services that were previously unprofitable, broadening appeal for human interaction and cognition. However, it will not necessarily draw forth labor in equivalent quantity, geography, remuneration and output access. A handful of top-tier specialists can be present alongside disappearing labor in routine intellectual work. The crucial issue is whether the new and expanded activities generate adequate labor demand to sustain the overall wage pool as a share of output. Thus, the available evidence tends to accept a conditional proposition. Output expansion will offset displacement where demand is elastic, competitive pressure transmits cost savings onto consumers, complementary stimulating investment exists and workers are able to move into less exposed or newly created tasks. Replacement will dominate rather than complement where demand is saturated, production is standardized, output cannot expand and adoption is mainly focused on cost reduction. There is no support for either universal complementarity or universal replacement.

3. Are Productivity Gains Passed Through as Wages, Retained as Profits, or Transferred to Technology Suppliers?

A productivity gain creates an economic surplus without indicating the ultimate beneficiary. Where an AI-enabled employee takes four hours for work that previously required eight, the saving may turn into a larger payroll, a lower customer price, a larger corporate margin, extra output, a technology-service payment, or a higher shareholding valuation. Its distribution hinges on labor bargaining, the degree of competition, the ownership structure, buyer power and the supply of complementarities. For the state, the ultimate distribution matters as much as the gross increase since each form of income hits the tax machinery in distinct ways. Workers gain where AI increases their marginal productivity and labor-market institutions convert those gains into income.

Customer-support evidence suggests that AI can transmit expertise and improve the output of inexperienced workers, perhaps expanding access to productive employment.[17] The EIB firm research finds higher wages among AI adopters, although some of the association can be explained by traits of adopting firms and their personnel. By contrast, Danish administrative data have no measurable average earnings increase in the initial two years of AI adoption, with no measurable change in recorded hours.[18] Overall, the evidence suggests wage pass-through is feasible but neither instantaneous nor assured. Bargaining power decides whether productivity results in pay. Workers receive a larger share where complementary skills are scarce, job opportunities are high, collective bargaining extends to innovation and it's still performance-related pay. Firms win a bigger slice where workers can't monitor improvement in productivity, job opportunities are limited, or AI frees up labor. The very same technology can thus both tighten performance pay gaps and widen internal income gaps of occupational groups. The worst-performing workers may get more productive, but at the same time, engineers, domain experts and managers who are able to put AI to work may take the lion's share of wage premiums.

Consumers realize benefits through lower prices, quicker delivery of services, a wider product range, or better service quality. Consumer returns can therefore lead to an increase in real living standards despite static nominal income. Consumer surplus, however, is not directly taxed.[19] It enters into revenue when lower prices release income for other taxable spending, or if improved services lead to a rise in the volume of taxable transactions. The fiscal capture from a large consumer return can then be weak. Adopting firms keep the gain if competition is sparse, or prices are sticky, or AI is embraced, which makes internal processes more effective, whose gains are not readily observed by customers. Retained gains might show up as higher operating margins, cash flow, or investment in further intangibles. Direct taxation on the domestic taxable profit may grow if the gain is implemented in an approximate proportion within the economy. However, the overall productivity advantage and once-off domestic taxable gain are not ipso facto the same. Businesses need to pay for the supply of models, software, cloud services and consulting and integration, as well as complementary capital. These costs imply that the size of the surplus left to the adopter will be conditional on the expenditures made by the adopter.

Technology suppliers form their own claims. There are many forms of purchase. AI is often bought on a subscription. It is bought on a usage-based contract or a license through cloud services, enterprise software and licensing. A local provider may enjoy having a higher revenue per employee, at the same time as it gives away most of the added value to an upstream supplier. If this supplier is foreign, the domestically paid sum reduces domestic taxable profit and generates supplier income abroad.[20] The country can still benefit in terms of wages, consumer surplus and, as long as downstream competitiveness exists, the figures used to measure national productivity can exaggerate the part of the surplus that can exist in the corporate tax structure. Supplier capture is affected by market structure. Convergence of models and complementary services in a commoditized market could hold down prices, shifting more of the surplus to adopters and consumers. While economies of scale, proprietary data, switching costs and distribution premiums tend to reinforce market power in favor of technology suppliers, early adopters might pay a premium for the limited services available in a nascent market and, as the phenomenon diffuses, advantages accruing to technology suppliers might diminish.

Shareholders are paid the residual value of expected future profits. Gains from AI may be capitalized into equity prices prior to accounting income or dividends. The taxation of this value will be very different from payroll. Wages are normally taxed on an ongoing basis by withholding and social contributions. Capital gains can be taxed on a realization basis, at a lower rate, in a different jurisdiction, or not at all in some institutional or tax-privileged arrangements.[21] A productivity improvement can therefore increase private wealth without delivering an equivalent current flow of revenues domestically. AI-complementary workers in short supply form an intermediate category. Their compensation is legally labor income, although part may reflect scarcity rents. The key fiscal inference is that governments do not tax productivity as an abstract concept. They tax the wages, profits, dividends, realized gains and consumption through which the surplus is allocated.

4. How Might AI Shift the Labour Share?

AI share is not an officially established 'national-accounts' concept; it's a useful heuristic. It refers to that part of value added[22] that flows to AI-related capital, intellectual property, technology suppliers and economic rents after the extraction of labor compensation. Labor share is defined as employee compensation (sometimes with the addition of an adjustment for the labor share of self‑employment), as a share of value added. The research issue is whether and, if so, where and how AI influences this share. Five mechanisms alter the direction. Augmentation creates more output per worker and maintains or increases labor share if the benefit occurs alongside proportional wage increases. Substitution decreases the need for labor, exerting downward demand pull on the share. Scale expansion may restore labor demand if the lower costs induce enough additional output. The creation of new tasks supports labor income, though the creation of activities where the worker maintains the comparative advantage. Concentration of rent enables either the investor or the upstream supplier to convert its creation of productivity into profit. It turns on the relative power of the mechanisms, not the technical exposure.

Figure 3. In high-income economies, GenAI exposure reaches 41 percent of women’s employment, compared with 28 percent of men’s.

There is as yet no reliable empirical measure indicating that generative AI has brought down the economy-wide labor share by a quantifiable observed amount. Adoption is recent and measurement remains inconsistent and labor shares are subject to sectoral composition and business cycle effects, housing income, self-employment and transnational accounts. The global share of labor income had already begun heading down even before the current generative systems became widespread in their use. The ILO estimates that the global labor-income share fell from 52.9 percent in 2019 to 52.3 percent in 2022 and remained broadly unchanged through 2024,[23] but this longer-term trend cannot be traced to generative AI; it merely shows that labor entered the AI era without an irrefutable claim on productivity. Model-based analysis provides a cautious indicator of what scale we might expect: Acemoglu's upper-bound scenario assumes a small increase in the share of capital and a fall in labor's share, assuming that the impact of AI on output exceeds the growth of wages in the economy.[24] The resulting shift is comparable to small fractions of a percentage point rather than the sharp decline as postulated at times in the public discussion.

Figure 4. Labor’s global income share fell 1.6 percentage points after 2004, with 40 percent of the decline occurring after 2019.

Labor-share pressures are likely to be more visible in particular sectors than in the aggregate. Administrative support jobs, hidden in well-often concentrated in protected sectors, have generative-AI exposure and may be more easily measured and improved than other work tasks, as they involve repetitive tasks, which are relatively cheap to automate.[25] For example, simple computing tasks like data entry, calendar planning, routine filing, accounting and copying may be affected proportionally. As the demand for this underlying service is likely to be fixed in each task, the source of the decline in employment in administrative jobs is likely to be primarily reduced pay bills, rather than increased output. They are one of the clearest early signs of a decline in labor share. Finance, insurance, accounting and legal services are somewhat more mixed. Text, data and pattern-recognition tasks generate far more exposure, yet regulation, fiduciary duties, model risk and client confidence preserve important human roles. Jobs may stay relatively stable while importance is added and profit-per-partnership or shareholder increases faster than costs. Labor share thus can fall without mass layoffs. Within a business, junior roles may shrink while senior staff and rare technical professionals benefit from extra income.

Figure 5. Occupations with similar average exposure differ markedly in how evenly exposure extends across their tasks.

Public services need a different measure. If a tax office, hospital, school, or local authority handles many more cases with the same number of staff, there's a chance that measured labor share will hardly fall, because output in the non-market is often based on input prices.[26] Productivity improvement might take the form of fewer queues, better quality or efficiency, not higher 'market' profit. Standards of national labor-share figures could underestimate salient fluctuations in productivity in the public sector. The other factor is firm heterogeneity. As shown by the EIB evidence, productivity growth is heavily biased towards medium and large data, capital, software and skilled- worker intensive adopters.[27] Such productivity growth is potentially accompanied by market share shifts away from less productive small companies and a resulting reduction in aggregate labor share because the less productive small companies are more payroll-intensive than they are value-added-intensive. On the other hand, new, less costly AI tools may remove at least some competition barriers for small companies, the self-employed and independent professionals, thus avoiding income concentration.

The falling labor share does not necessarily mean real wages are falling. Wages may rise with a smaller share of a faster-growing total. Consumers may gain via lower prices. The fiscal worry is that governments tax wages/employment contributions sooner and more smoothly than retained earnings, unrealized capital gains, or border-crossing supply of income-that is, a small shift in factor shares may lead to a much larger shift in revenue composition, destination and timing. The current evidence belies any assertion of a sharp collapse in the labor share, as well as the presumption that distributional impacts are minor. Nor can any aggregate shift caused by generative AI be established, but significant reallocations at the firm and sector levels seem feasible and could be underway already. The key variables are payrolls as a share of value added; wage growth relative to labor productivity; entry-level employment in new firms; domestic operating profits; payments for imported technology services; and the generation of new labor-intensive jobs. Exposure indices simply cannot determine the extent to which the labor share will decline.

5. Are the Resulting Wages and Profits Taxable Where Affected Workers and Customers Reside?

The most direct channel is personal income tax. If AI raises wages and employment, then taxable labor income increases. If firms cut hours, cut hiring, or automate, there are larger cutbacks in the base. Since top personal income tax schedules are progressive, distributional effects matter. Substantial redistributions to a small number of highly skilled specialists might generate large revenue, but may be offset by large layoffs among middle-income employees when allowances, thresholds and behavioral responses are taken into account. Employee and employer social contributions are especially vulnerable as they are formally linked directly to pay and used to pay for pensions, health care, unemployment insurance and other benefits.[28] The bottom line for contribution bases can be eroded even if aggregate national income rises if compensation is shifted from wages to profits. Contribution ceilings may aggravate the problem by providing tax shelter above a cap for highly paid workers without generating any further revenue. A broad-based increase in middle-income wages is monetarily different from a similar total increase concentrated among a handful. Payroll taxes are subject to the same logic. They are mostly tied to the place where the worker works and collected on an ongoing basis, so that they are, in effect, more stable and less mobile.

Corporate profits and capital income are more concentrated, more volatile and sometimes more globally mobile.[29] A government may indeed hold on to total revenue during an initial AI-driven expansion while substituting a large labor base for a narrow profit base and its distinctive cyclicality. When the lost payroll adopters keep the productivity gain as a domestically taxable profit, their corporate-income-tax receipts can offset the losses in payroll. Reduced labor costs, enhanced output and improved margins all raise the corporate obligations. But this offset is not automatic; taxable profit depends on the ways a firm can manage its deductions, losses, financing, depreciation, investment incentives and where it claims income. A few dominant corporations may dominate corporate receipts, a source of increased volatility and political liability. It also needs to be separated from payments to suppliers of the technology. Even a domestic firm can have had higher productivity in the home market but paid huge sums for foreign models, cloud services, software, integration, or IP, which eat into the profit that remains at home. The adopting country can continue to gain from wages, lower prices and higher competitiveness, but gross productivity growth will overstate the corporate income tax that can be levied by the government when so much has been exported to foreign suppliers.

Dividends and capital gains also offer another revenue stream. Increasing profits can be shared directly as dividends, or capital gains can increase to the extent investors anticipate future profits. Capture through capital gains is a function of the investor's residence and ownership, realization rules, exemptions and timing. Capital gains can be unrealized for years, can be expatriated, or can be held through pension funds and other tax-privileged vehicles and the requisite future tax receivable embedded in an equity valuation cannot fund today's unemployment benefits, retraining, or pensions in exactly the same manner as a monthly payroll withholding. The consumption tax only indirectly takes some of the gain. Rising real wages, dividends, or corporate profits may enable more household expenditures. Falling prices can raise real buying power, freeing income for other taxed purchases. However, surpluses are not themselves taxed. If AI lowers a taxed service price, with a fixed quantity, value-added tax revenue may fall. Spending the savings on a different good may restore overall VAT returns. This effect is driven by observed nominal expenditure, not uplift in consumer welfare. Public expenditure can increase during adjustment. Those displaced from exposed occupations may need unemployment insurance, income security, retraining, job-search assistance and mobility programs. Aggregate employment stability does not remove this cost, since disruption may be absorbed by individual occupations, geographic areas, or ages. An important national effect can have a profound local adjustment. Ineffective retraining, which produces costs without improvement in employment, should also be taken into account.

AI can also boost the public balance sheet through tax administration and public-service efficiency. Against a background of expanding data collection, tax authorities are increasingly harnessing AI and big data-driven systems to improve risk management, limit tax evasion and avoidance, select cases and provide services to taxpayers. OECD evidence indicates that a large majority of surveyed tax administrations by 2023 were either already exploiting or in the process of deploying AI.[30] Better targeting can drive compliance, increasing revenue and cutting collection costs; automation of routine administration can release officials to concentrate on more complex cases. These improvements will depend on quality data, supervisory human oversight, cybersecurity, legal protection and sophisticated procedures for challenging automated judgments. These benefits should be viewed as a potential fiscal offset, rather than a windfall. The three hypothetical 100 productivity-gain cases do show the importance of incidence. They are stylized accounting cases, not empirical estimates or forecasts. Assume that there is also a 20 percent tax on the additional wages, 10percent on the incremental domestic profit and the consumer-benefit component produces an illustrative VAT effect of 10 percent. The VAT assumption does not reflect the direct taxability of consumer surplus; it is just a simplified increase in taxable expenditure with regard to the benefit.

In Case A, €60 appears as wages, €30 as domestic profit and €10 as consumer benefit. Wage tax yields €12, profit tax €3 and VAT €1, producing total domestic revenue of €16. This is the widest base, widest in the sense that almost all of the gain ends up as domestic labor income. In Case B, €20 goes to wages, €60 to domestic profit and €20 to consumer benefit. The effect of wage taxes is €4, profit taxes €6, VAT effect €2, thus €12 in all. Corporate revenue is higher than in Case A, but not enough to offset the weaker labor base, assuming these tax rates. In the simple example of Case C, €20 goes to wages, €20 to domestic profit, €40 to payments to foreign AI suppliers and €20 to consumer benefit. If in this case we suppose no direct domestic tax collection on the foreign income outside the model, the wage is taxed at €4, the profit is taxed at €2 and the VAT effect is €2, totaling €8 in domestic revenue. The same €100 gross productivity advantage produces only one-half the revenue of Case A in this case because a relatively much greater share appears outside of the model's assumed domestic base.

These scenarios do not prove that AI will cut government revenues by 50 percent. They do not allow for progressive rates, social contributions, deductions, investment, supplier payroll, withholding taxes, 'trade' effects, behavioral responses and deferred taxation of capital gains. They accept the full €100 as available surplus, although real adoption implies complementary spending. Their purpose is narrower: calculating fiscal projections on the basis of GDP or productivity alone can be highly misleading when the distributional and 'jurisdictional' aspects of the gain are neglected. Current tax arrangements are therefore relevant to this composition problem. Personal-income taxes and social-security contributions combined constitute a significantly larger proportion of total revenues than do corporate-income taxes throughout the OECD.[31] Within the EU, taxes on labor (including social contributions) constitute roughly half of total tax revenue.[32] A continuing transfer of revenues from payrolls toward profits or capital gains or to foreign providers may harm the bases upon which social insurance and other public services are paid, even if GDP continues to grow.

Tax treatments might also shape the adoption path. In their model of a sub-optimally high automation rate, Acemoglu, Manera and Restrepo show that the US tax system taxed labor more heavily than capital over the course of the last century, thus incentivizing automation above the socially optimal level.[33] However, it must be noted that this study deals with automation as a whole, rather than production using generative AI specifically and reaches its conclusion on the basis of a set of assumptions made in the model. Nevertheless, it sends a clear message: tax policies should not unintentionally incentivize labor substitution in an artificially attractive light relative to augmentation, organizational change, or skills acquisition.

Figure 6. Personal income taxes and social contributions together account for nearly half of the OECD tax mix.

Measurement remains a central fiscal institution. Departments of finance, statistical offices, tax authorities and social-insurance agencies require linked data on firm-level adoption, payrolls, occupations, wages, sales, profits, imported services and the taxes they pay. Circumstances where firms have furloughed, increased wages, or transferred the proceeds abroad cannot be identified from aggregate exposure indices alone.[34] Within sector monitoring must compare observed developments with what technical potential permitted while monitoring flows into entry-level jobs, contribution density, a firm's labor share and intermediate supplier expenditure. Transition policies should be based on actual displacement, not on conjectural estimates of national employment figures. Unemployment insurance and income support are required where displacement takes place and retraining should be related to credible labor demand and assessed by subsequent employment and earnings. Short-term wage insurance may be more appropriate for a few mid-career workers than long classroom courses. Policies should be amplified where verified displacement increases and changed if labor demand resurges. Thus, AI may boost corporate revenue at the expense of wages, increase tax compliance at the expense of adjustment costs and improve consumer well-being at the expense of rent-sharing. The tax problem is not merely a single change in the base but one in the composition, timing, localization and jurisdiction of revenue.

Figure 7. Labour supplies more than half of EU tax revenue, making shifts away from payroll fiscally consequential.
6. Which Fiscal Structures Are Most Exposed?

The highest fiscal risk occurs where a number of vulnerabilities combine. High dependence on labor taxes alone is inadequate because employment and real wages could increase if augmentation succeeds. High occupational vulnerability alone will also be inadequate, because exposed jobs can be converted into other types of occupations. The greatest danger exists if all six of the following factors reinforce one another: high dependence on labor taxes; available substitutions of cognitively based employment; weak capital-income tax payments; low domestic ownership of capital stock; payments to foreign producers of technology; and rising expenditure pressure. Employers are hit directly by social-contribution systems as they fund them through payroll. The share of social-security contributions in total tax revenue in 2023 was over 40 percent in Czechia, Slovenia and Slovakia.[35] While these economies might not be the most exposed to generative AI in employment terms, especially where manufacturing is still relevant, the impact in fiscal terms is clear. A shift in the value-added composition from payrolls to profits or imported services undermines the funding base of pensions and social-insurance systems directly.

An analogous situation exists among countries exhibiting high tax wedges on formal employment. Belgium, Germany, France and Austria all register some of the highest totals of personal and social contributions borne by the typical worker of the OECD.[36] Again, a high tax wedge on employment does not necessarily point to a net employment effect of AI; it simply emphasizes that every euro taken away from the wage bill requires the outflow of much profit and capital income. If these are taxed less than current income forms, the tax effect of diversion increases. A similar sort of risk is represented by countries whose reliance on personal-income taxes, as opposed to social contributions, is especially pronounced. Denmark is just such an example, with significant levels of government revenues coming from the former while contributing less to the latter. Its highly digitized labor market entails significant technical exposure, but high levels of general consumption taxation, an efficient administrative bureaucracy and effective labor-market institutions are compensating factors. This sheds light on just why exposure should not be ranked within only a single tax category.

Figure 8. Belgium, Germany and France tax average labour costs substantially more heavily than the OECD average.

The highest amount of occupational exposure is likely to exist in service-intensive high-income economies. According to the ILO, 34 percent of employment in high-income states has at least some generative-AI exposure, relative to 11 percent in low-income states. In these economies, there is a fairly large prevalence of clerical, financial, professional, technical and administrative jobs, many of which have relatively high wages, meaning these economies could garner the most benefit from productivity but could also see the most significant movement away from broad labor compensation if rent concentration and substitution take hold.

Another distinct vulnerability arises if domestic ownership is weak. While a country can quickly utilize the AI to boost productivity and consumer benefits , payments for technology will be made to foreign owners, which will shrink the domestic operating surplus and needed capital income base and enhance "ownership income" in the homeland less than what would be reflected in the productivity figures. Firms, cloud providers, software and other intellectual properties have mostly foreign owners, though the business and capital stocks could grow less than the productivity data would indicate. The IMF scenario estimates help distinguish economies that produce and own important AI assets from those that only import these services, while acknowledging the scenario's nature.[37] Ownership should not be conflated with the physical location of the infrastructure. A state can host infrastructure and not be the owner of the most sophisticated models, software, or intangible assets and a domestic firm can own artificial intelligence assets while contracting for infrastructure services outside. In terms of fiscal incidence, the key variables are the recipient of the income, the contractual form of the transaction and the jurisdiction of the taxation of profit or capital income. An infrastructure nexus so detailed belongs to a different tax question.

In countries with only weak effective tax on capital and profits, exposure increases even with AI rents remaining within the economy. Changing from wages to retained profits, dividend income, or capital gain tax will lessen revenue where that source is either exempt, deferred or weakly enforced; and statutory corporate tax rates will be just half of the factor. Tax expenditures and old-style profits-for-years-with-no-tax arrangements, with institutional ownership and realization-based capital gains provisions, alter collection. Highly profitable economies can be susceptible in this way and multinational nodes can generate significant revenues through large, highly concentrated sets of firms, thus allowing profit taxes to offset a feeble labor force. However, reliance on such a limited and heavily internationalized pool is itself subject to fluctuations in the decisions of multinational enterprises. Ireland and Luxembourg should not be treated simply as weak-profit-tax jurisdictions. Their potential exposure to risk might stem from concentration and mobility rather than weakness.

All revenue sources are made more susceptible by aging and debt. In countries with growing pension, health care and long-term-care costs, a decline in contribution income in itself or an erratic tax portfolio will have harsher revenue consequences than the same or smaller downturn in payrolls for a more youthful, less indebted economy. The European Commission's aging projections are built-in scenarios based on what-if demographic and policy assumptions, but they clearly demonstrate the cost squeeze European countries face in future decades.[38] A small permanent weakening of payrolls in such a macroeconomic environment could be much more harmful than a temporary large disturbance to a younger population in a more lightly indebted country.

Cross-country fiscal evaluation, consequently, must involve many indicators. Countries should assess their reliance on labor taxation and social contributions, the skill composition of the formal sector, the true take-up and labor market responses, the effective taxation of profits and capital income, payments to foreign technology providers, domestic ownership of productive assets and the medium-term budgetary and debt constraints. There is no one measure of exposure. The existence of national stress tests can incorporate alternative distributions of incidence without assuming the ability to predict technological advances: in an all-wage scenario, an augmentation model may take full wage pass-through; in a high-corporate-profit model, a declining payroll and rising domestic corporate income model may be applied; in an imported-rent model, a significant share of suppliers receiving international payments can be incorporated. Each should apply the appropriate personal tax schedule, contribution system, effective corporate rates, consumption taxes and benefit expenditure schedule and be cast as conditional ranges that should be refined iteratively as administrative data improves.

Policy should match the diagnosed vulnerability. Payroll-dependent countries should monitor wage pass-through, contribution density and entry-level employment; weak capital income taxation requires broader effective capture, rather than a special AI levy; AI-service importers require better data on technology payments and a greater share of home value added. Heavily indebted and aging relies on contingency planning for adjustment costs and revenue volatility. Tax administrations can benefit from AI for ensuring compliance, while not compromising due process, transparency and human accountability. Pretax fiscal incidence is also affected by the provision of labor-market institutions. Institutions such as training, portable social insurance, efficient job matching, collective bargaining and competition can provide augmenting and scale-expanding effects. They do not just serve as reinforcing improvements after the shrinking revenue base has occurred. Instead, they influence whether higher productivity is revealed in increased wages, job creation, consumers’ gains or concentrated rents. The most durable fiscal cushion will not result from heavier taxation of a declining wage bill, but from a healthy economy where rising productivity is reflected in and taxed from widely spread, domestically generated income. Countries most exposed are those in which the gains will be incorporated into forms that the domestic fiscal system captures weakly, slowly, or not at all, while public obligations remain tied to employment, aging and social insurance.

7. Conclusion - When Does AI Expand Rather Than Erode the Tax Base?

Artificial intelligence does not mechanically erode the tax base. It shifts the tasks that produce income and can enhance productivity, wages, corporate profits, aggregate consumption and fiscal administration. Preliminary evidence suggests significant productivity gains in selected tasks and significant workplace exposure, but not a general job-shedding in the economy. The strategic risk is a divergence of productive capacity relative to domestic fiscal capture. An economy may gain productivity while shifting a broad, immediately taxed payroll base toward concentrated profits, deferred capital gains and revenue from foreign suppliers. High wage pass-through, competitive diffusion, new-task creation and growth of domestic enterprise can produce the opposite outcome as well. Governments must avoid complacency on the one hand and destructive technology-specific taxes on the other and should base intervention on measured rather than speculative displacement. Fiscal resilience will depend less on headline productivity improvement than on the extent to which the income generated by that productivity is distributed broadly and locally and ultimately able to be taxed.


This article was prepared as an independent research contribution following Professor Keith Lee’s presentation at the conference Inequalities in Longevity, held at Fondazione Giorgio Cini in Venice on 3–4 July 2026, and a subsequent substantive discussion with Federico Fubini of Corriere della Sera. The discussion helped sharpen the motivating questions concerning AI-driven productivity, the declining employment intensity of economic growth, and the resulting pressures on the public tax base.

The research, analysis, and conclusions were developed independently by the author. This publication is separate from the official conference proceedings and from the editorial coverage published by Corriere della Sera.

Unless expressly stated otherwise, this publication has not been commissioned, reviewed, or endorsed by Fondazione Giorgio Cini or Corriere della Sera. References to the conference and the subsequent discussion describe the intellectual context in which the research developed and do not imply co-authorship, institutional affiliation, formal collaboration, or partnership. The analysis, interpretations, and conclusions are those of the author and do not necessarily reflect the official positions of Fondazione Giorgio Cini, Corriere della Sera, the Swiss Institute of Artificial Intelligence (SIAI), or their respective affiliates.


References

[1] Noy, S. and Zhang, W. (2023) ‘Experimental evidence on the productivity effects of generative artificial intelligence’, Science, 381(6654), pp. 187–192.

[2, 15, 17] Brynjolfsson, E., Li, D. and Raymond, L.R. (2023) Generative AI at Work. NBER Working Paper No. 31161. Cambridge, MA: National Bureau of Economic Research.

[3] Dell’Acqua, F., McFowland III, E., Mollick, E.R., Lifshitz-Assaf, H., Kellogg, K.C., Rajendran, S., Krayer, L., Candelon, F. and Lakhani, K.R. (2023) Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge-Worker Productivity and Quality. Harvard Business School Working Paper No. 24-013. Boston, MA: Harvard Business School.

[4, 5, 6, 25] Gmyrek, P., Berg, J., Kamiński, K., Konopczyński, F., Ładna, A., Náfrádi, B., Rosłaniec, K. and Troszyński, M. (2025) Generative AI and Jobs: A Refined Global Index of Occupational Exposure. ILO Working Paper 140. Geneva: International Labour Organization.

[7, 24] Acemoglu, D. (2025) ‘The simple macroeconomics of AI’, Economic Policy, 40(121), pp. 13–58.

[8, 18] Humlum, A. and Vestergaard, E. (2025) Large Language Models, Small Labor Market Effects. NBER Working Paper No. 33777. Cambridge, MA: National Bureau of Economic Research.

[9, 13, 14, 34] Hampole, M., Papanikolaou, D., Schmidt, L.D.W. and Seegmiller, B. (2025) Artificial Intelligence and the Labor Market. NBER Working Paper No. 33509. Cambridge, MA: National Bureau of Economic Research.

[10] Autor, D.H. (2015) ‘Why are there still so many jobs? The history and future of workplace automation’, Journal of Economic Perspectives, 29(3), pp. 3–30.

[11, 12, 27] European Investment Bank (2026) EIB Group Investment Survey 2025/2026. Luxembourg: European Investment Bank.

[16] Acemoglu, D. and Restrepo, P. (2019) ‘Artificial intelligence, automation and work’, in Agrawal, A., Gans, J. and Goldfarb, A. (eds) The Economics of Artificial Intelligence: An Agenda. Chicago: University of Chicago Press, pp. 197–236.

[19] Brynjolfsson, E., Collis, A. and Eggers, F. (2019) ‘Using massive online choice experiments to measure changes in well-being’, Proceedings of the National Academy of Sciences, 116(15), pp. 7250–7255.

[20, 21, 28, 29, 37] International Monetary Fund (2024) Broadening the Gains from Generative AI: The Role of Fiscal Policies. Staff Discussion Note 2024/002. Washington, DC: International Monetary Fund.

[22] ILOSTAT (2024) Labour Income Share as a Percent of GDP: ILO Modelled Estimates and Methodological Description. Geneva: International Labour Organization.

[23] International Labour Organization (2024) World Employment and Social Outlook: September 2024 Update. Geneva: International Labour Organization.

[26] European Commission (2013) European System of Accounts: ESA 2010. Luxembourg: Publications Office of the European Union.

[30] Organisation for Economic Co-operation and Development (2024) Tax Administration 2024: Comparative Information on OECD and Other Advanced and Emerging Economies. Paris: OECD Publishing.

[31, 35] Organisation for Economic Co-operation and Development (2025) Revenue Statistics 2025. Paris: OECD Publishing.

[32] European Commission (2026) Data on Taxation Trends. Brussels: Directorate-General for Taxation and Customs Union.

[33] Acemoglu, D., Manera, A. and Restrepo, P. (2020) ‘Does the US tax code favor automation?’, Brookings Papers on Economic Activity, 2020(1), pp. 231–300.

[36] Organisation for Economic Co-operation and Development (2026) Taxing Wages 2026. Paris: OECD Publishing.

[38] European Commission (2024) 2024 Ageing Report: Economic and Budgetary Projections for the EU Member States, 2022–2070. European Economy Institutional Paper 279. Brussels: Directorate-General for Economic and Financial Affairs.

Corriere della Sera Features SIAI Analysis of AI and Fiscal Erosion

Corriere della Sera Features SIAI Analysis of AI and Fiscal Erosion

Picture

Member for

1 year 2 months
Real name
SIAI Editor
Bio
SIAI Editor

Corriere della Sera has featured commentary from Professor Keith Lee of the Swiss Institute of Artificial Intelligence in an article by Federico Fubini examining how artificial intelligence may reshape employment, labor-based taxation, and the fiscal capacity of advanced economies.

The article draws substantially on Professor Lee’s analysis of

  • highly AI-augmented workers,
  • the concentration of productivity gains among a relatively small group,
  • the transfer of economic value from labor income toward corporate profits and AI-related rents,
  • and the resulting pressure on labor-based tax systems.

The article also considers the difficulty governments may face in replacing income-tax and social-contribution revenues if AI adoption reduces employment in knowledge-intensive industries.

Note: Screenshot from Corriere della Sera coverage
Source: Tasse, la rivoluzione dell’intelligenza artificiale: chi le pagherà se i lavoratori sono già sempre di meno? | Corriere.it

Selected coverage in Italian:

Keith Lee, nativo della Corea del Sud, un master in finanza alla London School of Economics e un dottorato in finanza matematica alla Boston University, insegna le sue materie allo Swiss Institute of Artificial Intelligence. E ha iniziato a notare comportamenti nuovi fra i suoi studenti. Stanno emergendo quelli che lui definisce, «lavoratori superumani», persone in grado di assicurare da sole il lavoro per il quale prima sarebbero servite dieci o venti persone. O di comprimere enormemente i tempi per alcune forme di attività della conoscenza. 

Mi ha raccontato Keith: «L'intelligenza artificiale può consentire quella che descriverei come una produttività sovrumana in compiti specifici. Ho visto persone completare il lavoro diverse volte più velocemente utilizzando l'IA in modo efficace. In un compito recente, uno studente ha prodotto in otto ore un'analisi di 20-30 pagine con diversi grafici che normalmente avrebbe richiesto diversi giorni». Non che gli altri allievi fossero privi di esperienza - siamo in un corso universitario avanzato di AI - ma solo una in tutta la classe aveva capito come usare la nuova tecnologia per moltiplicare di molte volte la propria produttività.

Note: Screenshot from Corriere della Sera coverage
Source: Tasse, la rivoluzione dell’intelligenza artificiale: chi le pagherà se i lavoratori sono già sempre di meno? | Corriere.it

English translation of selected coverage:

Keith Lee, a South Korean native, holds a master's degree in finance from the London School of Economics and a doctorate in mathematical finance from Boston University, teaches his subjects at the Swiss Institute of Artificial Intelligence. And he began to notice new behaviors among his students. What he calls "superhuman workers" are emerging, people who can provide jobs on their own for which ten or twenty people would have been needed before. Or to enormously compress the time for some forms of knowledge activities.

Keith told me: "Artificial intelligence can enable what I would describe as superhuman productivity in specific tasks. I've seen people complete work several times faster using AI effectively. In a recent assignment, a student produced a 20-30 page analysis with several graphs in eight hours that would normally have taken several days." Not that the other students were inexperienced - we are in an advanced university course in AI - but only one in the whole class had understood how to use the new technology to multiply their productivity many times over.

A related Executive AI Brief expands on the discussion, including the relationship between AI-driven labor substitution, fiscal erosion, data-center taxation, and the geographic distribution of AI-generated value.

Theme in Corriere coverageSIAI follow-up research
AI-enabled labor productivityLabor Income, AI Rents and Fiscal Erosion
Mobility of AI-generated rentsData-Center Taxation and the Geography of AI Value
Tax treatment of AI profitsTaxing AI Profits in Europe
Europe’s strategic positionEurope’s AI Race and the Fiscal State

Further analysis by SIAI

Following the Venice conference at Fondazione Giorgio Cini and Professor Lee’s subsequent interview with Federico Fubini of Corriere della Sera, SIAI developed a four-part research series examining the fiscal consequences of AI-driven productivity growth.

The papers examine the same issue from the perspectives of labor income, AI-generated rents, data-center taxation, Euro tax policy, and Europe's broader AI strategy.

AI and Tax research series

Picture

Member for

1 year 2 months
Real name
SIAI Editor
Bio
SIAI Editor

AI and Tax: Who Captures the Productivity Dividend?

AI can create task-specific “superhuman productivity,” sharply compressing the time required for some knowledge work without implying that one person can replace ten complete jobs
he fiscal risk is less an immediate collapse of government revenue than a gradual shift from broadly taxed labour income toward concentrated profits, capital income and economic rents
Because those gains are more mobile across borders than payroll, governments will need stronger capital-income taxation, better measurement and international coordination rather than a blunt tax on AI itself

SIAI Presents Research on AI and Inequalities in Longevity at Fondazione Giorgio Cini

SIAI Presents Research on AI and Inequalities in Longevity at Fondazione Giorgio Cini

Picture

Member for

1 year 2 months
Real name
SIAI Editor
Bio
SIAI Editor

Professor Keith Lee of the Swiss Institute of Artificial Intelligence participated in the conference Inequalities in Longevity: Drivers, Frailties and Policy Responses to an Underestimated Challenge, held at Fondazione Giorgio Cini in Venice on 3–4 July 2026.

His presentation, Unemployed Growth and the Emerging Inequality of Longevity, examined how artificial intelligence may reshape employment, income distribution, access to healthcare technologies, and the ability of different socioeconomic groups to benefit from longer life expectancy.

The conference brought together international researchers and policy specialists to discuss the economic, medical, and social causes of unequal longevity, as well as potential policy responses.

Conference materials

Unemployed Growth and the Emerging Inequality of Longevity | Full presentation
Unemployed Growth and the Emerging Inequality of Longevity | 1-min Insight

Related SIAI Analysis

A related Executive AI Brief provides a more detailed account of the presentation, including the mechanisms through which AI-driven productivity growth and labor displacement may influence future inequalities in health and longevity.

Note: Longhena stairs, San Giorgio Cini at Fondazione Giorgio Cini conference for ‘Inequalities in Longevity’
Source: Inequalities in Longevity | Drivers, Frailties and Policy Responses to an Underestimated Challenge - Fondazione Giorgio Cini
Picture

Member for

1 year 2 months
Real name
SIAI Editor
Bio
SIAI Editor

Unemployed Growth and the Emerging Inequality of Longevity

AI can accelerate biological research, diagnosis, prevention and care, raising the prospect of longer and healthier lives
Its nearer-term labour effect may be unequal augmentation: a small group becomes dramatically more productive while others lose bargaining power or move into lower-quality work
Because work, income, wealth and access shape health, AI could widen longevity inequality unless its productivity and medical dividends are deliberately shared

From Feedback Loops to Causal Guardrails: Endogeneity in AI Systems

From Feedback Loops to Causal Guardrails: Endogeneity in AI Systems

Picture

Member for

1 year 10 months
Real name
Keith Lee
Bio
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.

His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.

He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.

Modified

Adaptive AI systems partly create the data from which they learn
Long memory can preserve hallucinations as easily as valid information
Causal identification must become a distinct layer of AI architecture

Artificial-intelligence systems are increasingly trained, evaluated, and updated through feedback. Language models learn from human rankings, recommendation systems learn from user engagement, autonomous agents learn from environmental rewards, and conversational systems use previous interactions to shape subsequent responses. These mechanisms allow models to adapt to users and improve beyond fixed supervised datasets, but they also create a fundamental statistical problem: the system influences the observations that are later used to evaluate and retrain it.

This feedback-generated data cannot always be treated as independent evidence. A model selects the content a user sees, observes the user’s reaction, and interprets that reaction as information about preference or quality. Yet the reaction is partly a consequence of the model’s own earlier decision. The system is simultaneously producing the treatment, influencing the environment, observing the outcome, and learning from the relationship among them. In econometrics, this problem belongs to the broader category of endogeneity.

Endogeneity is not resolved by increasing model size, adding more data, extending context windows, or collecting more human feedback. A larger system may estimate the endogenous relationship with greater precision while remaining mistaken about its causal meaning. The emerging challenge for AI architecture is therefore not only how to optimise models against feedback, but how to determine whether the feedback identifies the objective that developers actually intend to optimise.

AI Systems Are Becoming Their Own Data-Generating Processes

Traditional machine-learning pipelines often assume that training observations were generated before the model existed. Images, documents, transactions, or historical outcomes are collected, transformed into a dataset, and used to train a predictive function. The model may inherit biases from the dataset, but it does not necessarily determine which observations entered the original sample. The data-generating process remains largely external to the trained system.

Adaptive AI changes this relationship. A recommender system determines which videos, articles, products, or opinions are presented to a user. The user’s subsequent clicks, purchases, viewing time, and reactions are recorded as behavioural evidence. These observations then influence the next set of recommendations. The system is not merely discovering preferences that existed independently. It is shaping exposure, attention, familiarity, and sometimes the preference itself.

The same structure appears in conversational AI. The model selects an answer, the answer changes the user’s understanding of the issue, and the user writes the next prompt in response. Later interactions are therefore conditioned on a history partly authored by the model. When these conversations become training or evaluation data, the system learns from a behavioural environment that it helped construct. Without causal discipline, repeated interaction can be mistaken for independent confirmation.

RLHF Improves Behaviour, Not Identification

Reinforcement learning from human feedback has become a central method for aligning language models with human instructions and preferences. Human evaluators compare candidate responses, a reward model learns to predict those preferences, and the language model is optimised to generate outputs that receive higher predicted rewards. The procedure can substantially improve usability because it introduces direct information about how people judge model behaviour.

The reward, however, remains a proxy. Evaluators may prefer an answer because it is accurate, but they may also respond to confidence, fluency, familiarity, politeness, brevity, ideological compatibility, or emotional reassurance. These characteristics can be correlated with usefulness without being identical to it. The reward model therefore learns a composite signal whose latent components are not fully observed.

Once the policy is optimised against this signal, the distinction becomes consequential. The model may discover increasingly effective ways to generate approval without improving the underlying quality developers intended to measure. It can become more agreeable without becoming more correct, more confident without becoming more reliable, or more engaging without becoming more beneficial. Human feedback improves the behavioural surface of the system, but it does not automatically identify the causal relationship between response characteristics and true user welfare.

The Endogeneity of Human Preference Data

Human preference data are frequently treated as though they were labels attached to model outputs by independent observers. In reality, the model determines which outputs are available for evaluation. If a particular policy rarely produces a class of response, evaluators have little opportunity to assess it. The preference dataset is consequently selected through the behaviour of the model being trained.

The evaluator is also influenced by presentation. Response order, length, tone, framing, and the apparent confidence of the model can affect judgment. A persuasive but incorrect answer may receive a higher rating than a cautious but accurate one. If the evaluator lacks specialist knowledge, surface quality can substitute for factual validity. The resulting score is not a direct measurement of truth or usefulness; it is the outcome of an interaction among the response, the evaluator, the interface, and the surrounding informational environment.

The data become still more endogenous when feedback is implicit. Continued conversation, click-through rates, session length, or user retention may be interpreted as positive signals. Yet an emotionally validating, politically confirming, or sensational answer may increase engagement precisely because it reinforces a pre-existing bias. The system then learns that the response was desirable because it generated behaviour that the response itself helped produce.

Long Context Creates Persistent Error Processes

Longer context windows allow language models to retain names, facts, objectives, and earlier decisions across extended conversations. This improves continuity and makes agentic workflows more practical. It also expands the pathway through which an early hallucination can influence later reasoning. A false statement generated near the beginning of a thread may remain available as contextual evidence for thousands of subsequent tokens.

Once incorporated into the conversation, the error can reproduce itself. The model refers to the earlier statement, the user adopts its terminology, and later responses observe both the original error and the user’s dependent reaction. The false proposition acquires contextual frequency and internal consistency. A system that relies heavily on conversational history may then assign the proposition greater relevance because it appears repeatedly, even though every repetition descends from the same initial mistake.

This resembles a persistent error process in time-series analysis. The historical variable contains predictive information, but its predictive power does not establish its validity. A lagged statement can explain a later statement because the former was copied into the latter, not because either was independently supported by reality. Long-term memory therefore requires provenance and validation, not merely greater storage capacity.

Agentic AI Magnifies the Feedback Problem

Agentic AI systems do more than answer questions. They search for information, choose tools, write files, communicate with external systems, and update plans according to intermediate outcomes. Each action changes the environment from which the next observation is collected. The system’s decisions therefore become part of the causal process that generates its future data.

An early classification error can influence which sources the agent retrieves, which hypotheses it considers, and which actions it executes. The resulting environment may then appear to confirm the original classification because contradictory evidence was never collected. The agent can construct a path-dependent information set in which each later decision is conditioned on earlier, potentially biased selections.

The problem is particularly serious in multi-step workflows. A false entity resolution can lead to the wrong document, the wrong recipient, the wrong market variable, or the wrong policy interpretation. As the workflow expands, downstream outputs may remain internally coherent because they all share the same mistaken origin. Agentic reliability therefore depends not only on local model accuracy, but on mechanisms capable of detecting when the entire decision path rests on an endogenous or unverified premise.

Prediction Performance Can Conceal Causal Failure

Machine learning has traditionally rewarded predictive accuracy. If a model forecasts the next outcome successfully, it is considered useful regardless of whether it has identified the underlying mechanism. This standard remains appropriate for many applications, but it becomes inadequate when the model’s outputs influence the future observations being predicted.

A recommendation system may become highly accurate at predicting which content a user will consume after years of shaping that user’s exposure. The prediction is accurate, but the system may be predicting a preference that it helped create. A conversational model may accurately predict that an agreeable answer will sustain engagement, but the relationship does not show that agreement improves the user’s understanding or welfare.

The distinction becomes more important as AI systems move from prediction to intervention. An agent is not merely asked what will occur; it selects an action intended to produce an outcome. Causal identification is therefore necessary whenever developers want to know what would happen under a different action, policy, reward structure, or information regime. Predictive validation on historical feedback cannot answer this question by itself.

Instrumental Variables as a Causal Tool

Instrumental-variable methods are designed for settings in which an explanatory variable is correlated with unobserved causes of the outcome. A valid instrument changes the endogenous variable but does not affect the outcome through another direct pathway. By isolating variation that is plausibly independent of the hidden disturbance, the method can identify an effect that ordinary regression cannot recover.

In adaptive AI, an instrument might take the form of randomised exposure, assignment rules, interface variation, or controlled exploration that changes which response or action is selected without directly changing the evaluation criterion. For example, random variation in response order could help distinguish genuine preference from presentation effects. Randomly assigned model variants could help estimate how a change in behaviour affects engagement without relying only on observations produced by the existing policy.

The difficulty is that a valid instrument must be justified, not merely declared. If the instrument affects the outcome through attention, interpretation, or another unmeasured channel, the exclusion restriction fails. If it barely changes the action, it becomes weak. Instrumental variables are therefore not a universal software module that can be attached to every AI system. Their importance lies in the identification framework they impose: where did the variation come from, why is it independent, and what causal path does it represent?

Lagged Data Are Not Automatically Valid Instruments

Sequential AI systems naturally rely on historical variables. Reinforcement-learning algorithms use previous states, actions, rewards, and transitions to estimate future value. Conversational agents rely on earlier messages, retrieved documents, and prior tool outputs. Because these observations occurred before the current action, they may appear to be natural candidates for causal identification.

Econometrics provides a clear warning against this assumption. A variable does not become exogenous simply because it is lagged. If errors are serially correlated, a previous observation may remain correlated with the current disturbance. If the earlier variable was generated by the same policy now being evaluated, it may transmit policy-induced bias into the present period.

The validity of a lag depends on the data-generating process. Under some finite moving-average structures, sufficiently distant lags may become independent of the current shock. Under persistent autoregressive processes, the influence of a shock can continue through many periods. AI systems therefore need diagnostics for error persistence and policy dependence rather than a general rule that older context is safer context.

DQN and the Limits of Recursive Learning

Deep Q-Networks estimate the value of taking an action in a state by combining current rewards with the estimated value of future states. The Bellman recursion gives reinforcement learning its dynamic structure, while replay buffers and target networks help stabilise estimation. The method does not simply sum previous states, but it does depend on information generated through earlier interactions between the policy and the environment.

If the observed state contains all information necessary for future transitions and rewards, this framework can be well defined. In practice, however, environments are often partially observed. Hidden variables may affect both action selection and outcomes. A policy trained on historical transitions can then attribute reward differences to actions that were actually selected in response to unobserved conditions.

This is the reinforcement-learning counterpart of omitted-variable bias. The estimated value function may perform well under the behaviour policy that generated the data but fail when deployed under a new policy. The system learned an association embedded in the historical action-selection mechanism rather than the invariant effect of the action itself. Causal and instrumental-variable methods become relevant when developers need policy values that remain valid outside the original data-generating process.

A Causal Layer for RLHF

A more robust RLHF architecture would separate several quantities that are usually combined: immediate user approval, predicted reward, factual accuracy, policy compliance, and long-term benefit. These outcomes can be correlated, but they should not be treated as interchangeable. The system should therefore retain multiple evaluation channels rather than compressing every objective into one reward score.

Independent evaluation is essential. Responses optimised against a learned reward model should be tested by evaluators who did not generate the original preference data, by alternative models, and under prompt distributions not used during optimisation. The purpose is to determine whether reward improvements transfer outside the feedback process that created them.

Randomisation should also remain part of the system after deployment. If every response is selected deterministically by the current policy, the model will gradually observe only the consequences of its preferred behaviour. Controlled exploration permits counterfactual comparison and reduces the risk that the system becomes trapped inside a self-confirming policy. In this sense, randomisation is not wasted inefficiency; it is part of the causal measurement infrastructure.

Provenance as an Anti-Endogeneity Mechanism

Conversational and agentic systems need to distinguish among external evidence, user-provided claims, retrieved documents, model-generated inferences, and earlier model outputs. Without this distinction, every statement in the context appears as equivalent text, even though the reliability and causal origin of each statement differ substantially.

A provenance layer should record where each material claim originated and whether it has been independently verified. A statement retrieved from an authoritative database should not be treated in the same way as a hypothesis generated by the model several turns earlier. When a later conclusion depends heavily on model-originated content, the system should trigger verification or lower its confidence.

This mechanism helps interrupt endogenous error propagation. The objective is not to remove historical context, but to prevent the model’s own prior output from masquerading as new external evidence. Repetition should not increase confidence unless the repeated claims originate from genuinely independent sources.

Counterfactual Memory for Long Conversations

Current memory systems are designed primarily to improve continuity. They retrieve earlier information judged semantically relevant to the present request. A causally disciplined memory system must ask an additional question: what would the current reasoning look like if a particular remembered statement were removed or contradicted?

Counterfactual memory testing can identify whether one early claim dominates the later chain of reasoning. The system may generate an alternative response without the contested memory, compare the resulting conclusions, and determine whether the difference is material. If the conclusion changes substantially, the remembered claim deserves stronger verification.

This approach treats memory not as a passive archive but as an active source of possible confounding. A remembered statement may be useful, irrelevant, or harmful depending on its origin and its relationship to current uncertainty. Memory architecture should therefore support deletion, isolation, and competing histories rather than assuming that a single continuous thread provides the most reliable state representation.

Reward Models Need Causal Audits

Reward models are typically evaluated by how accurately they reproduce human preference rankings. This measures predictive agreement with the labels, but it does not reveal which features caused the rankings or whether those features align with the desired objective. A reward model can achieve high predictive performance by exploiting superficial signals.

Causal auditing would test how the reward changes when content is preserved but style is altered, when confidence language is removed, when answer length changes, or when ideological framing is varied. These interventions help determine whether the model rewards correctness, persuasion, verbosity, agreement, or another hidden characteristic.

The same procedure should be applied across evaluator groups and domains. A reward pattern learned from general users may not remain valid for scientific, medical, legal, or financial tasks. The architecture should allow different reward components to be measured and governed separately rather than assuming that one universal preference model represents human values.

Engagement Is an Endogenous Objective

Commercial AI systems often rely on engagement metrics because they are continuously observable. Session length, return frequency, click-through rates, and conversational depth appear to provide objective behavioural feedback. Yet these metrics are especially vulnerable to endogeneity because system behaviour directly shapes them.

An assistant that produces emotionally reassuring or controversial responses may generate longer interactions than one that provides a restrained, accurate answer. A recommender system that repeatedly exposes users to polarising content may observe increasing engagement and infer stronger preference. The system then amplifies the behaviour responsible for creating the metric.

Engagement should therefore be treated as an outcome requiring causal interpretation, not as a direct measure of value. Developers need to distinguish engagement caused by usefulness from engagement caused by dependency, anxiety, outrage, or confirmation. This requires long-horizon evaluation, controlled experimentation, and outcome measures external to the interaction itself.

Endogeneity in Multi-Agent Systems

The problem becomes more complex when multiple AI agents interact. One agent’s output becomes another agent’s input, and the resulting decision affects the environment observed by both. Errors, strategic behaviour, and shared training biases can circulate across the system.

If several agents rely on similar models or data sources, agreement among them does not provide independent confirmation. Their errors may be correlated because they share the same latent assumptions. A multi-agent debate can therefore create the appearance of consensus while reproducing a common training-data bias.

Robust multi-agent design requires diversity of models, independent evidence channels, and mechanisms that trace the causal origin of agreement. The system should distinguish convergence produced by separate information from convergence produced by shared architecture or copied intermediate conclusions. Otherwise, the number of agents increases apparent confidence without increasing identification.

From Model Evaluation to Process Evaluation

Conventional benchmarks evaluate the final answer. Endogeneity requires evaluation of the process that generated the answer. Developers must examine which data were selected, which alternatives were excluded, how earlier outputs affected later observations, and whether the evaluation signal was independent of the model’s behaviour.

This process-level view changes the meaning of reliability. A correct answer reached through a fragile, self-confirming chain should not receive the same confidence as a correct answer supported by independent sources and robust counterfactual checks. The former may fail sharply under small changes in context, while the latter reflects a more stable reasoning process.

AI evaluation should therefore include causal stress tests. Prompts can be reordered, misleading context inserted, earlier claims removed, and alternative evidence supplied. The objective is to determine whether the system tracks the underlying facts or merely follows the correlations embedded in its current informational path.

The Architecture of Causal Guardrails

A practical causal-guardrail layer would sit between model generation and consequential action. It would record data provenance, identify policy-dependent observations, detect persistent unverified claims, and assess whether the evidence supporting an action is independent of the system’s own prior outputs.

The layer would also manage experimental variation. Randomised response selection, alternate retrieval paths, shadow policies, and independent evaluator models could generate the counterfactual data required to distinguish correlation from intervention effects. Statistical components would test for serial dependence, distributional change, reward drift, and sensitivity to hidden confounding.

Not every interaction would require full causal analysis. The architecture should allocate stronger safeguards to high-impact decisions, long-running agentic tasks, and feedback-sensitive domains. The essential principle is that causal identification must become an explicit system function rather than an informal assumption hidden inside the training pipeline.

A New Division of Labour between AI and Econometrics

Computer science has developed the architectures, optimisation methods, and computational infrastructure required for modern AI. Econometrics has developed a mature framework for analysing systems in which choices, outcomes, and expectations are jointly determined. Adaptive AI now sits at the intersection of these traditions.

AI researchers do not need to replace machine learning with instrumental-variable regression. They need to recognise which tasks are purely predictive and which require causal interpretation. Econometricians, in turn, must adapt identification methods to high-dimensional, sequential, and policy-dependent environments where conventional linear models may be insufficient.

The emerging field will require both capabilities. Neural networks can approximate complex relationships, LLMs can interpret unstructured information, and reinforcement learning can optimise sequential decisions. Causal inference determines when those relationships can support intervention. Without that layer, more powerful models may simply become more efficient at learning from feedback loops they do not understand.

Toward Causally Disciplined Agentic AI

The next generation of AI systems will retain longer memories, execute more complex tasks, and interact with users over extended periods. These capabilities increase utility, but they also make endogenous feedback more persistent. Every action can alter the state, every output can shape the user, and every stored memory can influence future interpretation.

Causally disciplined systems should therefore treat their own past actions as possible sources of bias. They should verify claims that originated internally, preserve independent evaluation channels, and maintain enough experimental variation to estimate what would have happened under alternative actions. They should also recognise when the available data cannot identify a causal conclusion and communicate that limitation.

RLHF was an important step toward incorporating human judgment into AI behaviour. It should not be mistaken for the completion of alignment or causal validation. Human feedback itself is generated within a system of incentives, exposure, interpretation, and model influence. The next step is to build AI architectures capable of examining that feedback rather than merely optimising against it.

Picture

Member for

1 year 10 months
Real name
Keith Lee
Bio
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.

His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.

He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.

From Narratives to Prices: AI and the New Data Architecture of Prediction Markets

From Narratives to Prices: AI and the New Data Architecture of Prediction Markets

Picture

Member for

1 year 10 months
Real name
Keith Lee
Bio
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.

His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.

He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.

Modified

Digital markets convert previously unpriced expectations into observable data
LLMs connect news, discussion, and policy language to real-time price movements
Prediction markets may create measurable signals for political, scientific, and social risk

Financial markets have long provided researchers with a continuous record of human expectations. Prices, volumes, spreads, volatility, and order flow reveal how investors respond to earnings, interest rates, regulation, geopolitical events, and changing perceptions of risk. These data are imperfect and often difficult to interpret, but they offer something that most other areas of social inquiry lack: an observable sequence showing how collective beliefs change over time.

Many important forms of uncertainty have historically remained outside this structure. Political instability, regulatory intervention, scientific breakthroughs, institutional failure, technological disruption, and social conflict are frequently discussed through reports, interviews, and qualitative assessments. Analysts may describe a risk as increasing or declining, but they usually cannot observe a continuously updated market value attached to that assessment. The absence of a tradable asset means that beliefs remain dispersed across documents and conversations rather than consolidated into a common numerical signal.

Digital prediction markets are beginning to narrow this gap. By creating contracts around specific future events, they transform expectations that once existed only as language into observable bids, offers, and transaction prices. The resulting markets remain smaller and less mature than conventional financial exchanges, but they introduce a potentially important form of data infrastructure. They make it possible to observe not only what people believe about an uncertain event, but also when those beliefs change, how strongly participants disagree, and which information appears to move the market.

The Expansion of Observable Market Data

For much of modern economic history, market data referred primarily to financial assets, commodities, currencies, and derivatives. These markets generated sufficiently frequent transactions to support empirical analysis, risk modelling, and the development of pricing theory. Researchers could estimate distributions, examine volatility, identify regime changes, and test how new information affected asset values. Other forms of uncertainty remained difficult to quantify because no corresponding market existed.

The development of digital platforms has changed the economics of market creation. A market no longer requires a physical exchange, specialised dealer network, or substantial operational infrastructure. A platform can define a contract, register participants, record orders, match trades, and preserve the complete history of activity within a database. Once these functions become inexpensive, markets can be created for narrower and more specialised questions than traditional exchanges would support.

This enables the emergence of segmented markets. Instead of observing only broad political or economic conditions, researchers can construct markets around individual elections, policy decisions, research outcomes, climate records, disease declarations, court rulings, or technological milestones. Each market produces a small dataset, but thousands of such markets could collectively form a new layer of information about expectations across society. The importance of prediction markets may therefore lie less in any single forecast than in the infrastructure they create for recording beliefs that were previously invisible.

Source: Kalshi - Prediction Market for Trading the Future, June 10th, 2026 (EST)

From Qualitative Risk to Measurable Expectations

Political risk provides a useful example. Firms, investors, and governments routinely assess the probability of elections, sanctions, regulatory changes, political unrest, and international conflict. These assessments influence investment decisions and strategic planning, yet the underlying reasoning is usually contained in confidential reports, expert judgment, or narrative scenarios. Because the information is fragmented and expressed in incompatible forms, it is difficult to compare expectations across time or institutions.

A prediction market can impose a common structure on these beliefs. Participants must translate their judgments into positions linked to a defined event and deadline. The market price does not become an objective measure of political risk, but it creates an observable estimate generated through interaction among multiple participants. Researchers can then study how that estimate changes after speeches, polls, diplomatic developments, court decisions, or media reports.

With sufficient data, the value extends beyond the individual contract. Price reactions across related markets may reveal how participants connect events. A new regulation may change expectations not only for one company or industry, but also for elections, public spending, technological investment, or international relations. Prediction markets could therefore help reveal the implicit structure through which participants price political and institutional risk—an area that has traditionally resisted direct measurement.

Text Has Always Moved Markets

Market prices have never been determined by numerical information alone. Earnings announcements, central-bank statements, political speeches, newspaper articles, analyst reports, rumours, and public commentary all influence expectations. The difficulty has been that text is far more complex to analyse than prices. A price is already represented in a standard numerical format, while language must be interpreted in context.

Earlier forms of text analysis relied heavily on dictionaries, keyword counts, and manually labelled datasets. These methods were useful but limited. The meaning of a sentence could change with context, negation, technical vocabulary, or institutional setting. A word such as “risk” could indicate deterioration, prudent management, or merely a formal disclosure requirement. Researchers could process large volumes of text, but the resulting classifications often lost much of the meaning contained in the original documents.

Large language models materially expand this capability. They can identify claims, entities, causal relationships, uncertainty, disagreement, and changes in tone across large collections of documents. They can compare a new statement with earlier statements, distinguish expert analysis from speculative commentary, and organise text according to the events or markets to which it is relevant. This makes it increasingly feasible to connect the information environment directly to subsequent market behaviour.

Connecting Messages to Price Movements

A prediction-market platform can record the precise time at which orders are submitted and prices change. News and discussion platforms also preserve timestamps. When these datasets are integrated, researchers can begin examining how particular forms of language affect collective expectations. A newspaper article, government announcement, scientific paper, or community discussion can be linked to the price movements that followed.

The simplest analysis would measure whether a price increased or decreased after a relevant message appeared. A more sophisticated system would examine the size, speed, duration, and distribution of the response. Did prices move immediately, or only after the information was repeated by other sources? Did trading volume increase before the price changed? Did the initial movement reverse after expert criticism appeared? Did a small group of participants react first, followed by the wider market?

LLMs can help classify the textual event that preceded each reaction. They can identify whether the message introduced new evidence, repeated known information, expressed an opinion, challenged an existing consensus, or used unusually emotional language. The resulting dataset would allow researchers to study not merely whether news moves prices, but what types of messages move them, which participants respond, and how long the effects persist.

Beyond Sentiment Analysis

The term “sentiment analysis” is often used to describe the classification of text as positive, negative, or neutral. This framework is too limited for prediction markets. A statement can be pessimistic in tone while reducing uncertainty, or optimistic while providing little new information. What matters is not simply emotional direction, but how the message changes beliefs about the probability of a defined event.

An AI system designed for prediction-market analysis would therefore need to identify informational structure. It would distinguish new facts from interpretations, separate evidence from speculation, and recognise whether a statement supports or contradicts the conditions specified in a market contract. It would also need to estimate whether the information was already reflected in the price before publication.

This opens a broader research question: how does information become price? The answer may depend on source credibility, technical complexity, participant expertise, market liquidity, and the clarity of the contract. LLMs provide a way to represent and compare the textual inputs, while market data provide the behavioural response. Together, they create an empirical setting in which the transmission of information can be observed rather than assumed.

Market Prices as Behavioural Data

Prediction-market prices are frequently interpreted as forecasts, but their value as behavioural data may be equally important. A price movement can reflect a rational response to new evidence, but it can also reveal anxiety, herding, overconfidence, political identity, or excessive attention to dramatic events. These effects are often treated as noise from the perspective of forecasting. From the perspective of behavioural research, they are central observations.

The combination of price data and textual data makes it possible to separate some of these mechanisms. If a scientific report produces a gradual adjustment consistent with its empirical findings, the market may be processing information efficiently. If a sensational headline creates a large temporary movement that later reverses, the data may reveal an attention shock. If community discussion amplifies a movement without adding new evidence, the platform may be observing social contagion.

Repeated observations across many markets could allow researchers to identify recurring behavioural patterns. Certain topics may be especially vulnerable to fear, technological optimism, political loyalty, or media amplification. Certain participants may consistently react early and accurately, while others may follow momentum. These patterns would provide a richer account of collective judgment than either surveys or final market prices alone.

Detecting Changes in the Information Regime

Financial markets are often analysed in terms of regimes. A market may shift from low volatility to high volatility, from stable expectations to crisis conditions, or from fundamental valuation to speculative momentum. Prediction markets may exhibit comparable changes. A contract can remain largely inactive and then suddenly become the focus of intense trading after a major event.

The analytical task is to determine whether a movement represents ordinary updating or a transition into a different information regime. A temporary price change may disappear once uncertainty is resolved, while a persistent change may indicate that participants have adopted a new interpretation of the event. Volume, spreads, order concentration, textual intensity, and the diversity of participating accounts may all contribute to identifying such shifts.

AI can assist by combining these heterogeneous signals. Statistical models can detect distributional changes in market data, while LLMs can identify changes in the surrounding narrative. When both the numerical and textual environments shift simultaneously, researchers may have stronger evidence that the market has entered a new state. This integrated approach could be useful not only for prediction markets, but also for financial, political, and strategic risk analysis more broadly.

The Importance of Database Architecture

The analytical potential of prediction markets depends on data design from the beginning. A platform built only to display current prices will lose much of its scientific value. Research requires detailed records of orders, cancellations, trades, positions, participant histories, contract revisions, resolution decisions, and external information events. The database must preserve the sequence through which the market evolved.

Metadata are equally important. Each contract should have a clearly defined subject, deadline, resolution source, geographic scope, and event category. External documents should be linked to relevant markets with timestamps and source classifications. Participant privacy must be protected, but behavioural continuity should be preserved sufficiently to study calibration and learning over time.

This is where digital infrastructure becomes part of the research design. The database is not merely a technical support system for the platform. It determines which scientific questions can later be answered. A poorly structured system may produce visible prices while discarding the underlying process. A research-oriented system should treat every order, message, and revision as part of a longitudinal record of collective belief formation.

A Non-Monetary Market as Research Infrastructure

Real-money markets use financial incentives to encourage participation and penalise inaccurate confidence. They also attract gambling demand, expose participants to losses, and introduce legal and regulatory concerns. For a scientific institution, a non-monetary system may provide a more appropriate starting point.

Participants could receive equal virtual capital and accumulate reputation through calibrated forecasts and successful information contribution. Performance could be evaluated across many markets rather than through a single large wager. The platform could reward consistency, early incorporation of reliable information, and transparent reasoning. This would make the service less attractive to gambling users while preserving much of the market structure required for research.

The absence of real money would not remove every distortion. Participants could still trade carelessly, follow others, or attempt to manipulate rankings. These behaviours would themselves become subjects of analysis. The key advantage is that the institution could design the system around data quality, experimental control, and participant learning rather than transaction revenue.

The Emerging Research Opportunity

The convergence of prediction markets, modern databases, and large language models creates a research opportunity that did not previously exist at comparable cost. Digital markets can generate continuous behavioural data for narrowly defined forms of uncertainty. Databases can preserve the full history of how beliefs evolve. LLMs can organise the text that influences those beliefs and connect narrative changes to market responses.

This combination may help researchers approach questions that have remained resistant to measurement. How rapidly do people incorporate scientific evidence? Which media sources exert disproportionate influence? When does uncertainty become panic? How are political and regulatory risks translated into numerical expectations? Which participants possess genuine forecasting skill, and which merely benefit from favourable outcomes?

Prediction markets will not provide definitive answers to these questions by themselves. Their prices remain products of market design, participant composition, and imperfect information. Yet they can create something valuable: a structured empirical record where previously there were only scattered opinions. Once beliefs, messages, and price movements can be observed together, uncertainty becomes more amenable to scientific analysis.

Toward an SIAI Labs Research Platform

For SIAI Labs, the strategic opportunity is not to replicate an existing betting platform. It is to build an experimental environment in which market design, statistical modelling, behavioural analysis, and AI-based text interpretation can be studied together. A non-monetary prediction system could begin with a limited number of scientific, technological, economic, and policy questions and expand as its research methods become more reliable.

The platform could compare market forecasts with expert judgments, statistical models, and LLM-generated estimates. It could test how different information sources affect prices, identify regime changes, and examine whether participants improve through repeated forecasting. Over time, the resulting data could support research in political risk, scientific forecasting, strategic intelligence, and the economics of information.

The deeper significance lies in the creation of a new data category. Financial markets made asset expectations observable. Digital prediction markets may make event expectations observable. LLMs can then connect those expectations to the language through which society interprets uncertainty. The result is not a machine that predicts the future with certainty, but a research infrastructure that reveals how humans continuously attempt to price it.

Picture

Member for

1 year 10 months
Real name
Keith Lee
Bio
Keith Lee is Professor of AI and Finance at the Gordon School of Business, Swiss Institute of Artificial Intelligence (SIAI). His primary research lies in financial mathematics and AI-driven computational science, with a focus on quantitative modeling of complex economic and financial systems. His work integrates machine learning, stochastic modeling, and data-centric methods to study structural transformations in markets and institutions.

His recent work examines the broader socioeconomic consequences of artificial intelligence, including labor markets, public finance, demographic change, institutional adaptation, and the distributional effects of technological progress.

He holds a PhD in Mathematical Finance from Boston University, and previously earned an MSc in Finance and Economics from the London School of Economics. He completed his undergraduate studies in Economics at Seoul National University under the Korea Foundation for Advanced Studies scholarship program.