Generative AI at Work: Measured Use, Abandoned Effort
Authored on
Modified
Survey data show wide but shallow generative AI adoption Chat logs misplace tasks and ignore abandoned or corrected answers Future data collection should record what happened to results

Generative AI at work has already passed the testing stage. In more than 80 percent of occupations in the U.S., at least one in five workers uses it on the job, according to the Real-Time Population Survey, which covered nearly 14,000 workers aged 18 to 64 through May 2026. The number sounds like a success and is usually presented as one. Everyday experience has a different texture: a question answered comfortably and without substance, a second and a third message to correct the mistake, an hour lost, and in the end the work done by hand, as it would have been from the start. No line in these statistics records that last step. The distance between what the tool is said to do and what it does is small but constant, and that is where time and money are spent.
Generative AI at Work: Wide Reach, Shallow Depth
Two numbers from the same survey are enough to show the pattern. The survey asks each worker which of the ten most important tasks in an occupation the job includes and which of those AI regularly helps with. In more than 80 percent of occupations at least one in five workers uses the tool, and more than 40 percent of individual tasks have adoption rates above 20 percent. Fewer than 3 percent of tasks cross the 50 percent threshold, and none reach 70 percent. Even in the most popular tasks, at least three in ten of the workers who perform them do not use AI for them. The reading that emerges almost automatically is that this is an early-stage technology, waiting for its spread to progress. There is a second, less convenient version: part of the shallow adoption may be the imprint of people who tried, were disappointed, and went back to the old way. Neither reading is confirmed by the available data.
Another finding makes easy explanations difficult. Exposure scores, estimates of what work a model can take on in principle, explain up to half of the variation in adoption between occupations and tasks, but less than 10 percent among individual workers. Two people in the same position, with the same tasks and the same tool at their disposal, come to a different decision, and what the system can theoretically do explains little of this difference. This is where the first crack opens between what is said and what is done. The public debate speaks of possibilities as if they were a property of the job itself, while in practice they are judged in the office, question by question, by someone who has to decide whether the answer is worth keeping. Sometimes it is not worth it, and that decision, with its cost in time and money, does not appear in any of these tables.
What Conversation Logs Miss About the Worker
The companies that build the models publish analyses of their own conversations, classified into tasks from the U.S. Department of Labor's O*NET catalog. The method makes sense because it allows measurement on a scale that no survey can reach. The classification system, however, sees the text of the conversation and not the person who wrote it. Without knowledge of the occupation, a chat in which a professor asks for help with trends in a dataset is filed under a task that belongs to data analysts, and the wrong task then points to the wrong occupation. The consequences can be seen in the numbers. Across 332 intermediate work activities, the correlation between the survey and the conversation data is 0.34 for Anthropic, 0.10 for Microsoft and 0.11 for OpenAI. The chat sources agree no better with one another, with correlations of 0.08 to 0.38. A further difficulty is that a task's share of chats mixes how common the task is with how many of its workers use AI, and chat data hold no non-adopters that would allow the two to be separated.

The concentration of usage differs as well. In conversations, the largest single task absorbs 15 percent to 23 percent of the total, and the ten largest tasks absorb 46 percent to 61 percent. In the survey, the corresponding numbers are 4 percent and 22 percent. Editing takes up over 15 percent of OpenAI's work-related conversations, yet only 2.4 percent of U.S. workers are in occupations for which ONET lists editing as a task. Far more workers edit texts, of course, but ONET describes their jobs through a purpose, such as preparing reports, and the classifier picks up the activity while missing the purpose. The survey's top task, directing an organization's operations, belongs to occupations that cover 43 percent of the workforce and is almost invisible in the conversations. What the conversations measure as use is largely basic activity, such as writing and searching, cut off from the work that justifies it.
When Abandonment Counts as Use
The two methods, conversation logs and the survey, share a blind spot that has nothing to do with occupation. They measure whether AI was used, not whether the answer endured as the final product. In task-based classification, a conversation with a correct answer and a conversation with an invented answer, which the user read, re-asked twice and finally threw away, land on the same list as one use of the same work. In the survey, a worker who states that AI regularly helps with a task has no obvious way to state that some attempts left the help worse than its absence. The question concerns tasks where regular help exists, and efforts that cost time instead of saving it have no clear place in that formulation. Abandonment, in other words, subtracts nothing from measured use, while the effort itself adds to it.

The cost of this invisible side of generative AI at work is specific. A question that would be closed quickly with a reliable source becomes a long dialogue with a system that insists with certainty on the wrong number, date or reference, while the subscription or per-request fee runs as normal. Worse than a mistake that is eventually detected is the need to check every answer, because that turns the help into a second job. Each round of correction adds another prompt and another careful read of the answer. At some point the choice is between one more attempt and the work itself from the beginning, and the second sometimes wins. The claim rests on everyday experience rather than data, which is precisely the gap: none of the datasets discussed here counts the number of times the last step was taken by hand.
Measuring Outcomes, not Attempts
Public bodies have already begun to rely on these measures. The U.S. Bureau of Labor Statistics uses two conversation-based measures in the AI exposure categories that accompany its employment projections, and notes that these measures do not directly observe whether workers in an occupation used AI on the job. Measures of this kind are increasingly used as proxies for AI exposure, so the agency's caveat reaches every table built on them. In April 2026 the Workforce Transparency Act was introduced in the U.S. Senate. The bill would direct the Department of Labor, the Bureau of Labor Statistics and the Census Bureau to set up a process for collecting task-level data on AI use, with the backing of Anthropic, Google, Microsoft and OpenAI. Should the bill pass, the open matter is what will be asked. A question about whether the tool was used would produce another adoption number, while a question about what happened to the result, whether it was used as delivered, corrected or thrown away, would bring to light what is currently missing.
An obvious objection is that the models are improving quickly and the disappointment is already a thing of the past. That may be so, but the survey data run only through May 2026 and describe the tools of that period, and improvement remains an impression rather than a finding as long as it is not measured as the share of results kept. The survey has limits of its own, since it covers only the ten most important tasks in each occupation and cannot resolve small occupations, which is why surveyS and logs are best read as complements. For decision-makers in businesses that deploy generative AI at work the implication is practical: the number of logins or conversations says little about productivity, while the share of responses that reached the final deliverable says a great deal and can already be recorded internally, without waiting for any government statistics.
The headline figure will probably stay at more than 80 percent of occupations with at least one in five workers using generative AI at work, because that is what the surveys and the logs can count. The same survey also shows a top task, directing an organization's operations, that sits in occupations covering 43 percent of the workforce and barely registers in chat data. What became of the answers produced along the way is a separate matter, and none of the datasets discussed here holds it. A Bureau of Labor Statistics table built partly on conversation data, and a bill that would collect task-level usage figures, would inherit the same gap unless a question about outcomes is added. Whether the share of answers that survive to the final deliverable is small or large remains unknown, and an organization that records it will hold a number that no public statistic currently offers.
This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.