© 2026 FUTURE PROOF™
Skills & the Future of Work · AI Exposure

Which jobs will AI change? Ask about tasks.

“Will AI take my job?” is the question everyone asks and research cannot answer. “Which of my tasks can AI now do?” is the question research answers rather well — and the answers have a shape almost nobody’s intuition predicts. Four decades of task economics, and what the LLM exposure studies actually measured.

TL;DR

The finding: Jobs are bundles of tasks, and technology hits tasks, not jobs. That is why the famous “47% of jobs at risk” estimate shrank to roughly 9% when analysts measured at the task level. For large language models, the best-known exposure study estimates that around 80% of workers have at least a tenth of their tasks exposed, and about 19% have half or more. Exposure concentrates in higher-wage, higher-education work — the reverse of the automation wave before it.

The mechanism: Occupations survive by re-bundling. When a technology absorbs some tasks, the human share shifts toward what machines can’t do — and historically, tasks that can’t be substituted get more valuable, not less. Exposure is a map of where work will be rearranged, not a list of who will be unemployed.

The product: An exposure map is only useful if you know what your people can do today and what adjacent skills they can reach. Future Proof supplies that half: measured skills, mapped adjacencies, and reskilling pathways that target the re-bundled roles.

In this article

  1. 01The task framework
  2. 02The 47% scare and the task-level correction
  3. 03What LLM exposure actually measures
  4. 04How to read an exposure score without fooling yourself
  5. 05From exposure map to skills strategy
  6. 06What the evidence doesn’t show
  7. 07What this means for practice
© 2026 FUTURE PROOF™
The route. 7 sections, from “The task framework” to “What this means for practice”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Every wave of automation produces the same question — will it take the jobs? — and the same two unhelpful answers. The doomsayers say yes; the boosters say no. Both answers come with confidence, both argue at the level of whole occupations and whole economies, and both were refuted by the last five technology shifts.

The research tradition that actually made progress did so by refusing the question’s unit of analysis. Jobs, it turns out, are the wrong object to study. A job is a bundle of dozens of distinct tasks. Technologies arrive with sharp opinions about specific tasks and no opinion at all about job titles. Everything useful the field has learned follows from taking that seriously.

This article traces that idea from its origin to the present. It covers the task framework that reshaped labor economics, and the famous 47% estimate with the correction that cut it by four-fifths. It ends with the new wave of exposure studies mapping what large language models can reach. The destination is practical rather than prophetic. Exposure research, read correctly, is not a doom index, and was never designed as one. It is the best targeting map a reskilling strategy can buy — provided the reader knows exactly what the numbers do and do not measure, which is what the next few sections establish.

The task framework

The mental toolkit for thinking clearly about AI and work was built twenty years before ChatGPT, in a paper about spreadsheets and assembly lines. It remains the single most useful idea a workforce planner can borrow from economics.

The modern analysis begins with Autor, Levy and Murnane. In 2003, they asked precisely what computers could and could not do, and matched the answer against what workers in different occupations actually spend their time on. Computers, they argued, excel at routine tasks: work that can be fully described by explicit rules, whether mental (bookkeeping, filing) or manual (repeated assembly). They struggle with non-routine tasks that demand flexibility, judgment, or people skills. Tracking U.S. task content over three decades, they showed labor input shifting exactly as the model predicted. Routine task input fell as computing got cheap; non-routine analytic and interactive input rose (Autor, Levy & Murnane, 2003).

The framework explains what occupation-level thinking cannot, and its predictions have aged well across every technology since. Bank tellers survived the ATM by re-bundling toward relationship work. Radiologists absorbed computer-aided detection without vanishing. The general principle appears in Autor’s later synthesis of two centuries of automation worry: occupations rarely die of automation; they shed the automated tasks and concentrate on the rest. And because the remaining tasks are now scarcer relative to demand, they often command better pay (Autor, 2015).

Acemoglu and Restrepo formalized the full accounting. Technology displaces labor from some tasks. But it also reinstates labor by creating new tasks that did not exist — the web designer, the MRI technician, the data engineer. The net employment effect is a race between the two forces, not a one-way subtraction (Acemoglu & Restrepo, 2019).

The 47% scare and the task-level correction

The framework’s value is easiest to see in the episode where it was ignored. In 2013, Frey and Osborne asked machine-learning experts to rate how automatable 70 occupations were, then extended the ratings to 702 via a classifier. Their conclusion: 47% of U.S. employment sat in occupations at “high risk” of computerization (Frey & Osborne, 2017). The number detonated in public debate. It still echoes through boardroom slide decks a decade later.

The correction came from applying the field’s own founding insight. Arntz, Gregory and Zierahn redid the analysis at the task level, using survey data on what workers in each occupation individually do. The reframing collapsed the estimate to about 9% of jobs across OECD countries with a mostly automatable task profile. The reason: even “high-risk” occupations turn out to contain large shares of interaction, judgment, and problem-solving that the occupation-level label averaged away (Arntz, Gregory & Zierahn, 2016). Same technology assumptions, same economy — a five-fold difference, produced entirely by the unit of analysis. It is the cleanest lesson on method this literature owns: any claim about “jobs” that has not been broken into tasks is numerically untrustworthy.

The number

47% → ≈9% What happened to the share of jobs “at high risk” when the same question was re-asked at the task level — a five-fold collapse produced entirely by the unit of analysis (Frey & Osborne, 2017), (Arntz, Gregory & Zierahn, 2016).

Why does the larger number still circulate a decade after its correction? Partly because fear travels better than method. Partly because the correction is less quotable — “it depends on the task mix of each role” fits no headline. But firms that plan against the scary number make real errors in both directions. They over-prepare for wholesale job loss that the task data does not support. And they under-prepare for the pervasive partial re-bundling that it does.

0% 25% 50% 75% 100% “High-risk” jobs, occupation-level (2013) 47% Same question, task-level (OECD 2016) ≈9% Workers with ≥10% of tasks LLM-exposed ≈80% Workers with ≥50% of tasks LLM-exposed ≈19%LLM era — exposure, not automation (Eloundou et al.) Share of employment / workers © 2026 FUTURE PROOF™
Figure 1. Two eras, one lesson. The occupation-level 47% collapsed to ≈9% when measured at the task level. LLM exposure estimates are broad but describe task overlap — how much of a job the technology can touch — not predicted job loss. Estimates as reported; definitions differ across studies — see references. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What LLM exposure actually measures

With the task discipline in hand and its cautionary tale told, the LLM-era studies can be read for what they are. They are the most careful attempt yet to map a general-purpose technology onto the actual anatomy of work — published while the technology was still moving.

Large language models triggered a new generation of exposure studies, and the best known applied the task discipline from the start. Eloundou, Manning, Mishkin and Rock rated the tasks of every U.S. occupation. The test: could access to an LLM cut the time to complete the task by at least half, at equal quality? Their headline estimates: roughly 80% of the workforce has at least 10% of tasks exposed, and about 19% has 50% or more. And — the finding that inverted a generation of intuition — exposure rises with wage and education (Eloundou et al., 2024). The routine-task era hit the middle of the wage range; the language-model era reaches writing, analysis, programming, and communication — the task diet of the professional class.

The finding did not come from nowhere, and its agreement across independent methods is what gives it weight. Felten, Raj and Seamans had already built occupational AI-exposure indices. They linked specific AI capabilities — language modeling, image recognition, translation — to the abilities each occupation requires, and found white-collar analytic work most exposed to the language-centric ones (Felten, Raj & Seamans, 2021). Webb mined the overlap between AI patent texts and job-task descriptions, and reached the same inversion. Unlike software and robots before it, AI’s patent footprint points up the skill ladder (Webb, 2020).

Read the fine print, though, because the studies do so more carefully than the coverage they get. Exposure means the technology can plausibly touch the task — speed it up, draft it, check it. It does not distinguish augmentation from replacement; a fully “exposed” legal-research task may mean a paralegal doing five times the research, not a fifth of the paralegals. Brynjolfsson and Mitchell made the deeper point before the LLM wave: knowing what machine learning can do technically says little about what will make economic sense to deploy. That depends on error costs, on process redesign around the tool, and on where a task sits inside a workflow (Brynjolfsson & Mitchell, 2017).

Who the technology reaches, by wage level Task exposurelower-wage tasks mid higher-wage tasks software & robot era: peaks mid-wage LLM exposure: rises with wageGradients after Autor et al. (2003), Webb (2020), Eloundou et al. (2024) © 2026 FUTURE PROOF™
Figure 2. The inversion: routine-era automation hit the middle of the wage distribution, while language-model exposure rises with wage and education — the task diet of the professional class. Schematic after Autor et al. (2003), Webb (2020) and Eloundou et al. (2024); read the contrast, not the decimals. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

How to read an exposure score without fooling yourself

A worked example makes the reading discipline concrete — and shows why the same score supports three different futures. Suppose a financial analyst’s role breaks into ten tracked tasks. An exposure index rates six of them — data assembly, first-draft reporting, summaries, routine variance notes, chart production, meeting notes — as highly exposed. It is tempting to read that as “60% of this job is gone.” The task framework licenses three different readings, and history has produced all three.

Reading one: substitution with headcount effects — fewer analysts produce the same reporting. Reading two: substitution with demand effects. Reporting gets so cheap that the firm consumes far more of it, and analyst time shifts to interpreting the larger flow. This is the bank-teller path, where falling task cost expanded the surrounding business (Autor, 2015). Reading three: quality escalation. The same reports get produced faster, and the freed hours flow into the four unexposed tasks — stakeholder work, judgment calls, anomaly hunting, model scrutiny — which were always the binding constraint.

Which reading happens is not a property of the technology. It is a property of demand elasticity, error tolerance, and management choice (Brynjolfsson & Mitchell, 2017). An exposure score tells you where the fork in the road is. It does not tell you which branch your firm will take — that decision, and the skills to execute it, remain stubbornly human.

Why it matters

Read every exposure score as a fork, not a forecast. The same “60% exposed” role can end in fewer people doing the old volume, the same people doing far more of a newly cheap task, or the freed hours flowing into the judgment work that was always the bottleneck — and which branch materializes is a management choice, not a property of the model.

Tasks that cannot be substituted by automation are generally complemented by it. David Autor, Journal of Economic Perspectives, 2015

From exposure map to skills strategy

Everything to this point has been about reading the research correctly. The remaining question is what to do with it — and here the exposure literature turns from commentary into an operating tool.

Treated as a forecast of unemployment, exposure indices overreach. Treated as a planning tool, they are the most useful artifact this literature has produced. They tell a firm, occupation by occupation, where work is about to be re-bundled. That is exactly the information a skills strategy needs, and almost never has in time.

The logic runs in three steps. First, high-exposure task lists show which parts of each role will compress: the drafting, summarizing, boilerplate analysis, and first-pass code that models increasingly handle. Second, the complement principle shows what expands to fill the space: judgment over model output, problem framing, quality checking, client work, and the domain knowledge that lets someone notice a fluent answer is wrong (Autor, 2015). Third, the reinstatement principle says entirely new tasks will appear around the technology itself — evaluation, integration, oversight — and someone will have to staff them (Acemoglu & Restrepo, 2019). Every step is a training target. None of it is answered by the exposure score alone: you must know which of your people hold which skills, and which nearby skills each could reach fastest.

The inverted wage gradient carries one more strategic lesson, and HR planning has been slow to absorb it. In the routine era, reskilling programs pointed at the factory floor and the back office. This time the most exposed task portfolios sit in marketing, law, finance, software, and analysis. Those groups have rarely been the object of systematic reskilling, inside firms that have never had to re-bundle professional work at scale. The firms that treat that as a measurement problem now, before the re-bundling forces it, are buying their transition at a discount.

What the evidence doesn’t show

Exposure research is genuinely useful and easily oversold, sometimes in the same slide deck and occasionally in the same sentence. Five boundaries keep the planning honest:

  • Exposure is not adoption, and adoption is not displacement. The indices measure technical overlap between model capability and task content. Deployment depends on cost, liability, regulation, and workflow redesign — the economics Brynjolfsson and Mitchell flagged as the binding constraint (Brynjolfsson & Mitchell, 2017).
  • No employment outcomes yet. The LLM exposure indices are too young to have been validated against realized job or wage changes. The previous generation’s record urges humility: the 47% forecast is now old enough to grade, and the mass technological unemployment it implied has not appeared (Frey & Osborne, 2017) (Arntz, Gregory & Zierahn, 2016).
  • Expert capability ratings drift. Both the 2013 and 2023 exercises rest on judgments about what the technology can do — judgments that model progress can invalidate in either direction, and that GPT-4-assisted rating (used as a robustness check in the LLM study) makes newly recursive (Eloundou et al., 2024).
  • Task lists miss tacit work. Occupational databases record what jobs officially involve. The undocumented glue — relationship maintenance, error-catching, institutional memory — is precisely the hardest to automate and the least likely to appear in the data.
  • The net is not the gross. Displacement and reinstatement run simultaneously (Acemoglu & Restrepo, 2019); an economy-level “safe” verdict is fully compatible with wrenching transitions for specific occupations, firms, and people. Averages don’t retrain anyone.

Where the evidence stops

  1. 1Exposure is not adoption, and adoption is not displacement
  2. 2No employment outcomes yet
  3. 3Expert capability ratings drift
  4. 4Task lists miss tacit work
  5. 5The net is not the gross
© 2026 FUTURE PROOF™
The boundary. 5 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What this means for practice

For a workforce-planning team, the literature compresses to four instructions. Break roles into tasks before you decide — no role-level AI judgment survives contact with its own task list. Map exposure against your actual roles, not national averages; the same job title carries different task bundles in different firms. Target the complements — the judgment, checking, and people tasks that expand when drafting compresses. Treat them as trainable skills with measurable mastery, because they are. And measure the transition: re-bundling shows up first as skill gaps, and a firm that measures skills continuously sees the shift while it is still a training problem rather than a layoff program.

The cadence matters as much as the checklist. Exposure maps date quickly — each capability jump re-rates tasks that last year’s analysis called safe — so the exercise is a standing review, not a one-time audit. The durable asset is not any particular map. It is the muscle of knowing, at any moment, what your people can verifiably do and how fast they can learn the next adjacent thing. Firms that built that muscle for other reasons will find the AI transition merely demanding. Firms that manage skills by job title and annual review will read their exposure the way the 2013 audience read the 47% — too coarsely, and too late.

Applied at Future Proof

How Future Proof™ applies this.

Exposure indices tell you where work will be re-bundled; Future Proof supplies the other half of the decision. The knowledge map holds a live, measured picture of what each person can actually do — not what their job title implies — so exposure maps land on real skill inventories instead of org charts. Reskilling pathways are built along skill adjacencies, verified by assessment rather than attendance, and the analytics show transition readiness by team and role while it is still cheap to act on.

See the knowledge map
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above.

The evidence, by year

  • 2003Autor
  • 2015Autor
  • 2016Arntz
  • 2017Frey
  • 2017Brynjolfsson
  • 2019Acemoglu
  • 2020Webb
  • 2021Felten
  • 2024Eloundou
© 2026 FUTURE PROOF™
The evidence base. The 9 sources cited here span 2003–2024, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Autor, D.H., Levy, F., & Murnane, R.J. (2003). The Skill Content of Recent Technological Change: An Empirical Exploration. Quarterly Journal of Economics 118(4): 1279–1333. DOI
  2. Autor, D.H. (2015). Why Are There Still So Many Jobs? The History and Future of Workplace Automation. Journal of Economic Perspectives 29(3): 3–30. DOI
  3. Acemoglu, D., & Restrepo, P. (2019). Automation and New Tasks: How Technology Displaces and Reinstates Labor. Journal of Economic Perspectives 33(2): 3–30. DOI
  4. Frey, C.B., & Osborne, M.A. (2017). The future of employment: How susceptible are jobs to computerisation? Technological Forecasting and Social Change 114: 254–280. DOI
  5. Arntz, M., Gregory, T., & Zierahn, U. (2016). The Risk of Automation for Jobs in OECD Countries: A Comparative Analysis. OECD Social, Employment and Migration Working Papers No. 189. PDF
  6. Brynjolfsson, E., & Mitchell, T. (2017). What can machine learning do? Workforce implications. Science 358(6370): 1530–1534. DOI
  7. Felten, E., Raj, M., & Seamans, R. (2021). Occupational, industry, and geographic exposure to artificial intelligence: A novel dataset and its potential uses. Strategic Management Journal 42(12): 2195–2217. DOI
  8. Webb, M. (2020). The Impact of Artificial Intelligence on the Labor Market. Working paper, Stanford University. PDF
  9. Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science 384(6702): 1306–1308. PDF
Try the AI engine

Map your exposure to your actual skills.

Book a 20-minute demo. We’ll show you a live skills inventory against role task-bundles — and the reskilling pathways that target the complements, not the casualties.

9 citations Reviewed August 2026 Open peer review welcomed