© 2026 FUTURE PROOF™
Future of Work · AI at Work

AI at work: the first field experiments.

Beyond the punditry, generative AI’s workplace impact now has controlled evidence: randomized and quasi-experimental trials in real jobs. The gains are large, weirdly distributed — novices benefit most, experts sometimes get worse — and bounded by a frontier nobody can see. What the studies found, and what they mean for how organizations build skill.

TL;DR

The finding: The first controlled studies of generative AI in real work — writing tasks, customer support, consulting, software development — find large average productivity gains, roughly 15–55% depending on task. There is a consistent twist in who gains: the least-experienced workers gain most, which narrows performance gaps. But the gains hold only inside the technology’s capability frontier. On tasks just outside it, AI assistance made consultants measurably worse, because the tool is most persuasive exactly where it is wrong.

The mechanism: AI embeds top-performer patterns and hands them to everyone — instant expertise for the routine core, misleading confidence at the edges. The scarce human skills shift toward judgment: knowing when the output is wrong, and where the frontier runs today.

The product: Future Proof trains for the complement: the domain knowledge that lets people verify AI output, the judgment scenarios where the frontier bites, and the measurement that shows who can catch a confident wrong answer.

In this article

  1. 01The cornerstone experiments
  2. 02Reading the compression finding
  3. 03Not every novice, everywhere
  4. 04The frontier problem is a judgment problem
  5. 05What happens to the expertise pipeline
  6. 06What the evidence doesn’t show
  7. 07What this means for practice
© 2026 FUTURE PROOF™
The route. 7 sections, from “The cornerstone experiments” to “What this means for practice”. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

This literature matters for a reason beyond curiosity. The rollout decisions are being made now, in every firm, mostly on vendor claims and anecdote. A small body of truly controlled evidence exists. Its findings agree well enough to build policy on. The gap between what it shows and what the average rollout assumes is wide enough to be costly both ways: gains lost where the evidence says deploy, and quality harm where it says beware.

That gap has a history. No technology has drawn more workplace forecasts per unit of workplace evidence than generative AI. For its first two years, the debate ran on demos and surveys; strategy decks quoted projections at each other. What changed the standing of the debate was the arrival of real experiments: random assignment, real tasks, measured output. Economists ran them fast, because they saw that the labor question was an empirical one with parts that could be answered.

That speed deserves a note of thanks. Field experiments of this quality usually take years to design and publish. The first wave arrived within eighteen months of the technology itself: random assignment inside real firms, outcomes registered in advance, working papers out for scrutiny while the questions still mattered. Whatever the findings’ half-life, they set the bar for “evidence about AI at work”. Vendor claims should now be held to it.

The early canon is small, consistent, and strange in ways that matter greatly for workforce growth. This article walks its four cornerstone studies. It then draws the strategy conclusions that survive them — including one that lands squarely on this library’s home turf. What happens to learning and expertise when a machine hands everyone the answer?

The cornerstone experiments

Four studies anchor the canon, and each adds a different piece. A randomized task experiment sets the basic effect. A field rollout maps who gains. An adversarial design finds the boundary. The developer studies trace the lab-to-field slope. Read in order, they form something close to a complete first sketch.

Writing tasks. The first prominent randomized experiment gave college-educated professionals realistic writing jobs — press releases, delicate emails, analysis plans. Half got access to a chatbot assistant. The treated group finished about 40% faster, and their output was rated clearly higher in quality. The spread also narrowed: the weakest writers gained most, while strong writers mostly saved time (Noy & Zhang, 2023).

Customer support. The first major field study followed thousands of support agents at a software firm. The agents used an AI assistant trained on successful past conversations. Output rose about 14–15% on average — but the average hides the finding: novice and low-skill agents improved by around a third, while the most experienced agents gained little or nothing. The authors’ reading is the study’s legacy: the AI had bottled the tacit patterns of top performers and was handing them to everyone else. In effect, it compressed months of on-the-job learning into the tool. Agents with AI access also climbed the experience curve faster — two months with the assistant matched six without (Brynjolfsson, Li & Raymond, 2025).

The number

2 months ≈ 6 Support agents with the AI assistant reached in two months of tenure the performance that took unassisted agents six — the experience curve itself, compressed by borrowed patterns (Brynjolfsson, Li & Raymond, 2025).

Consulting: the jagged frontier. The most instructive design gave hundreds of BCG consultants a battery of tasks. Most sat inside current AI capability; one was carefully built to sit just outside it. Inside the frontier, AI-assisted consultants were far better — more tasks, faster, over 40% higher rated quality, with below-average performers again gaining most. Outside the frontier, the same help flipped sign: AI-assisted consultants were roughly 19 percentage points more likely to get the task wrong than unassisted colleagues, seduced by fluent, confident, wrong output. The authors named the shape “the jagged frontier”: capability borders that are real, invisible, and nowhere near where intuition draws them (Dell’Acqua et al., 2023).

Software development. The controlled task experiment on AI pair-programming found developers finishing a standard coding task about 56% faster with an AI assistant (Peng, Kalliamvakou, Cihon & Demirer, 2023). Later field experiments across thousands of engineers found smaller but reliable gains in merged work — landing in the low double digits. That is the lab-to-field discount every applied literature in this library would predict (Cui et al., 2024).

Where AI help lands — and where it flips sign Support agents — gain falls as experience risesNovice agents ≈34%Average agent ≈14%Expert agents ≈0 Consulting task — same tool, sign flips with difficultyInside frontier 40%+Outside frontier −19 pts−20 0 +20 +40 change vs unassisted control (metrics differ by study) © 2026 FUTURE PROOF™
Figure 1. Two cornerstone findings on one axis. Among support agents the gain shrinks as experience rises — roughly a third for novices, 14–15% on average, close to nothing for the most experienced (Brynjolfsson, Li & Raymond, 2025). The consulting experiment shows the same assistance adding over 40% to rated quality inside the capability frontier and costing about 19 percentage points of accuracy outside it (Dell’Acqua et al., 2023). Metrics differ by study, so read each bar against zero rather than against its neighbours. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

Reading the compression finding

A finding this steady across separate teams, tasks, and continents is rare in applied research. It demands a mechanism, not a shrug. Why should a general-purpose tool help the weakest workers most? Most tools in history did the opposite: they amplified skilled workers’ edge. The novice-uplift pattern is the most replicated feature of this literature, and its mechanism ties directly to this library’s expertise cluster.

What separates experienced from novice performers, our knowledge and deliberate-practice reviews argue, is a stored library of patterns — situations seen, responses that worked. The support-agent study makes the case that the AI assistant is that library, externalized. Trained on top performers’ conversations, it supplies the pattern the novice hasn’t yet stored, at the moment of need (Brynjolfsson et al., 2025). Experts gain little because they already own the patterns. Novices gain hugely because they are borrowing expertise on demand — a rental whose long-term terms nobody has yet seen.

For workforce strategy, the compression cuts two ways. The bright side: onboarding speeds up, performance floors rise, and the experience premium for routine competence shrinks. Firms can staff capably with greener teams. The shadow side is a question the trials flag but cannot yet answer. If the assistant performs the pattern, does the novice ever store it?

Borrowed expertise that is never made one’s own leaves workers hooked on the tool and helpless at the frontier. Early evidence from education sharpens the worry. Students who could lean on AI freely during practice did better while practicing and worse on the unassisted test afterward — the classic crutch signature this library’s desirable-difficulties review would predict. A tutor-style AI that scaffolded without answering avoided the harm (Bastani et al., 2024).

Not every novice, everywhere

The compression story has a telling counterexample that keeps it honest. A field experiment offered AI business advice to Kenyan small-business owners, and found the opposite spread. High performers turned the advice into measurable gains. Low performers — who asked about harder problems and got generic advice they could not adapt — did slightly worse (Otis et al., 2024). What squares the two results is task structure: support agents work inside a defined workflow, where the assistant’s suggestion is directly usable. Entrepreneurs must diagnose their own situation, pick what applies, and execute alone — a loose loop where the power to absorb advice, not access to it, is the binding constraint.

The pair of findings brackets the rollout question every firm faces. Where work is structured and the assistant delivers patterns ready to use, expect rising floors and compression. Where work requires diagnosing which advice applies — strategy, complex sales, architecture — expect the rich to get richer. Using the tool well there requires exactly the judgment it cannot supply. The same instrument spreads expertise in one setting and widens the gap in another. The difference is readable in advance from the task’s structure.

The frontier problem is a judgment problem

If the compression finding is the good news, the frontier finding is the warning label. The two must be held together: rollouts built on the first alone walk straight into the second. The BCG result reframes what “AI skills” means. The consultants who failed outside the frontier were not careless. They were rationally extending trust that had just been rewarded on a dozen inside-frontier tasks (Dell’Acqua et al., 2023). The tool’s fluency is constant while its accuracy is not, and nothing in the interface marks the border.

The scarce skill the experiments reveal is therefore not prompting technique, which the trials show workers pick up quickly. It is verification capacity. That means enough domain knowledge to check an answer on your own, calibrated confidence about when checking is required, and the discipline to do it against fluent output engineered to feel finished.

The catch

On the task built to sit just outside AI capability, assisted consultants were roughly 19 percentage points more likely to get it wrong than unassisted colleagues — the same tool, the same fluency, the opposite sign (Dell’Acqua et al., 2023). Budget for verification wherever you budget for licenses.

The consultants’ failure also shows why disclosure and disclaimers underperform as guardrails. Participants knew, in the abstract, that the tool could err — everyone knows. The knowledge did not survive contact with a dozen straight successes. Trust calibrates on experienced reliability, not on warnings. So the working defense must be procedural — checking built into the flow — rather than attitudinal, a reminder to be careful.

Every part of that checking skill is home territory for this library. Verification runs on the domain knowledge our background-knowledge review calls the substrate of critical thinking: you cannot check what you could not have outlined yourself. Boundary-sensing is calibration, trainable as our metacognition review describes. And the discipline is a habit built exactly the way our behavior cluster prescribes: practiced, cued, and reinforced. “AI literacy” programs built from tool tours miss all three. The trials suggest the durable curriculum is mostly the old one — deep domain knowledge plus calibrated judgment — pointed at a new failure mode that fluency was built to hide.

Average gains in the first controlled studies Different tasks report different metrics — read the sizes, not the decimalsCoding — lab task ≈56% fasterWriting tasks ≈40% fasterConsulting, in-frontier 40%+ rated qualitySupport agents ≈14–15% productivityCoding — field low double digitsOutside the capability frontier the sign flips: ≈19 points more wrong answers (Dell’Acqua et al., 2023) © 2026 FUTURE PROOF™
Figure 2. The cornerstone experiments’ headline gains, largest in controlled lab tasks and smaller in field deployments — and negative outside the frontier. Schematic after Peng et al. (2023), Noy & Zhang (2023), Dell’Acqua et al. (2023), Brynjolfsson et al. (2025) and Cui et al. (2024); read the contrast, not the decimals. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What happens to the expertise pipeline

The support-agent study holds one more result that deserves its own billing, because it touches the core subject of this library. Agents with AI access learned faster. The tool worked as an always-on coach whose hints were also worked examples (Brynjolfsson et al., 2025). Is that speed-up real learning, or just measured output with help? That is exactly the line the education experiments force. And it frames the deeper question the trials only gesture at: fields build experts by routing novices through exactly the routine work AI now absorbs.

The junior lawyer’s document review, the junior analyst’s first-draft models, the junior developer’s boilerplate — each was low-value output and high-value training at once. That repetition is how pattern libraries get built. A firm that automates the routine tier has not just changed its cost structure. It has quietly knocked out the bottom rungs of its own expertise ladder — while still needing the top rungs occupied a decade from now.

The answer cannot be to keep drudgery as a teaching tool — the economics forbid it. The answer the learning science suggests is to rebuild the training role of routine work on purpose. Use simulation-shaped practice to supply the reps the workflow no longer does, and unassisted sessions to force pattern storage rather than pattern borrowing. Gate promotion on shown solo skill, not assisted output. The firms that notice this early will run dual tracks — AI-boosted production and deliberately built training. Their rivals will discover, some years on, that the senior bench stopped refilling when the junior work disappeared.

The tool’s fluency is constant while its accuracy is not — and nothing in the interface marks the boundary. The jagged-frontier problem, after Dell’Acqua et al. (2023).

What the evidence doesn’t show

  • It doesn’t show economy-wide effects. These are task- and job-level experiments over months; aggregate productivity, employment, and wage implications remain contested territory the trials were never designed to settle.
  • It doesn’t freeze the frontier. Every capability boundary in these studies is dated; the strategic constant is that a frontier exists and moves, which argues for continuously refreshed judgment training rather than any fixed map (Dell’Acqua et al., 2023).
  • It doesn’t bless every deployment. The strong results come from deployments with structure — curated assistants, defined tasks, feedback loops. Handing licenses to a workforce is not the treatment the trials tested.
  • Novice uplift is not novice development. The trials measure performance with the tool; whether assisted novices are learning or leaning is the open question, and the early evidence says the answer depends on the assistance design (Bastani et al., 2024).

Where the evidence stops

  1. 1It doesn’t show economy-wide effects
  2. 2It doesn’t freeze the frontier
  3. 3It doesn’t bless every deployment
  4. 4Novice uplift is not novice development
© 2026 FUTURE PROOF™
The boundary. 4 limits this article draws around its own claims. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.

What this means for practice

The playbook the trials support is specific enough to write down. Deploy where the evidence deployed: high-volume, pattern-rich, inside-frontier tasks, with the assistant grounded in your own top performers’ work. That setup produced the support-agent gains, and it is the one most likely to reproduce them. Map your roles’ tasks against the frontier honestly — with pilots, not opinions. For tasks near or beyond it, install the guardrail the consultants lacked: required checking steps, second sources, and the cultural permission to distrust fluent output. Track error catches, not just speed — the metric that shows whether judgment is operating.

Measure the frontier locally rather than trusting anyone’s general map. Run periodic spot-audits: known-answer tasks — some inside, some outside current capability — flow through the assisted workflow, and you watch what gets caught. The audit prices your firm’s actual verification capacity. That is the number the BCG design showed matters most.

Then protect the learning loop the tools quietly threaten. Keep unassisted practice in the syllabus — recall without the copilot, the training-wheels-off sessions that reveal what was learned versus borrowed. Let tests tell the two apart, because with-tool output and without-tool skill now diverge by design. Invest hardest in the complement skills the trials keep pointing at: domain depth for checking, calibration for boundary-sensing, and judgment drills where confident-wrong output must be caught. The deepest lesson for L&D is almost ironic. The more able the assistant, the more valuable the knowledge in the human’s head — because that knowledge is now the only error-catching system in the loop, and the loop is trusted with weightier work every quarter.

Applied research

How Future Proof™ applies this: training the complement.

The platform builds exactly what the field experiments say stays scarce: domain knowledge deep enough to verify machine output, calibration sharp enough to sense the frontier, and judgment rehearsed on scenarios where fluent answers are wrong. Assessment separates with-tool performance from internalized capability — unassisted retrieval checks alongside applied work — so organizations can see who is learning with the assistant and who is merely leaning on it. The tools supply the patterns. We make sure your people still own enough of them to catch the confident mistake. That division of labour is the productivity dividend without the skill decay.

See complement-skills training
References

Selected papers.

This is not an exhaustive bibliography — these are the studies cited above. The full reading list is in the downloadable Science Library PDF.

The evidence, by year

  • 2015Autor
  • 2023Noy
  • 2023Dell’Acqua
  • 2023Peng
  • 2024Cui
  • 2024Bastani
  • 2024Otis
  • 2025Brynjolfsson
© 2026 FUTURE PROOF™
The evidence base. The 8 sources cited here span 2015–2025, oldest to newest. Figure © 2026 Future Proof™ — reuse permitted with attribution and a link.
  1. Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654): 187–192. PDF
  2. Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at work. Quarterly Journal of Economics 140(2): 889–942. PDF
  3. Dell’Acqua, F., McFowland, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K.R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School Working Paper 24-013. PDF
  4. Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. arXiv:2302.06590. PDF
  5. Cui, Z., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2024). The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers. SSRN working paper. PDF
  6. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. SSRN working paper (Wharton). PDF
  7. Otis, N., Clarke, R., Delecourt, S., Holtz, D., & Koning, R. (2024). The uneven impact of generative AI on entrepreneurial performance. SSRN working paper. PDF
  8. Autor, D. (2015). Why are there still so many jobs? The history and future of workplace automation. Journal of Economic Perspectives 29(3): 3–30. PDF
Try the AI engine

Build the skills the machines make scarce.

Book a 20-minute demo. We’ll show you verification-depth training, calibration measurement, and assessments that separate with-tool performance from owned capability.

8 citations Reviewed August 2026 Open peer review welcomed