If AI Does the Junior Work, Where Do Experts Come From

Author
Ravi Prajapati

AI is absorbing the entry-level work that built professional judgment. What the research actually shows about expertise, apprenticeship and the pipeline ahead.
For most of the modern professions, there was an unspoken exchange at the bottom of the org chart. The firm got cheap hands for work nobody senior wanted to do: the document review, the reconciliation, the test suite, the competitor deck. The junior got something that never appeared on an invoice. They got thousands of hours of low-stakes exposure to real problems, and permission to be wrong about them.
That second half of the exchange is now quietly being cancelled, and mostly by accident. Firms are automating the tasks. Almost nobody is deliberately replacing what those tasks were secretly doing.
The interesting question is not whether AI takes junior jobs. It is whether a profession can keep producing experts once the beginner stage has been optimized away.
The bottom rung is thinning. The cause is genuinely contested.
Start with the strongest data. Researchers at the Stanford Digital Economy Lab, using ADP payroll records covering millions of US workers, have tracked employment by age and AI exposure since ChatGPT's release. In their revised August 2026 paper, Erik Brynjolfsson, Bharat Chandar and Ruyu Chen report that employment among workers aged 22 to 25 in the most AI-exposed occupations sits roughly 19% below where it would be had it tracked their less-exposed peers. Workers with more experience show no comparable gap.
Three details matter more than the headline number.
First, the researchers find no evidence of widespread economy-wide displacement. This is a compositional shift, not a jobs apocalypse. Second, the gap operates through reduced hiring rather than increased firing, which is exactly the signature you would expect if firms were quietly deciding they need fewer beginners rather than deciding to remove the ones they have. Third, and most usefully, the declines concentrate in occupations where AI is used to automate tasks rather than augment them. Where humans and models collaborate, the effect is muted.
That third finding is the one operators should tattoo somewhere visible. It suggests the damage tracks how AI is deployed, not how capable it is.
Now the complication. In May 2026, Peter John Lambert and Yannick Schindler at the LSE Centre for Economic Performance published The Broken Ladder, analysing 243 million new hires and 407 million job postings across the US, UK, Canada and Australia from 2017 to 2025. Measured separately, generative-AI exposure and work-from-home exposure each predict a fall of roughly five percentage points in the junior share of new hires. Measured together, the remote-work effect holds and the generative-AI coefficient attenuates sharply, often to statistical insignificance.
Their proposed mechanism deserves attention: remote work raises the cost of supervision, monitoring and on-the-job learning, and those costs fall hardest on the people who need the most oversight.
So we have two credible research groups, two large datasets, and two different culprits.
Anyone telling you this is a settled AI story is overselling. But notice what the two explanations share. Whether the cause is a model absorbing the tasks or a distributed workforce making supervision expensive, the thing being destroyed is identical: proximate, supervised, failure-tolerant early work. The debate about the cause obscures a quiet agreement about the mechanism. That is why the rest of this argument survives either verdict.
The firm-level signals point the same direction, though they are reported rather than measured. UK graduate recruitment at the Big Four accounting firms fell sharply between 2022 and 2024, with KPMG's intake down around 29%, and the Financial Times has reported further cuts and executive commentary attributing part of this to AI adoption. Treat these as evidence that decision-makers believe the story, which has its own consequences, rather than as proof of the story.
Junior work was never just cheap labour
Here is what the productivity framing misses.
The traditional path to expertise runs: repetition, then mistakes, then feedback, then pattern recognition, then judgment. The economically worthless part of that sequence, from the firm's point of view, was always the mistakes. From the learner's point of view, the mistakes were the entire point.
Michael Polanyi's observation that we know more than we can tell is the reason this cannot be shortcut with documentation. Tacit knowledge, the sense that a contract clause is unusual before you can articulate why, that a number in a model is wrong before you have checked it, transfers through exposure and correction rather than instruction.
We have unusually good evidence for what happens when technology severs that exposure, and it does not come from AI at all.
In a two-year ethnography across five hospitals plus blinded interviews at 13 top-tier US teaching hospitals, published in Administrative Science Quarterly, Matt Beane studied how surgical residents learned robotic surgery. In open surgery, the trainee was physically necessary: someone had to hold the incision. Necessity forced participation, and participation produced skill. Robotic systems concentrated the work at a single console. Residents rotated through cases, logged hours and received certifications on schedule. On paper the training pipeline was intact.
In the operating room it was not. Beane found residents graduating licensed to operate while missing tacit competencies their predecessors had absorbed almost invisibly. The ones who did develop real skill got there through what he called shadow learning: specialising prematurely, rehearsing in simulators without proper supervision, and seeking out what he termed "undersupervised struggle."
Three things about that finding should worry anyone redesigning junior roles right now. The formal curriculum stayed intact while the learning collapsed. The problem was invisible in the credentialing data. And the successful adaptation was rule-breaking, individual, and unevenly distributed, which means it advantaged those already advantaged.
Why are entry-level jobs important for developing experts?
Entry-level work historically functioned as training infrastructure, not just cheap production. Repetitive junior tasks gave beginners high-volume exposure to real problems at a scale where errors were survivable and correctable. That cycle of attempt, failure and feedback is how tacit judgment forms. Remove the tasks and you remove the failures that generated the expertise.
Producing expert output is not the same as becoming an expert
This is the distinction the entire debate turns on, and there is now a clean experiment for it.
Hamsa Bastani, Osbert Bastani and colleagues ran a preregistered randomised controlled trial with nearly a thousand high school maths students in Turkey, published in PNAS in 2025. Three arms: no AI, a standard ChatGPT-style interface ("GPT Base"), and a version with prompts designed to protect learning ("GPT Tutor").
During assisted practice, the AI groups crushed the control group. GPT Base improved performance by 48%. GPT Tutor by 127%.
Then the researchers took the tools away and ran an unassisted exam. The GPT Base students scored 17% worse than students who had never had AI access at all. Not worse than they might have; worse than the control group. The students had used the model as a crutch, produced excellent work, and learned less than if they had struggled.
Two further details make this the most important study in this debate. First, the skill-gap narrowing that AI produced during practice did not persist once access was removed. Second, students' own assessments of how much the tools had helped them learn were overly optimistic.
That second point generalises uncomfortably. When METR ran a randomised controlled trial with 16 experienced open-source developers across 246 real tasks in their own repositories, the developers predicted AI would make them 24% faster, and after finishing still believed it had made them about 20% faster. Measured completion times showed they were 19% slower. (METR has since said it believes developers are likely more sped up with early-2026 tools, but that selection effects in its follow-up make the size of that improvement unreliable. Take the slowdown as a finding about one moment, and the perception gap as the durable lesson.)
Anthropic's own June 2026 Economic Index reports that 68% of surveyed users say they learn more when using AI, and that heavy delegators report learning at the same rate as everyone else. The report then adds the caveat most coverage dropped: these are self-assessments, and "skills can erode even as they become more valuable and as someone reports learning more, so the data do not rule out skill erosion."
Put the three together and you get a genuinely awkward conclusion for anyone running a team. Self-report is not a valid instrument here, and output quality is not either. Both read healthy in precisely the scenario where capability is being hollowed out. The dashboards will look fine for years.
Can AI help junior employees become experts faster?
Sometimes, and it depends almost entirely on deployment design rather than model capability. In the PNAS trial, the AI tutor built with learning safeguards produced no measurable learning harm, while the unrestricted chatbot left students 17% worse off on unassisted exams. Same underlying model, opposite outcomes. The variable was whether the tool gave answers or forced reasoning.
The Verification Paradox
Ask any executive what juniors will do now and you get a version of the same answer: they will supervise the AI. Review its output. Catch its errors. Add judgment on top.
It sounds obvious. It is close to backwards.
Reviewing AI output well requires you to know what a correct answer looks like without producing one, to spot the plausible-but-wrong result, and to sense that a confident, well-formatted, entirely fluent output is subtly off. That is not a beginner capability. That is the thing seniority is.
The BCG experiment makes the cost visible. Fabrizio Dell'Acqua and colleagues ran a preregistered study with 758 BCG consultants, about 7% of the firm's individual-contributor consultants. On tasks inside AI's capability frontier, consultants using GPT-4 completed 12.2% more tasks, worked 25.1% faster, and produced work rated around 40% higher in quality. On tasks that fell outside that frontier, consultants using AI performed 19 percentage points worse than those working without it.
The researchers' explanation is the part that matters here. The consultants who did badly tended to adopt the AI's output uncritically and interrogate it less. The frontier is jagged and unmarked. You cannot tell from the output which side of it you are on. Detecting that requires domain judgment.
So the proposed new junior role has a circular dependency at its centre:
To verify AI output, you need judgment. You used to acquire judgment by doing the work. The work is what AI now does.
This is not an argument that juniors cannot review AI output. It is an argument that doing so is a senior task being handed to juniors on the assumption it is a simple one, and that assigning it without redesigning how judgment gets built is how organisations end up with people who are fluent, productive, and unable to tell when they are wrong.
The AI Expertise Ladder
A framework for locating where the damage actually happens. Each rung answers a different question, and each was historically built by doing the rung below it repeatedly.
Rung | The question it answers | AI capability today | Where humans used to build it | Status |
|---|---|---|---|---|
1. Execution | Can I produce the thing? | High and rising | Repetition at volume | Largely absorbed |
2. Verification | Is this output correct? | Poor; models are confidently wrong | Catching your own errors after making them | Assigned to juniors, unsupported |
3. Diagnosis | Why did this fail? | Partial; good on known failure modes | Debugging your own mistakes | At risk |
4. Judgment | Which technically valid option is right here? | Weak; requires context AI lacks | Watching consequences play out over time | Human |
5. Stewardship | What do I own when this goes wrong? | None; accountability cannot be delegated | Carrying responsibility for real outcomes | Human |
The framework's point is not that AI stops at rung 2. It is that rung 1 was never the destination. It was the staircase. Organisations are automating the staircase while continuing to recruit for the upper floors, and treating the resulting gap as a hiring problem.
The Stanford finding gives this teeth. Employment declines concentrated where AI automates rather than augments. Automation removes rung 1 entirely. Augmentation keeps the human on rung 1 while raising the ceiling. Same technology, different pipeline consequences, and the difference is a deployment decision rather than a technological inevitability.
Anthropic's usage data suggests which way the default is drifting: the share of "directive" interactions, where a user hands over a whole task with minimal back-and-forth, rose from 27% to 39% over eight months. In its June 2026 report, the median Claude Code session producing a blog post contained a single human prompt, against 13 rounds of back-and-forth for the same output in chat.
What this looks like across five professions
Profession | Traditional junior work | What AI now does well | What that work was secretly teaching |
|---|---|---|---|
Software | Bug fixes, tests, documentation, small tickets | Generates all four competently | How this system actually behaves under stress; why the last person made that choice |
Law | Document review, legal research, first-draft contracts | Retrieval, summarisation, clause comparison | Volume-based intuition for what is standard, and therefore what is anomalous |
Accounting & audit | Reconciliation, sampling, working papers | Matching, anomaly flagging, schedule preparation | Feel for when a number is wrong before you can prove it |
Consulting | Research, benchmarking, first-cut analysis, slides | Fast and fluent across all of it | Recognising a weak analysis, because you built weak ones and were corrected |
Marketing | Copy variants, SEO research, campaign reporting | High-volume production, competitor scraping | Which messages fail, learned by shipping some that did |
The pattern holds across all five: the automated task was the cheapest available source of high-volume feedback in that profession. That is what is being removed, and no profession has replaced it yet.
The strongest case that this argument is wrong
This is the section the topic usually skips, so let me make the opposing case properly.
AI demonstrably helps novices most. Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,172 customer support agents in research published in the Quarterly Journal of Economics in 2025. Access to an AI assistant raised issues resolved per hour by about 15% on average, but the distribution is the story: less experienced and lower-skilled workers improved in both speed and quality, while the most experienced saw small speed gains and small quality declines. The authors argue the model disseminates the tacit best practices of the firm's strongest agents, and find suggestive evidence of genuine learning, with newer workers moving down the experience curve faster.
If AI can extract tacit knowledge from top performers and hand it to beginners, it could be the most scalable apprenticeship mechanism ever built. That is not a small possibility, and it is supported by a peer-reviewed study with a large sample.
Design solves at least part of the learning problem. The PNAS experiment did not just find harm. Its tutor arm, same model with learning safeguards, produced the largest performance gains of any group and no detectable learning penalty. The harm was a property of the interface, not the technology.
The causal case is weak. As noted, the LSE analysis finds remote work is a more robust predictor of the early-career hiring decline than generative AI exposure. If that holds, the remedy is managerial rather than technological, and considerably more tractable.
Practitioners closest to AI are not panicking about themselves. In Anthropic's survey of roughly 9,700 linked respondents, the people who delegate most heavily to AI are the most optimistic about their pay, security and job prospects, and 57% say AI has made their skills more valuable.
Where the counterargument runs out. Two places. First, the same Anthropic survey shows respondents are markedly more worried about others than themselves: over a third put the probability of a junior colleague losing their job in the next year above 60%, against 10% who rated their own job loss likely. The optimism is not evenly distributed down the ladder. Second, and more fundamentally, every optimistic finding above measures performance while the tool is available. The one study that removed the tool found a 17% deficit. Until we have workplace research that measures unassisted capability after a period of AI-assisted work, the optimistic evidence and the pessimistic evidence are not actually in conflict. They are measuring different things.
What companies should do before cutting junior roles
Not a maturity model. Four decisions with evidence behind each.
1. Track your automation-to-augmentation ratio, not your AI adoption rate. The Stanford data locates employment damage specifically where AI automates rather than augments. Adoption rate tells you nothing about pipeline risk. This ratio does, and you control it.
2. Measure unassisted capability at intervals. This is the only recommendation here that is genuinely uncomfortable, and it is the most important. Every perception-based instrument in the evidence base points the wrong way: students, developers and self-reporting professionals all overestimated their gains. If you never test people without the tools, you will not discover a capability gap until you need someone to handle an exception.
3. Redesign junior work around the failure, not the output. The learning input was never the task; it was the correction.
Old junior role | Redesigned junior role |
|---|---|
Produce first draft | Predict what the AI will get wrong, then run it |
Get corrected by a senior | Diagnose why the output failed, in writing |
Repeat until fluent | Defend a judgment call against a senior who disagrees |
Success = accepted output | Success = correctly identifying the one flawed output in five |
That last row is the operational core. It is directly testable, it cannot be faked with a good model, and it measures rung 2 rather than rung 1.
4. Fund the supervision, not just the headcount. If the LSE analysis is right and supervision cost is the real driver, then hiring juniors without funding proximate mentoring reproduces the problem at greater expense. Beane's surgical residents had the rotations. What they lacked was a role in the work that made their participation necessary.
What junior professionals should do
Do the work before you check the work. Attempt problems unassisted first, then compare against the model. The gap between your answer and its answer is the highest-density feedback available to you, and it disappears entirely if you prompt first.
Optimise for exception handling. The value is concentrating in the cases where the model fails. Deliberately seek the messy, contested, under-documented problems that will not be automated soon.
Get proximate. If remote-work supervision cost is a real driver of junior hiring reluctance, then physical or synchronous access to senior people is a career asset, not a lifestyle preference.
Build a verifiable track record of judgment. Not output volume. Anyone can produce volume now. Documented instances of catching something wrong are the scarce signal.
Treat shadow learning as the warning it is. Beane's residents who succeeded broke rules to do it, which means the adaptation favoured the confident and well-connected. Do not assume the informal path stays open.
So, where will senior experts come from?
The honest answer, for the next few years, is that they will come from the pipeline that already exists. Today's seniors were trained under the old regime. Firms will not feel this for a while, which is exactly why the incentive to act now is so weak. The productivity dividend arrives this quarter. The expertise deficit arrives in someone else's tenure.
After that, three futures, and we do not yet know which.
In the first, the pessimistic case holds: firms optimise the beginner stage away, tacit knowledge fails to transfer, and a decade from now the profession discovers it has plenty of fluent operators and very few people who can tell when the fluent output is wrong.
In the second, the Brynjolfsson result generalises: AI turns out to be an effective mechanism for transmitting expert practice, juniors move down the experience curve faster than any previous cohort, and the worry looks in hindsight like every previous panic about calculators and compilers.
In the third, and I suspect most likely, the outcome splits by organisation. Some firms will redesign junior work around verification and diagnosis, will actually measure unassisted capability, and will produce experts faster than before. Others will bank the productivity gain, report excellent numbers, and discover the hole a decade later when it is expensive and slow to fix. The technology will be identical in both. The deployment choice will not be.
Which returns to the thing that makes this genuinely different from previous automation waves. Calculators removed arithmetic but left mathematical reasoning intact. Compilers removed assembly but left system design intact. This is the first tool that competently performs the entry stage of expert work while leaving the expert stage to humans, and expertise has never been a set of separable layers. It has been a staircase.
Beane's surgical residents got their certifications on time. Every measurement said the training worked. That is the part worth sitting with: the mechanism failed silently, and the instruments that were supposed to detect it kept reporting success.
We are about to run that experiment across most of the professional economy at once. The least we could do is measure the right thing while it happens.
FAQ
Will AI replace entry-level jobs entirely?
Unlikely, but the composition is shifting. Stanford Digital Economy Lab research through June 2026 found employment for 22-to-25-year-olds in highly AI-exposed occupations about 19% below where it would otherwise be, with no comparable gap for experienced workers and no evidence of economy-wide displacement. The effect runs through reduced hiring rather than layoffs, and concentrates where AI automates rather than augments tasks.
Is AI definitely the cause of falling entry-level hiring?
No, and this is genuinely contested. A 2026 LSE Centre for Economic Performance paper analysing 243 million hires found that when AI exposure and remote-work exposure are tested together, the remote-work effect holds while the AI effect often becomes statistically insignificant. Both explanations point to the same underlying mechanism: the loss of proximate, supervised early-career work.
Does using AI make junior workers more skilled?
It makes their output better; whether it makes them more skilled is a separate question. A PNAS randomised trial found students using an unrestricted AI tutor performed 48% better during practice but 17% worse than a control group once the tool was removed. A version designed with learning safeguards eliminated that penalty, so deployment design appears to matter more than model capability.
Which junior roles are most exposed to AI?
Roles built on high-volume, well-specified production: document review and legal research, reconciliation and audit preparation, first-draft analysis and slide production in consulting, routine tickets and test-writing in software, and content and reporting work in marketing. The common factor is that the task is codifiable and produced at volume, which is also what made it useful for learning.
Could AI cause a shortage of senior experts later?
It is a plausible risk rather than an established forecast, and no dataset can confirm it yet because the lag is measured in years. The mechanism is documented elsewhere: Matt Beane's study of robotic surgery found residents completing certification on schedule while missing the tacit skills their predecessors acquired through hands-on participation.
How should companies train juniors in the AI era?
Shift junior work from producing output to interrogating it: predicting where a model will fail before running it, diagnosing why an output was wrong, and defending judgment calls. Critically, test unassisted capability periodically. Self-reports and output quality both read healthy in exactly the situation where underlying skill is eroding.
How can young professionals gain experience if AI does the junior work?
Attempt problems unassisted before consulting the model, since the gap between your answer and its answer is the highest-value feedback available. Seek out messy, under-documented, exception-heavy work that resists automation. Prioritise proximity to senior colleagues, and build a record of catching errors rather than producing volume.
Does this mean juniors should avoid using AI?
No. Avoidance forfeits real productivity gains and does not build judgment either. The evidence points toward sequencing rather than abstinence: struggle first, then compare against the AI, then reconcile the difference. The failure mode is using the model to skip the attempt, not using the model at all.
Comments (0)
No comments yet. Be the first to share your thoughts!