The performance-learning paradox in the age of AI — and what it means for how you develop your people
OECD research shows AI can boost task performance while employee skill stays flat. Here's what that performance-learning paradox means for how growing companies should design AI-enabled training and L&D.
On 19 January 2026, the OECD published a 247-page review of how generative AI is actually landing in classrooms and universities worldwide — not the marketing version, the evidence.
It was written for education ministries. But strip out the classroom language and what remains is one of the clearest evidence bases available anywhere on how people learn, or fail to learn, while working alongside AI. That distinction matters just as much on a sales floor or in a regulatory affairs team as it does in a lecture hall.
Researched and drafted with AI assistance; reviewed, fact-checked and edited by Primož Verbič.
In brief
- AI can improve how well someone performs a task while their actual retained skill stays flat or drops — OECD-cited research found a 17% performance gap once AI support was removed.
- The cause is “cognitive offloading”: handing a task to AI skips the diagnosis, evaluation and iteration steps that build lasting understanding.
- Well-configured, purpose-built AI tools help far more than generic chatbot access — and help newest employees most (a 9-point pass-rate gain for the least experienced users in one study).
- The strongest human-AI model is “augmentation” — iterative critique and refinement together — not replacing tasks outright.
- For L&D: sequence work as attempt-then-assist, prioritise AI access for new hires, and measure competence with delayed, tool-free checks.
The performance-learning paradox in the age of AI
The report’s central finding is a mismatch between performance and learning. Generative AI can make someone measurably better at a task in the moment while their underlying skill stays exactly where it was, or drops.
17% worse. A randomised trial with 1,000 Turkish high-school students found that those who practised maths with a general-purpose GPT-4 chatbot solved far more problems correctly during practice — but scored 17% worse than students who’d never used AI at all, once the chatbot was taken away for a closed-book exam.
Researchers call the mechanism behind this “cognitive offloading.” When a tool hands over a finished answer, people skip the steps that actually build understanding: diagnosing the problem, weighing options, iterating. A separate study using brain imaging found that students who wrote an essay with ChatGPT could barely recall their own content an hour later — only 12% could quote a line from it, against 89% of students who wrote unaided. Writing first and bringing in AI afterward preserved recall; leading with AI did not. Sequence, not just usage, decided the outcome.
The report isn’t anti-AI — it’s specific about what makes AI help rather than hurt. Tools configured around real pedagogical practice consistently outperformed raw, general-purpose chatbot access.
9-point gap. A Stanford-built tutoring assistant lifted student pass rates by an average of 4% — but the gain was uneven: +9 percentage points for the least experienced tutors, almost nothing for the most experienced ones. The better-designed the tool, the more it closed the gap for people who needed the most support.
On how AI should sit alongside human expertise, the report weighs three models: replacement, complementarity, and augmentation, where human and AI iteratively critique and refine each other’s output. Augmentation produced the strongest results and the least erosion of professional skill; simply automating tasks away was consistently the weakest option.
One more finding is worth carrying out of the classroom: AI-generated feedback now matches human feedback in accuracy and depth. But people trust it less and act on it less when they know a machine wrote it, even when it’s objectively just as good. The credibility gap doesn’t close just because the quality gap does.
What this means for upskilling your people
None of the above was written with a P&L in mind. But every finding maps directly onto how a growing company should — and shouldn’t — build AI into how its people develop.
-
Measure competence, not output speed. If your L&D metrics track how fast a new hire produces a first draft or a client email, you’re measuring the exact number the OECD data warns about — the one that can rise while real capability stays flat. Add a delayed check: can this person do the task without the tool a few weeks later? That’s the closed-book exam, applied to your team.
-
Point AI at your newest people first. The tutoring result — biggest gains for the least experienced, smallest for the most — is the single most useful data point here. A well-configured AI assistant for onboarding or objection-handling will do more for a new hire in month one than for your most senior specialist. Rollout priority should follow that curve, not seniority.
-
Set the order: attempt, then assist. Have people draft the proposal or the account plan before opening the AI tool — even a rough five-minute attempt. Bringing AI in afterward to refine preserves the thinking; leading with AI and cleaning up after doesn’t. This is a work-assignment policy, not a training slide.
-
Don’t hand out a chatbot licence and call it training. The gap between a generic AI seat and a tool configured around your company’s actual playbooks and language is the same gap the report finds between general-purpose and purpose-built educational AI. Treat an internal AI assistant as a designed product, not a procurement line item — map the workflows before you buy the seats.
-
Let AI draft feedback; keep a human name on it. Use AI to generate the first pass of a coaching note or a call review — it’s fast and it’s good. But have the manager read, adjust, and deliver it. People discount advice they know came from a machine, regardless of quality.
-
Teach people to work with AI, not just click into it. The strongest results in the report came from users who understood what the tool was doing — how to prompt, when to distrust a confident-sounding answer. That’s a trainable skill. Build it as deliberately as negotiation or product training, with augmentation as the target behaviour.
The report’s authors frame the challenge for policymakers as making sure AI is a learning partner, not a learning shortcut. It’s the same design question for any team investing in AI right now — and it’s the one we’re building directly into how we design AI adoption for clients.
Common questions
How quickly does skill erosion from AI use show up? Fast. The clearest evidence comes from a closed-book exam given immediately after a short practice period using AI, where the effect was already measurable. That’s why a delayed competence check should happen on a similar timescale — weeks, not an annual review cycle.
What’s the difference between giving employees an AI licence and AI literacy training? A licence is access; literacy is judgement — knowing when to trust an output, recognising a confident-sounding hallucination, and prompting effectively. The strongest results in the underlying research came from users who understood the tool’s limitations, not simply from users who had access to it.
Does this apply to non-technical or non-client-facing roles? Yes. The underlying mechanism, cognitive offloading, isn’t role-specific — it applies anywhere a tool can produce a finished output instead of supporting someone’s own reasoning. The source studies span maths practice, essay writing and tutoring, none of which are client-facing roles, so the pattern generalises well beyond knowledge work.
Source: OECD (2026), OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education, OECD Publishing, Paris.
Where's your constraint?
Five minutes tells you which of the five drivers is capping your growth.
Take the Scale Scorecard →