← Latest reporting

A skills-erosion survey should trigger workflow observation, not a critical-thinking course

IBM surveyed 1,500 CHROs and 8,800 employees and found concern about judgement, skills erosion and hidden verification work. Self-reported gaps should lead to observed task evidence and redesigned decision rights, not a generic training response.

Skills Demand and Labour MarketWork and Role Change
A bold magenta and teal flat print shows a magnifying lens examining proxy outputs beside a large manual override lever and a row of review cards.
Conceptual illustration generated with AI under editorial direction; it does not depict a real event.

What happened

An IBM Institute for Business Value study, conducted with Oxford Economics from April to June, surveyed 1,500 CHROs across 21 geographies and 23 industries and 8,800 employees across 28 countries. IBM reports that 71% of CHROs prioritised supervising, validating and overriding AI while 29% of employees prioritised judgement.

Why it matters

Survey differences identify a hypothesis about work design and recognition, not a measured skill deficit. Organizations need to observe where judgement occurs, whether overrides are safe and whether verification work is visible, supported and rewarded.

IBM's CHRO study, conducted with Oxford Economics from April to June, surveyed 1,500 senior workforce executives across 21 geographies and 23 industries and 8,800 full-time employees across 28 countries. IBM reports that 71% of CHROs called the ability to supervise, validate and override AI the most essential skill, while 29% of employees ranked judgement as important. Sixty per cent of employees worried about skills erosion.

HR Dive's independent summary highlights the same disconnect and the risk that AI-related work goes unseen. These are self-reported perceptions collected in cross-sectional surveys. They do not show that AI caused a measured decline in critical thinking, or that a course would restore performance.

Observe where judgement actually enters the workflow

Select a small set of AI-assisted processes and shadow real work with consent. Mark each point where a person frames the task, checks evidence, rejects an output, asks for another source, resolves a disagreement, escalates risk or accepts responsibility. Record the time, information available and consequence of the choice. Compare this map with the formal job description and performance metrics.

The likely gap is often not knowledge alone. A worker may know how to challenge an output but lack time, access to evidence, authority to override or a safe escalation route. A generic critical-thinking course cannot repair those conditions. Redesign the workflow so a reviewer can see source provenance, uncertainty and prior edits; set an explicit override right; and ensure production targets do not punish careful checking.

Make invisible verification visible

IBM reports that many executives see AI creating invisible work, while employees describe extra checking that may not be recognised. Measure it directly: minutes spent validating, duplicated searches, correction loops, handoffs, emotional load and after-hours recovery. Add the work to capacity planning and role expectations. If verification is essential to quality, it is production work, not discretionary diligence.

Use errors as learning material without turning them into individual blame. Review a sample of accepted and rejected AI outputs, classify failure modes and ask whether the person had the evidence and authority to act. Train against those cases. A short scenario on resolving conflicting sources or stopping an automated decision is more diagnostic than a broad course completion rate.

The counterargument is that observation is slow and intrusive. Bound it: sample two weeks, anonymise unnecessary personal data, involve worker representatives where appropriate and publish the measurement purpose. Pair observation with system logs, but do not infer judgement quality from clicks or time alone. The objective is to redesign the conditions for good decisions, not surveil individuals.

After redesign, test transfer rather than attendance. Give workers unfamiliar but realistic cases, allow them to use the same tools available in production and score whether they identify missing evidence, choose a safe override and escalate appropriately. Repeat later to detect decay. Compare teams with and without the redesigned workflow before attributing improvement to training. Course completion, confidence and quiz scores are useful diagnostics, but they are not substitutes for safe decisions in context.

Report results by task and consequence, not as one enterprise critical-thinking score. A reliable override in customer support does not establish judgement in hiring, finance or safety work. Each workflow needs its own evidence threshold, reviewer capacity and rollback signal. This keeps capability claims narrow enough to guide staffing and learning investment.

The immediate decision is therefore to delay a broad training purchase until the organization knows which judgement tasks are failing and why. Choose three workflows, measure verification and override behaviour, fix structural blockers and then target practice at observed gaps. The Skills Atlas can help name relevant capabilities, but evidence should come from work. A survey is a signal to investigate; a changed workflow with safer decisions is the outcome.