SEO title: Companies Are Rehiring People They Cut for AI. Here’s Why.
Meta description: Why are companies rehiring employees after AI-attributed layoffs? The answer is not just overhyped technology. It is a measurement failure between benchmark capability and real deployment-context capability.
The AI Layoff Boomerang Is Here
In September 2025, Starbucks deployed an AI-powered inventory management system across more than 11,000 North American stores. In evaluation, computer vision counted inventory up to eight times faster than manual methods, with accuracy approaching 99 percent. Nine months later, the deployment was discontinued entirely.
That story is now everywhere, in different costumes. Klarna cut roughly a thousand customer service roles after announcing its AI assistant could do the work of 700 agents, watched satisfaction fall, and started rehiring. Ford brought 350 experienced engineers back into quality work after automated inspection tools failed to catch what veteran judgment catches, and this year topped JD Power’s Initial Quality Study for the first time since 2010. IBM’s AskHR system resolves 94 percent of routine HR requests; the remaining 6 percent, the sensitive and ambiguous cases, still lands on humans, and the humans had to be there to catch it.
The pattern has a name now. Robert Half found that 32 percent of U.S. hiring managers who eliminated a role primarily because of AI later rehired for the same or a similar position. Forrester’s Predictions 2026 report found that 55 percent of employers regret AI-attributed layoffs, and Orgvue independently arrived at the same 55 percent figure in its own survey. Forrester goes further, predicting that half of AI-attributed layoffs will be quietly refilled, often offshore or at lower salaries.
The commentary around this boomerang mostly blames one of two villains: the technology, because AI is overhyped, or the executives, because leadership chased a trend. My research suggests both explanations miss the mechanism. What failed was not the AI and not, primarily, the judgment of the people buying it. What failed was the measurement.
The Gap Between Capable on Paper and Capable in Your Building
In my working paper series, Four Lenses of AI Measurement, the third paper, “Capable on Paper,” examines a distinction most procurement processes never make: the difference between general deployment capability and deployment-context capability.
General deployment capability is what a system can do under favorable, controlled conditions. It is what benchmarks measure and what vendor demonstrations show.
Deployment-context capability is what the same system delivers under your operational conditions, with your data, your staff, your edge cases, and your pace of change.
Organizations routinely assess the first and assume they have purchased the second. When a vendor reports 99 percent accuracy, that number describes a ceiling established under evaluation conditions. It is not a forecast of what you will experience. Yet workforce decisions, including layoffs, get made as if it were.
Five Ways AI Capability Degrades in Production
The gap between general capability and deployment-context capability is not random noise. It follows five identifiable degradation pathways.
1. Distribution Shift
The inputs at deployment differ from the inputs at evaluation. Across 11,000 stores, lighting, packaging, and shelf arrangements varied in ways no evaluation set captured.
2. Integration Degradation
The system’s reliability depends on everything it connects to. An inventory miscount does not stay an inventory miscount; it flows into ordering decisions downstream.
3. Scope Creep
The tasks demanded in production exceed the tasks demonstrated in evaluation. A system assessed on a limited product range gets asked to handle a full rotating inventory.
4. Temporal Degradation
The world changes while the system stays static. Seasonal and promotional items drift steadily away from the conditions under which capability was assessed.
5. Use-Condition Divergence
Real operators under real service pressure behave differently from the attentive evaluators the assessment assumed.
These pathways do not act in isolation. They compound. Scope creep widens distribution shift; integration propagates errors downstream; temporal decay accumulates quietly until it does not. Any assessment that scores these separately and averages them will miss the composite signal, which is precisely the signal that determines whether the deployment survives.
Rereading the Boomerang Cases Through This Lens
Now reread the boomerang cases through this lens.
- Klarna is use-condition divergence and scope creep: customers are not demo users, and “the work of 700 agents” turned out to include work never demonstrated.
- Ford is integration degradation: quality judgment is embedded in a web of institutional knowledge that inspection tooling cannot see.
- IBM’s 6 percent is scope: the demonstrated range covered routine requests, and the undemonstrated remainder was exactly where the human stakes were highest.
None of this required hindsight. Every one of these gaps was visible, in principle, before a single role was eliminated. That is the uncomfortable part, and the useful part.
What Decision-Makers Can Do Before the Next AI Commitment
The practical failure was not that these organizations lacked data. It is that no one asked the comparative question: under what exact conditions was this capability established, and how closely do those conditions match ours?
My paper proposes a structured pre-deployment diagnostic built around that question. It is deliberately not a predictive model. It requires no access to model internals and no operational data that does not yet exist. It asks non-technical decision-makers to assess divergence along each of the five pathways, flags where pathways are likely to compound, and routes findings to proportional actions before commitment.
One finding worth underlining: if a vendor cannot document the demonstrated range of the system, that absence is itself a finding, and a serious one.
Three Takeaways You Can Apply This Quarter
- Treat benchmark scores as ceilings, not forecasts. The 99 percent belongs to the evaluation environment, not to yours, until someone demonstrates otherwise.
- Change the procurement question. Not “can this system do X,” but “under what conditions was X established, and how closely do those conditions match our deployment context?” Vendors who can answer that question precisely are telling you something. So are vendors who cannot.
- Make workforce decisions on deployment-context capability only. If the capability has not been demonstrated under conditions resembling yours, for a sustained period, at your scale, then the layoff is a bet on a number that was never measured. Forrester’s 55 percent regret rate is what that bet looks like in aggregate.
The Boomerang Is One Symptom of a Larger Measurement Problem
Capability measurement is one of four lenses in this research. The same structural confusion shows up when organizations mistake ISO conformance for a legally defensible compliance posture, and when aggregate maturity scores mask governance dimensions that have quietly regressed.
Different instruments, same disease: measuring one thing and making decisions as if you had measured another.
I write about all four lenses, with the frameworks and the cases behind them, in my newsletter, The Measurement Brief, which you can subscribe to here: https://thebluenarwhal.beehiiv.com/.
If your organization is heading into an AI procurement decision and you want the vendor-evidence side of this in practical form, my book Procurement-Ready AI: The Vendor’s Evidence Playbook covers exactly what to ask for and what the answers should look like.
The full working paper, “Four Lenses of AI Measurement: Capable on Paper,” is available on SSRN: https://papers.ssrn.com/sol3/cf_dev/AbsByAuth.cfm?per_id=11297550.
The organizations rehiring today are not failures. They are early data points. The question for everyone else is whether you will measure before you commit, or after.
JM Wofford is the founder of The Blue Narwhal, an AI governance advisory practice, a professor of computer science, and the author of the Four Lenses of AI Measurement working paper series, drawing on over fifteen years of experience across enterprise technology and governance.