Ninety-five percent of enterprise AI pilots delivered no measurable P&L impact last year. That number, from MIT's Project NANDA report, has been quoted in every board deck and consultancy pitch for twelve months. What almost nobody has done in that time is change the metric they use to justify the next pilot.
The chart went up. The money did not.
I've sat through enough AI review meetings to recognize the pattern. The dashboard shows seats deployed, weekly active users, prompt volume, percent of employees who logged in at least once. These are activity metrics. They measure whether people touched the tool. They say nothing about whether the tool did any work.
McKinsey's 2026 State of AI survey found that 80% of AI users report improved personal productivity. Yet only 37% of organizations can attribute any EBIT impact to AI use, unchanged from the prior year. Eighty percent feel more productive. Thirty-seven percent can prove it on the books. That gap is the whole story.
The Vanity Denominator Problem
There's a name for this shape in operations. It's a vanity denominator. You pick a number that's easy to move, you push on it, and you point at the movement to justify the budget.
The 91% adoption number is a vanity denominator. So is the two-million-weekly-active-users figure. So is your internal Copilot deployment count. Every one of those can climb every quarter while the P&L line for AI stays flat, and the person defending the budget will still get a hearing.
What's not a vanity denominator is the outcome the pilot was supposed to change. Revenue per head. Cycle time. Margin. The work that actually got better, measured in terms the CFO already tracks.
Bill Schmarzo at TechTarget put it bluntly:
AI optimizes exactly what leaders tell it to optimize.
Bill Schmarzo
Most organizations don't have an AI strategy problem. They have an AI ROI measurement problem. They've pointed an extraordinarily capable optimization engine at metrics that were built to track activity, not to define business value.
Tokenmaxxing: When the Metric Becomes the Game
The most absurd proof of this arrived earlier this year. Amazon shut down an internal leaderboard called KiroRank after discovering employees were running AI agents on trivial tasks just to inflate their token consumption scores. The practice earned a name: tokenmaxxing.
Amazon isn't alone. Meta built a similar ranking system called Claudeonomics that tracked token usage across 85,000 employees. In a single 30-day window, usage exceeded 60 trillion tokens. The leaderboard was quietly taken down after it became public.
This is Goodhart's Law, accelerated. British economist Charles Goodhart articulated it in 1975: "When a measure becomes a target, it ceases to be a good measure." The moment you attach a reward or consequence to a metric, people optimize for the metric itself, not the outcome it was meant to represent.
Token consumption was, briefly, a half-decent signal. More tokens, more AI in the workflow, probably more value. Then companies tied it to leaderboards managers could see. The instant it became the target, it stopped measuring productivity and started measuring competitive anxiety. Engineers optimized the number, exactly as designed. The number just no longer pointed at anything real.
The Workflow Redesign Gap
Here's where the data gets interesting. McKinsey's QuantumBlack division found that the single strongest predictor of EBIT impact from AI isn't the model you choose or the vendor you buy from. It's whether you fundamentally redesigned the workflow the AI was supposed to run.
High performers are 2.8 times more likely to have done this. Only 21% of organizations have done it at all.
That 21% figure is the tell. It means roughly four in five organizations are layering AI on top of processes that were never rethought. They take the existing workflow, insert an AI step, and expect transformation. What they get is incremental improvement: a faster version of a process designed for humans working alone, not for humans and AI working together.

Layering AI onto an existing claims-handling workflow means the model drafts what a human would otherwise have typed. The shape of the process is unchanged. Redesigning the workflow means the model classifies, routes, and drafts in one pass, while people review exceptions and own the complex cases. The first approach trims minutes off each transaction. The second changes the throughput of the whole operation. That's where EBIT impact actually comes from.
The Baseline Nobody Captured
Ran Aroussi at VarOps identified another uncomfortable truth: most companies never measured their people on much of anything to begin with.
No baseline for how long a task should take. No tracked throughput for most knowledge work. No agreed definition of "done well." The work happened, it shipped, nobody clocked it.
So when AI shows up and someone asks "how much did this improve us?" there's nothing to compare against. Instead of admitting that, organizations start inventing brand-new metrics for processes they never tracked. Some invent the bar itself: new quality thresholds, new performance standards that didn't exist twelve months ago.
That's not measuring AI's impact. That's building a yardstick after the race. And it always ends up measuring whatever is easiest to count.
What Actually Moves the Needle
The MIT NANDA report didn't conclude that AI doesn't work. It concluded that enterprises don't work, at least not the way AI needs them to. The divide between winners and losers was, in the authors' words, "determined by approach," not by model quality or regulation.
The 5% that captured value share three traits:
They bought externally rather than building internally. External partnerships saw roughly twice the success rate of internal builds. The 67% vs. 33% spread suggests that buying capability, not building it, is the faster path to value.
They chose back-office friction over front-office visibility. More than half of enterprise AI budget went to sales and marketing, despite clearer returns appearing in back-office automation. The investment bias favors visible, top-line functions over high-ROI operations.
They measured workflow change, not license adoption. The successful pilots instrumented the process before the AI touched it, captured a baseline, and tracked whether the work actually got better.
The Measurement That Matters
Ariel Agor at Agor AI framed it precisely:
Measuring ROI on AI initiatives is not a spreadsheet exercise you run after a pilot ships. It is an architectural exercise you run before the pilot ships, because the only ROI a model can produce is the ROI the workflow was designed to capture.
Ariel Agor
Gartner's research found that only 29% of executives can confidently measure AI ROI today, while 79% already see productivity gains. The two numbers don't reconcile until you look at what they're measuring. Productivity is real inside the workflow. The measurement infrastructure to prove it on the P&L is not built.
The fix isn't complicated. It's just uncomfortable. Before the next pilot ships, answer three questions:
- What number on the P&L is this supposed to move? Not "productivity" or "efficiency." An actual line item the CFO already reports.
- What is that number today? If you don't know, you can't prove improvement. The baseline is unrecoverable after the fact.
- Who owns the outcome, not the adoption? Someone has to be accountable for whether the work got better, not whether people logged in.
Marketing is like dating. You don't propose on the first ad impression. And you don't declare AI success on the first login metric. The dashboard being green doesn't mean the relationship is working. It means someone showed up. What happens next is what actually counts.