Consultants beat controls where the AI was right, and underperformed where it was confidently wrong. They didn't lose competence. They lost the habit of pushing back.
BCG / Dell'AcquaMeasure the human, not just the model.
The quiet threat to work is not displacement. It is keeping the job and losing the capability, invisibly, because every proxy for capability measures output — and output is now the part the AI supplies. I built the instrument, ran it on myself for four months, and wrote up what it means.
Vermeer, Woman Holding a Balance, c. 1663 · the pans are empty — she is testing the instrument before trusting it
The problem
We measure the machine obsessively. We don't measure what it's doing to us.
Two years of work has gone into knowing whether the model is good: evals, observability, agent memory. The human side got its first instrument in July 2026 — Anthropic's Reflect, in beta — and it measures usage, not judgment. You can audit an agent's every step. You still cannot audit your own.
The evidence
This isn't speculation. Separate teams keep measuring it, and the signal is getting stronger.
Across separate studies the same shape appears: AI makes the work easier and the human less critical, and the loss outlasts the tool.
Less mental effort from LLM users than search-engine users completing the same task.
MIT Media LabPassive AI use lowers self-efficacy, ownership, and meaning, and the dip outlasts the AI use itself.
Nature · Lee et al.Losing one's edge ranked the 4th-biggest concern among 81,000 Claude users. 13,000 raised it; close to half already feel it.
Anthropic · 81k interviewsDelegation lowers learning: hand it over and comprehension drops. Ask the AI to explain, and the skill is kept.
Anthropic · coding RCTThe cleaner the output looks, the less the user checks context (−5.2pp) or questions the reasoning (−3.1pp).
Anthropic · Fluency IndexWhy now · 2026
The work has shifted from process to outcome, and the judgment moments are vanishing.
AI is becoming more relational, persistent, and accurate about you specifically, each one harder to push back on. The human's contact with the work narrows to three seams: the brief, the redirect, the sign-off. The people closest to it said the same thing in a single week.
“I no longer steer. I commission.”
“I don't prompt Claude anymore… my job is to write loops.”
“The loop doesn't know if you understand the work or are avoiding it. You do.”
The ownership account
People keep the job and lose the capability. That is the quiet threat to work.
Through mid-2026 the practitioner accounts converged: engineers finishing AI-heavy projects faster and knowing them less, unable to defend the decisions when someone asks why was this done? The decisions existed — they just left no trace anyone can retrieve at the moment of challenge.
“The craftsmen are tired. Very tired. The entire burden of review falls on the craftsman.”
“It becomes very mechanical and takes away what a lot of devs love about their job.”
“I am no developer anymore. I am the Product Owner with the technical skill they wish they had… plus QA engineer.”
Every one of these reports is a missing-record problem, and it reframes what the digest is: not just a mirror for judgment drift, but a record of title — the part of the work that is still yours, written down while it was warm, there when someone asks why. The full account is in the research document.
I ran it on myself
On the record, I push back about one time in four.
Before deciding anyone else has this problem, I pointed the instrument at my own work: more than two hundred of my real coding sessions and the nine-hundred-odd messages I sent inside them. One narrow question: what does the transcript show me doing when the machine hands me something finished?
Four moves show up in the record. About one time in four, I visibly push back on whether it's right. About as often, I steer how it reads. A third of the time I'm handing over the next task. And roughly one in seven leaves no visible check at all, with one honest caveat: a transcript only sees what I say out loud. It can't tell a silent, careful read from a genuine wave-through.
The shape is specific. The pushback that surfaces is mostly about how the AI writes, the voice and the approach, and less often, out loud, about whether it's true. That visible-judgment gap is the thing worth watching.
I take this seriously: I build with guardrails and care about using AI safely. I half-expected to come out of this looking fine. I didn't. No single session means much, but across all of them, the judgment that showed up on the record was thinner than I'd thought.
Call it an estimate; measure it a few ways and it shifts a point or two. And it only ever sees the visible half: silent reading leaves no mark.
The method
Quiet, local, pointed at you. The study runs on anyone.
Not a dashboard that nags, and not a score. A thin instrument that records the decisions behind your work and shows you the drift, the way a training log shows a runner their pace. It is a research method, not a product — everything needed to replicate the study on yourself is in the repo.
Capture, locally
A hook reads each session on your machine when it ends. Nothing identifiable leaves your device.
Reflect, privately
It records what you steered, what you checked, and what you let go. No score, no grade.
See the drift
Over time the mix and its trend become visible, at the one moment the loop can't watch itself: sign-off.
A personal tool helps the person who runs it. On its own it won't save the population. That has always taken a curriculum, a platform, or a professional standard.
Open questions
What's still unknown.
- 01
What's the right unit to watch? A single turn, a whole session, the irreversible commands, or the moment you accept a long run you never watched.
- 02
How much checking is silent? Reading a diff and saying nothing is real scrutiny that no transcript can see.
- 03
Does the gap compound over months, or do people correct once they can see it? Answering that needs far more history.
- 04
Does showing someone their own mix change how they work, or just hand them a number to game?
- 05
Is the right altitude the individual, the team's curriculum, or the platform itself? And how do you surface it without becoming the tool people switch off on day one?
- 06
What does hiring for judgment look like once output stops evidencing it? Fluency can be interviewed for. The contest rate can't, yet.
The platform has started building it. The judgment half is still open.
Reflect landed the same week that memo went live — tier one, on the chat surface, measuring usage. What stays unmeasured is the half this document is about: judgment on the working surface, the contest rate at sign-off, and what erosion means for hiring.
The research is ongoing. If you work on the measurement layer, on AI adoption, or on the same question from any side, I'd like to swap notes.
