AI doesn't make anyone a better or worse manager. It amplifies whoever they already were — and the leverage gap it opens between a leader and their team is made almost entirely of decisions the leader controls: who gets the tools, who has permission to experiment, and whose output becomes the baseline everyone else is measured against. This is about that gap: why it's structural rather than attitudinal, what it quietly costs about eighteen months out, and the three things to measure instead.
There's a scene playing out across a lot of leadership careers right now, and almost nobody talks about it. A leader gets real access to frontier AI tools. Within weeks their personal output jumps — agents running in parallel, work that used to take a week landing in an afternoon. The gain is genuine, and so is the feeling that comes with it. It's a dopamine hit. I've felt it, and most people using these tools seriously have.
The hit isn't the problem. The problem is what happens next, quietly. The designed essay is below, and the written version follows.
The bar moved. Nobody scheduled a meeting about it.
Once your own output jumps, your sense of what "normal" looks like quietly shifts. You don't notice it happening, because you only see your finished work — not the six dead-end prompts, the two restarts, the judgment calls that fifteen years of experience let you make in seconds. You end up judging by a highlight reel.
It's worth being precise about what the two tapes actually contain. The highlight reel is one line long: finished work, in an afternoon. The full tape is six dead-end prompts, two restarts, fifteen years of judgment spent in seconds, and then the finished work. The new sense of "normal" gets built from the first tape. The estimate it goes on to judge was built in the second.
Then, sooner or later, they read a team estimate and think: this could be done by Friday. Without any malice at all, the team is now being measured against a number that was never a team number. It was one person's number, produced with tools the team may not have, on a task where that person's experience was doing most of the quiet work.
Ask the leader, and they're just holding the bar. Ask the team, and the bar moved — and nobody scheduled a meeting about it.
The tell is usually the same: estimate reviews turn quietly tense, and it takes months to locate the source. Part of why it takes months is that nobody has language for it. The leader experiences it as the team getting slower, or more defensive about scope. The team experiences it as the leader becoming unrealistic. Both are describing the same event from opposite ends, and neither description names it. The estimates haven't changed. The yardstick has.
Call it the yardstick problem.
Is the gap about effort, or about structure?
It's tempting to blame the distance between a leader's output and the team's on effort, or on people being slow to adopt the tools. Usually it's neither. Three built-in differences do most of the work, and none of them is about how hard anyone is trying.
Access
The executive has the frontier tools and no procurement friction. The analyst has a locked-down license and a policy memo. Same company, different century.
Leverage
This is the tricky one, because the research points in two directions at once. The big field studies found AI narrows skill gaps: in the Harvard–BCG study, the bottom half of consultants improved 43% with AI versus 17% for the top half, and a study of five thousand support agents found new hires gained 34% while veterans barely moved. On work the tool is good at, AI is a leveler.
But the same support-agent study found the flip side: on tasks just past what the tool could handle, people who leaned on it were 19 points more likely to get the answer wrong. What separated people wasn't prompting skill. It was knowing what to trust. When the hard part is judging the output rather than producing it, experience wins — and a leader who assumes everyone gets their results from the same tools is missing that gap entirely.
AI levels the doing. It doesn't level the judging — and it punishes trusting it blindly.
Permission
Senior people experiment freely. Junior people often genuinely don't know what they're allowed to touch, so they use less, learn slower, and fall further behind a bar they never saw move. Ambiguity reads as risk to the person with the least standing to absorb it.
This one is nearly free to fix and almost never gets fixed, because from the top it doesn't look like a policy problem at all. From the leader's chair the rule is obvious: use good judgment. From three levels down, "use good judgment" about a tool nobody has explicitly sanctioned, on data nobody has explicitly cleared, is a coin flip with career consequences on one side. People resolve that by not flipping.
None of these is fixed by telling people to work harder. All of them are fixed by leadership decisions.
The second-order cost: the bench stops growing
There's a quieter failure mode that matters more. When a senior person can do a junior's work in ten minutes, the tempting move is to stop assigning it. Just do it. Ship faster.
This is already happening at the level of the labour market. Stanford's Digital Economy Lab, using payroll data covering millions of workers, found that jobs for early-career workers in the most AI-exposed fields have fallen roughly 13% relative to their peers since late 2022 — while experienced workers in the same fields show no comparable drop. The economy is already running the experiment: keep the seniors, skip the juniors.
Inside a single team, the same trade is invisible for about a year. The dashboard shows throughput up and cycle time down, because the seniors are shipping at machine speed — and every one of those numbers is real. What no dashboard shows is that the bench stopped growing. No metric turns red when a team's capacity to produce its next senior person quietly goes to zero, which is exactly why that decision gets made by default rather than on purpose. The bill arrives roughly eighteen months later, when nobody below the leader can own anything without them.
Run a team this way and you've traded next year's capability for this quarter's speed. The bench doesn't disappear all at once. It just stops growing.
Delegation was never only about efficiency. It's how the next generation of judgment gets built.
That's still true. It's just easier than ever to forget.
Amplification cuts both ways
Here's the part that gets missed. AI doesn't make anyone a better or worse manager on its own. It amplifies whoever they already were.
A manager who already measured people against themselves now does it with a superhuman baseline. A manager who already hoarded the interesting work now hoards it at machine speed. But a manager who already thought in terms of team capability now has more leverage than ever to distribute — and those teams are pulling away from everyone else's.
The same amplifier, pointed at two different people, produces two different companies. Which is why this is worth being deliberate about rather than letting it resolve itself.
Bad managers are getting worse. Good ones are getting better. The technology is neutral about which one you become. It is not gentle about it.
What should leaders measure instead?
Three replacements, none of them complicated.
Swap the baseline: the leader's personal output → team throughput. A leader's output is an input to the system, not the standard for it. The moment your own afternoon becomes the reference point, every estimate below you is being graded against a number that was never comparable.
Swap the order: expectations first → tools and permission first. Nobody gets held to an AI-era standard until they have AI-era access — and explicit permission to experiment, and to fail while learning. Raising the bar before handing out the tools isn't ambition. It's a tax on the people with the least leverage to pay it.
Swap the signal: shipping alone → shipping plus learning. If the work is getting done but nobody is getting more capable, that's not a team. It's a queue.
That last one is the hardest to instrument, and it doesn't need to be elaborate. It's usually one question asked at the end of a cycle: who can do something now that they couldn't do at the start of it? If the honest answer is nobody, the throughput number is borrowing against a balance nobody is tracking.
Don't benchmark the gap. Close it.
The gap between a leader's leverage and the team's is real, it's growing, and it's mostly made of decisions leaders control. So the job isn't to measure people against the gap. The job is to close it.
Where this could be wrong: if tool access genuinely commoditizes — everyone on a frontier model, no procurement lag, permission settled by default — then two of the three differences collapse on their own and the gap narrows without anyone managing it. That's a plausible two-year outcome, and it would make the access and permission arguments here temporary. It wouldn't touch the third one. Knowing what to trust doesn't arrive with a license, and it's the difference that compounds.
The best leaders in this era won't be the ones who move fastest. They'll be the ones whose teams do.
