A principal-level developer in Microsoft's coding-agent study described what they did while the agent wrote code: reviewed other engineers' pull requests and read design documents. A small, specific image, but it reveals the shape of a much larger reorganization.
When agents handle routine operations within scoped authority, the human role doesn't disappear. It reconstitutes itself around maintaining the conditions under which delegation remains safe. The boundaries of the agent's authority have to be defined. Escalation paths for when the agent encounters something outside those boundaries have to be designed. Policies governing what the agent can and cannot do have to be authored. And the agent's output has to be reviewed with enough context to catch the cases where the artifact looks correct but isn't. Every one of those tasks requires judgment that the agent cannot provide, because the agent is the thing being judged.
I'd call this review-shaped work: policy authoring, exception handling, escalation design, review sampling. None of it is new. Managers have always done versions of it. The volume and velocity have changed, but so has something more structural. Traditional management review happens within shared context. The manager often did the work themselves once, knows the domain intimately, can sense when something is off before they can articulate why. Review-shaped work increasingly asks people to evaluate artifacts produced by a process they didn't participate in and may not fully understand. When a coding agent produces 24% more pull requests, someone reviews 24% more pull requests. Each one was generated by a process that left no trail of human reasoning to follow.
The temptation is to treat review as a lighter version of doing. It is heavier. Review requires the ability to evaluate an artifact you didn't produce, under time pressure that structurally rewards approval over scrutiny. When a human signs off on an agent's output without the context to evaluate it, the authority to reject it, or the ability to recover from a bad approval, that is governance theater. We have built remarkably effective systems for producing the appearance of oversight at scale.
McKinsey's 2025 survey found that among organizations reporting meaningful business impact from AI, respondents were nearly three times as likely to have fundamentally redesigned individual workflows. Deploying the agent is the technology decision. Redesigning the human work around it is the organizational one.
McKinsey's 2025 survey offers a useful signal here, with appropriate caveats about self-reported data from roughly 2,000 respondents. Among organizations attributing 5% or more of EBIT to AI (about 6% of the sample), respondents were nearly three times as likely to have fundamentally redesigned individual workflows. This was one of the strongest contributors to meaningful business impact in their analysis. The sample is small and the data is self-reported, but the direction is suggestive: organizations reporting value aren't the ones that deployed agents alongside existing processes and hoped for the best. They rebuilt the human work around the agent's output. They treated the organizational redesign as the actual project.
That rebuilding is where things get difficult. Deploying an agent is a technology decision. Redesigning review workflows, escalation authority, and exception-handling processes is organizational work, which is slower, more political, and less exciting. It doesn't demo well. It requires negotiating authority, redefining roles, and confronting the question of who is actually empowered to override the system. Most organizations would rather ship a feature.
The Taobao experiment illustrates what happens when you skip this step. The agent was fast and contained. But the escalation design assumed that handing a frustrated customer to a human would repair the interaction. It often didn't, because the handoff arrived after negative sentiment had accumulated and the human lacked the context to recover. Workers in emotionally escalated chats sent fewer messages, sought less information, offered fewer solutions. The workflow around the agent hadn't been rebuilt to match what the agent actually produced. The technology worked. The institution hadn't caught up.
The clerk model concentrates human responsibility into fewer, higher-stakes moments of judgment. That can be a sound design. But only if the humans in those moments have what they need: context, authority, and the organizational backing to say no when saying no is the right call. Without those, you haven't automated the routine. You've constructed a system where the most consequential decisions look like rubber stamps on someone's morning queue. Someone always pays for that arrangement. It is never the people who chose to defer the redesign.

