Practitioner's Corner

Practitioner's Corner

The Queue After the Queue

When a telecom deployed a conversational agent, daily call volume to human agents barely moved. But the average call got longer. The system was resolving cases and making the human operation slower at the same time. Automating routine work strips out the predictable middle and leaves behind a concentrated residue of ambiguity, judgment calls, and emotional difficulty. Organizations celebrating an 80% automation rate may be running a human operation that works less well than the one it replaced.
The Queue After the Queue
When a telecom deployed a conversational agent, daily call volume to human agents barely moved. But the average call got longer. The system was resolving cases and making the human operation slower at the same time. Automating routine work strips out the predictable middle and leaves behind a concentrated residue of ambiguity, judgment calls, and emotional difficulty. Organizations celebrating an 80% automation rate may be running a human operation that works less well than the one it replaced.

Alexandre Drouin and the Benchmark as a Theory of Work

Alexandre Drouin's WorkArena benchmark tests AI agents on enterprise tasks inside live ServiceNow software: filtering incident lists, ordering configured laptops, reading dashboard values. Every included task shares one property — it ends in a system state that a program can verify. That criterion is also a claim about work. Tasks whose outcomes live in database fields can be benchmarked; tasks whose value depends on interpretation, conversation, or judgment that never resolves into a checkable field cannot. Drouin's design choices show where that line falls more clearly than his scores do.

Alexandre Drouin and the Benchmark as a Theory of Work
Alexandre Drouin's WorkArena benchmark tests AI agents on enterprise tasks inside live ServiceNow software: filtering incident lists, ordering configured laptops, reading dashboard values. Every included task shares one property — it ends in a system state that a program can verify. That criterion is also a claim about work. Tasks whose outcomes live in database fields can be benchmarked; tasks whose value depends on interpretation, conversation, or judgment that never resolves into a checkable field cannot. Drouin's design choices show where that line falls more clearly than his scores do.

The Emotional Residual

A field experiment on Alibaba's Taobao tested what happens when an AI agent handles customer service and routes hard cases to humans. With 647 workers and 680,000 chats, the results split cleanly. Technical escalations — cases the agent couldn't resolve — preserved service quality. Emotional escalations — already-frustrated customers — produced 40% longer resolution times, ratings down nearly a full point, and recontact rates up six percentage points.
The interesting finding is where the quality collapsed. Workers handling emotional escalations maintained empathy scores comparable to those on technical cases. But they sent fewer messages, sought less information, offered fewer solutions. The authors suggest learned helplessness: workers receiving pre-frustrated customers may anticipate failure and reduce effort accordingly.
The study ran seventeen days. Seventeen days was enough to shift behavior. The question organizations building these systems should be asking is what a year does — not to service metrics, but to the people generating them.
External Reading








