A field experiment on Alibaba's Taobao tested what happens when an AI agent handles customer service and routes hard cases to humans. With 647 workers and 680,000 chats, the results split cleanly. Technical escalations — cases the agent couldn't resolve — preserved service quality. Emotional escalations — already-frustrated customers — produced 40% longer resolution times, ratings down nearly a full point, and recontact rates up six percentage points.
The interesting finding is where the quality collapsed. Workers handling emotional escalations maintained empathy scores comparable to those on technical cases. But they sent fewer messages, sought less information, offered fewer solutions. The authors suggest learned helplessness: workers receiving pre-frustrated customers may anticipate failure and reduce effort accordingly.
The study ran seventeen days. Seventeen days was enough to shift behavior. The question organizations building these systems should be asking is what a year does — not to service metrics, but to the people generating them.
The experiment: 647 customer-service workers, 680,676 chats on Alibaba's Taobao, August 2024. Treatment workers supervised an AI agent on eligible chats while personally handling ineligible ones.
Emotional escalation outcomes (vs. matched human-handled chats): +40.8% duration. −0.928 rating points on a 5-point scale. +6pp seven-day recontact rate.
Technical escalation outcomes: Quality preserved. Duration +19.1%.
Worker engagement, emotional vs. technical: 8.9 vs. 10.0 human messages per chat. 43% vs. 65% human share of chat rounds.
Timing: Worker-initiated escalations happened earlier and produced smaller quality drops than algorithm-triggered ones.
Learned helplessness: The authors' interpretation, inferred from behavioral patterns. Not directly tested. No data on long-run effects on worker skill, morale, or retention.

