The week after Labor Day is a fair moment to reconsider what "productive" means. A Stanford study of nearly 250,000 AI conversations found that high-stakes tasks averaged 12.4 turns — almost double the 6.5 for routine work. As stakes rose, users modified AI output more, accepted it unedited less, and challenged the model's reasoning successfully over 80% of the time. People used the exchange to develop their own thinking.
Enterprise AI is still evaluated mostly by time saved. For routine tasks, that captures real value. But the most consequential collaboration — longer sessions, more friction, more human intervention — registers, by that metric, as inefficiency. What organizations choose to measure will determine what kind of human-AI work survives.
12.4 vs. 6.5 — average turns per conversation, high-stakes vs. ephemeral tasks
56% of actionable conversations rated consequential or higher; only 20% were ephemeral
60% of users in consequential work modified AI output rather than accepting it directly
~50% of conversations hit friction — but 78.7% of users attempted recovery rather than abandoning
72% of conversations were human-led, with the person retaining primary responsibility
Source: Shao et al., Stanford, Aug 2026. Single-provider data (Claude.ai); behavioral classification, not outcome measurement.

