I learned production at seventeen, at 2 AM, getting paged for absolutely nothing interesting. Disk filled up. A cron job climbing on top of itself. A log rotation that, as it turned out, had never once rotated a log. None of that was hard work. All of it was mine, and somewhere in the second year of that tedium I started being able to hear a machine going wrong the way you hear a car going wrong three blocks before it quits.
Nobody handed me that as a curriculum. It was just the shitty end of the job.
You never learned a city from a map. You learn a city by getting sent out on errands, botching them, at night, again, until the streets stop being information and turn into knowledge.
Which is why the productivity research keeps catching on something in me and won't let go.
Start with the finding everybody hates. METR randomized 16 experienced developers across 246 tasks in codebases they'd been living inside for years. With AI assistance they were 19% slower. They had forecast a 24% speedup. Afterward, they still believed they'd gotten one.
Now flip the telescope. A large observational study of coding agents at Microsoft found roughly 24% more merged pull requests per engineer per day. The tenure cut is the interesting part: the biggest estimated lift, +83%, landed on people with under a year at the company. It's not a tidy slope. Long-tenured veterans did fine too, and the authors are careful to say the newcomer number may be tangled up with ordinary onboarding. The silhouette is still there.
Same silhouette outside software. A field experiment with 647 support workers at Taobao routed standardized chats to an agent, which cut them 16.8% shorter. Then a conversation went emotionally sideways, a human stepped in late, and satisfaction fell nearly a full point on a five-point scale.
The routine gets eaten. The hard gets concentrated. And the routine is the only place anybody has ever learned to survive the hard.
That's a pattern, not a mechanism. Mechanism has been tested head-on maybe a handful of times, and the cleanest test sits somewhere you wouldn't think to look. About a thousand Turkish high-school students got unrestricted chatbot access. Practice scores up 48%. Those same kids then came in 17% below controls on a closed-book exam. One school, one subject, tested immediately, and the authors say flatly that this is not workplace deskilling. Fine. It's still assistance succeeding while learning quietly declined to happen.
Here's the part I can't set down. Some professions sat down and decided out loud that the reps were non-negotiable. A US general surgery resident logs 850 major cases. 250 of them before the third year, 200 as chief. The accrediting body's argument is exposure and nothing fancier: you don't get to operate alone until enough of it has passed through your own hands.
Software has never counted anything. No case log, no hours requirement, no floor. So if the volume of formative tedium drops by half, there is no instrument anywhere that reports it. It surfaces a decade later as a cohort with gorgeous output histories and no ear for a system going wrong. And nobody in the room who'd think to suspect, the way I once did, a disk-shaped problem wearing a network-shaped costume, because the pages that taught me that shape now resolve themselves before anyone's phone buzzes.
This publication has been here before, in Issue 33 and Issue 34. Both of them pivoted to what to do about it. I'm not going there.
The experts we've got already paid for their mental models in full, and nobody can repossess them. The narrower and uglier question is what the ramp looks like for the engineer who started this January. She's been handed a hell of a map. I have no idea whether anyone is ever going to send her out on an errand.
-
The assistant taken away: Three randomized experiments with 1,222 American adults working fractions and reading comprehension found that the previously assisted group solved fewer problems and skipped more of them once the help disappeared.
-
Borrowed skill, not learned skill: Nearly 500 BCG consultants beat their unassisted peers by as much as 49 percentage points on technical work outside their skill set, then showed no detectable knowledge gain on related questions after the tool was withdrawn.
-
Measurement is getting harder: METR's own follow-up reports that some developers now refuse to participate in studies that would make them work without AI, which quietly degrades our ability to run the experienced-worker comparison again.
-
Why medicine counts cases: The accreditor's explanation of case-log minimums describes them as expert-consensus exposure floors sitting inside a competency system — roughly the instrument software has never bothered to build.

