Market Pulse

Market Pulse

The Infrastructure for Arguing Back

Microsoft's coding-agent study reports a 24% increase in merged pull requests across tens of thousands of engineers, then the authors quietly concede that merged PRs are an output proxy, not delivered value. That concession matters more than the number. It points toward a structural question the usual conversation about agent adoption keeps skating past: the predictor of durable adoption may not be task complexity or model capability, but whether a domain already has organizational machinery for catching mistakes, contesting outputs, and reversing course before the cost lands on whoever is least equipped to absorb it.

The Infrastructure for Arguing Back
Microsoft's coding-agent study reports a 24% increase in merged pull requests across tens of thousands of engineers, then the authors quietly concede that merged PRs are an output proxy, not delivered value. That concession matters more than the number. It points toward a structural question the usual conversation about agent adoption keeps skating past: the predictor of durable adoption may not be task complexity or model capability, but whether a domain already has organizational machinery for catching mistakes, contesting outputs, and reversing course before the cost lands on whoever is least equipped to absorb it.
Coding Agent Evidence
Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI
Agents gain traction where artifacts already flow through diffs, tests, CI, and rollback, making machine-generated work institutionally legible.
The authors say no, and that distinction matters: review systems make agent work visible, but visibility and quality are different things.
Coding Agent Evidence
The Shift to Agentic AI: Evidence from Codex
Users encoding repeatable workflows look a lot like people building review-shaped contracts for agent output, not just prompting ad hoc.
The heaviest concurrent-agent figures skew toward OpenAI's own teams; external enterprise adoption remains thinner and less even.
The Unreviewed Number

The EntSQL benchmark asked text-to-SQL systems to answer enterprise questions requiring domain knowledge: metric definitions, reporting conventions, organizational rules. The best evaluated English-input system reached 15.9% accuracy. Set that aside for a moment and consider what the other 84% looks like once it ships.
In coding, wrong output hits a diff, a test suite, a reviewer. Distrust already has a workflow. In analytics, the query produces rows. Rows become a number, and the number becomes a sentence in a deck. Somewhere between the SQL and the sentence, the metric definition, the filters, the caveats all fall away. Nobody later argues with the query. They argue with the conclusion, long after anyone could reconstruct whether "active customers" included trials, paused accounts, or subsidiaries.
A bad database mutation tends to be loud. A bad number can be quiet and durable. It looks precise enough to travel.
Primary Sources




Past Articles

Visa's new agentic commerce rules hold cardholders responsible for agent-initiated payments "as if the Cardholder initia...

Researchers at UC Davis watched seven commercial browsing agents navigate an instrumented website this spring. A behavio...

Across the agent ecosystem, multiple players are simultaneously shipping the same category of development: admin console...

Companies with governance tooling deploy twelve times more AI projects to production. Only 4 of 13 frontier-autonomy age...
