Practitioner's Corner

Practitioner's Corner

What Aman Gupta's Agents Don't Decide

Aman Gupta's team at Nubank built an agent architecture for over a hundred million customers largely by subtraction: deterministic sequences collapsed into single tools, toolsets held to five or fifteen options, routines translated from human procedure. The model decides only what couldn't be engineered away in advance. That's a precise account of where the model's authority ends. What the architecture doesn't yet specify is what a person receives when a live case comes back to them, because the scratchpad that keeps the agent moving isn't the artifact someone needs to pick up where it stopped.

What Aman Gupta's Agents Don't Decide
Aman Gupta's team at Nubank built an agent architecture for over a hundred million customers largely by subtraction: deterministic sequences collapsed into single tools, toolsets held to five or fifteen options, routines translated from human procedure. The model decides only what couldn't be engineered away in advance. That's a precise account of where the model's authority ends. What the architecture doesn't yet specify is what a person receives when a live case comes back to them, because the scratchpad that keeps the agent moving isn't the artifact someone needs to pick up where it stopped.
The Intervention Horizon as a Design Parameter

Agent architectures put a human approval step in the design and call it a safety measure. Approval only works if the person arrives before the outcome is locked in, with enough context to understand what they're approving and enough time to choose otherwise. Past that point, it records consent and offers no alternative. The last moment a human can still change what happens next deserves the same specification you'd write for latency, capacity, or failure behavior. Most systems never define it.
The Intervention Horizon as a Design Parameter
Agent architectures put a human approval step in the design and call it a safety measure. Approval only works if the person arrives before the outcome is locked in, with enough context to understand what they're approving and enough time to choose otherwise. Past that point, it records consent and offers no alternative. The last moment a human can still change what happens next deserves the same specification you'd write for latency, capacity, or failure behavior. Most systems never define it.


The Queue Used to Be a Curriculum — A Compliance Analyst on What Agent Handoffs Actually Feel Like
CONTINUE READINGEvidence Sidebar

After working alongside AI polyp-detection tools, endoscopists at four Polish centers saw their unassisted adenoma detection rate fall from 28.4% to 22.4%. The retrospective design leaves room for confounders — workload shifts, temporal changes. But the direction is consistent with what a 2025 meta-analysis of procedural skills quantified more precisely: accuracy-based skills lose roughly half their acquisition gains within 6.5 months of nonuse. Intermittent practice was a significant moderator, which means the decay responds to countermeasures — if someone designs them in.
Automated driving research adds a useful complication. Across 129 studies, urgency produced faster takeover reactions, but reaction speed and takeover quality turned out to have different determinants. Prior experience and adequate time budgets improved quality. Urgency alone did not.
What connects these findings across colonoscopy suites, driving simulators, and skill-retention labs is a structural fact about how competence works: it requires representative exposure and periodic unaided practice. Without those, the erosion is slow enough that nobody notices the accumulating deficit until something goes wrong.
Reading List


Past Articles

When an AI agent canceled a gym booking in Australia earlier this year, the only record of what happened was screenshots...

A database recovery log carries its own interpretation. A multi-agent failure trace does not: in one recent study, six e...

I carried a pager for years attached to monitoring that rolled everything into averages at write time. Any question I ha...

Most agents in production stop somewhere around ten steps and hand the work to a person. Easy to read that as the models...

