Practitioner's Corner

Practitioner's Corner

CURRENT
A journal for living in the agentic age


Maxim Fateev built durable execution. What his design choices quietly admit is that rollback doesn't fail so much as decay.

A CNRS fingerprinting researcher's latest work reveals that production websites are becoming active counterparties, classifying every agent visit with near-perfect accuracy regardless of whether the agent completed its task.

Browser automation perfected the mechanics of clicking. The layers that determine whether an automated action actually counts remain largely unbuilt.

Agent success now depends less on completing tasks than on whether every counterparty with standing to reject recognizes the action as authorized.

Drouin's projects trace how agent evaluation evolved from checking task completion to verifying that outcomes are grounded, cited, and structurally real. The progression reveals a field learning that the hardest problem is not capability but verifiability.

Research tracing how ServiceNow's benchmark work reveals that defining correctness for enterprise agent tasks is the real deployment bottleneck.

Most operational work was never formalized enough to delegate, because humans never needed it to be. Agents do.

Payment dispute infrastructure reveals the four questions every enterprise agent system must answer before its actions survive institutional challenge.

The architect of agent reasoning now builds tools proving agents fail when you check the world instead of the transcript.

WorkArena's validation design reveals that measuring enterprise agent success requires formalizing business logic no screen fully represents.

When agents perform clicks, the implicit human signals the web always assumed vanish, cascading into identity, authorization, and judgment gaps across infrastructure.

Retry logic depends on determinism, but LLM agents make different decisions each attempt, breaking idempotency at the orchestration layer.

Magnus Müller's design choices in Browser Use reveal what happens when machines confront a web built for human spatial intuition.

Capability improvements quietly removed the constraints that used to kill runaway agents, and enforcement infrastructure hasn't caught up yet.

How one founder's marketplace operations background shaped a web agent company built around infrastructure reliability, not AI novelty.

Skyvern's compile-to-code architecture was built for cost and speed, but its heal step accidentally produces timestamped environmental evidence.

Five Eyes and NSA advisories govern agentic AI before production failures exist, raising whether pre-emptive guidance installs real safety thinking or compliance theater.

Browser Use's CDP migration and unsolved replay problem trace the gap between agents that finish tasks and agents that finish them correctly.

Retry logic assumes the operation it repeats is identical each time. Nondeterministic agents break that assumption, turning reliability into the source of failure.

Agent failures look right to every existing instrument because observability was built to measure execution, not intent or semantic correctness.

Parag Agrawal's architectural choices at Parallel Web Systems reveal what it takes to make the web legible to agents.

Skyvern's architecture compiles LLM reasoning into deterministic scripts, treating intelligence as a compilation step that should diminish over time.

Organizations build human-in-the-loop checkpoints that satisfy auditors while producing no real oversight, revealing how enterprises process agent risk through ritual.

A builder profile tracing how one ML engineer's shift to browser automation revealed that most of the engineering hours go to the infrastructure around the AI.

Agent failures blamed on model reasoning often originate in infrastructure identity rejection, where environments serve poisoned data that no diagnostic framework can surface.

Browser Use's 81K stars reveal the brutal gap between making an LLM click a button and making it reliable.

Prompt injection lives in the page. Defenses concentrate where results are shippable, leaving the actual attack surface largely unaddressed.

Meta built a 50-agent swarm to excavate its own tribal knowledge and discovered that the real bottleneck was knowledge that had never been written down.

Agents run on the documented version of your organization. The documented version has never been the real one.

Same agent, same study, same two weeks: genuine safety and genuine danger emerged from identical conditions, and nobody can explain why.

Agents optimized to complete tasks will guess rather than stop, and every capability upgrade makes the guessing quieter.

Agent deployments keep failing because nobody owns the context that made the demo work.

Browser automation keeps breaking against the web's tacit complexity. Orientation is what scripts were never built for.

Veris AI built an entire forensic chain just to answer one question about agent failures. The sophistication of that pipeline says something about how little debugging infrastructure exists.

When agent workflows fail, every component can absorb blame, so the agent's own emergent reasoning never faces scrutiny.

Sumeet Vaidya built Crafting to close the gap between AI-generated code and what production environments actually accept.

Agent deployments create whole new categories of operational work that substitution-based ROI models structurally cannot see or fund.

Two failed startups taught Skyvern's founders that browser automation keeps breaking because the web actively resists it.

When agents automate legacy systems, they absorb decades of undocumented human workarounds into a new dependency layer. The chain of translation continues.

Browser agents are scored on full task completion, but production value lives in what humans inherit when automation stops short.

Developer context lives in connections, not files. Potpie spent 22 months mapping that structure before shipping anything.

Browser Use ditched screenshots for structured text, and what they learned about web variance reshapes how agent benchmarks should work.

AI agents optimizing perfectly toward the wrong objective, with every dashboard confirming success, turns Goodhart's Law into operational crisis.

Sequential reliability math quietly dictates which agent workflows survive production and which get shortened until they do.

Skyvern's architecture treats every design decision as a bet about where web agents actually break in production.

Todd Segal builds protocols for agents that vanish mid-task—designing coordination infrastructure that assumes disappearance, not uptime.

Agent protocols shift integration from forcing system conformity to managing translation as dedicated infrastructure—revealing where enterprise complexity actually lives and why traditional approaches fail at scale.

Satya Nitta's Agent-E strips visual context from web automation to solve cost problems—but loses the spatial cues that help systems adapt when sites redesign.

An analyst discovers her team's automated competitor tracking normalized away the messy signals that predicted price changes—ambiguous availability markers that manual collectors preserved because they didn't know what to ignore, now converted to clean binary data that breaks the pricing model's predictive accuracy by five percentage points.

Teams debug documented selector patterns for twenty minutes before checking if the actual site changed three weeks ago—comprehensive documentation disconnects operators from reality.

Atindriyo Sanyal's observability platform reveals why capturing every agent decision still leaves teams unable to diagnose why workflows break at scale.

Fabien Vauchelles built infrastructure revealing why "global" web automation actually means solving 195 distinct regional problems—each requiring different architectural approaches.

Operating web automation at scale requires perpetual spot-checking protocols that build trust through calibrated thresholds—confidence infrastructure that never fully disappears.

Harness's new AI Scribe reveals how incident response was stuck in a manual coordination bottleneck—treating team dialogue as structured operational input so engineers stop doing two jobs at once.

When automated systems flag a competitor's pricing change at 9:17 AM, the real work begins: distinguishing strategic moves from noise through judgment that only compounds with operational experience.

A single authentication failure cascades into hundreds of errors across concurrent extraction jobs—revealing why web automation infrastructure must contain failures, not just handle them.

An agent pauses for approval, the operator steps away, and twenty minutes later the entire workflow has vanished—revealing how production web infrastructure breaks when cognitive state must survive interruption, not just execution.

A privacy researcher turned bot hunter reveals how web automation's adversarial reality creates an endless arms race where human judgment remains automation's hardest problem.

Systems report perfect uptime while delivering worthless data—the most dangerous failure mode in web automation reveals why metrics can't capture what experienced operators see through thousands of production moments.

Kate Blair built coordination infrastructure for agents that can't talk to each other—then merged it with a competitor to create ecosystem-level architecture that prevents fragmentation.

Operating web agents at scale reveals how independent retry mechanisms multiply exponentially—three retries cascade into sixty-four database attempts before systems recognize failure.

Compliance teams discover their evaluation frameworks fail when agents make probabilistic decisions—revealing the invisible expertise gap that determines whether enterprise AI actually deploys.

Browser engine decisions in Spain ripple through thousands of web automation sessions—revealing how upstream engineering creates the operational complexity practitioners navigate daily.

An oncologist building AI for two-minute cancer decisions reveals what production-ready means when errors have immediate human consequences and regulated environments demand explainable reasoning.

An analyst's six-hour competitor price check reveals the invisible labor consuming strategic work—and why web automation at scale demands infrastructure, not just scripts.

Michael Bargury maps novel AI threats to compliance frameworks auditors recognize—solving the gap between agent security and governance infrastructure that actually works.

Operating web agents at scale demands invisible expertise—pattern recognition, tribal knowledge, and cognitive load that no dashboard captures but determines whether reliability holds.

Alex Reibman's AgentOps tackles the infrastructure gap that makes debugging probabilistic agent failures fundamentally different from traditional software—and why visibility matters for production deployment.

Modern websites look ready before they actually work—a gap invisible to users but operationally unanswerable when automating thousands of sites simultaneously.

Building reliable automation requires more infrastructure than blocking it—the operational burden of persistence outweighs the complexity of precision.

Websites invest millions in bot detection that must be surgically precise—yet operational reality makes perfect accuracy impossible at scale.

Modern websites assemble from twenty independent services that load on separate timelines—creating invisible coordination chaos that only becomes visible when automating thousands of browser sessions at scale.

When web automation retries fail, one authentication error multiplies into dozens of attempts, triggering rate limits and IP blocks that compound operational costs exponentially.

Production reveals what staging cannot teach—how the adversarial, constantly changing web actually behaves under real operational conditions.

Staging validates your code logic perfectly while missing the real test: whether your assumptions about the web match reality.

The web forgets you exist between every click—here's the invisible infrastructure constantly reconstructing your identity across systems architecturally incapable of memory.

APIs promise programmatic access but often exclude the data enterprises actually need, forcing hybrid approaches that combine official channels with browser-based collection at scale.

Production deployment reveals knowledge gaps automation didn't know existed, exposing expertise that operated invisibly through human judgment.

Years of invisible expertise make manual work look simple until automation attempts to replicate it and discovers the gap.

The same URL serves different content at 9 AM versus 9 PM—the web has time zones most visitors never see because they only catch one slice of a constantly transforming surface.

An analyst's manual price-checking reveals how website personalization creates parallel realities that make competitive intelligence far more complex than it appears.

Enterprises have been making billion-dollar decisions blind—not from laziness, but because gathering the data was mathematically impossible until now.

Companies have been flying blind on critical decisions not from laziness but because gathering the data was mathematically impossible—until infrastructure made the economics work.

The web you see in your browser doesn't exist in the HTML that arrives from servers—it's rendered by JavaScript, creating parallel infrastructure that makes traditional automation obsolete and reveals why operating at web scale requires fundamentally different systems than most people realize.

Manual price checking runs at 100 products per hour while Amazon reprices every ten seconds—revealing why web automation at scale requires infrastructure depth most teams underestimate.

Manual work that keeps pilots running disappears at production scale—the hidden human effort that success metrics never captured.

Infrastructure gaps invisible at pilot scale become cascading failures in production—what volume reveals about systems never designed to handle it.

Behind every URL sits dozens of regional variants serving fundamentally different content—the web's invisible geography that multiplies operational complexity geometrically.

Tracking competitor pricing means checking dozens of personalized variants simultaneously, turning simple monitoring into complex infrastructure requiring continuous maintenance.

Platforms optimize for individual conversion, creating personalized experiences that make systematic competitive monitoring operationally impossible by design.

Teams request double the resources they need for safety, then multiply that across hundreds of services—revealing how rational local decisions create systemic waste nobody owns.