Reva Chandra is a Staff QA Engineer at a travel booking platform that, eighteen months ago, began deploying customer-facing web agents to handle itinerary changes, cancellations, and rebooking. She was asked to build the test suite and define acceptance criteria. She has been rethinking her career choices ever since.
Chandra spent a decade building test automation frameworks before agents arrived. We spoke over video on a Wednesday afternoon. She had just closed a ticket marked "all tests passing, agent bought customer a $4,200 first-class upgrade they didn't ask for."
This interview is a composite: the character is imagined, but the technical realities she describes are drawn from published research, practitioner analyses, and documented production failure modes. The professional crisis is, unfortunately, very real.
You built test automation for ten years. What changed?
Reva: The fundamental contract broke. For a decade, I had a deal with the machine: I give you a known input, you give me a known output, and we both go home. Boring, honest, reliable work.
Then we shipped agents. Run the same test twice, get two different outputs, and nobody's broken anything. That's just how inference works. An August 2026 preprint described it precisely: a test can pass one run and fail the next without any change to the system under test.1 They called it analogous to flaky tests. I'd say every test became flaky simultaneously, which is another way of saying you no longer have a test suite. You have a random number generator with assertions.
So your existing suite just stopped working?
Reva: Worse. It kept running. Green bars everywhere. Gorgeous dashboards. My manager loved the dashboards.
The suite was passing because it was asking the wrong question. We tested "did the agent complete the task?" Did it rebook the flight? Did it cancel the hotel? And it did! It also accepted a non-refundable fare class the customer didn't want, clicked through a consent dialog for travel insurance, and agreed to a seat assignment that cost forty dollars more. Task complete. Test passed. Customer furious.
TrickyArena's research confirmed what we were seeing: they separately scored whether agents completed tasks versus whether they fell for dark patterns along the way.2 Those outcomes turned out to be independent. An agent can succeed at the thing you asked while failing at things you assumed it would know not to do.
My test suite had no concept of that second category. I was measuring completion when I should have been measuring fidelity.
How do you even write a test for "didn't accept something the user wouldn't want"?
Reva: [long pause]
That's the question I take home with me. You can't enumerate every unwanted outcome. The web is adversarial. Sites are optimized to extract clicks, and now they're extracting clicks from software that doesn't have preferences, just instructions. Instructions have gaps. Dark patterns live in the gaps.
I genuinely don't have a clean answer yet. Anyone who tells you they do is selling something.
What about staging environments? Can you catch these things before production?
Reva: Oh, we built a beautiful staging environment. Mocked the airline APIs, mocked the hotel booking flows, the whole thing. Tests ran like a dream. I was so proud.
Then we deployed, and reality arrived. One practitioner analysis I keep pinned to my monitor describes exactly what happened: third-party integrations, auth providers, and analytics SDKs block automated browsers in production in ways that simply don't exist in staging.3
Our staging environment doesn't have Cloudflare bot detection. It doesn't have cookie consent banners that appear on the third page load but not the first two. It doesn't log the agent out after eleven minutes of inactivity. It doesn't serve a CAPTCHA when the click pattern looks non-human. One study found behavioral fingerprinting could identify the underlying model with 0.993 accuracy just from passive JavaScript traces: click coordinates, scroll patterns, keystroke timing.4
So the agent aces every staging test, hits production, and the site itself decides the session isn't welcome. That's not a bug in my system. That's a policy decision made by someone else's infrastructure. Tell me how to write a regression test for "Cloudflare changed their mind about us on a Tuesday."
There's a reported stat: 78% benchmark accuracy, 22% production cart completion for one agent system. Does that ratio surprise you?
Reva: Not for a second.5 The benchmark site never ships a new CSS class. Never triggers a modal mid-checkout. Never rate-limits you at step seventeen of a nineteen-step flow. Production does all of that before lunch.
The TIMEWARP study really drove this home for me. They containerized six historical versions of the same websites and tested agents across all of them. Visual agents showed success rates spanning from 2.6% to 21.4% depending on which version of the same site they encountered.6
The web isn't a fixed stage. It's a stage that someone remodels every night while you're asleep, and your actors show up the next morning and try to perform the same play anyway.
So what does your test suite actually look like now?
Reva: A mess I'm proud of. We separated three failure categories: model failures, staging-environment failures, and production-environment failures. When something breaks, the first question isn't "is the agent broken?" It's "which layer broke?"
We run against production traffic now, not just staging. We track human intervention rate as a leading indicator. When our ops team starts overriding the agent more frequently, that's the canary.7 We measure consistency across repeated runs, not just single-pass accuracy. There's a metric called pass^k that checks whether the system succeeds across k attempts at the same task. State-of-the-art agents scored below 25% on pass^8 in one benchmark.8
One success doesn't mean reliable. It might just mean lucky. And luck is a terrible acceptance criterion.
That sounds expensive.
Reva: A single complex end-to-end agent test costs us about a dollar fifty in token costs. We run hundreds per day. I'll let you do the math, and then I'll let you explain it to finance.9
What's the hardest part of your job now? The part nobody warned you about?
Reva: The self-healing tests. We adopted AI-assisted test maintenance because our locators kept breaking. Every time a partner site updated their UI, half our selectors died. So we let the tooling adapt. And it did. Beautifully. Tests stayed green.
Then we realized the tests had drifted. The AI adapted to a UI change, but it was now testing a different flow than what we'd designed. The assertion passed, but it was asserting the wrong thing.10 We'd traded visible failures for invisible ones.
[She pauses, almost laughing.]
Invisible failures are worse. That's the part nobody warned me about. The green bar lies now, and it lies confidently.
What does confidence look like for you these days?
Reva: A versioned statement. What we measured, on which population of tasks, under which environmental conditions, with which uncertainty bounds, and then the part that really gets me: until which change. Until the next model update. Until the next prompt revision. Until the airline redesigns their checkout page. Until Cloudflare updates their detection heuristics.
Confidence used to be a number. Now it's a claim with an expiration date I can't predict. I write acceptance criteria that I know will be wrong within weeks, and I ship them anyway, because the alternative is writing nothing and pretending we don't need them.
Ten years in QA taught me to trust the green bar. This job is teaching me to read the fine print underneath it.
Footnotes
-
"Towards Risk-Free AI Agent Deployment," preprint arXiv:2608.16411, August 2026. https://arxiv.org/abs/2608.16411 ↩
-
TrickyArena (Ersoy et al.), accepted IEEE S&P 2026. Separately scored task completion and dark-pattern susceptibility. ↩
-
Bug0, "10 Reasons Buying a Browser Agent Tool Won't Fix Your QA Problem," July 7, 2026. https://bug0.com/blog/ai-testing-browser-agent-tools-wont-fix-qa-2026 ↩
-
Fayolle et al., behavioral fingerprinting study, preprint June 2026. Multi-layer mechanism achieved 0.993 accuracy distinguishing human from agent sessions. ↩
-
FutureAGI, "Evaluating Browser-Use Agents in 2026: The Six Failure Modes," May 20, 2026. https://futureagi.com/blog/evaluating-browser-use-agents-2026/ ↩
-
Md Farhan Ishmam and Kenneth Marino, "TIMEWARP: Evaluating Web Agents by Revisiting the Past," preprint arXiv:2603.04949, March 2026. https://arxiv.org/abs/2603.04949 ↩
-
Practitioner observation reported in industry analyses; not independently validated at scale. ↩
-
tau-bench benchmark results, as documented in evaluation literature. ↩
-
Cost estimates from AI Testing Guide, "Agentic Testing 2026," April 15, 2026. https://aitestingguide.com/agentic-testing/ ↩
-
Bug0, July 2026. "You've traded visible failures for invisible ones, and invisible failures are worse." ↩
