Vera Loopback is the Exception Resolution Manager at a mid-size online home and lifestyle retailer. The company is large enough to process tens of thousands of orders monthly but small enough that she is, as she puts it, "the entire exception department." For five years, she has handled the return cases that automated systems route past: wrong variants, listing disputes, items that technically match the SKU but not the customer's expectation. More recently, her queue has started filling with something new. Returns from purchases where the buyer never visited the product page at all.
We spoke over video. Behind her, a whiteboard was covered in what appeared to be a color-coded taxonomy of return reason codes, several circled in red with question marks. She'd been expecting us to ask about operational efficiency. We didn't.
Vera Loopback is a fictional character. Her exception queue is not.
Your job title says "resolution." What are you actually resolving?
Vera: The official answer is escalated return cases. The real answer is that I resolve the cases where the reason code is wrong. That's the common thread. The automated system handles the clean ones — bought it, don't want it, here's your label, here's your refund timeline. My queue is everything where the customer's actual complaint doesn't fit the dropdown menu.
"Item not as described" — one reason code covering about fifteen different realities. The throw pillow that photographs as terracotta but arrives looking like salmon. The "premium stainless steel" water bottle that weighs nothing. The size chart that's technically correct in centimeters but wrong for how anyone in that market actually understands "large." Each of those is a different organizational failure. They all arrive in my queue wearing the same label.
You've described your cases as containing intelligence. What does that mean concretely?
Vera: Last quarter I processed maybe three hundred cases involving products from one supplier. Different products, different categories. Kitchen organizers, storage bins, a desk lamp. Unrelated items. But the pattern was identical: listed dimensions were consistently about 15% larger than what arrived. Not wrong enough to be a shipping error. Wrong enough that customers felt misled.
That's a supplier-quality signal. Procurement should know. The listing team should know. The category manager should know. But my system has no field for "escalate to procurement." I have "resolved," "refunded," "exchanged," and "denied." Four buttons. So I close the case, and the intelligence stays in my head until the next batch arrives.
How much of your queue carries that kind of cross-functional signal?
Vera: Most of it, honestly. The routine stuff is gone — automation handles that beautifully. What's left is the interesting cases, which sounds like a compliment until you realize "interesting" means "expensive and slow to resolve." My cost-per-case has gone up every quarter for two years. Not because I've gotten worse at my job. Because the easy cases left, and nobody recalibrated what "good performance" looks like when your queue is 100% hard problems.1
Let's talk about the new category — returns from agent-mediated purchases.
Vera: This is the one that keeps me up at night. We're starting to see orders where — and I had to piece this together myself, because there's no flag for it — an AI shopping agent placed the order on behalf of the customer. The customer told some purchasing agent what they wanted, and the agent selected a specific product and bought it.
Customer opens the box. "This isn't what I wanted." I look at the case and the product matches the listing perfectly. The listing is accurate. Fulfillment was correct. Payment cleared. Everything worked. Except the customer never saw the product page. Their agent saw it, or saw whatever structured data it could parse from it, and made a choice the customer wouldn't have made themselves.2
How do you even code that return?
Vera: You don't. That's the whole problem. "Item not as described" — no, it was described fine. "Wrong item sent" — no, we sent exactly what was ordered. "Changed my mind" — they didn't change their mind, they never made up their mind. Their agent did.
I've been writing "other" and putting notes in the free-text field that I'm pretty sure no one reads. It's like shouting into a suggestion box that's been bolted shut.
What's the most common version of this?
Vera: Variant mismatches. The agent picks the right product but the wrong version of it. Wrong scent, wrong colorway, wrong size within a size range. I had one — a customer's agent ordered protein powder, same brand the customer always buys, but it grabbed the stevia-sweetened version instead of the monk fruit one. Same product line, same GTIN family, technically a valid substitution by any reasonable standard. Customer can't stand stevia.3
The agent wasn't wrong in any way our systems can identify. It was wrong in a way that only the customer's palate knows about. And there's no data field for palate.
Visa recently defined "agentic transactions" as purchases completed without direct cardholder-merchant interaction. Does that framing match what you're seeing?
Vera: It matches the structure. But the framing is about payment authorization — did the consumer intend this purchase? My problem sits one layer deeper. The consumer intended a purchase. They authorized the agent to buy protein powder. The agent bought protein powder. The question is whether authorizing the category is the same as authorizing the specific selection.4
I've started calling it "mandate depreciation." The instruction was valid when issued. By the time it became a specific product in a specific box, something was lost in translation. Not fraud, not error. Drift between what someone meant and what got executed.
Your KPIs measure cases resolved and cost per resolution. What would you measure if you could?
Vera: [long pause]
Signal routed. Cases where the intelligence in the exception actually reached the function that could act on it. Right now that number is effectively zero, because there's no routing mechanism. I close the case, refund the customer, and the listing that caused the problem stays exactly the same until it generates enough returns to show up in someone else's dashboard. Which might take months.
The research backs this up. The tools writing product titles and bullet points have literally never seen a return reason.5 The optimization pipeline and the returns pipeline don't talk to each other. I'm the only person who crosses that gap, and I do it in my head, with no tools, and nobody's asked me to.
That sounds like it could make someone bitter.
Vera: It makes me specific. There's a difference. I'm not angry that nobody listens. I'm precise about what the silence costs. Every unrouted signal is a listing that stays wrong, a supplier that stays unchecked, a product that keeps coming back. I can put a dollar figure on it. Twenty to thirty dollars per return in direct processing costs, but the real cost is the next fifty returns from the same listing that nobody fixed.6
Where does this go? As agent-mediated purchasing scales, does your queue get worse?
Vera: It gets different. And denser. The W3C and GS1 are convening a workshop next month on how to distinguish "finding a product" from "identifying the exact physical variant the user intended."7 That's my Tuesday. That's what I do every day. The standards bodies are just now asking the question I've been answering case by case for two years.
If agent-mediated purchasing scales — and it will — then the exception queue becomes the place where you learn what agents get wrong. Not in the abstract. In the specific, case-by-case, this-customer-got-stevia-instead-of-monk-fruit way. Someone has to read those cases. Someone has to notice the patterns. Right now that someone is me, and my job description says I'm supposed to make myself unnecessary.
The automation is supposed to eliminate my role. But the automation is generating a new category of cases that only a human can interpret. I'm not going anywhere. I'm just getting more expensive and harder to justify on a spreadsheet. Which, if you think about it, makes me the most accurate leading indicator my company has. They just haven't built the dashboard for it yet.
After we ended the call, Vera sent a follow-up email with a single attachment: a spreadsheet cross-referencing return reason codes with supplier IDs across six months of exception cases. "Nobody asked for this," she wrote. "I made it anyway. Someday there'll be a field for it."
Footnotes
-
Automated systems can handle 70–80% of standard return requests; the remaining 15–25% requiring human judgment are proportionally more complex and costly. See Heeya, "How to Handle Returns & Refunds with an AI Chatbot," May 2026. https://heeya.fr/en/blog/handle-returns-refunds-with-ai-chatbot-2026 ↩
-
Visa defines an Agentic Transaction as "an electronic-commerce transaction undertaken by an Agentic Payment Provider on behalf of a cardholder... completed without direct cardholder-merchant interaction." Associated Press reporting on Visa's framework noted that when an AI agent acts on a consumer's behalf, existing law does not clearly establish whether the consumer authorized the specific purchase or merely the general act of shopping. ↩
-
GPT-4o's accuracy on the APOLLO preference-recovery benchmark — which tests whether an agent can correctly infer user preferences from interaction history — was 51.16%. Better than random. Not good enough for stevia vs. monk fruit. ↩
-
Chargeflow, "Agentic Commerce Regulation 2026: What Merchants Must Know," July 2026. https://www.chargeflow.io/blog/agentic-commerce-regulation-what-merchants-need-to-know ↩
-
Avenirlabs reported that 38% of Amazon returns in 2025 cited "quality not as expected" or "item not as described" — complaints tracing to the listing itself — yet listing optimization tools operate on a completely separate data pipeline from returns intelligence. https://www.avenirlabs.com/2026/05/05/the-blind-spot-in-ai-ecommerce-tools-why-returns-data-never-reaches-your-listing-optimizer/ ↩
-
NRF / Happy Returns data establishes that processing a single return costs a retailer roughly $20–30 once shipping, inspection, restocking, and support are counted. Average e-commerce return rates run 20–30% across most product categories. ↩
-
W3C/GS1 Workshop on E-Commerce for Humans and AI Agents, scheduled September 8–9, 2026. https://www.w3.org/2026/ecommerce-agents/ ↩
