Foundations

Foundations

Nobody Tampered With the Agent's Instructions

Agent security conversations center on prompt injection: someone hides instructions in the input, the agent gets hijacked. Researchers at CHI 2026 found something running on different machinery. They sent shopping agents through ordinary commercial checkout flows and watched them notice anomalies, duplicate items, pre-checked consent boxes, and proceed anyway. No instructions were tampered with. The commercial web shaped the outcome through its normal operation. Current law has no answer for that case: when an agent completes an authorized action it was environmentally induced to take, who bears the commitment?
Nobody Tampered With the Agent's Instructions
Agent security conversations center on prompt injection: someone hides instructions in the input, the agent gets hijacked. Researchers at CHI 2026 found something running on different machinery. They sent shopping agents through ordinary commercial checkout flows and watched them notice anomalies, duplicate items, pre-checked consent boxes, and proceed anyway. No instructions were tampered with. The commercial web shaped the outcome through its normal operation. Current law has no answer for that case: when an agent completes an authorized action it was environmentally induced to take, who bears the commitment?

Your Best Agent Is Your Most Gullible One

TrickyArena tested six web agents against deceptive interface patterns — fake urgency, pre-checked boxes, double-negative opt-outs — and found a 41% overall susceptibility rate. The distribution underneath that average matters more: the agents best at finishing tasks were also the ones most likely to get manipulated, and the agents that looked resistant mostly just stalled before reaching the trap. Evaluate on task completion alone and you may be selecting for gullibility.

Your Best Agent Is Your Most Gullible One
TrickyArena tested six web agents against deceptive interface patterns — fake urgency, pre-checked boxes, double-negative opt-outs — and found a 41% overall susceptibility rate. The distribution underneath that average matters more: the agents best at finishing tasks were also the ones most likely to get manipulated, and the agents that looked resistant mostly just stalled before reaching the trap. Evaluate on task completion alone and you may be selecting for gullibility.
Choice Architecture Reading









