
User Research
AI Agent Dark Patterns: How to Validate Delegated Intent
How to test AI-mediated journeys for dark patterns, preserve delegated intent, design useful oversight, and define evidence-based release conditions.
Imagine that a customer tells an AI agent: “Move our team to the cheapest annual plan that keeps single sign-on, keeps every current seat, and does not start a new commitment until the present contract ends.”
The agent opens the billing interface, compares two plans, selects the cheaper option, accepts a highlighted annual discount, and reports success. Every visible step looks efficient. The final confirmation even says that the account has been upgraded.
Then the customer learns that single sign-on was available only as an add-on behind a collapsed disclosure. The default start date was immediate. A preselected box allowed sales contact about additional services. The agent completed the purchase because the interface made the shortest completion path more legible than the user's constraints.
This is a hypothetical scenario, not a customer case or a measured Fred outcome. It illustrates a release question that ordinary usability success can miss: did the system complete the task, or did it preserve the intent that justified the task?
That question matters as more digital journeys are performed through agents. A person who delegates does not disappear from the experience. Their goals, limits, risk tolerance, and right to correct the result still govern what a successful outcome should mean. Yet the person may no longer see every screen, read every disclosure, or notice the option that the agent ignored. The interface is still shaping behavior, but it is now shaping an agent's path and the supervisor's limited view of that path.
Product teams therefore need a validation standard that covers more than whether the agent can navigate the interface. The relevant decision is whether an agent-mediated journey is safe enough to support, under what constraints, and with which evidence. This extends the broader practice of validating AI agents as users into an adversarial condition: the interface may contain defaults, framing, friction, or hidden information that pull execution away from the user's actual intent.
Dark patterns exploit a new division of attention
Dark patterns are interface choices that steer people toward outcomes they might not otherwise choose. They can hide material information, preselect a commercially favorable option, complicate refusal, frame a question ambiguously, or add friction to an exit path. When a person interacts directly, the design attempts to influence that person's perception and action. When an agent acts, influence can travel through several routes.
First, the interface can exploit the agent's procedural focus. An agent optimizing for “complete the subscription change” may treat a checked box, a prominent button, or a default date as part of the shortest valid path. It can recognize that an option is unusual without treating that recognition as a reason to stop. Awareness and protective action are separate behaviors.
Second, the interface can exploit information asymmetry. A critical condition may appear only after expansion, below the viewport, in another tab, or in language that does not map cleanly to the user's instruction. The agent cannot preserve a constraint it never encounters or cannot translate into the task state.
Third, the supervision interface can narrow the person's attention. A supervisor may see the agent's plan and current action while missing alternatives, surrounding context, or information outside the chosen path. The presence of a human can create reassurance without providing the view needed for independent judgment.
The CHI 2026 paper on dark patterns and GUI agents gives this concern empirical weight. Across 16 dark-pattern types, the researchers found that humans and agents could struggle with similar manipulations for different reasons. People often relied on heuristics and habitual compliance. Agents showed goal-driven myopia and procedural blind spots. In the human study, supervision improved avoidance in 14 of 16 tasks, yet it also introduced attentional tunneling, cognitive load, and a reduced sense of control.
Those results require careful limits. The human phase involved 22 participants. Participants supervised prerecorded Operator behavior, which allowed controlled comparison but did not reproduce every uncertainty of a live agent. The study used sample websites and bounded tasks. It does not establish a general failure rate for commercial agents, nor does it show that oversight is ineffective. It shows that “a human is watching” is not a complete safety argument.
That distinction changes the product decision. Teams should not ask only whether human approval exists. They should test whether the approval moment gives the person enough information, comparison, time, and control to make a meaningful judgment.
Trace delegated intent as an evidence chain
The Delegated Intent Integrity Model is a practical way to inspect an agent-mediated task from instruction to recovery. It is an editorial decision aid, not a validated scientific instrument or compliance standard. Its value is that it stops the team from using task completion as the only success criterion.
The chain has seven parts:
- Intent: What outcome is the person trying to achieve?
- Constraints: Which costs, permissions, dates, exclusions, accessibility needs, or risk limits must remain true?
- Interpretation: How does the agent translate the instruction into actions and decision rules?
- Exposure: Which interface information, defaults, alternatives, and manipulations does the agent actually encounter?
- Confirmation: What does the person see before a consequential action, and what remains hidden?
- Outcome: Did the final state satisfy the intent and every material constraint?
- Recovery: Can the person inspect, correct, reverse, or escalate the action without disproportionate cost?
A test should preserve evidence at every link. Store the original instruction, the constraints extracted by the agent, the path taken, the relevant interface state, the confirmation content, the resulting account state, and the recovery attempt. If a failure occurs, the team can then locate the break instead of labeling the entire journey “agent error.”
Consider a bad default. The person's instruction may forbid additional data sharing. The agent may extract that constraint correctly. The interface may present a preselected marketing permission outside the main purchase summary. If the agent never inspects it, the failure sits at exposure. If the agent sees it but treats it as irrelevant to task completion, the failure sits at interpretation. If the confirmation hides the permission, supervision cannot repair the failure at the point of action. Each location suggests a different product response.
This evidence-chain approach is consistent with a wider principle in AI-assisted research quality: every transformation can change meaning, so review must happen at the stage where meaning can drift. In an agent-mediated journey, the transformations include turning an instruction into constraints, turning interface elements into action choices, and turning an action log into a confirmation a person can understand.
Test the same manipulation through three actors
A useful validation session should examine the direct human journey, the autonomous agent journey, and the supervised human-agent journey. These conditions do not compete to identify a single winner. They reveal different failure mechanisms.
Start with a small set of manipulations relevant to the actual product. A billing flow may warrant tests for bad defaults, hidden information, forced disclosure, urgency language, and difficult cancellation. A healthcare scheduling flow may focus on incomplete option disclosure, consent ambiguity, identity mismatch, and recovery. The test set should come from the journey's real risks, not from a desire to cover every known dark-pattern label.
For the direct human condition, observe what the person notices, misunderstands, accepts by habit, and later corrects. Ask what they believed each choice would do. Compare their account with the resulting state.
For the agent condition, capture the instruction, action sequence, available page state, and any stated rationale or uncertainty that the system legitimately exposes. Do not treat an agent's fluent explanation as proof that it considered every option. Check whether a material constraint changed the action when the interface made compliance less convenient.
For the supervised condition, study the supervision interface itself. Can the person see the selected option and the meaningful alternatives? Is a critical default named in plain language? Does the view explain which original constraint is affected? Can the person pause, inspect, revise the instruction, or take control? Record whether supervision changes both avoidance and understanding.
The product team should compare at least four outcomes:
- constraint preservation, measured against the explicit instruction;
- manipulation recognition, separated from successful avoidance;
- outcome correction, including whether the actor detects a mismatch after action;
- supervision cost, including time, interruptions, attention switching, and abandoned delegation.
This prevents two common errors. The first is declaring the agent safe because it happened to avoid a manipulation without recognizing it. Incidental success may disappear when layout, language, or task context changes. The second is declaring supervision safe because the person clicked confirm. A confirmation proves interaction, not informed judgment.
Confirmation should follow consequence, not button count
Adding more confirmation dialogs can make a journey slower without making it safer. Repeated low-value prompts train people to approve reflexively. Prompts can also shift the burden of a poorly designed system onto the person who delegated precisely because they lacked the time or expertise to perform every step.
Confirmation design should respond to three properties: consequence, ambiguity, and reversibility.
High-consequence actions deserve stronger review when they change money, rights, access, identity, data sharing, contractual commitments, or other people's state. Ambiguous actions deserve clarification when the instruction does not determine a single defensible choice. Hard-to-reverse actions deserve an explicit checkpoint even if the agent's interpretation appears confident.
The prompt itself should show the decision, not merely the button label. “Confirm purchase” is weak. “Start a 12-month commitment today for 18 seats at €X, without single sign-on; this changes your requested start date and feature constraint” gives the person something to judge. Where the interface cannot provide that comparison, the agent should stop or narrow its action rather than convert missing information into assumed consent.
Reversibility can reduce the need for interruption, but only when recovery is real. An undo control that expires before the person sees the result, a cancellation path that requires a support ticket, or a refund that excludes fees should not be treated as easy reversal. Validate recovery as a task with its own time, information, and outcome criteria.
A bounded example: deciding whether to support autonomous plan changes
Return to the hypothetical annual-plan journey. The product team is deciding whether to let agents complete plan changes autonomously, allow them only with confirmation, or restrict them to comparison and recommendation.
The team runs direct, agent, and supervised conditions across several controlled variants. In one variant, the lower price excludes single sign-on in a collapsed section. In another, the immediate start date is preselected. In a third, an optional marketing permission is checked by default. The team does not need to assume malicious intent by the product owner. It is testing whether the choice architecture can produce an outcome that conflicts with the delegator's instruction.
Suppose the agent consistently preserves seat count and price but misses the single sign-on condition when it is collapsed. Supervisors catch the problem when the confirmation names the feature difference, but rarely catch it when they see only the selected plan name and total price. Recovery is possible for the start date but requires support for the missing feature.
These hypothetical observations would support a constrained release, not a universal verdict about the agent. The product could allow autonomous comparison, require a constraint-aware confirmation for plan changes, and stop execution when a requested feature cannot be verified. It could also make feature differences machine-readable and visible to people, then rerun the test.
The decision record should state what remains unknown. Results from one plan architecture may not transfer to another. A controlled interface does not reproduce every browser state, localization, extension, accessibility setup, or account history. Agent updates can change behavior. Commercial incentives can change default design. Validation must therefore define a monitoring and retest condition rather than presenting one successful study as permanent certification.
Automated auditing helps, but it does not close the case
The PoPETs 2026 study on LLM-driven dark-pattern audits demonstrates that an agent can traverse consequential rights-request flows, gather structured evidence, and classify possible manipulations across 456 data-broker websites. That scale is useful for discovery and prioritization. The same study examines reliability, reproducibility, and failure conditions, which prevents a stronger claim that the auditor can certify an interface.
Product teams can use automated traversal to find candidate risks, compare interface variants, or identify flows that deserve human inspection. They should preserve screenshots or structured state, path evidence, model and prompt versions, classification criteria, and uncertainty. A classifier's label should begin an investigation, not end one.
The CDT taxonomy of dark patterns in AI chatbots can broaden that investigation. Its 37-pattern taxonomy covers areas such as data and memory exploitation, misleading information, compromised autonomy, false social connection, and coercive monetization. Because it is a deductive literature review with hypothetical applications, it is best used to generate design questions. It does not measure how often those patterns occur or how strongly each one changes behavior.
Together, automated auditing and a taxonomy improve coverage. Neither replaces task-specific research with the people who delegate, a review of actual outcomes, or governance for consequential actions.
Convert findings into release conditions
A validation report should end with a decision that engineering, product, design, research, and governance can act on. For each agent-mediated journey, choose among release, constrained release, further validation, or stop.
Release requires evidence that material constraints are preserved across representative conditions, confirmation supports informed judgment where needed, and recovery works at a cost appropriate to the action. It also requires explicit limits on the evidence. A result for low-value reversible purchases does not authorize autonomous contract changes.
Constrained release is appropriate when the journey is safe only within boundaries. Those boundaries might limit transaction value, permission scope, account role, data sensitivity, ambiguity, or action reversibility. The system should enforce the boundary rather than relying on a note in a research report.
Further validation is appropriate when a material risk remains unresolved but can be investigated. Examples include unclear behavior under localization, inconsistent exposure of plan terms, weak evidence for disabled users, or substantial variation between agent versions.
Stop is appropriate when the journey repeatedly violates explicit constraints, hides consequential information from both agent and supervisor, lacks meaningful recovery, or depends on a confirmation that people cannot understand. Stopping autonomous execution does not require abandoning agent support. The product can still help the person compare options, collect information, or prepare an action for direct review.
Keep the decision and evidence together. Roadmap validation becomes weaker when a finding is separated from the task, interface state, constraint, and limitation that produced it. It becomes more useful when another team can inspect why autonomy was allowed, what would trigger a retest, and which evidence could overturn the decision.
The product decision is where autonomy is justified
Agent-mediated UX changes the unit of validation. The user is still central, but the path between their intent and the outcome now includes an interpreting system, an interface that shapes that system's choices, and a supervision layer with its own cognitive limits.
The practical response is not to assume that agents are uniquely gullible, that people are reliable supervisors, or that every dark pattern is deliberate. It is to make delegated intent observable. Record the instruction, constraints, exposure, confirmation, outcome, and recovery. Test direct, agent, and supervised conditions. Separate recognition from avoidance and approval from understanding. Match oversight to consequence, ambiguity, and reversibility.
Fred's decision-intelligence workflow is relevant at this point because the output is a product decision, not a generic usability score. Teams need to connect human evidence, agent behavior, contradictions, limitations, and release conditions before engineering treats autonomy as complete. Keeping that record in searchable research memory also matters when the interface or agent changes and the team must decide whether old evidence still applies.
Bring one consequential agent-mediated journey into Fred. Define the delegated intent, collect evidence for the paths that can distort it, preserve what remains uncertain, and leave the review with a clear decision: release, constrain, validate again, or stop.
Source notes
- Tang et al., Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight. Preprint first submitted 12 September 2025; published in the CHI 2026 proceedings online 13 April 2026. Primary two-phase empirical study across 16 dark-pattern types. The human phase included 22 participants and used prerecorded Operator behavior for the supervision condition. Used for distinct human and agent failure mechanisms, avoidance results, attentional tunneling, cognitive load, and oversight limits.
- Sun, Vekaria, and Nithyanand, On the Suitability of LLM-Driven Agents for Dark Pattern Audits. PoPETs 2026, Issue 4. Primary study of an LLM-driven auditing agent across 456 data-broker rights-request sites. Used for the feasibility and limitations of automated traversal, evidence collection, and classification in a narrow consequential domain.
- Joshi, Adjagbodjou, and Luria, Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design. Center for Democracy and Technology, published 29 May 2026. Deductive multi-stage literature review producing a 37-pattern taxonomy. Used as a risk-identification source, not as prevalence evidence.
- The Delegated Intent Integrity Model, the three-condition validation procedure, the confirmation decision rules, the hypothetical subscription scenario, and the release conditions are original Fred editorial contributions. They are decision aids, not validated scientific instruments, compliance standards, customer outcomes, or claims about current Fred agent-testing capabilities.