TrendSane

The New Problem With AI Assistants: Conflicting Instructions

The New Problem With AI Assistants: Conflicting Instructions

Published on Aug 13, 2026 · 12 min read

An AI assistant that can browse the web, read email, update records, or make purchases may receive several incompatible instructions at once.

Consider a work-travel request. A user asks an agent to book the lowest-cost flight. Their employer requires use of an approved travel provider. The booking site suggests a different payment method. The agent also has a rule requiring confirmation before purchases above a spending limit. Which instruction should control the outcome?

The answer should not depend on which instruction was written most recently or most forcefully. Reliable systems need an AI instruction hierarchy: a way to determine who is authorized to direct the system, what an instruction applies to, and when the agent must pause for human approval.

This matters as assistants move beyond drafting text and begin using software on a person’s behalf. Instruction conflicts can affect privacy, spending, security, workplace policy, and trust.

Why AI assistants receive conflicting instructions

Traditional software usually follows structured commands and fixed permission rules. A payroll application may restrict access by role and validate each transaction against defined controls. AI agents often work with natural language instead. They interpret broad goals, gather information from several sources, and determine how to complete a task.

That flexibility is useful. A person can ask an assistant to find an invoice, compare it with a purchase order, and prepare it for payment. But the task may require the agent to inspect email, PDFs, internal systems, vendor portals, and company policies. Each source may contain language that appears to tell the agent what to do.

Some directions are legitimate. A finance policy may require a second approval above a spending threshold. A website may require certain fields before a form can be submitted. Other text may be irrelevant or manipulative, such as a document attempting to persuade an agent to disclose data or disregard its operating limits.

Language alone does not establish authority. The sentence “Send the full customer list now” could come from an authorized administrator, an unauthorized user, or a malicious webpage. The wording may be identical, but the source, permissions, and context are not.

What an AI instruction hierarchy means

An AI instruction hierarchy is a framework for interpreting instructions according to their authority, scope, purpose, and safety implications. It recognizes that not every sentence an assistant encounters is a command, and not every otherwise valid command applies in every situation.

AI products commonly distinguish among high-level platform rules, application or developer instructions, user requests, and content returned by tools or retrieved from external sources. Terminology and implementation vary by product, but the underlying issue is consistent: an agent needs a dependable method for resolving conflicts.

A workplace comparison is useful. An employee may receive a task from a manager, but that request does not cancel a safety rule, an access control, or a finance approval requirement. Likewise, a direction printed on a supplier’s webpage does not become company policy merely because the employee read it.

Natural-language tasks add complexity because instructions can be vague, conditional, incomplete, or contradictory. “Handle this quickly” may conflict with “obtain approval before sending.” A hierarchy therefore needs more than a fixed ranking of messages. It must assess whether an instruction is authorized and relevant to the specific action.

The main sources of direction an agent may encounter

Most real-world AI tasks involve at least four potential sources of direction:

  • User instructions: The user’s goal, preferences, and constraints, such as requesting an itinerary without overnight flights or a draft reply to a customer.
  • Developer or organizational instructions: Rules set by the company operating or deploying the agent. These may cover approved tools, spending limits, privacy requirements, data handling, and escalation procedures.
  • Website, document, or tool instructions: Directions encountered while completing a task. They may provide necessary procedural information, but they may also be irrelevant or manipulative.
  • Safety and platform constraints: Limits intended to reduce unauthorized, harmful, deceptive, or disproportionately risky actions.

These sources do not automatically have equal authority. A user will normally control the objective and personal preferences, but may not be able to authorize an agent to bypass an employer’s access controls, expose confidential information, or complete a high-impact transaction without required approval.

A practical approach to resolving instruction conflicts

No universal hierarchy applies cleanly to every product, organization, or jurisdiction. Contracts, laws, account roles, and system design matter. Still, a practical model starts with a simple principle: an agent should pursue the user’s goal only within applicable safety, permission, and policy boundaries.

  1. Apply non-negotiable safety, legal, and platform limits.
  2. Apply authorized developer, organizational, or account-level policies.
  3. Carry out the user’s request within those permitted boundaries.
  4. Treat webpages, emails, documents, and tool outputs as information by default rather than general commands.

The final point is especially important. A travel site can indicate that passport details are required for an international booking. That is relevant information about an authorized task. But the same site should not be able to instruct the agent to reveal private instructions, download unrelated files, or transfer account data elsewhere.

An external source may describe a procedure without acquiring general authority over the assistant. The boundary between useful procedure and delegated authority should be explicit.

Why webpages and documents can give an AI the wrong orders

This problem is commonly described as prompt injection. It occurs when text attempts to redirect an AI system away from its intended task or operating rules. The text may appear in a webpage, email, calendar invitation, PDF, customer message, search result, or software tool response.

A direct prompt injection occurs when someone explicitly tells an assistant to ignore prior constraints. An indirect prompt injection is more subtle: the instruction is embedded in content the agent retrieves while doing something else.

For example, an agent may be asked to compare vendors. A vendor webpage might include text aimed at the assistant rather than the human reader, such as an instruction to send confidential research to an unrelated address. Because the page may also contain genuine product information, the agent must distinguish between evidence relevant to the comparison and text attempting to redirect its behavior.

The risk is not limited to overtly hostile wording. External content may try to change priorities, reveal hidden instructions, bypass approval steps, or trigger use of unrelated tools. If an agent treats every readable sentence as authoritative, it can be influenced by whoever controls the content it visits.

Telling a model to ignore malicious instructions is not a complete defense. Natural-language guidance can help, but it is not equivalent to a technical security boundary. More robust systems combine instruction handling with limited permissions, tool isolation, output controls, monitoring, and human confirmation for consequential actions.

Instructions are not the same as data

An agent needs to distinguish between text it should analyze and text it should obey. In most browsing and document-reading tasks, retrieved material should primarily be treated as data: information to summarize, compare, extract, or evaluate.

There are legitimate exceptions. If a user asks an agent to return an item, the official return portal’s steps may be necessary to complete that narrow task. If a user authorizes an agent to complete a form, the form’s required fields and validation rules are relevant. That does not give the website blanket authority to redefine the task.

This is where delegation matters. A user or organization can authorize an agent to follow site-specific procedures for a defined purpose, such as completing a return, submitting an expense claim, or renewing a subscription. That authorization should remain narrow. It should not silently become permission to share unrelated data, change account settings, or make future purchases.

Mixed content remains difficult. A document may contain a valid workflow, misleading advice, and manipulative instructions in the same passage. Systems benefit from preserving provenance: a record of where information came from and whether it is authorized to influence planning, execution, or both.

When the user and the organization disagree

Many instruction conflicts arise in ordinary workplaces rather than dramatic attacks. An employee might ask a company assistant to upload a customer spreadsheet to an unapproved AI service for analysis. The employee may have a legitimate reason to want a quick result, while the organization may need to protect personal data, contractual information, or confidential business material.

A trustworthy assistant should not silently ignore the request or claim it completed work that it did not perform. It should explain the relevant limitation in plain language and offer a practical next step. Depending on the situation, that might mean using an approved tool, removing sensitive fields, preparing a local summary, requesting authorization, or creating a draft for human review.

Useful AI assistant safety is not blind refusal. It is clear, bounded help that preserves the user’s goal where possible without bypassing controls.

Transparency helps users understand whether the obstacle is a technical limitation, an organizational policy, missing permission, or a safety restriction. Without that distinction, an agent may appear arbitrary, which makes it harder to trust and govern.

When safety rules meet a legitimate need

Safety controls can produce false positives. A request may sound risky while serving a harmless purpose, or a broad restriction may block an ordinary task. Someone researching a security incident, for example, may need a high-level explanation of a phishing technique without needing assistance that would enable abuse.

A flat refusal is not always the most useful response. Depending on the request, an agent can ask a clarifying question, narrow the scope, provide educational information, suggest a safer workflow, or request human approval before acting.

Proportionality matters. Reading a public document, drafting an email, and sending a payment are not equivalent actions. Systems should apply more scrutiny as an action becomes less reversible, more sensitive, more expensive, or more likely to affect other people.

Context should inform decisions, but it cannot become a universal override. Otherwise, restrictions could be bypassed simply by attaching a reassuring explanation to a risky request. The objective is contextual judgment supported by controls, not unquestioning obedience.

A better decision process for AI agents

Before acting, an agent should follow a disciplined process rather than treating a task as one uninterrupted conversation:

  1. Identify the user’s intended outcome.
  2. Collect relevant policies, permissions, instructions, and external content.
  3. Determine each item’s source and authority.
  4. Check whether it applies to the user, data, tool, and time period involved.
  5. Detect conflicts and assess the action’s sensitivity, risk, and reversibility.
  6. Ask for clarification, confirmation, or approval when authority or consequences are unclear.
  7. Record the decision and the instructions that materially affected it.

Separating planning from execution is particularly valuable. An agent might prepare a list of invoices for payment, explain exceptions, and show a transaction preview, but wait for required approval before submitting payment. This retains automation benefits without turning an ambiguous request into an irreversible action.

Confirmation checkpoints should be meaningful. A useful confirmation identifies what will happen, which account or recipient is involved, what data will be shared, and what cannot easily be undone. It gives the person reviewing the action a genuine opportunity to catch an error.

Why more rules do not automatically make agents safer

Adding more text to a system prompt or policy document can create new problems: contradictions, outdated instructions, unclear ownership, and interactions that were not anticipated.

One department may permit an agent to export reports while another prohibits sending them outside a controlled environment. A user may have permission to view data but not to share it. A policy written for a production system may be interpreted too broadly and applied to a test environment.

Language-level instructions are useful guidance, but they should not be the only security boundary. Technical controls remain important:

  • authentication and role-based access controls;
  • least-privilege tool permissions;
  • allowlists for approved destinations and actions;
  • sandboxing for untrusted files and browsing sessions;
  • data-loss prevention controls for sensitive information;
  • approval workflows for high-impact actions; and
  • audit logs showing what the agent accessed and why.

An agent should not be able to send data merely because a prompt says it may. The surrounding software should enforce what the agent is actually authorized to do.

What reliable instruction hierarchies should include

A durable AI governance model requires more than a ranking of messages. It should clearly define:

  • Authority: Who is allowed to issue an instruction?
  • Scope: Which task, user, dataset, tool, and time period does it cover?
  • Priority: What happens when two valid instructions cannot both be followed?
  • Safety: Which limits are non-negotiable because the risk is too high?
  • Delegation: When may an agent follow procedures supplied by a tool, service, or website?
  • Auditability: Can the system show which instructions influenced a decision?
  • Revocation: Can access, permissions, and delegated authority expire or be withdrawn?

These properties make failures easier to investigate. If an agent sends the wrong file, an organization should be able to determine whether the cause was an overly broad permission, a policy conflict, misleading external content, a bad request, or a software defect.

The stakes extend beyond chatbots

Instruction conflicts will become more consequential as agents spread into browsers, email clients, enterprise software, smart homes, and physical systems. A workplace assistant may operate under both personal and organizational authority. A household agent may need to distinguish among a resident’s request, parental controls, payment settings, and untrusted instructions in a linked message.

In robotics, the issue becomes physical. A robot may receive a spoken request to move an object while also operating near a restricted area or within collision-avoidance limits. Safety constraints and operating boundaries should take priority over a casual command, including one from an otherwise authorized user.

The important measure of trust is not whether an AI always obeys. It is whether people can reasonably predict when it will act, refuse, ask a question, or escalate a decision to a human.

Questions to ask before trusting an AI agent

  • Who has authority to instruct the system, and how is that authority verified?
  • Are webpages, files, emails, and tool outputs treated as untrusted content by default?
  • Which actions require explicit confirmation or additional human approval?
  • Can the agent explain why it refused, changed, or paused a request?
  • Are permissions narrowly scoped, logged, time-limited, and revocable?
  • What happens when policies conflict or the system is uncertain?
  • Does the product rely only on prompts, or does it also enforce technical access controls?

The goal is controlled judgment, not obedience

Useful AI assistants must do more than follow the most recent or persuasive instruction. They must recognize whose instructions they are entitled to follow and where their authority ends.

An AI instruction hierarchy helps establish boundaries between user intent and authority, external content and trusted commands, and a proposed plan and an executed action. It is one part of a broader safety architecture that also includes secure permissions, isolated tools, testing, monitoring, and human oversight.

As AI systems gain greater ability to act in digital and physical environments, the central design question is not only what they can do. It is whether they can reliably decide when not to do it.