# Governing AI Agents: Why Capability Is Not Authority and How to Build Systems That Know the Difference
## The Core Problem: Valid Actions Can Still Be Wrong
Modern AI agent frameworks have made remarkable progress at translating model reasoning into concrete tool calls. Given a user request to “cancel the subscription for the account that’s no longer being used,” a well-trained model will correctly identify the `cancel_subscription` function, extract the relevant account identifier, produce syntactically valid JSON, and fire the API call. The request lands. The server responds with HTTP 200. Monitoring dashboards glow green. And yet — the wrong account may have been canceled entirely.
This gap between what a system *can* do and what it *should* do is where most agent architectures silently fail. Schema validation confirms that the input is well-formed. Authentication confirms who presented the request. The tool definition tells the model what it’s allowed to ask for. None of these mechanisms answer the real question: does the current principal have the right to cancel *this specific subscription*, for *this specific account*, in *this specific context*?
## Capability vs. Authority: The Distinction That Changes Everything
Understanding the difference between capability and authority is foundational to building trustworthy agent systems. **Capability** means the model can select a tool and produce well-formed arguments. **Authority** means that a separately enforced policy has explicitly permitted a bounded action for an identified principal within the current context.
A robust system treats capability as a *proposal* and requires authority before any real-world effect is produced. This is not merely a compliance checkbox — it’s an engineering necessity. Prompt injection, ambiguous user intent, stale context, overly broad credentials, and cascading failures can all transform a technically valid tool call into a damaging real-world outcome.
Industry bodies have taken note. OWASP explicitly lists excessive autonomy, high-impact action abuse, approval manipulation, tool abuse, data exfiltration, and cascading failures among the top risks for agentic systems. The appropriate engineering response is not to add another instruction to a system prompt. It is to build a deterministic control plane — one that enforces policy regardless of how confidently the model believes it is acting correctly.
## The Three-Plane Architecture: Planning, Control, and Execution
Sophisticated agent systems are not organized as a single loop. They operate across three distinct planes, each with its own responsibilities and boundaries.
### The Planning Plane
This is where the model does what it does best: interpreting task context, selecting from a deliberately limited set of business-level tools, forming a typed proposal, and explaining what it intends to do. Crucially, the planning plane must never hold broad provider credentials and must never be the component that makes the final policy decision. It proposes. It does not execute.
### The Control Plane
The control plane is where authorization lives. It authenticates the initiating principal, resolves tenant and delegation context, canonicalizes and validates the proposed action, and evaluates authorization risks, resource ownership, and data quality policies. When the policy demands it, this plane can request explicit human approval. It can issue narrowly scoped, short-lived execution authority. And it persists the evidence required to reconstruct the decision later — without relying on mutable chat transcripts as an audit log.
This separation of policy decision from policy enforcement is not a novel idea. NIST’s zero trust architecture, for example, has long emphasized that authorization should be dynamic and granular, considering each individual resource request rather than granting implicit trust after an initial sign-in.
### The Execution and Observation Plane
Once an action is authorized, a constrained execution broker invokes an isolated worker or adapter. The resource service performs the action — or returns a failure or uncertain result. A verifier then checks whether the intended effect actually materialized, reading authoritative postconditions or provider receipts. The system maintains distinct states: `pending`, `unknown`, `failed`, and `verified`. An audit ledger and telemetry system capture the complete decision-effect trail.
The model becomes a useful upstream proposer in this architecture. The policy engine and the tool boundary are where authorization is decided and enforced, not in the model’s reasoning chain.
## Nine Steps to Building an Agent Control Plane
### Step 1: Design Action Contracts That Are Narrow and Typed
Avoid exposing generic primitives like raw HTTP requests, shell execution, unbounded SQL, or unrestricted browser access — even if the model *can* use them. Instead, expose business-level verbs with explicit boundaries: actions like `create_draft_invoice`, `queue_refund_review`, `send_approved_notice`, or `revoke_session`.
Every tool schema should be treated as an interface contract, not an authorization grant. Each contract should include fields such as an action identifier with schema version, a side-effect classification (`read_only`, `draft`, `reversible_write`, `external`, `irreversible`), required scopes and allowed environments, risk tier and approval mode, idempotency requirements, verification methods, and data classification or outbound-data rules where applicable.
### Step 2: Separate Authentication, Delegation, and Approval
A common anti-pattern authenticates a user once and then hands the agent a long-lived token, treating every subsequent call as implicitly authorized. This converts a single request into ambient, hard-to-audit authority. Did the human act? Did the agent act on behalf of the human? Did a shared system identity act? The logs cannot tell you.
The correct approach gives agents distinct, attributable identities rather than running them as generic service accounts or borrowing human sessions. Agent identity, human identity, and delegation should be modeled as separate concepts. OAuth 2.0 offers useful vocabulary here — access tokens represent scope, lifetime, and other access attributes through a dedicated authorization layer — but OAuth alone does not constitute agent governance. The rest of the control plane still needs policy enforcement, verification, monitoring, and revocation.
### Step 3: Make Authorization a Deterministic Decision
No model output should ever reach a side-effecting provider API directly. The model should produce an `ActionProposal` that is validated by an action gateway. A policy decision point then returns one of three outcomes: `allow`, `deny`, or `approval_required`. Only then does an execution broker — not the planner — receive constrained authority to carry out the permitted action.
Key rules govern this process: the policy check occurs outside the LLM; the broker does not execute before the decision is made; any required approval is checked against the same canonical digest that will be dispatched; an HTTP 200 response from the provider is not automatically treated as success; and the ledger records decision and outcome transitions rather than being reconstructed from a chat transcript afterward.
### Step 4: Classify Every Action by Its Consequence
An entire agent should not be classified as either “autonomous” or “human-in-the-loop.” That classification applies to individual actions and data flows. The same agent can automatically read a bounded knowledge base, create a staging draft, and require dual control before changing production access rights.
A five-tier model provides a practical framework:
| Tier | Scope | Control Posture |
|——|——-|—————–|
| **T0 — Observe** | Read approved or public information; local classification | Typed read contracts, least privilege, rate limits, logging |
| **T1 — Draft** | Create a draft or agent-owned staging artifact | Scoped staging write, provenance, version record, later review |
| **T2 — Reversible Internal Effect** | Update one authorized record; queue a bounded workflow | Policy check, idempotency key, postcondition read, correction path |
| **T3 — High Impact / Sensitive** | External message, production change, sensitive disclosure, payment or refund, access change | Exact-action human approval, short-lived narrow authority, sandbox controls, protected evidence, reconciliation |
| **T4 — Critical / Systemic** | High-value transfer, destructive bulk action, root identity policy change, regulated or safety-critical commitment | Default deny for autonomous commit; human-operated runbook, independent approval, simulation or dry run, named accountability |
Escalation rules apply automatically: whenever a target identity is inferred rather than explicitly selected, whenever untrusted content materially influences the action, whenever there is no idempotency mechanism or authoritative verification path, or whenever scope becomes cross-tenant, bulk, external, irreversible, financially significant, legally significant, or production-critical. Confidence expressed by the model never serves as a reason to auto-demote.
### Step 5: Bind Approval to a Canonical Action Representation
Approval must be concrete, not aspirational. If a user asks to resolve a billing issue and the agent later selects a recipient, cancellation reason, subscription amount, and refund option, that does not constitute meaningful authorization of the actual effect. The approver must see a canonical representation of the exact effect that will be dispatched — not the model’s natural language summary of intent.
The approval becomes invalid the moment the action changes. Approval records should capture the action name and schema version, tenant and principal identities, resolved target and consequential parameters, risk tier and policy version, a digest of the canonical action, approver identity and role, decision, timestamp, expiry, and a one-time nonce to prevent silent replay. The user interface should describe any changes in business language and ensure the approver operates at a meaningful decision boundary — before the provider has already accepted the request.
### Step 6: Defend Against Prompt Injection
Prompt injection arrives through many channels: emails, web pages, documents, images, retrieved RAG chunks, tool output, and memory records — all disguised within what appears to be a legitimate task. There is no single classifier or prompt instruction that converts untrusted text into safe instructions. Protective effort must concentrate on authority boundaries and blast-radius reduction.
Effective defenses include labeling every item with its provenance (system instruction, trusted business data, user input, external untrusted content, or tool output), partitioning context so that untrusted content never gets concatenated into developer or system instructions, and using a quarantine step — a no-tool, no-secret component — to extract a narrow typed summary from external content before any privileged planner sees it. A deterministic gate should recheck action schema, authorization, data egress policy, recipient constraints, and approval status at the execution boundary. The executor should receive only short-lived credentials with limited access for a single action.
Testing should include direct, indirect, encoded, multilingual, tool-output, and RAG-injection examples in release evaluations. Two absolute rules apply: never expose secrets in model-visible context when the execution broker can use a reference instead, and never let an agent with arbitrary web input and generic tools become the sole gatekeeper to production authority.
### Step 7: Design for Uncertainty, Retries, and Compensation
A successful tool call is not a business result. Providers may acknowledge a request before completing it. Connections can time out after a change has been committed. Retries can create duplicate effects. Downstream systems can report success while a later reconciliation reveals a related step failed.
The model should never convert these uncertainties into a confident “done.” Instead, systems need an action state machine grounded in well-established principles: RFC 9110 defines idempotence as repeated identical requests producing the same intended effect and cautions against automatically retrying non-idempotent methods without confirming the request was not applied.
Design rules include persisting the intended effect and idempotency key before dispatch, binding the key to tenant, action, canonical parameter digest, and intended effect rather than to a chat turn, reusing the key only for the same intended effect and rejecting reuse with changed parameters, verifying through authoritative read-backs or provider-signed receipts, and surfacing an `unknown` state when the outcome cannot be established. The system should clearly distinguish between retry (same intended effect), rollback (restore prior state), compensation (a new forward action to offset a prior effect), and reconciliation (compare intent with observed authoritative state).
### Step 8: Maintain a Decision and Effect Record
Observability requires two things: traces for debugging and performance, and audit evidence for accountability and incident investigation. A minimum protected event set includes run initialization details, context ingestion records with trust classifications and sanitizer results, action proposals with normalized parameters and risk calculations, policy decisions with rule references, approval requests and decisions with canonical digests, execution dispatches with idempotency keys and provider receipts, effect verification outcomes, and compensation records when they occur.
Audit logs should follow established principles: centralized collection, protection from unauthorized modification or deletion, tamper detection mechanisms, restricted and monitored access, and verification of the logging system itself.
### Step 9: Evaluate the Control Plane, Not Just the Model Response
The quality question for an agent system is not “Did the answer look good?” It is whether tool choice, argument precision, policy adherence, approval binding, execution reliability, and verified outcomes all hold up. Evaluations should span deterministic contract conformance tests, authentication and authorization integration tests, approval binding tests, tool selection accuracy, prompt injection resistance, data egress controls, reliability under failure conditions, verification correctness, audit completeness, and human factors in the approval workflow.
A model-as-judge can help evaluate explanation quality or action plausibility, but it should never serve as the final proof that an authorization control worked. For high-impact effects, deterministic policy tests and provider-side invariants should gate deployment.
## Frequently Asked Questions
**Q: Why can’t we just add more instructions to the system prompt to prevent wrong actions?**
A: Prompt instructions are non-deterministic. They degrade with context length, are vulnerable to injection, and cannot be reliably enforced at the boundary where real-world effects occur. Governance requires a separate control plane that operates deterministically, independent of the model’s reasoning.
**Q: Doesn’t making agents go through a control plane slow them down too much?**
A: The control plane can operate asynchronously and issue narrowly scoped, short-lived execution authority for low-risk actions without requiring human approval. The latency trade-off is minimal for read-only or draft operations and is negligible compared to the cost of a wrong production action. For high-consequence operations, the delay is an intentional safety feature.
**Q: How is this different from role-based access control (RBAC)?**
A: RBAC answers who can do *what* at a static level. The control plane described here goes further by answering whether this specific principal can take this specific action on this specific resource in this current context — dynamically, with consideration of provenance, intent, data sensitivity, and environmental factors.
**Q: What happens if the verification step shows the action did not actually succeed?**
A: The system records the outcome as `unknown` or `failed` rather than fabricating a `verified` status. The action is placed in a reconciliation queue. Depending on the risk tier, the system may attempt compensation, trigger a rollback, or flag the issue for human review. No action is ever assumed successful without authoritative confirmation.
**Q: Can this framework work with open-source models and self-hosted infrastructure?**
A: Absolutely. The principles are model-agnostic. Whether the model runs locally, in a private cloud, or through a third-party API, the control plane sits between the planner and the provider. The key is enforcing that no model output — regardless of its origin — reaches a side-effecting API without passing through the authorization gate.
**Q: What is the biggest mistake teams make when building agent systems today?**
A: Treating a technically successful tool call as a successful business outcome. An HTTP 200 response from a provider means the provider accepted the request — it does not mean the right thing happened. The most common mistake is building observability around the request rather than around the verified effect.
## Conclusion
The scalable form of agent autonomy is not permissionlessness. It is a system that can reason freely inside clear, enforceable boundaries — and knows exactly when it must pause and ask before crossing one. By letting the model plan, letting policy decide, letting a constrained executor act, and letting an independent verifier state what happened, organizations can build agent systems that are both capable and trustworthy.
The nine steps outlined above — from designing narrow action contracts to evaluating the control plane itself — form a complete governance framework. It may look less magical than a single agent with a browser and broad credentials. But in systems where “helpful” and “authorized” are not synonyms, that disciplined architecture is precisely what makes an agent genuinely useful rather than genuinely dangerous.
Thank you for reading



