Why Bearer Tokens Fail
AI agents can plan, delegate, and call tools at machine speed, but a bearer token grants access based on possession, not intent. Once an agent inherits a broad credential, a prompt injection, compromised tool, or mistaken plan can turn that credential into a path across systems. Zero trust changes the test: verify every action, not just the agent’s initial login. Establish a distinct identity for each agent and task, then evaluate the requested operation against its purpose, context, risk, and current policy.
Also worth reading: How Are You Monitoring AI Agents in Production? · How Do Enterprises Secure AI Agents in Production Without Slowing Down Innovation? · How Should You Evaluate AI Agents Before Production Deployment in 2026?
Before production, teams need enforceable least privilege, short-lived credentials, tool-level authorization, and rapid revocation. They also need isolation between agents and workloads, clear approval gates for consequential actions, and durable records linking decisions to identities and policies. Emerging efforts such as Peon’s Rust runtime with Casbin and OWASP-aligned AgentSign point toward policy enforcement built for agent workflows, but a framework alone is not a security model. CISOs should test failure modes, including delegation and changing context, and prove that controls work under pressure. The goal is not to slow agents indiscriminately; it is to make every capability explicit, bounded, and auditable before granting autonomy.
Designing Identity-First Agent Architecture
Zero-trust AI agents cannot become production-ready by attaching API keys to prompts and calling a bearer token sufficient protection. As Anthropic’s zero-trust test, Pap’s model, Peon’s Rust runtime, and AgentSign’s OWASP-aligned engine all suggest, the agent itself must become a verifiable, short-lived identity. Every model call, retrieval request, tool invocation, and data transfer needs continuous authorization based on user, device, task, environment, risk, and requested resource—not merely possession of a credential.
That identity-first architecture changes engineering, not just policy. Casbin-backed policy decisions can constrain agents at runtime; scoped credentials eliminate reusable secrets; egress controls limit lateral movement; and tamper-evident logs show who acted, under which delegation, and why. High-impact actions still require human approval, revocation must propagate immediately, and agents should receive no more privilege than the current task needs. The result is not a safer chatbot wrapper but a constrained actor that can be authenticated, authorized, observed, and stopped. For a practical architectural review, agustin-otegui.com helps teams align AI speed with trust before deployment.
Enforcing Least Privilege by Action
Zero-trust AI agents cannot reach production with broad credentials and a promise that the model will “follow instructions.” Static bearer tokens fail because they are replayable, overly durable, and blind to who or what is acting. Every tool call, data read, message send, and code change instead requires a verifiable workload identity, short-lived authorization, contextual policy decisions, and an auditable denial path. Least privilege must be enforced outside the model, not merely described in prompts or requested in system messages.
Recent work from Anthropic, Pap, Peon, AgentSign, Forbes, and Morphisec converges on the same test: can policy remain effective when an agent is compromised? A production runtime should evaluate identity, resource, action, environment, and risk on every request, using deny-by-default rules and scoped credentials. A Rust runtime with Casbin can make those decisions explicit, while OWASP-aligned engines such as AgentSign can expose and govern them. High-impact operations may still require human approval. As an AI Architectural Consultant at agustin-otegui.com, I see the priority as simple: keep enforcement independent, observable, and faster than permission expansion.
Adding Visibility and Runtime Evidence
Zero-trust AI agents must move beyond perimeter assumptions and static prompts. Anthropic’s framing exposes a core failure: a bearer token proves possession, not identity, intent, authority, or current trust. Before production, platforms need continuous verification, least-privilege tools, scoped credentials, approval gates, tamper-evident logs, and runtime policy evaluation. Pap and the Rust-based Peon show the direction: enforce authorization beside agent execution, with Casbin helping isolate capabilities rather than giving models broad trusted sessions.
Production also requires evidence operators can use. AgentSign’s OWASP-aligned work suggests security cannot be an API wrapper around opaque orchestration. Teams should correlate identity, prompts, tool calls, retrieved data, policy decisions, and outputs in traceable run records, while supporting rapid revocation and ownership. Morphisec’s CISO perspective reinforces the point: AI speed matters only when controls remain observable and enforceable. As with single-vendor SASE, consolidation can simplify operations but also conceal gaps and concentrate risk. The production test is whether every consequential action can be authenticated, authorized, explained, and stopped.
A Practical Zero-Trust Rollout
Zero-trust AI agents need more than a renamed API key. As Anthropic’s zero-trust guidance and Pap’s work emphasize, a bearer token fails once it can be copied, replayed, or overprivileged. Before production, every agent should receive a short-lived, audience-bound identity, with permissions scoped to tools, resources, environments, and actions. AgentSign’s OWASP-aligned approach and Peon’s Rust runtime, using Casbin, point toward an enforcement layer rather than another policy document. Runtime authorization must evaluate each call, constrain delegation, redact context, log decisions, and fail closed without strangling autonomy.
The rollout should connect agent policy to the controls used for human and workload access. At agustin-otegui.com, I frame this as an architecture problem: establish identity, enforce least privilege continuously, contain lateral movement, and preserve evidence across models and frameworks. Morphisec’s CISO research and Forbes coverage reinforce the urgency, while the single-vendor SASE analogy exposes the risk of treating one platform as a budget shortcut instead of a security boundary. Production readiness depends on measurable controls, adversarial testing, rapid revocation, and tested recovery—not simply labeling an agent “zero trust.”
Agent Security Control Comparison
| Production Control | Current Failure Mode | Required Change |
|---|---|---|
| Agent identity | Shared API keys and bearer tokens enable replay, impersonation, and excessive privilege | Issue per-agent workload identities with short-lived, audience-bound credentials and automated rotation |
| Authorization policy | Static RBAC cannot account for task intent, context, or changing risk | Enforce default-deny, least-privilege ABAC at every action using an engine such as Casbin |
| Tool and data access | Broad MCP, API, or database permissions allow lateral movement and destructive actions | Apply scoped capabilities, egress allowlists, sandboxing, filtering, and approval thresholds |
| Runtime assurance | Perimeter-only controls provide little visibility into autonomous decision loops | Add continuous policy evaluation, tamper-evident audit trails, kill switches, and OWASP-aligned monitoring |