# How Should You Design Runtime Agent Security Architecture in 2026?

Savannah Jenkins · September 27, 2026

> Direct Answer: Treat the Agent as an Untrusted Program Runtime agent security architecture is the set of controls applied while an AI agent is...

## Direct Answer: Treat the Agent as an Untrusted Program

Runtime agent security architecture is the set of controls applied while an AI agent is reasoning, calling tools, reading data, and changing systems—not only before or after execution. The direct answer is to place a policy-enforcing control plane between the model and every consequential action, then give that plane explicit identity, context, audit, and emergency-stop mechanisms. The agent should receive least-privilege credentials, tool calls should be evaluated against deterministic policies, sensitive outputs should be inspected before release, and high-impact actions should require independent approval. This model assumes that the model can be manipulated, that instructions can conflict, and that an apparently harmless tool response can contain hostile content. It also recognizes that runtime security is not a replacement for ordinary application security, IAM, data protection, or secure model deployment.

**Also worth reading:** [How Do Enterprise Security Teams Handle Agentic AI Threat Modeling in Modern System Architecture?](https://agustin-otegui.com/knowledge/how_do_enterprise_security_teams_handle_agentic_ai_threat_modeling_in_modern_system_architecture.php) · [How Should Modern Enterprises Design an AI Financial Valuation Architecture to Drive Long-Term Value?](https://agustin-otegui.com/knowledge/how_should_modern_enterprises_design_an_ai_financial_valuation_architecture_to_drive_long-term_value.php) · [How do you design a zero trust architecture for agentic AI systems?](https://agustin-otegui.com/knowledge/how_do_you_design_a_zero_trust_architecture_for_agentic_ai_systems.php)

A useful architecture usually has four boundaries: the model boundary, the orchestration boundary, the tool boundary, and the infrastructure boundary. At the model boundary, prompts and retrieved content are labeled by trust level. At the orchestration boundary, the runtime controls planning, state, budgets, retries, and delegation. At the tool boundary, each action is authorized based on the agent’s identity, task, target, data classification, and risk. At the infrastructure boundary, operating-system, network, and runtime controls limit what a compromised process can do. The aim is not to trust the agent because it passed a pre-deployment evaluation, but to constrain it continuously because behavior can change with context, tool output, and user instructions.

## How Runtime Control Works During an Agent Session

The runtime begins by issuing a short-lived, workload-specific identity rather than handing the agent a general-purpose API key. It records the user, tenant, agent version, objective, allowed data domains, and permitted tools in a signed execution envelope. Every tool request must match that envelope, and authorization should be reevaluated when the target, parameters, volume, or sensitivity changes. For example, a support agent permitted to read five order records should not automatically gain permission to export an entire customer table or issue a refund above a stated limit. Policy decisions should be explainable enough for an auditor to answer who acted, under which authorization, and through which control.

Open Policy Agent, the policy technology used by projects such as Cupcake, is one practical enforcement mechanism for tool and orchestration decisions. Policies can permit or deny actions using attributes such as role, environment, resource, requested operation, data classification, time, and transaction value. Deterministic rules are especially useful for actions that cannot tolerate probabilistic judgment, such as production writes, payments, privilege changes, or bulk exports. The model may propose an action, but it should not be the final authority when a hard business or security rule exists. A production architecture would normally combine policy-as-code, scoped credentials, egress filtering, content inspection, and complete trace collection rather than treating a policy engine as the whole solution.

## A Reference Architecture for AI Agents

A logical request path can move through an API gateway, agent gateway, policy decision point, tool proxy, and protected execution environment. Before the model runs, the gateway authenticates the caller, validates the requested model and tools, and establishes quotas. The agent gateway then maintains session state and distinguishes trusted system instructions from untrusted documents, web pages, email, code, and tool results. A policy decision point evaluates each proposed tool call, while a tool proxy supplies temporary credentials and attaches transaction identifiers to downstream requests. Sandboxing, egress controls, and data-loss prevention remain active even when a tool appears safe or has been approved previously.

The architecture should also separate planning from execution. A planner may produce a proposed sequence, but deterministic workflow code can decide which steps are allowed and whether human confirmation is needed. This separation reduces the chance that a prompt injection embedded in retrieved material silently changes the agent’s objective. State should be immutable or append-only where practical, with signed decisions stored alongside prompts, tool schemas, policy versions, and outputs. The runtime should cap tool-call counts, execution time, token use, fan-out, and cost; sensible starting points might be 20 tool calls per task for a low-risk workflow, no unrestricted subprocess creation for general agents, and a zero-tolerance default for production credential access. Exact thresholds should come from task testing, not arbitrary industry rules.

## Comparison of Runtime Security Approaches

There is no single product category called an agent runtime gateway, and vendors describe overlapping capabilities differently. Some controls sit in a gateway, some in an agent orchestration framework, some in a sandbox, and some in cloud-native infrastructure. A buying decision should compare enforceable boundaries and failure behavior rather than marketing labels.

| Feature | Gateway or Policy Layer | Sandboxed Execution Environment |
| --- | --- | --- |
| Primary role | Authorizes tool calls, identity, context, and policy decisions | Constrains code, files, processes, and resource consumption |
| Best-known decision | Should this action proceed? | What can this process safely touch? |
| Strength | Centralized, auditable policy across models and tools | Limits damage after code or agent instructions are compromised |
| Common weakness | Policy mistakes can overauthorize actions | Does not by itself understand business intent or user authorization |
| Human review | Useful for high-risk or novel actions | Usually reserved for deliberate escape or elevated operations |
| Typical cost | Often usage-based, per request or per decision | Usually based on compute, storage, and observability consumption |
| Required complement | Sandboxing, IAM, logging, and DLP | Gateway policy, scoped identity, egress controls, and audit |

Gateway-based controls are attractive when an organization already operates API security, IAM, or policy-as-code. Sandboxing is necessary for code execution, but a sandbox cannot infer that a database query is inappropriate merely because the process is technically contained. The strongest option usually uses both, with a third layer of human or workflow approval for irreversible actions. Kernel-level and eBPF-based monitoring can add visibility and enforcement at lower levels, but it should complement—not obscure—the application-level decision trail.

## Practical Implementation Steps and Initial Thresholds

Start with one bounded workflow and inventory every capability the agent can exercise, including reading files, browsing sites, sending messages, executing code, calling APIs, and spawning other agents. Classify tools by consequence: read-only internal data, external disclosure, reversible mutation, and irreversible or regulated action. Begin with read-only tools and shadow mode, where the runtime records what the agent would do without executing it. This creates a baseline and reveals normal call volume, data access patterns, and failure modes without granting broad production permissions.

Next, define deny-by-default policy and conduct adversarial tests before deployment. Include direct prompt injection, indirect injection in retrieved documents, poisoned tool output, credential theft, cross-tenant access, data exfiltration through URLs, encoded payloads, and attempts to bypass approval by splitting a large action into many small ones. Use a threshold such as a maximum of 10 sensitive records for a standard support task, one approval for a payment, and immediate termination for access to a production secret. These numbers are design examples, not universal standards; teams should adjust them according to data sensitivity, model performance, and business tolerance for false positives.

Then introduce staged enforcement: log and alert, block clear violations, require approval for high-impact actions, and automate only after measured precision is acceptable. Every override should have an owner, reason, expiration time, and audit record. Track security metrics such as policy denial rate, approval rate, blocked exfiltration attempts, tool-call latency, unauthorized-action rate, false-positive rate, mean time to revoke credentials, and percentage of high-risk actions with a complete trace. A practical initial target might be 100% credential coverage and 100% traceability for privileged actions, because those are binary architectural requirements, not aspirational percentages.

## Common Security Mistakes and Design Traps

The most common mistake is treating prompt instructions as an authorization system. A model can follow a malicious instruction hidden in a web page, spreadsheet, issue, or email, and a system prompt is not a security boundary. Another mistake is giving one agent a shared service account across users or tenants. Even if the orchestration layer behaves correctly, a confused or compromised process can reuse that identity outside intended context. Credentials should be short-lived, audience-bound, and selected by the tool proxy after authorization, with no raw secret returned to the model.

Teams also make the mistake of logging prompts without protecting the logs. Traces may contain credentials, personal data, source code, or confidential business information, so retention, encryption, redaction, and access policy matter. Another trap is evaluating only whether the final answer contains a secret while missing exfiltration through tool parameters, DNS, image URLs, or outbound messages. Conversely, indiscriminate content inspection can create privacy, latency, and reliability problems. Inspection should be risk-based, with stronger controls for secrets, regulated records, and high-volume or unusual destinations.

A subtler error is assuming that a benchmark score predicts authorization correctness under hostile conditions. A runtime must be tested against the actual system context, including tool descriptions, retrieved text, retries, delegated agents, and changing permissions. Do not allow a child agent to have broader access than its parent merely because it is faster to configure. Avoid circular trust between the model, orchestrator, and policy engine, and do not permit an agent to modify its own policy, audit trail, approval rule, or credential scope. Finally, do not deploy an elaborate security platform before defining the smallest set of actions that genuinely need to be blocked.

## When to Act, and What It May Cost

Act during design for agents that can write code, access internal systems, handle regulated data, execute transactions, or delegate to other agents. Waiting until after an incident is especially risky because runtime policy touches credentials, workflows, logging, and service ownership; retrofitting it under urgency often produces a brittle gateway that teams bypass. A smaller pilot can still begin within a few weeks, but production-grade assurance may take several months because it requires data classification, identity design, threat modeling, red-team scenarios, and operational ownership.

Costs depend heavily on the implementation route. Open-source policy engines, orchestration frameworks, and sandboxes can reduce software licensing costs, but engineering, cloud compute, storage, observability, incident response, and security review remain real expenses. Commercial API gateways and runtime security products may use per-request, per-decision, per-host, or annual subscription pricing; cloud-native services often add charges for model calls, protected compute, log ingestion, and network inspection. A small internal agent with modest volume may cost tens to hundreds of dollars monthly in infrastructure before labor, while a high-volume regulated deployment can reach thousands or more as traces and compute grow.

The correct investment is proportional to consequence, not to the novelty of AI. If a failed agent can only draft a response for human review, a gateway with logging and simple restrictions may be enough. If it can deploy code, move money, alter customer records, or retrieve confidential data, the architecture should include isolation, strong identity, policy enforcement, human approval, continuous monitoring, and tested revocation. A budget should include the cost of evaluating false positives and manual approvals; a control that blocks legitimate work at a 5% rate may be more expensive than a narrower rule that blocks the actual abuse path.

## A Practical Decision Framework for 2026

The key architectural question is not “Does the model produce a good answer?” but “What is the maximum credible damage if this agent is wrong or manipulated?” For a low-impact drafting agent, isolate it, restrict network access, and keep the final action with a person. For an internal operations agent, add workload identity, tool-level authorization, tenant boundaries, transaction logging, and approval for writes. For an agent that operates production infrastructure, use a dedicated ephemeral environment, narrowly approved tools, no standing credentials, independent policy control, immutable audit events, and a kill switch tested at least quarterly. For multi-agent systems, propagate authority downward and prevent a delegate from accumulating permissions across tasks.

By 2026, runtime security is converging with API security, cloud workload protection, identity governance, and AI governance, but those fields do not automatically solve agent-specific problems. Agents generate dynamic action sequences, interpret untrusted content, and can combine otherwise legitimate tools into an illegitimate outcome. Meta’s reported use of a kernel-level sentinel for Muse illustrates the direction toward deeper runtime protection, while NVIDIA’s work on agent infrastructure and in-silicon security points toward controls closer to hardware. Those developments may improve defense in depth, yet they do not remove the need for application-level authorization and business context.

A defensible rollout therefore begins with inventory, classification, and a narrow pilot; progresses through shadow enforcement, adversarial testing, scoped credentials, sandboxing, and human approval; and ends with measured operating thresholds. The architecture should be reviewed whenever a new model, tool, data source, agent role, or delegation path is added. Success is not the absence of blocked requests, but evidence that the runtime prevents unacceptable actions, preserves legitimate work, and gives operators a fast, understandable account of every consequential decision.

## Quick answers

### Is runtime agent security the same as a firewall?

No. A firewall controls network paths, while an agent runtime evaluates identity, intent, context, tool parameters, and data sensitivity for each action. It may still need firewalls, egress controls, and sandboxing underneath it, but a network rule cannot determine whether a particular database update is authorized.

### What is the minimum viable runtime security design?

A useful minimum is a deny-by-default tool gateway, short-lived scoped credentials, untrusted-content labeling, audit logs, and human approval for high-impact actions. Add sandboxing and egress restrictions when the agent executes code or retrieves sensitive data. The exact design should follow the highest credible consequence of a compromised agent.

### Can Open Policy Agent secure an AI agent?

Open Policy Agent can evaluate tool calls and orchestration decisions using policies based on identity, resource, action, and context. It does not provide identity issuance, sandboxing, secret redaction, or complete audit storage by itself. Those controls must be supplied by the surrounding runtime.

### How do you stop an agent from exfiltrating data?

Use allowlisted destinations, data classification, parameter inspection, scoped credentials, and controls on outbound messages, URLs, DNS, and file transfers. Monitor unusual volume or novel destinations, and block actions that combine sensitive reads with unapproved external writes. No single prompt instruction is sufficient because tool output can itself be hostile.

### Should every agent action require human approval?

No, because constant approval can make an agent slow, expensive, and operationally fragile. Automate low-risk, reversible actions with clear limits, and reserve approval for irreversible, regulated, privileged, or unusually novel actions. Measure false positives and approval latency so the approval boundary improves over time.

Canonical: https://agustin-otegui.com/knowledge/how_should_you_design_runtime_agent_security_architecture_in_2026.php
Markdown: https://agustin-otegui.com/knowledge/how_should_you_design_runtime_agent_security_architecture_in_2026.php/index.md
