Direct Answer: Treat the Agent as an Untrusted Dynamic Identity
A defensible runtime agent security architecture treats every AI agent as a temporary, non-human identity whose permissions must be issued, constrained, observed, and revoked in real time. Static application security is insufficient because an agent can interpret natural-language instructions, select tools, generate code, move data, and change its next action without a fixed predeclared transaction path. The control objective is not merely to block known attacks; it is to limit the maximum damage an agent can cause under unexpected instructions, compromised context, malicious tool output, identity theft, or model error.
Also worth reading: How Should an AI Architect Design an MCP Gateway Architecture for Enterprise Security and Scale? · How should enterprises architect security for Model Context Protocol deployments in 2026? · What Is a Runtime Security Wrapper for Autonomous AI Agents, and When Do You Need One?
The architecture should combine identity and policy enforcement, tool gateways, sandboxed execution, network controls, data-loss prevention, runtime telemetry, and an independent audit trail. Open Policy Agent, or OPA, can evaluate policy decisions expressed as code, while eBPF-based monitoring can observe activity at the operating-system and network layers without requiring every application to be modified. Neither product solves the entire problem: OPA is primarily a decision and policy layer, whereas eBPF is a low-level telemetry and enforcement mechanism. A separate control plane should correlate the agent's identity, prompt or task provenance, intended action, data sensitivity, destination, and accumulated risk before allowing a consequential operation.
For most enterprises, the practical pattern is deny by default, grant short-lived access, require human approval for defined high-impact actions, and continuously evaluate behavior. A useful starting threshold is to permit read-only operations automatically while requiring approval for external writes, privileged commands, production access, regulated-data transfers, or financial transactions. Exact thresholds should depend on the agent's role and business impact rather than on a universal percentage or security score. The important principle is that authorization must be attached to a specific action and context, not simply to a broad statement that the agent is allowed to “use tools.”
How Agent Runtime Risk Differs from Conventional API Security
Conventional APIs normally execute endpoints selected by a client, while an agent uses an LLM to decide which endpoint, command, query, or file operation should occur next. That decision process is probabilistic and influenced by system instructions, retrieved documents, user messages, memory, tool descriptions, and observations returned during execution. An attacker may therefore manipulate the context rather than exploit a fixed endpoint directly, causing the model to invoke a legitimate tool with dangerous parameters. Runtime security must account for both the immediate request and the chain of decisions that produced it.
The main threat classes are prompt injection, tool abuse, excessive agency, identity and credential compromise, data exfiltration, supply-chain manipulation, and unauthorized lateral movement. Prompt injection can arrive through a web page, email, PDF, issue ticket, database record, or tool response. Tool abuse includes invoking a permitted tool outside its intended purpose, chaining tools to bypass a control, or requesting excessive records. A system that permits read_file may still become dangerous if the agent can read secrets, recursively inspect unrelated directories, encode results, and send them to an attacker-controlled host.
Controls should therefore bind permissions to workload identity, task scope, tool, target resource, environment, data classification, and time. For example, a research agent working on a public policy issue might receive read access to a 100-document public corpus for 30 minutes, but no access to source code, customer records, or outbound destinations outside an approved research domain. If the workflow later moves to a private dataset or a production system, it should obtain a new authorization rather than inherit broader access. This limits the blast radius of both successful prompt injection and ordinary model hallucination.
Reference Architecture and Enforcement Boundaries
A practical runtime agent security architecture has six functional layers: workload identity, policy and risk evaluation, execution isolation, tool mediation, data and network enforcement, and observability with response. These layers need a shared event model, but they should not be collapsed into one overloaded security product. Identity services issue short-lived credentials and map the agent to a human sponsor, service account, agent definition, and permitted task. Policy services consume the identity, requested action, resource attributes, current session history, and environmental signals before returning allow, deny, or step-up decisions.
Tool gateways are especially important because they turn abstract model output into controlled operations. Instead of allowing an agent to issue unrestricted HTTP, shell, database, or filesystem calls, each tool should have a typed contract, parameter validation, destination restrictions, result filtering, rate limits, and an audit record. Shell commands should run inside ephemeral containers, virtual machines, or microVMs with a read-only base image, non-root user, restricted system calls, and no host credentials. Network access should use a default-deny egress model, with approved DNS destinations, proxies, and service identities replacing unrestricted internet connectivity.
eBPF technology can support kernel-level visibility into process execution, file access, socket activity, and signals that may be difficult to interpret from application logs alone. It can also enforce selected kernel actions, although it does not reliably reveal whether an operation was caused by malicious instructions, an authorized user, or a model mistake. Application-level events remain necessary for recording the task, tool call, policy result, approval, and outcome. The defense-in-depth principle is stronger when both layers participate, while teams should recognize that kernel observability can generate high volumes, sensitive data, and operational complexity if retention and filtering are poorly designed.
Policy, Approval, and Anomaly-Based Decision Making
Runtime protection should evaluate actions continuously because an agent's risk changes as it moves through a workflow. A harmless document summary may precede access to a restricted file, a credentialed API call, a code deployment, or an outbound transfer. A static role cannot describe every acceptable sequence, so policies can use session state, cumulative action counts, sensitive data access, novelty of destinations, and prior human approvals. A small denial-of-service control should also cap tool calls, tokens, runtime duration, fan-out, and spend per task.
OPA-style policy-as-code is useful for explicit, testable authorization rules and can separate policy decisions from tool implementations. Rules can deny access to production secrets, require step-up authentication for privileged operations, constrain database queries to approved schemas, or require a change ticket before deployment. Its flexibility does not eliminate policy-design risk: broad labels such as “trusted context” or “safe tool” can conceal ambiguous assumptions. Every policy should be versioned, unit-tested against expected and adversarial cases, associated with an owner, and reviewed when the agent model, tool contract, or data environment changes.
Anomaly detection should complement deterministic rules rather than replace them. Useful signals include a sudden increase in tool calls, first-time access to a sensitive repository, repeated failed authorization attempts, unusual process ancestry, communication with a newly registered domain, or access patterns outside the agent's normal task envelope. Threshold selection should start from a baseline rather than a fabricated industry norm. For a new agent, monitor in report-only mode for roughly 2 to 4 weeks, establish expected tool and destination distributions, then introduce conservative blocking for confirmed high-risk patterns. Statistical anomalies are leads for investigation, not proof of compromise, because legitimate tasks can be unusual.
Comparison of Runtime Security Approaches
No single category covers identity, application behavior, infrastructure activity, and governance. The most credible architecture combines approaches, but the balance depends on existing controls, deployment model, agent autonomy, and the sensitivity of reachable systems. A low-code agent using only public data may justify a lighter architecture than a production coding agent with credentials and write access, yet even public-data agents can face resource abuse and malicious retrieval content.
| Feature | OPA and Policy-as-Code | eBPF Runtime Monitoring | Sandboxed Tool Gateway | Traditional IAM and Network Controls |
|---|---|---|---|---|
| Primary control | Explicit authorization and contextual decisions | Process, file, and network visibility or enforcement | Constrained execution of model-selected actions | Identity, secrets, service access, and segmentation |
| Deployment focus | Application or sidecar decision layer | Hosts, containers, VMs, or kernels | Agent runtime and connected tool services | Cloud, platform, and enterprise infrastructure |
| Strength | Fast policy iteration and auditability | Broad, low-level telemetry | Direct control over tools and data flows | Mature governance and credential management |
| Main limitation | Cannot infer business intent alone | High-volume signals and limited semantic context | Requires well-defined tool contracts | Often static and unaware of agent reasoning chains |
| Typical role | Deny, allow, or require approval | Detect anomalies and enforce selected OS actions | Validate, filter, rate-limit, and record actions | Establish identity and network boundaries |
| Best combination | Policy engine plus gateway and telemetry | Telemetry plus identity and application context | Gateway plus sandbox and egress policy | Foundation beneath all agent components |
Practical Implementation Steps and Test Program
Begin with an inventory of every model, agent definition, tool, credential, dataset, destination, and privileged action. Classify capabilities by confidentiality, integrity, availability, safety, and financial impact, then identify the human or business process accountable for each one. Remove unused tools and dormant credentials before designing enforcement, because every retained capability expands the attack surface. For a new deployment, a reasonable pilot is limited to 5 to 10 low-risk tools, a small set of sandboxed users, and a maximum of 2 weeks of access before a formal security review.
Implement short-lived workload identity and deny-by-default network access next. Place each tool behind a gateway with schema validation, argument limits, response filtering, idempotency controls, and audit identifiers. Run untrusted code separately from control-plane software, use non-root identities, remove package managers from production images where possible, and mount only task-specific data. Secrets should be brokered per operation rather than placed in prompts, environment variables that can be dumped, or general agent memory.
Testing must include unit tests for policy, integration tests for identity and tools, red-team scenarios for indirect prompt injection, and operational tests for alert routing and revocation. Measure both security and engineering outcomes: unauthorized tool calls blocked, percentage of high-impact actions requiring approval, median approval latency, false-positive rate, policy evaluation latency, and mean time to revoke credentials. A target such as 100% of production deployments having an owner, 100% of privileged calls being attributable to a workload identity, and less than 1% of routine low-risk actions requiring manual review are reasonable internal objectives, not universal industry benchmarks.
Common Mistakes and Cost Tradeoffs
The most common mistake is assuming that a system prompt can serve as a security boundary. Prompt rules can improve behavior, but they are advisory content processed by a probabilistic model and should be backed by authorization outside the model. Another mistake is giving one broadly scoped service account to all agents, because this destroys attribution and makes least-privilege review impractical. A third error is adding a security dashboard without an enforcement path; if detections cannot revoke a session, block a tool, quarantine output, or isolate a host, the system is primarily observational.
Teams also underinvest in data minimization. Blocking outbound traffic does not help if every tool returns 10 megabytes of unnecessary context that may contain secrets. Responses should be truncated, structured, classified, and stripped of fields that the next task does not require. Kernel telemetry without retention limits can create privacy, storage, and cost concerns, while automatic human approval for every low-risk read destroys usability and encourages users to bypass the system. Risk-based step-up should be reserved for genuinely consequential actions.
Costs vary by deployment and cannot be responsibly reduced to a universal monthly figure. Open-source components such as OPA, eBPF libraries, Dapr, and sandbox runtimes may be available at no license fee, but engineering, compute, policy maintenance, telemetry storage, incident response, and model consumption still carry cost. A hosted AI platform may add subscription and usage fees while reducing integration effort, whereas on-premises policy and runtime controls can improve control but demand specialized staff. Small teams should first budget for identity, egress filtering, sandboxing, logging, and a few hard rules; larger organizations should add continuous telemetry, policy simulation, dedicated red teaming, and segmented response capabilities.
When to Act and How to Decide the Appropriate Level of Control
Act immediately when an agent can access secrets, execute code, modify production systems, communicate externally, make financial decisions, or handle regulated or personal information. These capabilities turn model mistakes or prompt injection into security incidents, and prompt-only safeguards are particularly weak in this setting. Even a read-only agent deserves basic controls when it can browse untrusted content, process sensitive data, consume paid APIs, or run indefinitely, because those designs remain vulnerable to exfiltration, denial of service, and unintended disclosure.
Lower-risk prototypes do not require the same depth as an autonomous production agent, but they still need an owner, bounded credentials, explicit data retention, tool restrictions, and an off switch. Before launch, require evidence that privileged actions are attributable, denied by default, logged, tested, and revocable within minutes. Revisit the design when tools or models change, autonomy increases, new data sources are connected, or the agent crosses a network or organizational boundary. A material capability expansion should trigger renewed threat modeling just as a significant change to a conventional distributed service would.
A sensible maturity sequence is basic identity and sandboxing, centralized tool mediation and egress policy, contextual authorization and approvals, then anomaly detection and automated response. The final stage should not be automatic for every incident; high-impact response actions need confidence thresholds, rollback mechanisms, and human escalation. Measure containment time, blocked high-risk attempts, approval burden, and confirmed incidents over at least 90 days before claiming that the architecture works. Security maturity comes from tested feedback loops, not from the number of products connected to the agent.