# How Should Enterprises Control AI Agents Without Slowing Down Innovation?

Savannah Jenkins · October 2, 2026

> The Direct Answer Enterprises should control AI agents through a risk-based governance system that connects identity, authorization, model oversight...

## The Direct Answer

Enterprises should control AI agents through a risk-based governance system that connects identity, authorization, model oversight, tool access, runtime monitoring, human approval, and evidence collection. The objective is not to prevent agents from acting, but to define which actions each agent may take, under what conditions, with what data, and subject to which stop conditions. Governance works best when it is embedded in the platforms through which agents are built and deployed, rather than added later as a separate compliance review.

**Also worth reading:** [What Is an Agent Security Control Plane in 2026, and How Should Enterprises Choose One?](https://agustin-otegui.com/knowledge/what_is_an_agent_security_control_plane_in_2026_and_how_should_enterprises_choose_one.php) · [How Can Enterprises Cut Hybrid LLM Costs Without Sacrificing Reliability in 2026?](https://agustin-otegui.com/knowledge/how_can_enterprises_cut_hybrid_llm_costs_without_sacrificing_reliability_in_2026.php) · [How Should Enterprises Design Sovereign AI Architecture for Control, Resilience, and Scale in 2026?](https://agustin-otegui.com/knowledge/how_should_enterprises_design_sovereign_ai_architecture_for_control_resilience_and_scale_in_2026.php)

A useful starting threshold is simple: if an agent can change a customer account, move money, modify production infrastructure, send external communications, or access sensitive personal information, it should be treated as a privileged automation component. Low-risk work, such as summarizing a public document or drafting an internal response, can normally proceed with lighter controls. Medium-risk work should use restricted data, constrained tools, and post-action review; high-risk work should require explicit human approval before execution.

There is no single enterprise product that provides all of these functions. Some organizations use an agent control plane, others use API management, identity governance, policy-as-code, model operations, and security observability together. The architecture should be modular, but ownership must be clear. Without a named control owner, even a technically capable platform will produce inconsistent decisions, inaccessible audit records, and unclear accountability.

## Why Traditional AI Governance Is Not Enough

Conventional AI governance usually concentrates on model development: training data, evaluation results, bias testing, documentation, and approval for release. Agents introduce a different operational problem. A relatively capable model may produce an acceptable answer while invoking the wrong system, acting through an over-privileged identity, or continuing a task after conditions have changed. The risk therefore sits in the agent’s combination of model, instructions, memory, tools, credentials, and environment.

The distinction can be expressed as a chain of authority. The model proposes an action; the orchestration layer interprets the plan; a tool adapter sends a request; an API or infrastructure system executes it; and a monitoring system evaluates the result. Governance must inspect every link, not merely screen the prompt or output. For example, content filtering cannot determine whether a sales agent is allowed to issue a $250,000 discount, and model accuracy cannot prove that a procurement agent used an approved supplier.

This is why vendors described in 2025 and 2026 research are placing agent controls closer to infrastructure. Reports from CIO Dive, the Boston Consulting Group, Deloitte, Bain, IBM, NVIDIA, Red Hat, SAP, and other organizations frame enterprise agents as operational actors requiring orchestration, observability, and policy enforcement. Their proposed solutions differ, but the common requirement is consistent: a central policy model must influence real actions. Governance that produces a report after an incident has occurred is useful for accountability, but it does not prevent the incident itself.

## A Reference Architecture for Governed Agent Operations

The first layer is an inventory that identifies every agent, its business owner, technical owner, purpose, model, data sources, tools, users, and current autonomy level. A practical inventory does not need hundreds of fields. It does need enough information to answer who created the agent, who can alter its instructions, which identities it uses, and what happens when it fails. Each record should also carry a revision number so that governance evidence corresponds to the version that actually ran.

The second layer is identity. Agents should receive dedicated, non-human identities rather than reuse an employee’s broad access token. These identities need short-lived credentials, restricted roles, and rotation intervals tied to risk. A research assistant that reads approved public sources should not inherit the same permissions as an operations agent that can restart production services. Where agents collaborate, service-to-service identities should establish which agent is acting and which agent requested the action.

The third layer is a policy decision point placed between the agent and its tools. Policies can evaluate user, agent role, data classification, environment, action type, transaction value, location, time, and confidence. For example, a policy might permit a procurement agent to draft a purchase order below $10,000, require finance approval from $10,000 to $100,000, and block autonomous execution above $100,000. Thresholds should reflect the organization’s actual loss exposure rather than copied defaults.

The fourth layer is runtime telemetry. The system should record prompts or approved prompt references, model and tool versions, policy decisions, retrieved data, tool calls, outputs, approvals, and state changes. Logs should protect secrets and sensitive content while preserving enough evidence to reconstruct an action. High-volume telemetry may require aggregation, but security events, denied requests, privileged actions, and changes to an agent’s configuration should be retained in searchable form.

## Policy and Approval Models Compared

Organizations commonly choose among full autonomy, human approval, and graduated or risk-tiered autonomy. Full autonomy offers speed but concentrates risk in the agent’s ability to contain errors. Mandatory approval improves human oversight, although excessive review can make agents slower than conventional software and may train reviewers to approve without reading. A risk-tiered model is usually more practical because it reserves scarce human attention for consequential decisions.

| Feature | Full autonomy | Approval for every action | Risk-tiered autonomy |
| --- | --- | --- | --- |
| Best suited work | Low-impact, easily reversible tasks | Regulated or highly sensitive workflows | Mixed portfolios of agents and tools |
| Decision speed | Highest after deployment | Lowest | Moderate to high |
| Human role | Exception handling | Approves every action | Approves defined high-risk thresholds |
| Main weakness | Errors can scale rapidly | Review fatigue and excessive latency | More complex policy design |
| Evidence requirement | Continuous monitoring | Approval record and action trace | Tier assignment, policy version, and event trace |
| Typical cost profile | High expected loss or high monitoring cost | Highest labor cost | Balanced operating and control cost |

Human-in-the-loop approval should be meaningful. A reviewer needs the intended action, affected records, expected outcome, relevant evidence, and a concise reason for escalation. Asking a person to approve an unexplained plan shifts the governance problem rather than solving it. Automation can also sample low-risk actions, route unusual ones for review, and immediately escalate attempts to bypass policy.
An agent may need deterministic rules even when the underlying task uses machine learning. A claims agent can be creative in summarizing a case, but regulatory and payment rules should be encoded as explicit constraints. Combining probabilistic models with deterministic policy engines provides a clearer separation: the model handles uncertainty in language and planning, while the policy layer handles enforceable limits. Neither component is sufficient alone.

## Practical Steps for a 90-Day Implementation

The first 30 days should establish visibility and ownership. Identify the 10 to 20 agents with the greatest access, business value, or regulatory exposure, and assign an accountable owner to each. Review their models, prompts, tools, credentials, logs, and autonomy. Remove unused credentials, disable dormant agents, and document every production action the agent can take without human intervention. The result should be a ranked inventory, not an aspirational policy statement.

Days 31 through 60 should establish enforceable controls. Create dedicated identities, replace broad API keys with short-lived credentials, and introduce a policy decision point for sensitive tools. Define at least three action tiers and set explicit thresholds for data access, financial transactions, customer communication, infrastructure changes, and record modification. Test denied actions as well as approved ones; a control that has never rejected unauthorized work has not been demonstrated.

Days 61 through 90 should prove the system in a controlled environment. Select one workflow, such as vendor due diligence or customer-support resolution, and compare governed and manual performance. Measure completion time, human-review rate, policy denials, false approvals, rollback frequency, and total cost per completed case. Establish incident procedures that can disable an agent, revoke its credentials, stop active tasks, preserve evidence, and notify the responsible owner. After 90 days, expand only when the measured controls work and the organization accepts the operating cost.

A useful pilot target is not a universal percentage, but a concrete service objective. For example, an organization might require a 95% audit-sampling rate for privileged actions, a 100% approval rate for transactions above $50,000, and credential rotation every 24 hours. It might also target a policy-evaluation latency below 100 milliseconds for routine calls, with a slower path for complex approvals. Such targets should be tested against the workload rather than presented as industry benchmarks.

## Common Mistakes That Produce False Assurance

The first mistake is treating a governance dashboard as a control. Dashboards can reveal configuration and activity, but they do not stop an unauthorized tool call unless a policy engine can block it. The second is assuming that a sandbox makes an agent safe. Sandboxes reduce blast radius, yet they may still expose secrets, create excessive resource consumption, or connect to weakly governed internal services. The third is allowing agents to share the credentials of the developer who built them.

Another common error is measuring only model quality. An agent can have a high task-success rate and still create unacceptable risk through excessive permissions. Evaluation should include unauthorized-action attempts, policy bypass, prompt injection, data leakage, incorrect tool selection, approval fatigue, and recovery behavior. It should also test whether the agent follows a denial decision rather than repeatedly trying another route.

Organizations also err by reviewing every prompt and action manually. This approach does not scale and encourages superficial approval. Conversely, removing people from every decision is unsafe for high-impact actions. The better design is to automate evidence collection, use deterministic rules for repeatable decisions, and reserve human judgment for uncertain or consequential cases.

Finally, governance cannot be static. An agent’s behavior can change when its model, instructions, memory, tool schemas, data sources, or permissions change. A production configuration should have a version, owner, test result, and approval record. If any material component changes, the organization should know whether the existing evaluation remains valid or must be rerun.

## Cost, Pricing, and Build-versus-Buy Decisions

The largest cost is frequently operational rather than the license fee. An enterprise may need an identity system, API gateway, policy engine, logging platform, data catalog, evaluation service, incident process, and staff with both security and AI expertise. Training reviewers, storing detailed traces, integrating systems, and maintaining policies can exceed the subscription price of an agent platform. Cost analysis should therefore include people, infrastructure, integration, evaluation, and expected incident loss.

Open-source projects can reduce license expense, but their use does not make governance free. The organization remains responsible for hosting, upgrades, security patches, policy development, integrations, and 24-hour operations. Some research projects described in 2026 use Python libraries or OPA-style policy enforcement, while vendors are packaging control planes and runtime governance into commercial platforms. These are useful options, but no advertised product should be treated as complete without a technical and procurement review.

A build-versus-buy decision should consider control ownership and technical fit. Buy when the provider already supports the required identities, tools, regions, retention rules, audit exports, and service levels, and when switching costs are acceptable. Build when the workflow has unique controls, the organization has mature platform engineering, and the business value justifies ongoing maintenance. A hybrid approach is common: use an existing identity provider and API gateway, add a dedicated agent policy layer, and select an observability product for runtime traces.

Pricing should be evaluated per workflow, agent, user, tool call, environment, or governed transaction depending on the vendor. A per-seat price may look inexpensive while leaving machine identities and API usage uncounted. A per-action model can become costly for high-volume agents. Ask for annual and three-year costs, overage rules, support tiers, data-retention charges, and the cost of adding non-human identities. Do not infer a specific market price from open-source availability; costs vary widely by deployment and integration requirements.

## When to Act and Who Should Lead

An organization should act before deploying an agent in production, but it does not need to build a comprehensive control plane before the first experiment. The immediate priority is proportionate containment: dedicated credentials, restricted tools, logging, a human owner, and a shutdown path. A customer-facing agent that can alter account settings requires more formal review than an internal summarization prototype, and an agent that can execute financial transactions requires the strongest review of all.

The program should be jointly led by an AI architecture function and security or risk leadership, with business-process owners accountable for outcomes. Identity, platform engineering, data governance, legal, compliance, and internal audit contribute controls, but assigning the program solely to a model-development team will leave operational gaps. Procurement should be involved early because vendor claims about identity, auditability, model support, and data handling must be tested against enterprise requirements.

By October 2026, the market is moving toward open control planes, API-layer controls, infrastructure-level governance, and integrated runtime monitoring. That direction is sensible, but market terminology should not obscure a basic question: can the enterprise demonstrably stop an unauthorized action before it causes harm? If the answer is yes, and the evidence can be produced later for investigation, governance is becoming operational. If the answer is no, the organization may have visibility without control, or a control plane without enforceable policy.

## Quick answers

### What is the fastest way to improve AI agent governance?

Start with an inventory of production agents, dedicated non-human identities, restricted tool permissions, centralized logs, and a tested shutdown procedure. Prioritize agents that can move money, change customer records, access sensitive data, or modify infrastructure. A full control plane can follow, but basic containment should be in place before expanding autonomy.

### Do open-source agent control planes eliminate governance costs?

No. Open-source software can reduce license fees, but hosting, integration, upgrades, policy engineering, evaluation, security operations, and skilled staff still create substantial cost. The best choice depends on the organization’s technical maturity, required integrations, and ability to maintain the platform.

### How should an enterprise decide which agent actions require human approval?

Use thresholds based on financial value, reversibility, data sensitivity, regulatory exposure, affected population, and potential operational impact. For example, an agent might act below $10,000, require approval from $10,000 to $100,000, and be blocked above $100,000. Organizations should adjust those figures to their actual loss exposure and controls.

### Can API gateways provide complete AI agent governance?

API gateways are important enforcement points because agents commonly act through APIs. They can apply authentication, authorization, rate limits, and some policy decisions, but they do not automatically understand the agent’s business purpose, retrieved context, model behavior, memory, or downstream effects. Complete governance usually combines gateways with identity, policy, observability, evaluation, and human accountability.

### What is the best first production use case for governed agents?

Choose a workflow with measurable value, bounded tools, recoverable actions, and clear owners. Vendor research, internal document processing, or customer-support resolution can work when permissions are limited. Avoid beginning with irreversible actions at large scale, because failures are harder to detect and correct.

Canonical: https://agustin-otegui.com/knowledge/how_should_enterprises_control_ai_agents_without_slowing_down_innovation.php
Markdown: https://agustin-otegui.com/knowledge/how_should_enterprises_control_ai_agents_without_slowing_down_innovation.php/index.md
