# What Security Controls Should an LLM Gateway Have in 2026?

Savannah Jenkins · September 28, 2026

> What Are LLM Gateway Security Controls? LLM gateway security controls are the authentication, authorization, policy, content-safety, data-protection...

## What Are LLM Gateway Security Controls?

LLM gateway security controls are the authentication, authorization, policy, content-safety, data-protection, observability, and cost-management mechanisms placed between AI applications or agents and the models, tools, and data they access. A gateway may expose one OpenAI-compatible API while routing requests to several model providers, but its central security role is stronger than protocol translation: it creates a centrally governed control point for identities, model access, prompts, outputs, tool calls, and spending. The exact control set matters because a compromised gateway can expose API credentials, private prompts, connected databases, and agent actions that already possess access to production systems. Microsoft’s 2026 discussion of AI infrastructure as a target and Wiz’s research on authentication bypass through LiteLLM show why the gateway itself must be treated as a security boundary rather than an ordinary API proxy. An effective control plane should therefore answer four questions for every request: who sent it, what may they access, which models or tools may be used, and what actions or costs are permitted. No single feature is sufficient; the useful unit of protection is a policy decision made before model invocation and, where necessary, again before a tool executes.

**Also worth reading:** [What Are the Best MCP Enterprise Security Controls for Production AI Agents?](https://agustin-otegui.com/knowledge/what_are_the_best_mcp_enterprise_security_controls_for_production_ai_agents.php) · [What are AI agent security controls and how do they protect autonomous systems?](https://agustin-otegui.com/knowledge/what_are_ai_agent_security_controls_and_how_do_they_protect_autonomous_systems.php) · [How Should an AI Architect Design an MCP Gateway Architecture for Enterprise Security and Scale?](https://agustin-otegui.com/knowledge/how_should_an_ai_architect_design_an_mcp_gateway_architecture_for_enterprise_security_and_scale.php)

## Authentication, Authorization, and Tenant Isolation

Every request should reach a verified workload identity, not merely an API key embedded in client code. Strong controls include short-lived credentials, OAuth 2.0 or mutual TLS for service-to-service traffic, role-based and attribute-based authorization, and tenant identifiers derived from trusted identity claims rather than arbitrary request headers. Admin and policy-management endpoints should require phishing-resistant multi-factor authentication, while read-only analytics should be separated from configuration and credential operations. Authorization policies can restrict callers by model, provider, data classification, tool, geographic region, budget, and approved purpose. A customer service agent might be allowed to use model A with a retrieval tool but denied shell execution, direct database write access, or transfer to an unapproved provider. Tenant isolation should be enforced at routing, cache, log, vector-store, and tracing layers; an isolated prompt path that still shares keys or cached results is not actually isolated. Default-deny access is safer than assuming new agents are trustworthy, and emergency policy revisions should expire automatically rather than leaving an emergency exception in production indefinitely.

## Policy Enforcement for Prompts, Responses, and Agent Actions

Policy enforcement should occur before a request is sent to a model and before an agent invokes a tool. Input controls can detect or block secrets, prohibited personal data, prompt-injection patterns, excessive input length, unsupported file types, and requests aimed at changing the system policy. Output controls can scan for credentials, regulated data, disallowed topics, unsafe code, and policy violations, although detectors should be calibrated because false positives can interrupt legitimate work. More important for agentic systems is action-level policy: an agent may be permitted to draft a refund but not issue it, read a ticket but not alter billing, or query a database with SELECT but not DELETE or UPDATE. Tool schemas should use strict parameter validation, approved endpoints, least-privilege service accounts, egress allowlists, and transaction limits. Some gateways also add human approval for consequential actions, but approval prompts are not a complete control unless they identify the exact tool, target, and parameters being approved. Security policy must cover both direct user prompts and instructions inserted through retrieved documents, web pages, emails, or other untrusted content.

## Data Protection, Secret Handling, and Provider Privacy

The gateway should minimize sensitive data before transmission to any external model and should never log credentials by default. Transport encryption with TLS 1.2 or later is the baseline, while providers should be selected according to contractual retention, training-use, regional-processing, and compliance requirements. Tokenization, redaction, data-loss prevention, prompt filtering, and purpose-based data routing can reduce exposure, but each creates an engineering tradeoff. Aggressive redaction can remove fields required for a task, while retaining too much data can turn centralized logging into a secondary data breach. Sensitive prompts and outputs should have configurable retention periods—for example, 0 to 7 days for detailed production traces and 30 to 90 days for aggregate operational metrics—subject to legal and debugging needs. Secrets belong in a dedicated vault or secrets manager, should be rotated at least every 90 days for ordinary credentials and immediately after suspected exposure, and should never be passed through model context when a direct authenticated call is possible. Provider credentials must be isolated by tenant or security domain so a compromised integration cannot automatically expose every downstream account.

## Rate Limits, Quotas, Budgets, and Denial-of-Service Controls

LLM gateways need conventional API defenses as well as AI-specific consumption controls. Per-user, per-tenant, per-model, and per-tool rate limits can prevent one application from exhausting shared capacity, while concurrency caps control the number of simultaneously executing or streaming requests. Token-based limits are more meaningful for LLM use than request counts alone: a short classification request and a 100,000-token retrieval task should not consume the same allowance. A practical policy might permit 60 requests per minute per service identity, 20 concurrent streams, 100,000 input tokens per minute, and a daily department budget expressed in both dollars and model-specific units. Administrators need hard ceilings, soft alerts at 50%, 75%, and 90%, and a kill switch that blocks new calls without disrupting already authorized transactions. These controls mitigate runaway agents, scraping, retry storms, and credential theft, but they do not replace capacity planning or provider quotas. Circuit breakers should also detect abnormal error rates, latency, or model-routing behavior and temporarily remove a failing destination.

## Logging, Monitoring, Detection, and Incident Response

Security monitoring must correlate identity, prompt, model, tool, data-source, policy, latency, token, and cost events instead of recording only a request ID and response. Logs should be tamper-resistant, time-synchronized, access-controlled, and exported to the organization’s security information and event management platform. High-signal detections include impossible travel, token replay, sudden increases in tool calls, access to a previously unused provider, repeated policy denials, unusual document-retrieval patterns, and attempts to alter system prompts or administrative routes. Security teams should set a measurable detection target, such as alerting within 15 minutes for confirmed privileged credential misuse and beginning triage within 60 minutes for high-risk agent actions. Dashboards should break down denied requests, model failures, data-policy blocks, tool errors, and spend by tenant so that a rise in cost is not mistaken for a security incident. Incident playbooks must define how to revoke gateway tokens, disable one tool or tenant, rotate provider keys, stop outbound traffic, preserve evidence, and distinguish a malicious request from a client defect.

## Gateway Options and Comparison

There is no universally superior category. An open-source gateway such as LiteLLM can offer routing flexibility, OpenAI-compatible interfaces, and cost controls, but deployment, hardening, upgrades, and monitoring remain the adopting organization’s responsibility. A commercial enterprise AI gateway may provide managed identity, policy integration, support, and preconfigured data controls, usually at the cost of vendor dependence and potentially higher per-token or platform pricing. A cloud-provider gateway can simplify private connectivity and billing inside one ecosystem, but portability and policy consistency may be weaker. A custom gateway offers exact control, yet its maintenance burden is high and the resulting security bugs remain the customer’s problem. Security teams should score gateways on control depth, auditability, failure behavior, and total cost rather than accepting a feature checklist at face value.

| Feature | Open-source gateway | Commercial enterprise gateway | Direct model API |
| --- | --- | --- | --- |
| Core control point | Self-managed routing and policy layer | Managed identity, policy, monitoring, and support | Provider-specific endpoint and quota controls |
| Data and key ownership | Customer controls infrastructure and secrets | Shared responsibility governed by contract and configuration | Customer sends data directly to the selected provider |
| Typical cost | Software may be free; infrastructure and engineering are not | Platform, usage, support, and possible enterprise minimums | Usually pay-per-token with no separate gateway platform |
| Multi-provider routing | Often broad and customizable | Commonly available, depending on product tier | Requires separate integrations and controls |
| Best security fit | Regulated teams able to operate Kubernetes or cloud infrastructure | Organizations wanting managed controls and faster deployment | Low-volume, low-risk workloads with simple requirements |
| Main weakness | Misconfiguration and unsupported operations create exposure | Lock-in, contract limits, and premium pricing | No centralized cross-provider security or policy layer |

The comparison is intentionally categorical rather than naming vendors as universally safe. Open-source and commercial systems can both be secure or insecure, depending on version, configuration, integration, and operating discipline. Direct provider APIs are not automatically insecure, but they force the customer to reproduce authentication, data filtering, spend controls, logging, and incident response across every model and region. A sensible architecture may place only low-risk internal traffic directly on provider APIs while routing sensitive or agentic workloads through a hardened gateway. For a small workload, the gateway’s operating cost may exceed the model cost; for a production platform, that cost is buying a single policy boundary across many applications.

## Deployment Practices, Timing, and Cost Tradeoffs

A gateway should be introduced before production AI use when multiple applications share provider keys, users can access sensitive data, agents can call tools, or third-party prompts reach retrieval systems. Organizations with one experimental assistant, no persistent data, and no privileged tools may begin with provider-native controls, but should set a review date and trigger—often within 30 to 90 days—for gateway adoption when usage or autonomy grows. Implementation normally takes four stages: inventory models and tools, define identities and policy, test deny and revoke behavior, then route production traffic gradually. Start with 5% of traffic, expand to 25% and 50%, and reach 100% only after error rate, latency, policy decisions, and cost remain within agreed thresholds. A reasonable initial policy objective is to block 100% of unauthenticated production calls and 100% of direct long-lived key access, while reviewing false-positive rates weekly during rollout. Costs vary by platform and contract, so avoid invented universal prices; budget instead for gateway compute, logs, vector and security telemetry, model usage, engineering labor, and ongoing testing. The cheapest option is not the one with the lowest license fee, but the one whose security and operating costs are predictable and proportionate to the harm it prevents.

## Quick answers

### Is an LLM gateway the same thing as an API gateway?

No. A conventional API gateway handles transport-level tasks such as authentication, rate limiting, routing, and observability. An LLM gateway adds model-specific controls such as token and spend budgets, prompt policies, provider routing, tool authorization, content inspection, and agent action governance.

### Which LLM gateway security control is most important?

Strong workload identity with least-privilege authorization is the foundation. Without it, an attacker may impersonate an application and reach models, data, or tools before content filtering has anything useful to evaluate.

### Can prompt filtering alone protect an agentic AI system?

No. Prompt filtering can reduce obvious abuse, but indirect prompt injection may arrive through web pages, documents, email, or tool results. Agent systems also require strict tool permissions, parameter validation, network restrictions, and approval for consequential actions.

### How much should an LLM gateway cost?

There is no standard price because gateways may be open source, self-hosted, cloud-hosted, or priced per request, token, seat, or enterprise contract. Compare total operating cost, including engineering, infrastructure, logging, support, and model usage, rather than comparing license fees alone.

### When should a company replace direct provider access with a gateway?

Adopt a gateway when multiple applications or agents share credentials, sensitive data crosses model providers, or costs and permissions need centralized governance. For a low-risk prototype, direct access can be reasonable, provided the organization documents limits and has a defined migration trigger.

Canonical: https://agustin-otegui.com/knowledge/what_security_controls_should_an_llm_gateway_have_in_2026.php
Markdown: https://agustin-otegui.com/knowledge/what_security_controls_should_an_llm_gateway_have_in_2026.php/index.md
