# How Should Organizations Perform RAG Threat Modeling in 2026?

Savannah Jenkins · October 2, 2026

> What RAG Threat Modeling Actually Means RAG threat modeling is the structured analysis of how an attacker could abuse a retrieval-augmented generation...

## What RAG Threat Modeling Actually Means

RAG threat modeling is the structured analysis of how an attacker could abuse a retrieval-augmented generation system, its data sources, models, prompts, tools, and users. It is not a single penetration test and it is not equivalent to conventional application security. The purpose is to identify assets, trust boundaries, attack paths, likely failure modes, and acceptable controls before production deployment or after a major architecture change. A RAG system should be treated as a distributed application with probabilistic behavior, rather than as an ordinary chatbot attached to a database. The retrieved text can influence instructions, expose private information, manipulate generated answers, or cause downstream actions when an agent is connected to business systems. A useful model therefore asks both conventional security questions, such as “Can an unauthorized user read this document?”, and AI-specific questions, such as “Can retrieved content cause the model to ignore its system policy or invoke a tool?”. The threat model should be documented, tested, owned by a cross-functional team, and revisited whenever the data, model, embedding pipeline, permissions, or agent capabilities change.

**Also worth reading:** [How Do Enterprise Security Teams Handle Agentic AI Threat Modeling in Modern System Architecture?](https://agustin-otegui.com/knowledge/how_do_enterprise_security_teams_handle_agentic_ai_threat_modeling_in_modern_system_architecture.php) · [How Should Organizations Govern Identity and Permissions for Autonomous AI Agents?](https://agustin-otegui.com/knowledge/how_should_organizations_govern_identity_and_permissions_for_autonomous_ai_agents.php) · [What is Agent Permission Architecture in 2026 and how should organizations implement it?](https://agustin-otegui.com/knowledge/what_is_agent_permission_architecture_in_2026_and_how_should_organizations_implement_it.php)

## Why RAG Changes the Security Problem

RAG reduces hallucination in some use cases by grounding responses in retrieved material, but retrieval does not make the system inherently secure. It introduces a new data path: user input is used to select documents, those documents are inserted into a model context, and the model interprets both the user request and the retrieved content. An attacker may control a document, compromise a connected data source, manipulate search rankings, or create misleading content that looks authoritative. Prompt injection is especially important because instructions embedded in retrieved pages can compete with system instructions. RAG and fine-tuning do not eliminate that risk; they alter where instructions may appear and how much context the model can process. A poisoned retrieval result can also produce an indirect attack even when the user never typed a malicious prompt. The security objective is therefore not simply to block known phrases. It is to constrain what can be retrieved, who can retrieve it, how retrieved content is represented to the model, and what the model is permitted to do after generation. A RAG design that retrieves from a customer-uploaded document repository, for example, needs a different threat model from one that searches a small, curated internal knowledge base.

## The Main Assets, Actors, and Trust Boundaries

Start by identifying the assets that matter. These commonly include confidential documents, embeddings, prompts, conversation history, retrieval indexes, model weights, API credentials, tool permissions, generated answers, audit logs, and business actions such as refunds, ticket closures, or code changes. The actors may include ordinary users, malicious users, compromised employees, external document contributors, attackers targeting the embedding database, model providers, plugin developers, and third-party software suppliers. Trust boundaries exist between the user and the application, the application and the retriever, the retriever and the vector store, the vector store and the model, the model and external tools, and the model provider and the organization. Permission boundaries must be mapped separately from semantic boundaries. A system may correctly authenticate a user but still retrieve documents that the user should not see because the search index lacks document-level authorization enforcement. Threat modeling should explicitly test horizontal access, such as one customer seeing another customer’s records, as well as vertical access, such as a normal employee retrieving executive material. The model’s context window is not a security boundary. Neither is an embedding similarity score.

## A Practical Threat-Modeling Method

A workable process begins with an architecture diagram and a data-flow inventory. Record every ingestion source, transformation, storage layer, retrieval mechanism, prompt template, model call, tool, output destination, and administrative interface. For each flow, specify authentication, authorization, tenant isolation, data classification, retention, logging, and failure behavior. Next, define abuse cases and attack paths. Examples include direct prompt injection, indirect prompt injection through retrieved documents, poisoned documents, retrieval denial of service, cross-tenant leakage, sensitive-data exfiltration through citations, unsafe tool use, malicious administrators, insecure plugin access, and denial of service caused by oversized documents or expensive queries. Assign a likelihood and impact using a simple scale, such as 1 to 5, and prioritize paths that combine high impact with realistic access. Then design controls, define measurable tests, and assign an owner and review date. The review should include red-team prompts, adversarial documents, authorization tests, data-flow inspection, and negative tests that verify the system refuses or safely handles malicious content. The output should be a living threat model, not a one-time presentation delivered to executives.

## Controls That Matter Most

The strongest controls reduce the amount of authority available to untrusted content. Begin with source and tenant authorization at retrieval time, rather than relying on the application interface to hide results. Sanitize and classify documents before indexing them, remove active content such as scripts or embedded instructions where possible, and preserve provenance so that the system can identify which document contributed to an answer. Treat retrieved text as untrusted data and place clear boundaries around it, while recognizing that delimiters alone do not reliably stop prompt injection. Limit model tools to the minimum permissions required, use allowlists for destinations and actions, require approval for high-impact operations, and prevent generated text from directly changing security policy. Apply rate limits, query-size limits, retrieval timeouts, token budgets, and cost controls to reduce denial-of-service exposure. Encrypt sensitive data in transit and at rest, protect embeddings and indexes as production data stores, rotate credentials, and log retrieval events, policy decisions, tool calls, and administrative changes. Monitoring should distinguish failed access control from harmless malformed input. A detection rule that flags every unusual prompt will become noisy, while a rule that records a sequence of unauthorized retrieval followed by a sensitive tool call can provide a more actionable signal.

## Comparing RAG Security Approaches

Organizations often compare RAG with fine-tuning, guardrail models, and conventional access control. These techniques can be combined, but they solve different problems. Fine-tuning can improve task behavior and reduce some prompt dependence, yet it does not provide reliable document-level authorization and does not guarantee resistance to prompt injection. A guardrail model or output filter can catch some unsafe generations, but it may miss indirect attacks, leaks, or context-dependent instructions. RAG improves grounding and makes source material available, but the retrieval layer expands the attack surface. A small, curated RAG corpus is generally easier to govern than an open ingestion pipeline, while a large enterprise knowledge base may require stronger filtering, lineage, and monitoring.

| Feature | RAG with controlled retrieval | Fine-tuning or guardrails alone | Conventional application security only |
| --- | --- | --- | --- |
| Fresh knowledge | Can update indexed material without retraining | Knowledge may become stale until retraining | Does not solve model behavior |
| Source traceability | Possible when citations and provenance are designed in | Usually weaker or absent | Tracks requests, not semantic retrieval |
| Prompt-injection defense | Requires data, retrieval, and model controls | Can detect some cases but cannot guarantee prevention | Does not directly address prompt injection |
| Authorization | Must be enforced per document and tenant | Usually cannot replace source access controls | Strong for APIs and databases, incomplete for RAG semantics |
| Operating cost | Indexing, embedding, retrieval, storage, and monitoring | Training, inference, filtering, and evaluation costs | Mature tooling, but limited AI-specific coverage |
| Best use | Knowledge access with controlled source material | Task adaptation or output filtering | Identity, network, API, and data-store protection |

The correct choice is usually layered. Conventional security remains necessary, but RAG-specific controls must be added rather than assumed from an existing web application firewall or identity platform.

## Common Mistakes in RAG Threat Modeling

One common mistake is treating the model as the only component being secured. In many incidents, the weakest point is the retrieval index, a document ingestion service, a misconfigured metadata filter, or a tool credential. Another mistake is testing only direct prompt injection, such as asking the model to ignore its instructions, while ignoring indirect attacks placed in a PDF, web page, support ticket, or shared knowledge-base article. Teams also frequently assume that vector similarity is equivalent to permission. Similarity can rank a document as relevant, but relevance does not establish that the current user may access it. Other errors include exposing retrieved chunks in logs, returning private source text when the model is asked to repeat its context, failing to separate training data from production records, and allowing an agent to execute arbitrary commands after a weak safety check. A useful review challenge is to ask whether every generated claim can be traced to an authorized source and whether every tool action can be explained and reversed. If the answer is no, the system is not ready for the intended use.

## When to Act and How to Prioritize

Threat modeling should begin during design, before a RAG system is connected to customers or internal business actions. It should be repeated before production launch, after changing the model or embedding provider, when adding new data sources, when introducing agentic tools, after an incident or near miss, and at least periodically thereafter. A reasonable initial threshold is to complete a basic model before any external deployment and a detailed, cross-functional threat model before granting write access to systems. Organizations can prioritize by exposure: customer-facing systems, regulated data, multi-tenant repositories, financial or healthcare information, and autonomous actions deserve earlier review than an internal prototype using synthetic documents. For a low-risk prototype with no external ingestion and no tools, a concise model may be sufficient. For a production agent that can issue refunds or modify records, the review should include abuse-case testing, independent authorization checks, human approval for high-impact actions, and incident-response exercises. Cost is driven more by architecture and process than by the threat-modeling document itself. A small team can start with a one-day workshop and a diagram, but remediation may require weeks of engineering work. Costs include security engineering time, red-team tests, monitoring, index isolation, model evaluation, access-management changes, and ongoing review. Expensive controls are justified when the potential impact is high; they are less rational when the system has no sensitive data, no external writers, no tools, and a tightly limited scope.

## What a Mature 2026 Practice Looks Like

A mature program connects threat modeling to evidence. Security teams maintain a current architecture inventory, test suites for direct and indirect prompt injection, tenant-isolation tests, retrieval-quality safety tests, and tool-use authorization tests. Product teams record accepted risks and have a named owner for every high-impact finding. Engineers measure useful outcomes, such as the percentage of cross-tenant retrieval tests that are blocked, the time required to revoke a source, the number of unauthorized tool-call attempts detected, and the proportion of answers with verifiable provenance. These metrics are more meaningful than claiming that the system is “secure” because a model refused one example. Leaders should fund remediation according to business impact and be skeptical of inflated security claims, including the headline figure reported in the research context that 88% of organizations have been hit by AI-agent security incidents. That figure may describe a survey, incident definition, or sample that is not directly comparable across organizations, so it should be cited precisely rather than used as a universal probability. The defensible conclusion is narrower: AI-agent and RAG security incidents are sufficiently common to justify disciplined testing, but each architecture requires its own evidence. The best program treats RAG threat modeling as an engineering discipline that evolves with the system, not as a certification or product purchase that guarantees safety.

## Quick answers

### Is RAG more secure than fine-tuning?

No single method is inherently more secure. RAG can improve source traceability and allow controlled updates, but it adds ingestion, retrieval, authorization, and prompt-injection risks. Fine-tuning and guardrails can help with behavior or output filtering, but they do not replace document-level access control or tool authorization.

### Can prompt injection be completely prevented in RAG?

Complete prevention is not a realistic engineering guarantee today. Teams can reduce risk by treating retrieved content as untrusted, limiting tools and permissions, separating instructions from data, testing direct and indirect attacks, and requiring approval for consequential actions. RAG and fine-tuning alone do not eliminate prompt injection.

### How do you test for cross-tenant data leakage in RAG?

Create documents with distinguishable tenant markers, index them under separate tenants, and query the system as multiple users and roles. Verify that unauthorized content is absent from retrieval results, model context, citations, logs, and generated responses. Test both ordinary queries and adversarial prompts designed to reveal hidden context.

### How much does a RAG threat-modeling engagement cost?

There is no standard price. A small design review may take roughly one day and involve internal staff, while a production assessment involving red-team testing, architecture review, and remediation planning can take weeks. Cost depends primarily on data sensitivity, tenant complexity, model and tool integrations, and the required evidence.

### Does RAG threat modeling apply to agentic AI systems?

Yes, especially when the model can call tools, modify records, send messages, execute code, or retrieve information from external systems. Agentic systems add decision-to-action risks, so the model must be tested together with permissions, approval gates, audit logs, and rollback mechanisms.

Canonical: https://agustin-otegui.com/knowledge/how_should_organizations_perform_rag_threat_modeling_in_2026.php
Markdown: https://agustin-otegui.com/knowledge/how_should_organizations_perform_rag_threat_modeling_in_2026.php/index.md
