01. The Real Threat Model
The failure modes that matter in practice are specific and recurring: a chatbot answering from documents the current user shouldn't see, a crafted instruction hidden inside retrieved content changing its behavior, or a confidently wrong answer being treated as a binding commitment by the person reading it.
02. Specific Risks
- Retrieval permission gaps: content scoped incorrectly so an authenticated user can retrieve information from a different account or role than their own.
- Prompt injection via retrieved content: crafted text inside a document or ticket designed to override the system's intended behavior once retrieved and included in context.
- Hallucinated commitments: a fluent, wrong answer about a policy, refund, or contract term that a customer reasonably treats as authoritative.
- Transcript data exposure: conversation logs storing more sensitive content than the retention and access policy accounts for.
- Indirect data exfiltration: a user prompting the bot to summarize or repeat back retrieved content it was only meant to reference internally, effectively using the chatbot as an unintended export path for documents it can see.
Retrieval permission gaps are the most common finding in practice, and the most misunderstood. Teams frequently scope retrieval at the index level, one index per customer or tier, and assume that settles it. It doesn't cover role changes, offboarded accounts still holding valid sessions, or a support agent's elevated view being reused by a customer-facing bot instance that shares the same retrieval index. Permission scoping has to be checked at query time against the current identity, not assumed from how the index was originally partitioned.
03. Defenses That Work
- Scope retrieval to the authenticated identity's actual permissions at query time, not as a post-response filter.
- Treat all retrieved content and user input as untrusted, the same posture used for any external input to a production system.
- Require explicit confirmation, and human approval above a threshold, for any binding commitment or irreversible action.
- Apply the same data retention and access policy to transcripts as to the source systems they reference.
- Attach citations to every substantive answer, making it possible to verify what the answer was actually grounded in.
04. What to Ask a Vendor
- How is retrieval scoped to the current user's actual permissions?
- How is retrieved content treated, trusted or untrusted, when constructing a response?
- Which actions require explicit human approval regardless of answer confidence?
- What's the retention and access policy for conversation transcripts?
See our AI chatbot guide for how these controls fit into the wider retrieval and synthesis pipeline.
05. Frequently Asked
Can a chatbot be tricked into revealing information it shouldn't?
Yes, if retrieval isn't scoped to the authenticated user's actual permissions. The chatbot should only be able to retrieve content the logged-in identity is entitled to see, checked at retrieval time, not just filtered from the response afterward.
What is prompt injection through retrieved content?
If a document in the knowledge base contains text crafted to look like an instruction, such as a support ticket with embedded text saying 'ignore previous instructions', a chatbot that treats retrieved content as trusted can be manipulated into unintended behavior. Treating all retrieved content as untrusted input is the defense.
Can a chatbot make a binding commitment it shouldn't?
Not if designed correctly. Refund approvals, contract terms, and other binding commitments should require explicit confirmation and, for anything above a threshold, human approval, regardless of how confident the synthesized answer sounds.
Does encrypting transcripts at rest cover the security requirement?
It covers one requirement, not the whole picture. Encryption protects data if storage is breached, but it does nothing about a retrieval query returning content from the wrong account, or a transcript retained longer than the source system it references. Both are more common causes of exposure than a storage-layer breach.
Cloudz Computing scopes retrieval to authenticated permissions and treats every retrieved passage as untrusted input by default.
Explore the AI Chatbots solution →
Request a private consultation