Access Revoked. Why Can Your RAG App Still Answer?
At 09:00, a user asks an internal assistant about Project Cedar. The retriever checks access, finds the document, and the model produces an accurate answer. At 09:05, the user loses access to that project. At 09:06, th
At 09:00, a user asks an internal assistant about Project Cedar. The retriever checks access, finds the document, and the model produces an accurate answer.
At 09:05, the user loses access to that project.
At 09:06, the same question returns the same answer. The request never reaches the retriever: the answer cache handles it.
Every retrieval check can be correct while the application still serves restricted information.
This is a hypothetical design scenario, not a report of a production incident. It exposes a question worth asking during a RAG review: which copies of source information remain usable after permission changes?
Follow the information past retrieval
A document can leave several descendants:
Source document
-> indexed chunks
-> retrieved context
-> generated answer
-> answer cache
-> conversation history or summary
-> exported report
A permission check on the first retrieval does not automatically protect those descendants.
A citation is also an incomplete dependency record. A generated answer may use a chunk without citing it. Removing the link later does not remove what the answer learned from that chunk.
For application-controlled artifacts, record the sources supplied to generation, including context inherited from earlier turns. Treat the full supplied set as dependencies unless you have a sound way to establish a smaller set. This is conservative: it may invalidate some answers unnecessarily, but avoids treating a model's citation choices as an access-control decision.
Specify what revocation means
Before selecting a cache TTL, decide the product contract.
Does removing a user from a project stop new retrieval only? Does it also stop replay of stored answers and reuse of old conversation context? What happens to requests already running?
For this example, assume the requirement is: once revocation becomes effective in the application's authorization system, subsequent requests must not serve or reuse protected content from that project.
Previously downloaded files and information already seen by the user cannot be recalled. Historical conversations may have a separate retention and access policy. State that policy explicitly.
For in-flight generation, choose and document a rule: authorize against a request-start snapshot, or require authorization again before delivery. A final check reduces exposure but still leaves a check-to-delivery race unless authorization changes and delivery are coordinated. Do not promise instantaneous revocation without implementing the consistency needed to support it.
Treat a cache hit as another read
OWASP recommends validating authorization on every request and denying access by default. Those principles apply when a cache supplies the response too. OWASP Authorization Cheat Sheet.
A useful answer-cache record contains:
| Field | Purpose |
|---|---|
| Tenant and authorization scope | Prevent accidental reuse across security boundaries |
| Query and generation configuration | Identify the requested computation |
| Source IDs and content revisions | Track which inputs produced the answer |
| Permission revision or epoch | Detect relevant authorization changes |
| Expiration | Bound retention and ordinary staleness |
A user ID in the key separates users, but the same user's permissions can change. A permission epoch helps only if every relevant policy, membership, and document-ACL change advances it, and the request obtains a sufficiently current value. Hashing an old group list does not make it current.
Here is control-flow pseudocode, not a complete security implementation:
authenticate request and establish tenant
find candidate cached answer within that tenant
if candidate exists:
check its complete source dependencies
against current authorization
if dependencies are known and access is allowed:
serve under the chosen delivery policy
otherwise:
do not serve the cached body
retrieve currently authorized context
generate a new answer
record all input dependencies
apply the chosen delivery policy
An unavailable authorization service must not silently turn a denied or unknown decision into permission to serve an old answer. A safe outcome may be an explicit temporary failure. A system can also answer from independently authorized public content, provided it does not reuse the restricted context.
The index has its own clock
Refreshing the answer cache does not fix stale permissions in the index.
Azure AI Search's query-time ACL/RBAC documentation, currently marked preview, explains that ACL freshness depends on ingestion. Its documented update paths include reingestion for custom push ingestion and resynchronization for particular indexer scenarios. Microsoft: Query-time ACL and RBAC enforcement.
That creates separate delays to measure: identity/group resolution, permission propagation into the index, authorization-cache freshness, and invalidation of derived answers.
A five-minute answer TTL does not establish a five-minute revocation bound if the authorization data used to regenerate that answer remains stale for longer.
For stricter requirements, evaluate whether candidate results need a check against an authoritative permission service before their content reaches the model. The right design depends on the source's consistency guarantees and the latency budget.
Test the warm path and the resumed conversation
Use synthetic documents with a distinctive marker, two test users, and controlled permission changes. Make the authorization state observable so the test knows when revocation is effective.
| Experiment | Expected result under this example's contract |
|---|---|
| Warm the answer cache, revoke access, repeat the question | Cached protected body is not served |
| Ask a paraphrase that hits a semantic cache | Same authorization check applies |
| Revoke access, then resume an old conversation | Restricted chunks and derived summaries are not reused as context |
| Keep the search index deliberately stale | Unauthorized context does not reach generation |
| Make authorization unavailable | No fallback to the previously allowed answer |
| Leave a control user authorized | The control user can still retrieve and answer |
Inspect the retrieved context and outgoing model payload, not just the final response. A model declining to repeat the marker does not prove it never received the protected text. Check stored summaries and citation metadata too.
For a revoked source inside a mixed answer, dropping one citation is insufficient. Regenerate from permitted inputs or withhold the derived artifact. Without trustworthy dependency tracking, rebuilding the conversation context may be safer than attempting selective cleanup.
The architecture review question is concrete: after a permission change, which component stops the next cached answer, follow-up turn, and export from using the old access decision?
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.