Enterprise RAG: permissions and tests before polished answers
Preserving document permissions in retrieval, handling prompt injection and measuring retrieval separately from the quality of generated answers.

Every question has a user and a set of permitted sources
In an example company assistant, two people ask about the terms of the same deal. One can read a public product sheet; the other can also access an internal commercial proposal. The correct answer may therefore differ. Retrieving relevant documents is not enough: the system needs to know whom it is answering, which sources it may use and when it should acknowledge missing evidence.
RAG combines retrieving context with generating an answer. Evaluate these components separately when designing the system. If the assistant never retrieves the necessary document, a better language model cannot reliably replace the missing information. Even with correct evidence, generation may interpret it incorrectly. The following design is illustrative and does not describe the architecture or measured results of a particular client project.
Carry permissions through to every searchable chunk
When splitting a document into chunks, retain its source ID, version and the metadata needed to make an access decision. OWASP warns that permissions must also be enforced during retrieval. Microsoft's documentation describes identity-based filtering as one implementation option. The server builds these filters from verified identity; the user cannot substitute a tenant ID or group list of their own.
Apply the boundary before returning search results and check again before sending context to the model. A chunk with missing or invalid access metadata is not public by default. A reranker that uses another model to judge relevance must not receive unauthorised text either. This pseudocode simplifies the components and is not a complete implementation. A real system also needs to handle policy failures, permission revisions and earlier conversation content.
principal = authenticate(request)
scope = access_policy.read_scope(principal)
hits = search(
query=request.question,
tenant=scope.tenant,
allowed_documents=scope.document_ids
)
context = access_policy.require_current_access(principal, hits)
ranked_context = rerank_only_authorised(context)
answer = generate(question=request.question, context=ranked_context)
access_policy.require_current_access(principal, ranked_context)
return validate_answer(answer, allowed_sources=ranked_context)Permission changes affect the index, cache and conversation
A document can be indexed correctly and have its access changed later. Until the index receives an update, stale metadata could allow disclosure. For sensitive content, check the current policy against a trusted authority or use a verified permission revision with an agreed freshness requirement. Deny access when the policy is unavailable. An authorisation service outage must not silently widen permissions.
An answer cache has the same problem. Text prepared for a manager does not automatically belong to a colleague asking a similar question. Cache keys and reuse conditions must respect organisation boundaries, access scope and source versions. Revocation must invalidate affected results and account for them in conversation memory. Do not simply put an entire old transcript back into the next request if it contains evidence the user can no longer access.
Documents may carry an attacker's instructions
Even an authorised document may contain text that tries to change the assistant's behaviour. OWASP describes indirect prompt injection through retrieved evidence. Connecting a knowledge base does not remove this threat. Separate application instructions from quoted data and limit sources, context size and available tools. A sentence in a prompt telling the model to ignore outside instructions is not a security boundary.
An assistant designed to answer questions should not gain write access to orders merely because it could propose an update. If actions are added later, the application must check identity, permission and arguments for each operation independently. Input and output filters may catch some attacks, but they do not guarantee resistance to every variation. The design needs to support refusal, human review and a traceable account of which sources influenced an answer.
A test set needs expected evidence as well as questions
Build tasks from the actual information workflow without publishing confidential documents. Each test includes a question, user role, authorised sources, expected evidence and answer criteria. Include cases where the correct document does not exist, has conflicting versions or is inaccessible to the user. Refusal can be the right result for such a test.
The table illustrates a proposed test set. A person familiar with the workflow and documents selects the expected sources. For significant claims, record the supporting passage rather than only a filename. Reserve some tasks for independent evaluation so retrieval changes are not merely tuned to questions developers already know. A document update may require changing its expected result.
| Example case | Expected behaviour | Check |
|---|---|---|
| Ordinary technical question | Answer grounded in a valid version | Content and supporting passage |
| Question about another team | No unauthorised context | Policy and retrieved context |
| Missing evidence | Acknowledge insufficient information | No invented claim |
| Conflicting versions | Apply validity rules or flag the conflict | Versions and dates |
| Instructions inside a document | No permission change or unintended action | Output and tool calls |
Measure retrieval separately from generation
For retrieval, measure how many expected relevant sources appear in the results. If a test expects three specific documents and the system retrieves two, the corresponding document-ID recall is two thirds. Ragas describes ID-based context recall as well as coverage of claims in a reference answer. These are different measurements, so name the unit you use.
Keep the access scope, questions and retrieved chunk limit unchanged when comparing systems. Otherwise the difference may not come from the new retriever. Result count has a cost: more context increases transfer and model work and can introduce distracting text. Report security cases separately. A high average quality score for routine answers must not hide a single case of retrieving another user's data.
A citation does not prove a claim is correct
Check whether specific claims are grounded in the supplied context, whether the answer addresses the question and whether cited sources support the conclusion. An assistant can cite a real document while attributing a statement to it that it never contained. Application code can verify that citation IDs belong to authorised retrieved sources. The relationship between a claim and its evidence requires further evaluation.
An automated judge can help assess a larger test set, but compare its decisions with human assessments on a sample. The Ragas research distinguishes the quality of retrieved context from the final answer; a particular score does not have a universal threshold for practical use. Define errors that should block your workflow and errors that merely reduce usefulness. That distinction matters particularly for approvals or personal information.
A pilot should make failures diagnosable
During a pilot, record the versions of the index, source material, retrieval configuration and generation settings. You can then connect an outcome to the system decisions that produced it. Handle question and answer content according to its sensitivity; an operational log should not automatically copy all company documents. Also track response time, cost and requests where the assistant lacked adequate evidence.
The following checks are proposed tests, not the results of a client security audit. Add a reproducible case to the test set after fixing a failure. Before changing the model, repeat the same tasks with the same data and permissions. This lets you distinguish changes in writing style from changes in factual quality or information disclosure.
- Revoke document access during an in-flight request and check the system's behaviour before it returns the answer.
- Ask the same question as two roles and two organisations; inspect context, citations and cached results.
- Insert an adversarial instruction into an authorised test document and verify that it cannot grant permission or execute an operation.
- Replace a document with a new version, then delete it; check invalidation of its chunks and cached answers.
- Disable the authorisation service; the system must reject unsupported disclosure rather than continue with broader access.
Sources and documentation
For implementation, consult the documentation for the version you use.
Put the topic into practice.
Related project: CBC Slovakia
