esedark
Menu
Back to notes

RAG / business AI / knowledge / governance

RAG for business: permissions, sources and pilot scope

Before commissioning a chat interface for your documents, agree who can read each source, when it stops being valid and which tests will make the pilot worth accepting.

Retrieval-augmented generation, or RAG, searches an approved knowledge collection and gives the relevant passages to a language model before it answers. It does not retrain the model or guarantee truth. It creates a controlled path from a user question to source evidence, which can improve freshness, traceability and domain relevance.

How a RAG system works

An ingestion process reads documents, preserves metadata and permissions, splits text into useful units and creates searchable representations. At query time, retrieval finds candidates, optional reranking selects the strongest evidence and the model answers under explicit instructions. The response should cite its sources and abstain when evidence is insufficient.

When RAG creates business value

Good cases include internal policy assistants, technical support, product documentation, sales enablement and contract or procedure discovery. RAG fits when knowledge changes often, lives across many documents and needs source-level access controls. Start with one expensive search-heavy workflow and a measurable target such as resolution time or answer accuracy.

If answers will reach customers, evaluate resolution and handoff as well as retrieval. These customer service AI pilot acceptance tests show what to check when a policy is missing, an API fails or someone asks to speak to a person.

When not to use it

Do not add RAG to deterministic calculations, structured database reports or processes where a normal search interface works better. It is also a poor starting point when documents are obsolete, ownership is unclear or permissions cannot be represented. For high-impact decisions, use human review and authoritative systems rather than treating generated text as approval.

What makes a source current?

I would start with a small inventory: authoritative source, content owner, approved version, effective date, next review and permitted groups. An indexing timestamp tells you when a document was processed. It does not establish that its contents remain valid. An overdue review should trigger the agreed removal or human review policy.

Consider a hypothetical assistant for support procedures. When a new policy version appears, its owner must decide when it replaces the previous one. An old PDF ranking higher in search should not make that decision. If two approved sources conflict, the system should identify the conflict and send it to the owner.

Decisions I would document before the pilot
ChangeExpected behaviourOwner
New versionActivate the approved version and exclude its superseded chunks from answers about the current procedure.Procedure owner and integration team.
Expired or conflicting sourceStop using it to give current instructions and request review.Content owner.
Deleted documentRemove affected chunks and cached answers; confirm that replaying an old import cannot bring them back.Ingestion and cache maintainers.
Revoked accessDeny subsequent queries and check history, citations and caches that could expose restricted content again.Identity owner and application team.
Stopped connectorExpose the delay and suspend affected answers once the agreed limit is exceeded.Operations owner.

Agree separate tolerances for content updates and permission propagation. If revocation must be immediate, a nightly copy of permissions is insufficient: check current authorization or block access until it can be verified. The Azure AI Search documentation on changes and deletions illustrates why detecting edits and detecting deletions are different mechanisms. Verify your specific connector.

How to verify RAG permissions

The server should obtain the authenticated identity and enforce its permissions before sending passages to the model. Hiding a citation in the interface does not remove text already supplied in context. Asking the model to ignore restricted documents is insufficient. Filters must cover every chunk and tenant, denying access when authorization cannot be verified.

Microsoft's document-level access guide distinguishes filters from identity mechanisms whose availability depends on the source and version; some capabilities are in preview. I would not assume that installing a search engine automatically inherits every permission. I would require tests with test users, group changes and revoked access.

Include cached answers, previous conversations, source links and diagnostic logs in the review. They must respect current access when read or reused. You cannot make someone forget text they have already read, but you can prevent the system from serving it again after revocation. Limit data sent to providers and define retention and deletion for each copy.

Six acceptance tests for a RAG pilot

These are proposed tests for the support example, not results from a delivered project. I would prepare questions and expected evidence before development, including requests the system should refuse.

  1. Superseded version: the same question cites the current policy after the agreed update window; the trace records the version used.
  2. Expired source: without other valid evidence, explain that current information is unavailable and route the question for review.
  3. Deleted document: it disappears from search, model context and answer caches; replaying an old import does not restore it.
  4. Unauthorized user: identical questions from two identities receive only evidence permitted for each; private titles and snippets do not leak.
  5. Revoked permission: repeating a query and opening an earlier conversation do not serve the restricted content again.
  6. Instructions inside a source: a document asking the assistant to bypass permissions or take action remains data, without extending access or activating tools.

Record retrieval of the correct source, evidence supporting the answer and the user's ability to resolve the question separately. A good average answer score does not offset an information leak. An authorization failure should block expansion until it is fixed and affected tests pass again. For customer outcomes and human handoff, use the support tests linked above.

What to include in the budget and maintenance scope

Separate source preparation, identity integration, ingestion, evaluation and the initial application from recurring search, generation, reprocessing, document review and incident handling. File count alone does not establish effort: formats, change frequency, permissions and questions matter. Ask which tasks your team will own and what happens when a connector changes.

The guide to comparing software quotes helps turn this into deliverables and exclusions. Adslyfy shows execution and failure evidence in ad monitoring: a reference for operational visibility, not a RAG case study or an accuracy promise.

Common mistakes

  • indexing every file without an owner or quality filter
  • ignoring document permissions and tenant boundaries
  • splitting content with one arbitrary chunk size
  • testing only polished demo questions
  • accepting fluent answers without source support
  • using a larger model to hide poor retrieval
  • never refreshing, deleting or versioning content
  • launching without feedback, evaluation and fallback paths

Practical RAG checklist

  • choose one workflow and baseline its current cost
  • inventory sources, owners, sensitivity and update frequency
  • preserve metadata, permissions and document versions
  • build representative questions with expected evidence
  • measure retrieval separately from answer quality
  • require citations and an insufficient-evidence response
  • test leakage, prompt injection and cross-tenant access
  • monitor latency, cost, failed queries and user feedback
  • define refresh and deletion service levels
  • keep human review for consequential decisions

When hiring a technical person makes sense

Hire an AI engineer or technical lead when data crosses multiple systems, access rules matter, evaluation is unclear or the prototype must become a supported product. The work is mostly integration, retrieval quality, security and operations—not merely selecting a model. See how to evaluate an internal assistant’s answers and citations and when agents are appropriate.

Scope the first delivery

If you need answers grounded in internal documentation, I can review whether RAG fits through my automation and applied AI services. To discuss a pilot, bring an anonymized document sample, the people allowed to read it, a useful question and a request that should be refused. We can use those to define scope, ownership and acceptance criteria.