Pidoku

An Enterprise RAG Assistant

Expert 1h Difficulty 4/5 Lesson 02 of 03

Prerequisites Context and Retrieval, The Data Layer, Trust Boundaries in an AI System

The idea in one minute#

The system: an assistant that answers employees’ questions from 20 million internal documents — wikis, tickets, shared drives, chat — and cites its sources. It looks like a search box with a model on the end. The design work is almost entirely in two places that have nothing to do with the model: permissions (a user must never receive an answer built from a document they cannot open) and freshness and quality of the index (the model can only be as right as what is retrieved). The model is the easy part.

Step 1 — requirements#

RequirementDecision
Users30,000 employees; 3,000 active at peak
Corpus20 M documents, 6 source systems, ~150 M chunks; 2% change per day
Unit of workOne question answered with citations
Quality≥ 90% of answers judged correct and supported; “I could not find this” when unsupported
LatencyFirst token under 2 s; full answer under 12 s
PermissionsSource-system access controls enforced exactly, including group changes within 15 minutes
FreshnessA changed document searchable within 15 minutes
DataEverything stays in the company’s cloud; hosted model allowed under zero retention
DeletionA deleted document is unsearchable within 15 minutes

Step 2 — numbers#

query path
  questions at peak        3,000 users × 0.4 / min ÷ 60        = 20 / s
  model calls per question rewrite (small) + answer (large)     = 2
  answer call              10,000 in (3,500 cached), 450 out
  retrieval budget         rewrite 300 ms + search 80 ms + rerank 150 ms = ~0.5 s before the answer call

index
  vectors                  150 M × 1,024 dims × 4 bytes         = 614 GB raw
  with scalar quantization (1 byte/dim) + graph index           ≈ 200 GB in memory, originals on disk for rescoring
  embedding throughput     2% of 150 M = 3 M chunks/day         ≈ 35 chunks/s average, bursts of 2,000/s on bulk edits
  initial build            150 M chunks at 3,000 chunks/s per GPU ≈ 14 GPU-hours

The query side is modest: twenty questions a second. The index is the heavy component — hundreds of gigabytes in memory, kept fresh continuously, with permissions attached to every chunk.

Step 3 — architecture#

flowchart TB
  subgraph INGEST["Ingestion: continuous"]
    direction LR
    SRC[":confluence: <b>Sources</b><br/><small>wiki, tickets, drive, chat</small>"] --> CONN[":i-route: <b>Connectors</b><br/><small>change feeds + ACLs</small>"]
    CONN --> BUS[":apachekafka: <b>Change stream</b>"]
    BUS --> PARSE[":i-funnel: <b>Parse, chunk, scan</b>"]
    PARSE --> EMB[":vllm: <b>Embedding model</b>"]
    EMB --> IDX[(":qdrant: <b>Hybrid index</b><br/><small>vector + keyword + ACL</small>")]
    CONN --> ACL[(":postgresql: <b>Identity and ACL store</b><br/><small>users, groups, doc permissions</small>")]
  end
  subgraph QUERY["Query: per question"]
    direction LR
    U[":i-user: <b>Employee</b>"] --> GW[":envoyproxy: <b>AI gateway</b>"]
    GW --> ORC[":i-workflow: <b>Answer service</b>"]
    ORC --> RW[":vllm: <b>Rewrite</b><br/><small>small model</small>"]
    RW --> SR[":i-search: <b>Search</b><br/><small>filter: user's groups</small>"]
    SR --> RR[":vllm: <b>Rerank</b>"]
    RR --> ANS[":anthropic: <b>Answer model</b><br/><small>cite or decline</small>"]
    ANS --> CHK[":i-shield-check: <b>Citation check,<br/>output sanitiser</b>"]
  end
  IDX --> SR
  ACL --> SR
  CHK --> U
  ORC -.-> TEL[":opentelemetry: <b>Traces, feedback</b>"] --> EVAL[":i-scale: <b>Evaluation</b>"]
  class SRC,U neutral
  class CONN,BUS,GW,SR queue
  class PARSE,CHK warn
  class EMB,RW,RR,ANS,ORC compute
  class IDX,ACL memory
  class TEL,EVAL neutral

Ingestion#

  • Connectors read each source’s change feed — created, updated, deleted, permissions changed — rather than re-crawling. They emit the document and its access-control list.
  • Parse and chunk by structure, attaching the heading path, source URL, author, updated time and ACL to every chunk. Scan for secrets and for instruction-like text; quarantine matches for review rather than indexing them.
  • Embed with a self-hosted embedding model; the model name and version are stored with each vector.
  • Index in a hybrid store. Deletes and permission changes take the same path and the same 15-minute target as edits.

Permissions#

This is the part to get exactly right. Two workable designs:

DesignHowTrade-off
Filter at query timeEach chunk stores the groups allowed to read it; the query carries the user’s groups as a filter evaluated inside the indexOne index; group membership changes take effect immediately; filters must be efficient for users in thousands of groups
Check at read timeRetrieve more candidates, then ask the source system or an ACL store whether this user may read eachExact and current; adds latency; needs a fast permission service

Use the first for speed and the second as a verification step on the final handful of chunks before they enter the context. Never rely on the model to withhold a document it has been shown: if a chunk is in the context, treat it as disclosed.

Two subtler leaks to close: answer caches must be keyed by the permission set, not just the question; and citations must not reveal titles of documents the user cannot open.

Query#

  1. The gateway authenticates and attaches the user’s identity.
  2. A small model rewrites the question using the conversation, and may produce two or three searches.
  3. Hybrid search with the user’s group filter returns 100 candidates; a reranker keeps 8.
  4. The answer model receives instructions, the chunks with source labels, and the question. It must cite a label for each claim or say it could not find the answer.
  5. A citation check confirms each cited label was in the context, and the output sanitiser removes links and images that point outside trusted hosts.
  6. The trace and any user feedback go to the evaluation pipeline.

For questions that need several hops — “compare our leave policy with last year’s” — the answer service escalates to an agentic retrieval loop in which the model issues further searches, with a step limit. It is used only when a classifier flags the question as multi-part, because it triples cost and latency.

Model choices#

CallModelWhy
Rewrite, classifySmall self-hostedHigh volume, simple, latency-critical
Embed, rerankSmall self-hostedVolume; data stays inside
AnswerHosted frontier model, with a self-hosted large model as fallbackFaithfulness to sources is the quality bar; evaluated per release

Step 4 — failure#

FailureBehaviour
Nothing relevant retrievedThe answer says so and suggests where to look; it does not improvise
Index lagging behind sourcesA freshness metric per source; a banner when lag exceeds the target; stale answers carry their document date
Connector broken for one sourceThat source is marked degraded; others continue
Permission store unavailableFail closed: no answer rather than an unfiltered one
Answer model slowFirst-token deadline, then fallback route
Embedding model upgradedBuild a second index in parallel, evaluate recall, switch, then drop the old one
A document contains injected instructionsIngestion scan and quarantine; the answer path has no tools and a sanitised output, so the worst case is a wrong answer, not an action

Step 5 — cost#

answer calls   20/s × (6,500 × $3 + 3,500 × $0.30 + 450 × $15) / 1e6
             = 20 × $0.0273                                   ≈ $0.55 / s  → ~$15,700 per 8-hour day
small models   rewrite, rerank, embed on ~6 shared GPUs        ≈ $500 / day
index          ~200 GB memory × 3 replicas + storage           ≈ $300 / day
                                                                 ≈ $0.03 per question

The answer call is 95% of the bill. Levers: retrieve fewer, better chunks (8 instead of 20 halves the input); cache the instruction prefix; route simple factual questions to the self-hosted model where evaluation shows it suffices; cap answer length.

Step 6 — security#

ItemIn this design
Untrusted sourcesEvery document: any employee, and in ticket and chat sources any customer or outsider, can write text that will be retrieved
SinksText and citations rendered to the user. No tools.
TrifectaPrivate data ✓, untrusted content ✓, outbound channel: only via rendered links and images → closed by the output sanitiser and a strict content-security policy
PermissionsEnforced in the search filter and re-verified before context assembly; fail closed
PoisoningKnown-writer sources ranked above open ones; ingestion scanning; ability to remove a document and its chunks everywhere within minutes
Leakage through cachesNo cross-user semantic cache; exact caches keyed by permission set
AuditEach answer records the user, the chunks shown and their sources

Keeping this system tool-less is a deliberate security decision: an assistant that can only answer can be wrong, but it cannot be made to act. The day someone asks to add “and let it file tickets”, the design moves into agent territory and the trifecta analysis must be redone.

Evaluation#

LayerMetricGate
RetrievalRecall@8 on 500 labelled question–document pairsNo regression
AnswerFaithfulness and correctness by a validated judge; citation accuracy by code≥ 90%, no regression
DecliningShare of unanswerable questions correctly declined≥ 95%
PermissionsA fixed suite of “user A asks about document only B can read”100%, always
Latency and costp95 first token; cost per questionWithin budget

The permission suite is a must-pass gate on every change to connectors, index or query code.

Remember this#

  • In enterprise RAG the hard parts are permissions and index freshness, not the model.
  • Filter by permission inside the search and verify again before the context; fail closed.
  • Deletes and permission changes are first-class change events with the same deadline as edits.
  • The answer model must cite or decline; code verifies the citations.
  • A tool-less assistant has a small blast radius. Adding tools changes the threat model.

Try it#

  1. Recompute the cost if retrieval passes 20 chunks instead of 8. What happens to quality?
  2. A user is removed from a group at 10:00. Trace exactly when they stop receiving answers built from that group’s documents in each permission design.
  3. Design the test that proves the citation list never reveals a forbidden document’s title.

Check yourself#

  1. Why is “the model was told not to reveal it” not a permission control?
  2. What should the system do when the permission store is unavailable, and why?
  3. Why does upgrading the embedding model require building a second index?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom