The idea in one minute#
The system: an assistant that answers employees’ questions from 20 million internal documents — wikis, tickets, shared drives, chat — and cites its sources. It looks like a search box with a model on the end. The design work is almost entirely in two places that have nothing to do with the model: permissions (a user must never receive an answer built from a document they cannot open) and freshness and quality of the index (the model can only be as right as what is retrieved). The model is the easy part.
Step 1 — requirements#
| Requirement | Decision |
|---|---|
| Users | 30,000 employees; 3,000 active at peak |
| Corpus | 20 M documents, 6 source systems, ~150 M chunks; 2% change per day |
| Unit of work | One question answered with citations |
| Quality | ≥ 90% of answers judged correct and supported; “I could not find this” when unsupported |
| Latency | First token under 2 s; full answer under 12 s |
| Permissions | Source-system access controls enforced exactly, including group changes within 15 minutes |
| Freshness | A changed document searchable within 15 minutes |
| Data | Everything stays in the company’s cloud; hosted model allowed under zero retention |
| Deletion | A deleted document is unsearchable within 15 minutes |
Step 2 — numbers#
query path
questions at peak 3,000 users × 0.4 / min ÷ 60 = 20 / s
model calls per question rewrite (small) + answer (large) = 2
answer call 10,000 in (3,500 cached), 450 out
retrieval budget rewrite 300 ms + search 80 ms + rerank 150 ms = ~0.5 s before the answer call
index
vectors 150 M × 1,024 dims × 4 bytes = 614 GB raw
with scalar quantization (1 byte/dim) + graph index ≈ 200 GB in memory, originals on disk for rescoring
embedding throughput 2% of 150 M = 3 M chunks/day ≈ 35 chunks/s average, bursts of 2,000/s on bulk edits
initial build 150 M chunks at 3,000 chunks/s per GPU ≈ 14 GPU-hoursThe query side is modest: twenty questions a second. The index is the heavy component — hundreds of gigabytes in memory, kept fresh continuously, with permissions attached to every chunk.
Step 3 — architecture#
flowchart TB
subgraph INGEST["Ingestion: continuous"]
direction LR
SRC[":confluence: <b>Sources</b><br/><small>wiki, tickets, drive, chat</small>"] --> CONN[":i-route: <b>Connectors</b><br/><small>change feeds + ACLs</small>"]
CONN --> BUS[":apachekafka: <b>Change stream</b>"]
BUS --> PARSE[":i-funnel: <b>Parse, chunk, scan</b>"]
PARSE --> EMB[":vllm: <b>Embedding model</b>"]
EMB --> IDX[(":qdrant: <b>Hybrid index</b><br/><small>vector + keyword + ACL</small>")]
CONN --> ACL[(":postgresql: <b>Identity and ACL store</b><br/><small>users, groups, doc permissions</small>")]
end
subgraph QUERY["Query: per question"]
direction LR
U[":i-user: <b>Employee</b>"] --> GW[":envoyproxy: <b>AI gateway</b>"]
GW --> ORC[":i-workflow: <b>Answer service</b>"]
ORC --> RW[":vllm: <b>Rewrite</b><br/><small>small model</small>"]
RW --> SR[":i-search: <b>Search</b><br/><small>filter: user's groups</small>"]
SR --> RR[":vllm: <b>Rerank</b>"]
RR --> ANS[":anthropic: <b>Answer model</b><br/><small>cite or decline</small>"]
ANS --> CHK[":i-shield-check: <b>Citation check,<br/>output sanitiser</b>"]
end
IDX --> SR
ACL --> SR
CHK --> U
ORC -.-> TEL[":opentelemetry: <b>Traces, feedback</b>"] --> EVAL[":i-scale: <b>Evaluation</b>"]
class SRC,U neutral
class CONN,BUS,GW,SR queue
class PARSE,CHK warn
class EMB,RW,RR,ANS,ORC compute
class IDX,ACL memory
class TEL,EVAL neutralIngestion#
- Connectors read each source’s change feed — created, updated, deleted, permissions changed — rather than re-crawling. They emit the document and its access-control list.
- Parse and chunk by structure, attaching the heading path, source URL, author, updated time and ACL to every chunk. Scan for secrets and for instruction-like text; quarantine matches for review rather than indexing them.
- Embed with a self-hosted embedding model; the model name and version are stored with each vector.
- Index in a hybrid store. Deletes and permission changes take the same path and the same 15-minute target as edits.
Permissions#
This is the part to get exactly right. Two workable designs:
| Design | How | Trade-off |
|---|---|---|
| Filter at query time | Each chunk stores the groups allowed to read it; the query carries the user’s groups as a filter evaluated inside the index | One index; group membership changes take effect immediately; filters must be efficient for users in thousands of groups |
| Check at read time | Retrieve more candidates, then ask the source system or an ACL store whether this user may read each | Exact and current; adds latency; needs a fast permission service |
Use the first for speed and the second as a verification step on the final handful of chunks before they enter the context. Never rely on the model to withhold a document it has been shown: if a chunk is in the context, treat it as disclosed.
Two subtler leaks to close: answer caches must be keyed by the permission set, not just the question; and citations must not reveal titles of documents the user cannot open.
Query#
- The gateway authenticates and attaches the user’s identity.
- A small model rewrites the question using the conversation, and may produce two or three searches.
- Hybrid search with the user’s group filter returns 100 candidates; a reranker keeps 8.
- The answer model receives instructions, the chunks with source labels, and the question. It must cite a label for each claim or say it could not find the answer.
- A citation check confirms each cited label was in the context, and the output sanitiser removes links and images that point outside trusted hosts.
- The trace and any user feedback go to the evaluation pipeline.
For questions that need several hops — “compare our leave policy with last year’s” — the answer service escalates to an agentic retrieval loop in which the model issues further searches, with a step limit. It is used only when a classifier flags the question as multi-part, because it triples cost and latency.
Model choices#
| Call | Model | Why |
|---|---|---|
| Rewrite, classify | Small self-hosted | High volume, simple, latency-critical |
| Embed, rerank | Small self-hosted | Volume; data stays inside |
| Answer | Hosted frontier model, with a self-hosted large model as fallback | Faithfulness to sources is the quality bar; evaluated per release |
Step 4 — failure#
| Failure | Behaviour |
|---|---|
| Nothing relevant retrieved | The answer says so and suggests where to look; it does not improvise |
| Index lagging behind sources | A freshness metric per source; a banner when lag exceeds the target; stale answers carry their document date |
| Connector broken for one source | That source is marked degraded; others continue |
| Permission store unavailable | Fail closed: no answer rather than an unfiltered one |
| Answer model slow | First-token deadline, then fallback route |
| Embedding model upgraded | Build a second index in parallel, evaluate recall, switch, then drop the old one |
| A document contains injected instructions | Ingestion scan and quarantine; the answer path has no tools and a sanitised output, so the worst case is a wrong answer, not an action |
Step 5 — cost#
answer calls 20/s × (6,500 × $3 + 3,500 × $0.30 + 450 × $15) / 1e6
= 20 × $0.0273 ≈ $0.55 / s → ~$15,700 per 8-hour day
small models rewrite, rerank, embed on ~6 shared GPUs ≈ $500 / day
index ~200 GB memory × 3 replicas + storage ≈ $300 / day
≈ $0.03 per questionThe answer call is 95% of the bill. Levers: retrieve fewer, better chunks (8 instead of 20 halves the input); cache the instruction prefix; route simple factual questions to the self-hosted model where evaluation shows it suffices; cap answer length.
Step 6 — security#
| Item | In this design |
|---|---|
| Untrusted sources | Every document: any employee, and in ticket and chat sources any customer or outsider, can write text that will be retrieved |
| Sinks | Text and citations rendered to the user. No tools. |
| Trifecta | Private data ✓, untrusted content ✓, outbound channel: only via rendered links and images → closed by the output sanitiser and a strict content-security policy |
| Permissions | Enforced in the search filter and re-verified before context assembly; fail closed |
| Poisoning | Known-writer sources ranked above open ones; ingestion scanning; ability to remove a document and its chunks everywhere within minutes |
| Leakage through caches | No cross-user semantic cache; exact caches keyed by permission set |
| Audit | Each answer records the user, the chunks shown and their sources |
Keeping this system tool-less is a deliberate security decision: an assistant that can only answer can be wrong, but it cannot be made to act. The day someone asks to add “and let it file tickets”, the design moves into agent territory and the trifecta analysis must be redone.
Evaluation#
| Layer | Metric | Gate |
|---|---|---|
| Retrieval | Recall@8 on 500 labelled question–document pairs | No regression |
| Answer | Faithfulness and correctness by a validated judge; citation accuracy by code | ≥ 90%, no regression |
| Declining | Share of unanswerable questions correctly declined | ≥ 95% |
| Permissions | A fixed suite of “user A asks about document only B can read” | 100%, always |
| Latency and cost | p95 first token; cost per question | Within budget |
The permission suite is a must-pass gate on every change to connectors, index or query code.
Remember this#
- In enterprise RAG the hard parts are permissions and index freshness, not the model.
- Filter by permission inside the search and verify again before the context; fail closed.
- Deletes and permission changes are first-class change events with the same deadline as edits.
- The answer model must cite or decline; code verifies the citations.
- A tool-less assistant has a small blast radius. Adding tools changes the threat model.
Try it#
- Recompute the cost if retrieval passes 20 chunks instead of 8. What happens to quality?
- A user is removed from a group at 10:00. Trace exactly when they stop receiving answers built from that group’s documents in each permission design.
- Design the test that proves the citation list never reveals a forbidden document’s title.
Check yourself#
- Why is “the model was told not to reveal it” not a permission control?
- What should the system do when the permission store is unavailable, and why?
- Why does upgrading the embedding model require building a second index?