The idea in one minute#
A model knows only what is in its context window on this call. Context engineering is the work of deciding, on every call, what goes into that window: instructions, the conversation, retrieved documents, tool definitions and results. Retrieval is the part that fetches documents the model was never trained on. Quality problems that look like “the model is not smart enough” are, far more often, “the model was not shown the right thing”.
A retrieval pipeline has two halves: ingestion, which runs offline and turns documents into searchable pieces, and query, which runs on every request and must find the right pieces in tens of milliseconds.
A picture#
flowchart LR
subgraph ING["Ingestion: offline"]
direction LR
S[":i-file-text: <b>Sources</b><br/><small>docs, wikis, tickets</small>"] --> P[":i-funnel: <b>Parse and chunk</b><br/><small>keep structure, add metadata</small>"]
P --> E[":huggingface: <b>Embed</b><br/><small>text to vector</small>"]
E --> IX[(":qdrant: <b>Index</b><br/><small>vectors + keywords + ACLs</small>")]
end
subgraph QRY["Query: every request"]
direction LR
Q[":i-user: <b>Question</b>"] --> RW[":i-search: <b>Rewrite</b><br/><small>resolve 'it', split, expand</small>"]
RW --> H[":i-merge: <b>Hybrid search</b><br/><small>vector + keyword, filter by ACL</small>"]
H --> RR[":i-list-checks: <b>Rerank</b><br/><small>100 candidates to 8</small>"]
RR --> CB[":i-layers: <b>Context builder</b><br/><small>budget, order, cite</small>"]
CB --> M[":i-brain: <b>Model</b>"]
end
IX --> H
class S,Q neutral
class P,RW queue
class E,RR,M compute
class IX memory
class H,CB ioHow it really works#
The context window is a budget#
Treat the window as a fixed budget and allocate it on purpose.
a 200,000-token window, allocated for a support assistant
system instructions and policies 3,000 stable → cacheable
tool definitions 4,000 stable → cacheable
retrieved documents 12,000 per request
conversation so far 8,000 grows
the user's message 300
reserved for the answer 2,000
------
used 29,300 (most of the window stays empty, on purpose)Two facts drive the layout. First, more is not better: models attend less reliably to the middle of a very long context, and every token is paid for on every call. Second, order decides cost: providers and engines cache a prompt’s unchanged prefix, so stable content goes first and anything that changes per request goes last.
Ingestion#
Parse. Keep structure: headings, tables, code blocks, page numbers. A table flattened into a paragraph cannot be retrieved as a table.
Chunk. Split into pieces of a few hundred tokens along natural boundaries — sections, paragraphs, functions — and attach the path to each (“Handbook > Leave > Parental leave”). A chunk must make sense alone, because alone is how the model will see it. A widely used refinement is to prepend a sentence of generated context to each chunk before embedding it.
Embed. An embedding model maps each chunk to a vector; similar meaning gives nearby vectors. Choose it by retrieval benchmark results on text like yours, and record its name and version with every vector — changing the embedding model means re-embedding everything.
Index. Store vector, text, metadata and access-control labels together. Permissions are enforced as a filter inside the search, never by trimming results afterwards.
Query#
Rewrite. “What about for contractors?” retrieves nothing useful on its own. A small model rewrites the question using the conversation, and may split it into several searches.
Hybrid search. Vector search finds meaning; keyword search (BM25) finds exact terms —
error codes, product names, identifiers. Run both and merge the rankings. Reciprocal rank
fusion merges without needing comparable scores: each document scores 1 / (60 + rank) in
each list, summed.
Rerank. A cross-encoder reads the question and each candidate together and scores relevance far more accurately than vector distance. It is too slow for a million chunks and fine for a hundred: retrieve 100, keep 8.
Build the context. Put the best chunks in, each with a source label, and instruct the model to cite the labels. Citations are what let a user — and an evaluation — check the answer.
Retrieval is not only vectors#
| Need | Better tool than a vector index |
|---|---|
| Exact facts about entities and relations | SQL, or a knowledge graph |
| Fresh data: inventory, account balance | An API call, as a tool |
| Code | Symbol search and grep, by an agent |
| A handful of documents | Put them all in the context and cache them |
Agentic retrieval turns the pipeline inside out: the model is given search as a tool and decides what to look up, reads the result, and searches again. It handles multi-step questions that a single search cannot, at the cost of more calls. In 2026 it is the default for coding agents and research assistants; a fixed pipeline remains better for high-volume, low-latency question answering.
How it fails#
| Symptom | Usual cause | Fix |
|---|---|---|
| Right document exists, wrong one retrieved | Vector-only search missing exact terms | Hybrid search, reranker |
| Answer is confidently wrong | Nothing relevant retrieved and the model filled the gap | Detect empty retrieval; allow “I do not know” |
| Answer mixes two policies | Chunks lost their headings | Structure-aware chunking with paths |
| User sees a document they should not | Permission checked after retrieval | Filter inside the query by ACL |
| Quality dropped after an “upgrade” | Embedding model changed, index not rebuilt | Version embeddings; rebuild and re-evaluate |
| A document changes the assistant’s behaviour | Prompt injection in retrieved text | Treat retrieved text as untrusted: see Prompt Injection |
Measure retrieval separately from generation. Recall@k — is the right chunk among the top k? — tells you whether to fix search or fix the prompt.
What fills the box in 2026#
| Layer | Options |
|---|---|
| Vector and hybrid index | Postgres with pgvector (start here if you already run Postgres), Qdrant, Milvus, OpenSearch, Elasticsearch, LanceDB |
| Embedding and reranking models | Open-weights models served by your inference layer, or a hosted embedding API |
| Parsing | Document-layout models for PDFs and scans; plain parsers for Markdown and HTML |
| Frameworks | LlamaIndex, LangChain, or a few hundred lines of your own |
Code#
Reciprocal rank fusion, the merge step of hybrid search.
// rrf.go — merge a keyword ranking and a vector ranking with reciprocal rank fusion.
package main
import (
"fmt"
"sort"
)
// fuse scores each document 1/(k+rank) in every list it appears in, and sums.
func fuse(k float64, lists ...[]string) []string {
score := map[string]float64{}
for _, list := range lists {
for rank, doc := range list {
score[doc] += 1 / (k + float64(rank+1))
}
}
docs := make([]string, 0, len(score))
for d := range score {
docs = append(docs, d)
}
sort.Slice(docs, func(i, j int) bool {
if score[docs[i]] != score[docs[j]] {
return score[docs[i]] > score[docs[j]]
}
return docs[i] < docs[j]
})
for _, d := range docs {
fmt.Printf(" %-22s %.4f\n", d, score[d])
}
return docs
}
func main() {
// Query: "error E4012 when exporting parental leave report"
keyword := []string{"E4012-troubleshooting", "export-errors", "release-notes-9.2", "leave-reports"}
vector := []string{"leave-reports", "parental-leave-policy", "E4012-troubleshooting", "report-builder"}
fmt.Println("fused ranking:")
top := fuse(60, keyword, vector)
fmt.Println("\nsend to the reranker:", top[:3])
}The page that names the error code and the page about leave reports both rise to the top; a page that only one method liked falls behind. Neither list alone would have ordered them so.
Remember this#
- The context window is a budget. Allocate it; do not fill it.
- Stable content first, changing content last, so the prefix can be cached.
- Hybrid search plus a reranker is the dependable baseline.
- Permissions are a filter inside the search.
- Retrieved text is untrusted input.
Try it#
- Run
rrf.go, then changekto 1. What changes, and why is 60 the usual choice? - Take a document you know well and chunk it by hand. Which chunks would be meaningless without their heading path?
- For a system you use, name one question best answered by vector search, one by SQL and one by a live API call.
Check yourself#
- Why does prompt order affect cost?
- What does a reranker do that vector search cannot, and why is it not used for the first pass?
- Why must access control be applied inside the search rather than after it?