Pidoku

Context and Retrieval

Basic 55 min Difficulty 3/5 Lesson 02 of 06

Prerequisites What a Model Changes

The idea in one minute#

A model knows only what is in its context window on this call. Context engineering is the work of deciding, on every call, what goes into that window: instructions, the conversation, retrieved documents, tool definitions and results. Retrieval is the part that fetches documents the model was never trained on. Quality problems that look like “the model is not smart enough” are, far more often, “the model was not shown the right thing”.

A retrieval pipeline has two halves: ingestion, which runs offline and turns documents into searchable pieces, and query, which runs on every request and must find the right pieces in tens of milliseconds.

A picture#

flowchart LR
  subgraph ING["Ingestion: offline"]
    direction LR
    S[":i-file-text: <b>Sources</b><br/><small>docs, wikis, tickets</small>"] --> P[":i-funnel: <b>Parse and chunk</b><br/><small>keep structure, add metadata</small>"]
    P --> E[":huggingface: <b>Embed</b><br/><small>text to vector</small>"]
    E --> IX[(":qdrant: <b>Index</b><br/><small>vectors + keywords + ACLs</small>")]
  end
  subgraph QRY["Query: every request"]
    direction LR
    Q[":i-user: <b>Question</b>"] --> RW[":i-search: <b>Rewrite</b><br/><small>resolve 'it', split, expand</small>"]
    RW --> H[":i-merge: <b>Hybrid search</b><br/><small>vector + keyword, filter by ACL</small>"]
    H --> RR[":i-list-checks: <b>Rerank</b><br/><small>100 candidates to 8</small>"]
    RR --> CB[":i-layers: <b>Context builder</b><br/><small>budget, order, cite</small>"]
    CB --> M[":i-brain: <b>Model</b>"]
  end
  IX --> H
  class S,Q neutral
  class P,RW queue
  class E,RR,M compute
  class IX memory
  class H,CB io

How it really works#

The context window is a budget#

Treat the window as a fixed budget and allocate it on purpose.

a 200,000-token window, allocated for a support assistant

system instructions and policies     3,000    stable   → cacheable
tool definitions                     4,000    stable   → cacheable
retrieved documents                 12,000    per request
conversation so far                  8,000    grows
the user's message                     300
reserved for the answer              2,000
                                    ------
used                                29,300    (most of the window stays empty, on purpose)

Two facts drive the layout. First, more is not better: models attend less reliably to the middle of a very long context, and every token is paid for on every call. Second, order decides cost: providers and engines cache a prompt’s unchanged prefix, so stable content goes first and anything that changes per request goes last.

Ingestion#

Parse. Keep structure: headings, tables, code blocks, page numbers. A table flattened into a paragraph cannot be retrieved as a table.

Chunk. Split into pieces of a few hundred tokens along natural boundaries — sections, paragraphs, functions — and attach the path to each (“Handbook > Leave > Parental leave”). A chunk must make sense alone, because alone is how the model will see it. A widely used refinement is to prepend a sentence of generated context to each chunk before embedding it.

Embed. An embedding model maps each chunk to a vector; similar meaning gives nearby vectors. Choose it by retrieval benchmark results on text like yours, and record its name and version with every vector — changing the embedding model means re-embedding everything.

Index. Store vector, text, metadata and access-control labels together. Permissions are enforced as a filter inside the search, never by trimming results afterwards.

Query#

Rewrite. “What about for contractors?” retrieves nothing useful on its own. A small model rewrites the question using the conversation, and may split it into several searches.

Hybrid search. Vector search finds meaning; keyword search (BM25) finds exact terms — error codes, product names, identifiers. Run both and merge the rankings. Reciprocal rank fusion merges without needing comparable scores: each document scores 1 / (60 + rank) in each list, summed.

Rerank. A cross-encoder reads the question and each candidate together and scores relevance far more accurately than vector distance. It is too slow for a million chunks and fine for a hundred: retrieve 100, keep 8.

Build the context. Put the best chunks in, each with a source label, and instruct the model to cite the labels. Citations are what let a user — and an evaluation — check the answer.

Retrieval is not only vectors#

NeedBetter tool than a vector index
Exact facts about entities and relationsSQL, or a knowledge graph
Fresh data: inventory, account balanceAn API call, as a tool
CodeSymbol search and grep, by an agent
A handful of documentsPut them all in the context and cache them

Agentic retrieval turns the pipeline inside out: the model is given search as a tool and decides what to look up, reads the result, and searches again. It handles multi-step questions that a single search cannot, at the cost of more calls. In 2026 it is the default for coding agents and research assistants; a fixed pipeline remains better for high-volume, low-latency question answering.

How it fails#

SymptomUsual causeFix
Right document exists, wrong one retrievedVector-only search missing exact termsHybrid search, reranker
Answer is confidently wrongNothing relevant retrieved and the model filled the gapDetect empty retrieval; allow “I do not know”
Answer mixes two policiesChunks lost their headingsStructure-aware chunking with paths
User sees a document they should notPermission checked after retrievalFilter inside the query by ACL
Quality dropped after an “upgrade”Embedding model changed, index not rebuiltVersion embeddings; rebuild and re-evaluate
A document changes the assistant’s behaviourPrompt injection in retrieved textTreat retrieved text as untrusted: see Prompt Injection

Measure retrieval separately from generation. Recall@k — is the right chunk among the top k? — tells you whether to fix search or fix the prompt.

What fills the box in 2026#

LayerOptions
Vector and hybrid indexPostgres with pgvector (start here if you already run Postgres), Qdrant, Milvus, OpenSearch, Elasticsearch, LanceDB
Embedding and reranking modelsOpen-weights models served by your inference layer, or a hosted embedding API
ParsingDocument-layout models for PDFs and scans; plain parsers for Markdown and HTML
FrameworksLlamaIndex, LangChain, or a few hundred lines of your own

Code#

Reciprocal rank fusion, the merge step of hybrid search.

Go
// rrf.go — merge a keyword ranking and a vector ranking with reciprocal rank fusion.
package main

import (
	"fmt"
	"sort"
)

// fuse scores each document 1/(k+rank) in every list it appears in, and sums.
func fuse(k float64, lists ...[]string) []string {
	score := map[string]float64{}
	for _, list := range lists {
		for rank, doc := range list {
			score[doc] += 1 / (k + float64(rank+1))
		}
	}
	docs := make([]string, 0, len(score))
	for d := range score {
		docs = append(docs, d)
	}
	sort.Slice(docs, func(i, j int) bool {
		if score[docs[i]] != score[docs[j]] {
			return score[docs[i]] > score[docs[j]]
		}
		return docs[i] < docs[j]
	})
	for _, d := range docs {
		fmt.Printf("  %-22s %.4f\n", d, score[d])
	}
	return docs
}

func main() {
	// Query: "error E4012 when exporting parental leave report"
	keyword := []string{"E4012-troubleshooting", "export-errors", "release-notes-9.2", "leave-reports"}
	vector := []string{"leave-reports", "parental-leave-policy", "E4012-troubleshooting", "report-builder"}

	fmt.Println("fused ranking:")
	top := fuse(60, keyword, vector)
	fmt.Println("\nsend to the reranker:", top[:3])
}

The page that names the error code and the page about leave reports both rise to the top; a page that only one method liked falls behind. Neither list alone would have ordered them so.

Remember this#

  • The context window is a budget. Allocate it; do not fill it.
  • Stable content first, changing content last, so the prefix can be cached.
  • Hybrid search plus a reranker is the dependable baseline.
  • Permissions are a filter inside the search.
  • Retrieved text is untrusted input.

Try it#

  1. Run rrf.go, then change k to 1. What changes, and why is 60 the usual choice?
  2. Take a document you know well and chunk it by hand. Which chunks would be meaningless without their heading path?
  3. For a system you use, name one question best answered by vector search, one by SQL and one by a live API call.

Check yourself#

  1. Why does prompt order affect cost?
  2. What does a reranker do that vector search cannot, and why is it not used for the first pass?
  3. Why must access control be applied inside the search rather than after it?

↑↓ navigate↵ openesc close

drag to pan · scroll to zoom