The idea in one minute#
GPUs get the attention, but an AI platform stores six different kinds of data, and each has a natural home. Model artifacts go in a registry. Source documents go in object storage. Vectors and keyword indexes go in a search store. Sessions, tasks and memory go in a transactional database and a fast key-value store. Events — every request, token count and trace — go in a stream and a column store. Evaluation data is versioned alongside code. Most data-layer mistakes are one of these kinds living in another’s home.
A picture#
flowchart LR
subgraph IN["Sources"]
DOC[":i-file-text: <b>Documents</b>"]
HF[":huggingface: <b>Model hubs</b>"]
end
DOC --> OBJ[(":minio: <b>Object storage</b><br/><small>raw + parsed documents</small>")]
OBJ --> PIPE[":apacheairflow: <b>Ingestion jobs</b><br/><small>parse, chunk, embed</small>"]
PIPE --> VEC[(":qdrant: <b>Search index</b><br/><small>vectors, keywords, ACLs</small>")]
HF --> SCAN[":i-shield-check: <b>Scan and sign</b>"] --> REG[(":harbor: <b>Model registry</b><br/><small>OCI artifacts</small>")]
REG --> SRV[":vllm: <b>Serving</b>"]
VEC --> APP[":i-bot: <b>Agents and apps</b>"]
APP --> PG[(":postgresql: <b>Postgres</b><br/><small>tasks, memory, metadata</small>")]
APP --> RD[(":redis: <b>Redis</b><br/><small>sessions, caches, limits</small>")]
APP --> BUS[":apachekafka: <b>Event stream</b>"]
SRV --> BUS
BUS --> CH[(":clickhouse: <b>Column store</b><br/><small>usage, traces, evals</small>")]
class DOC,HF neutral
class OBJ,VEC,REG,PG,RD,CH memory
class PIPE,BUS queue
class SCAN warn
class SRV,APP computeHow it really works#
The six kinds#
| Data | Shape | Home | Why there |
|---|---|---|---|
| Model artifacts | Tens of GB, immutable, versioned | OCI registry, backed by object storage | Content-addressed, signable, cacheable on nodes |
| Source documents | Files of any type | Object storage | Cheap, durable; the source of truth an index can be rebuilt from |
| Search index | Vectors + text + metadata | Vector or hybrid search store | Approximate nearest-neighbour search with filters |
| Operational state | Sessions, task checkpoints, memory, tenants | Postgres; Redis for hot, short-lived items | Transactions and durability |
| Events and telemetry | Append-only, high volume | Stream, then column store | Cheap to write, fast to aggregate |
| Evaluation datasets | Small, curated, versioned | Git or a dataset store with versions | Must be reproducible and reviewable |
Model artifacts#
Treat weights like container images, because operationally they are: large immutable blobs that many nodes pull. Storing them as OCI artifacts gives content addressing, signatures, node-level caching and the same access control as images. The ingest path from a public hub runs through a gate: scan for unsafe serialisation formats, record provenance, sign, and only then promote to the registry that production pulls from. That gate is a security control — see Models and Supply Chain.
The search index#
Sizing is arithmetic:
10 million chunks × 1,024 dimensions × 4 bytes (float32) = 41 GB of raw vectors
+ graph index (HNSW) overhead, roughly 1.5× ≈ 60 GB in memory
with scalar quantization to 1 byte per dimension ≈ 10 GB + index
with binary quantization + rescoring from disk ≈ 1.3 GB + indexDesign points:
- One index or many? Per-tenant collections give hard isolation and simple deletion, and cost more. A shared collection with a mandatory tenant filter scales better and makes the filter a security control that must never be omitted.
- Filtering. Access-control lists, tenant and freshness are filters applied inside the approximate search; check that your store does this efficiently, because post-filtering silently returns too few results.
- Freshness. Decide how soon a changed document must be searchable. Minutes needs a streaming ingestion path; a day allows a nightly batch.
- Deletion. When a source document is deleted or a user invokes their right to erasure, the chunks and vectors must go too. Keep the mapping from source to chunks.
- Rebuild. The index is derived data. You will rebuild it when the embedding model or the chunking changes, so keep the pipeline runnable and the source documents intact.
If you already operate Postgres, pgvector is a sound first choice up to millions of vectors
and avoids a new system. Dedicated stores earn their place at larger scale or with demanding
filters.
Operational state#
Postgres holds what must not be lost: tenants and keys, agent task checkpoints, memory records, tool-approval decisions, audit trails. Redis holds what is hot and may be lost: session buffers, rate-limit counters, exact-match caches. The classic mistake is keeping the only copy of an hours-long agent task in Redis without persistence.
Events and the feedback loop#
Every model call produces a usage event at the gateway and spans from the agent runtime. Send them through a stream to a column store, and four consumers appear:
- Billing and showback — tokens and cost per tenant.
- Capacity planning — demand by model over time.
- Evaluation — sampled traces become judged samples and new test cases.
- Security — the audit trail of what each agent did with whose authority.
Prompts and completions are customer data. Store content separately from metadata, off by default, with short retention and audited access; see Designing observability for an inference platform.
Data governance in one table#
| Question | The design must answer |
|---|---|
| Residency | In which region may this tenant’s documents, vectors and prompts live? |
| Retention | How long is each kind kept, and what deletes it? |
| Lineage | Which source produced this chunk; which data trained this adapter? |
| Access | Who — and which agent, on whose behalf — may read it? |
| Training use | May these conversations be used for fine-tuning or evaluation? |
What fills the box in 2026#
| Need | Options |
|---|---|
| Object storage | S3-compatible: cloud buckets, MinIO, Ceph |
| Registry | Harbor or a cloud OCI registry; MLflow for model metadata and lineage |
| Search index | Postgres + pgvector, Qdrant, Milvus, OpenSearch, Elasticsearch |
| Operational state | Postgres; Redis or Valkey |
| Stream and analytics | Kafka or NATS; ClickHouse |
| Pipelines | Airflow, Spark, Ray Data, or queue-driven workers |
Remember this#
- Six kinds of data, six homes. Do not keep durable state in a cache or documents only in an index.
- Weights are OCI artifacts behind a scan-and-sign gate.
- The search index is derived: keep the sources and the pipeline to rebuild it.
- Tenant and permission filters run inside the search.
- Events feed billing, capacity, evaluation and audit; content is stored separately and sparingly.
Try it#
- Size the index for 50 million chunks at 768 dimensions, raw and with scalar quantization.
- A user deletes a document. Trace every store in the diagram that must change.
- For each of the five governance questions, write the answer for a system you know.
Check yourself#
- Why should model weights go through a registry rather than being downloaded by each pod?
- What goes wrong when permission filtering is applied after the vector search?
- Which stores can be rebuilt from others, and which cannot?