The idea in one minute#
A model is a dependency that runs with access to your data and GPUs, so it gets the treatment any dependency gets — and a little more, because a model file can contain executable code and its behaviour cannot be read from its contents. The controls form a pipeline: acquire from a known publisher, verify integrity and signature, scan for unsafe content, record what it is in an AI bill of materials, sign it yourself, and promote it to an internal registry. Production loads only from that registry, only by digest. The same pipeline applies to every other AI-specific dependency: adapters, tokenizers, prompts, skills and tools.
A picture#
flowchart LR
HUB[":huggingface: <b>Public hub or vendor</b>"] --> Q[":i-box: <b>Quarantine</b><br/><small>isolated, no credentials</small>"]
Q --> V[":sigstore: <b>Verify</b><br/><small>publisher signature, digest</small>"]
V --> SC[":i-search: <b>Scan</b><br/><small>format, unsafe ops,<br/>licence</small>"]
SC --> EV[":i-scale: <b>Evaluate</b><br/><small>quality + safety probes</small>"]
EV --> BOM[":cyclonedx: <b>AI-BOM</b><br/><small>model, data, lineage</small>"]
BOM --> SIGN[":sigstore: <b>Sign</b><br/><small>your key, OMS format</small>"]
SIGN --> REG[(":harbor: <b>Internal registry</b><br/><small>OCI artifact, by digest</small>")]
REG --> ADM[":kyverno: <b>Admission policy</b><br/><small>signed + from registry only</small>"]
ADM --> SRV[":vllm: <b>Serving</b>"]
SC -.->|"fails"| REJ[":i-ban: <b>Reject</b>"]
class HUB neutral
class Q warn
class V,SIGN,ADM queue
class SC,EV compute
class BOM io
class REG memory
class SRV compute
class REJ warnHow it really works#
Step 1 — acquire deliberately#
- Choose models from identified publishers; check the organisation, not only the name. Lookalike repository names are common.
- Pin a revision — a commit hash or digest — never a floating name like
mainorlatest. - Check the licence and usage terms; some forbid commercial use or specific applications.
- Download into a quarantine environment: isolated, with no cloud credentials and no route to production.
Step 2 — prefer formats that cannot execute#
| Format | Can run code on load? | Guidance |
|---|---|---|
safetensors | No — tensors and metadata only | The default; require it where available |
GGUF | No arbitrary code; parsers have had bugs | Acceptable; keep loaders patched |
| ONNX | Graph of known operators; custom operators are a risk | Acceptable with operator allowlists |
Python pickle and formats built on it (.pt, .bin, .pkl, .joblib) | Yes | Avoid; if unavoidable, scan and load only in a sandbox |
| Repositories requiring “trust remote code” | Yes, by design | Review the code like any dependency, or do not use |
Converting a model to safetensors once, in quarantine, removes the largest class of
malicious-model risk for everyone downstream.
Step 3 — scan#
Scanners inspect model files for unsafe serialisation operations, suspicious imports, embedded payloads and known-bad hashes. Major hubs scan uploads, and open-source scanners can run in your own pipeline. Scanning catches code in the file. It does not catch a behavioural backdoor in the weights, which is ordinary numbers.
For that, the available measures are weaker and layered: provenance (a publisher you have reason to trust, a documented training process), behavioural evaluation including probes for trigger patterns, comparison with a reference model, and — the one that actually bounds the damage — runtime controls that limit what any model’s output can cause.
Step 4 — record: the AI bill of materials#
An AI-BOM (or ML-BOM) describes a model the way an SBOM describes software:
model name, version, digest of every file, format, architecture
origin publisher, source URL, revision
lineage base model; fine-tuning runs; adapters
data dataset identifiers and versions used for training and evaluation
software framework and library versions needed to run it
licence model and data licences
evaluation results of quality and safety evaluations, with dates
signatures who signed itCycloneDX (ML-BOM, since version 1.5) and SPDX 3 (AI and Dataset profiles) are the two standard formats. The BOM answers incident questions quickly: “which deployments contain this base model?”, “which models were trained on that dataset?”
Step 5 — sign and verify#
A hash proves a file has not changed; a signature proves who vouched for it. The OpenSSF Model Signing (OMS) specification defines a detached signature over a manifest of all the files in a model — weights, configuration, tokenizer — so the whole artifact is verified as one unit. It works with ordinary keys, certificate chains, or keyless signing through Sigstore, where the signature is tied to a workload or person identity and recorded in a transparency log. Several large model hubs and vendor catalogues have adopted it.
Two signatures matter:
- the publisher’s, verified when you acquire the model;
- yours, applied after your own scanning and evaluation, and the only one production trusts.
A caution the specification’s own authors stress: a signature says the artifact is intact and names its signer. It says nothing about whether the model is safe or how it was trained. Integrity is not provenance, and neither is behaviour.
Step 6 — registry and admission#
Store approved models as OCI artifacts in an internal registry, exactly as container images are stored: content-addressed, access-controlled, cacheable on nodes, replicated between regions. Then enforce, at deployment time, a policy with three clauses:
a model may be loaded in production only if
1 it is pulled from the internal registry (not from the internet at start-up)
2 it is referenced by digest (not by a mutable tag)
3 it carries a valid signature from our signing identityOn Kubernetes this is an admission policy (Kyverno, OPA Gatekeeper or a validating admission policy) plus removing internet egress from serving pods, so that a misconfiguration cannot quietly download weights from elsewhere.
The same pipeline for everything else#
| Component | Acquire | Verify and scan | Pin | Registry |
|---|---|---|---|---|
| LoRA adapters, tokenizers, configs | As for models | Same formats and scanners | Digest | With the model |
| Datasets | Recorded provenance | Hash manifest; content scans | Version | Dataset store |
| Prompts, policies, routing rules | Written in-house | Code review; evaluation gate | Git commit | Version control |
| Skills, rules files, prompt packs | Known author | Read them — they are instructions | Commit or digest | Internal catalogue |
| MCP servers and tools | Known publisher | Code and description review; sandboxed trial | Version and description hash | Approved catalogue behind the tool gateway |
| Frameworks and SDKs | Package registry | SCA scanning; lockfiles; provenance attestations | Lockfile | Internal mirror |
| Container images | Trusted base | Image scanning; signatures | Digest | Internal registry |
For build integrity overall, SLSA levels describe how much assurance a build pipeline provides — from “there is a build record” to “the build is isolated and its provenance is unforgeable” — and the same idea extends to training runs: a signed record of code, data and base model that produced these weights.
What fills the box in 2026#
| Need | Options |
|---|---|
| Model scanning | Open-source scanners for unsafe serialisation; hub-side scanning; commercial model-security products |
| Signing | OpenSSF model-signing tooling; Sigstore (cosign) for OCI artifacts |
| AI-BOM | CycloneDX ML-BOM generators; SPDX 3 AI profile tooling |
| Registry | Harbor or a cloud OCI registry; MLflow or a model catalogue for metadata |
| Admission | Kyverno, OPA Gatekeeper, validating admission policies |
Code#
Verify a model directory against a signed-off manifest before loading: every file present, no extra files, every hash matching, and no forbidden format.
// verify.go — check model files against an approved manifest and a format policy.
package main
import (
"crypto/sha256"
"encoding/hex"
"fmt"
"path/filepath"
"sort"
)
func hash(b []byte) string {
sum := sha256.Sum256(b)
return hex.EncodeToString(sum[:])
}
var forbidden = map[string]bool{".pkl": true, ".pt": true, ".bin": true, ".joblib": true}
// verify compares the files found on disk with the approved manifest.
func verify(manifest map[string]string, files map[string][]byte) []string {
var problems []string
names := make([]string, 0, len(files))
for n := range files {
names = append(names, n)
}
sort.Strings(names)
for _, n := range names {
want, ok := manifest[n]
switch {
case forbidden[filepath.Ext(n)]:
problems = append(problems, n+": format can execute code on load")
case !ok:
problems = append(problems, n+": not in the approved manifest")
case hash(files[n]) != want:
problems = append(problems, n+": content differs from the approved version")
}
}
for n := range manifest {
if _, ok := files[n]; !ok {
problems = append(problems, n+": missing")
}
}
sort.Strings(problems)
return problems
}
func main() {
approved := map[string][]byte{
"model.safetensors": []byte("weights-v1"),
"config.json": []byte(`{"layers": 32}`),
"tokenizer.json": []byte(`{"vocab": 50000}`),
}
manifest := map[string]string{} // in production this manifest is itself signed
for n, b := range approved {
manifest[n] = hash(b)
}
onDisk := map[string][]byte{
"model.safetensors": []byte("weights-v1"),
"config.json": []byte(`{"layers": 32, "auto_map": "custom_code.py"}`),
"tokenizer.json": []byte(`{"vocab": 50000}`),
"extra_weights.bin": []byte("?"),
}
problems := verify(manifest, onDisk)
if len(problems) == 0 {
fmt.Println("verified: safe to load")
return
}
fmt.Println("REFUSING TO LOAD:")
for _, p := range problems {
fmt.Println(" -", p)
}
}Remember this#
- A model is a dependency with code-execution potential and unreadable behaviour.
- Quarantine, verify, scan, evaluate, record, sign, promote — then load only from your registry, by digest, with your signature.
- Prefer weights-only formats;
pickle-based formats run code. - A signature proves integrity and signer, not safety. Scanning finds code, not backdoors.
- The same pipeline covers adapters, datasets, prompts, skills, tools and packages.
Try it#
- Run
verify.go. Removeextra_weights.binfromonDisk; what remains, and why is a changedconfig.jsonworth blocking on? - List the model files a service of yours loads at start-up. Where do they come from, and are they referenced by digest?
- Draft the AI-BOM fields for one model you run. Which can you not fill in?
Check yourself#
- What does scanning a model file catch, and what can it not catch?
- What does a valid signature on a model tell you, and what does it not?
- Why should serving pods have no internet egress?