The model gate
Why urna refuses a text query embedded by a different model than the corpus, what model_hash covers, and where the check runs or does not.
A vector search with the wrong query model still returns numbers. The cosine arithmetic is valid, the dimensions may even match, and the ranking is noise. urna records a fingerprint of the embedding model in every corpus, model_hash, and the CLI compares it with the fingerprint the query embedder reports before it searches. A mismatch is an error, never a silently bad answer.
Three hashes, one gate
A build through urna build keys the embeddings of each model by three hashes. Only one of them takes part in the query-time gate.
| Hash | What it identifies | Where it lives | Used for |
|---|---|---|---|
model_hash | the model: the files and settings that decide what vector a text becomes | the .urna manifest, and the build's <name>.manifest.json | the query-time gate |
corpus_input_hash | the input rows: each row's text, source image, label and chunker_version, in row order | the .urna provenance section, and <name>.manifest.json | keying the build's embedding cache |
embedding_recipe_hash | how the model was driven: query and document modes, image settings, normalize, dtype, device class | <name>.manifest.json only | keying the build's embedding cache |
Together the three form the key of the shared embedding cache, so a rebuild reuses vectors only when the model, the input and the recipe are all unchanged. See Reproducible builds.
At query time only model_hash is checked. The recipe is not carried to the query side: ask and retrieve build the query embedder from the preset's defaults, not from overrides in the build spec. Where a recipe setting also enters model_hash (the dtype policy of a sentence-transformers model, for example), the gate catches the difference; where it does not, nothing does.
What model_hash covers
model_hash is sha256: followed by 64 hex digits, computed by the embedder over a canonical JSON fingerprint. What goes into the fingerprint depends on the embedder:
| Embedder | Fingerprint inputs |
|---|---|
potion (the default, minishlab/potion-base-8M/v1) | the bytes of config.json, tokenizer.json and model.safetensors, the tokenizer settings, mean pooling, normalize, the dimension (256) and the float32 policy |
| sentence-transformers presets (jina, wemm) | the weight, tokenizer and processor files, the model-repo code files, pooling, normalize and the dtype policy |
| open_clip presets (clip, siglip2) | the model id and pretrained tag, the preprocessing transform, a digest of every weight tensor, and L2 normalization |
a sentence-transformers model fingerprinted with python/model_fingerprint.py (checkout) | up to ten config, tokenizer, pooling and weight files, the dimension and normalize |
The vendored potion table gives sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98. A different table gives a different hash, and corpora built with the old one stop answering. That is the gate doing its job.
For sentence-transformers presets the dtype policy follows the device: bfloat16 on CUDA, float16 on MPS, float32 on CPU. The same weights built on one device class and queried on another report different hashes unless URNA_ST_DTYPE (or the spec's dtype) pins the same dtype on both sides.
What the CLI checks at query time
urna ask, urna retrieve, urna search-text and the ask tab of urna tui run the query embedder, parse its JSON and check, in order:
- The reported
embedding_modelequals the manifestembedding_model. - The reported
embedding_dimand the vector length equal the manifestembedding_dim. - The manifest
model_hashis not the placeholder,sha256:followed by 64 zeros. - The reported
model_hashequals the manifestmodel_hash.
Each failure exits 1 before the search runs. A real mismatch on the quickstart corpus:
Error: model_hash mismatch: corpus was built with sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98, embedder reports sha256:abababababababababababababababababababababababababababababababab
fingerprint reported by embedder: {"fake":true}
hint: --model-path PATH to point at the exact snapshot, or rebuild the corpus with the model you intend to use.ask and retrieve offer no way around the gate. search-text --skip-model-hash-check turns off checks 3 and 4 and keeps 1 and 2. The protocol and the exact messages are in Query embedder protocol.
Verbs that take a raw vector (search, search-ann, search-graph) never see a model, so they have no gate. search-space checks a named space's own model_hash only when you pass --expect-model-hash.
Opt-in in Python
In the Python API the gate is off unless you ask for it.
| Call | Checks model_hash |
|---|---|
UrnaFile.retrieve(query, k, ..., expected_model_hash=None) | only when expected_model_hash is given |
UrnaFile.search_space(name, query, k, expected_model_hash=None) | only when expected_model_hash is given; it compares against that space's hash |
UrnaFile.search, search_ann, search_hybrid, search_graph | never |
Pass the hash your query embedder reports. With the wheel installed as urna[embed]:
import urna
from urna.embed_potion import potion_embedder
db = urna.open("corpus.urna")
emb = potion_embedder()
query = emb.embed_texts(["can I use this offline"])[0]
hits = db.retrieve(query, 3, expected_model_hash=emb.model_hash())On a mismatch, retrieve raises before it parses the query:
ValueError: model_hash mismatch: the query was embedded with ..., but the corpus was built with .... Results would be cosine-valid but semantically wrong. Pass expected_model_hash=None to bypass this check.The comparison is a string compare: there is no name or dim layer in Python, although the runtime still rejects a vector of the wrong length. The FastAPI, Flask and notebook examples in the repository call retrieve without the hash, so the gate is off there. See Use urna from Python.
Placeholders are accepted on write
urna.build checks that model_hash has the sha256: plus 64 hex shape. It does not refuse the all-zero placeholder, so a file can be written, validated and opened with it. The refusal happens only at the CLI's query gate, with this message:
manifest carries the legacy placeholder model_hash (sha256:0000000000000000000000000000000000000000000000000000000000000000). Rebuild this corpus with a real fingerprint, or pass --skip-model-hash-check to proceed at your own risk.A placeholder corpus cannot be asked
A corpus written with the placeholder model_hash passes urna validate and opens in Python, but urna ask and urna retrieve always refuse it, and in Python retrieve with expected_model_hash refuses it too. Only search-text --skip-model-hash-check, the raw-vector verbs and Python calls without expected_model_hash reach it. Pass a real model_hash when you build. See Known limits.
What the gate does not prove
The gate compares two claims: the hash the builder recorded and the hash the query embedder reports. It does not check that the stored vectors were produced by the model the manifest names; urna.build trusts the vectors and the hash it is given. And a hash identifies a model version, it does not vouch for it. For how far the file's own hashes go, see Citations and hashes and Security.
Search paths and the exact rerank
The exact, HNSW, hybrid, graph and named-space search paths in urna, how ask and retrieve pick one, and why every returned score is a real cosine.
Offline by construction
Where urna touches the network and where it cannot: the runtime has no network stack, queries embed offline, and only installers and setup download.