docsv0.5.1

The model gate

Why urna refuses a text query embedded by a different model than the corpus, what model_hash covers, and where the check runs or does not.

A vector search with the wrong query model still returns numbers. The cosine arithmetic is valid, the dimensions may even match, and the ranking is noise. urna records a fingerprint of the embedding model in every corpus, model_hash, and the CLI compares it with the fingerprint the query embedder reports before it searches. A mismatch is an error, never a silently bad answer.

Three hashes, one gate

A build through urna build keys the embeddings of each model by three hashes. Only one of them takes part in the query-time gate.

HashWhat it identifiesWhere it livesUsed for
model_hashthe model: the files and settings that decide what vector a text becomesthe .urna manifest, and the build's <name>.manifest.jsonthe query-time gate
corpus_input_hashthe input rows: each row's text, source image, label and chunker_version, in row orderthe .urna provenance section, and <name>.manifest.jsonkeying the build's embedding cache
embedding_recipe_hashhow the model was driven: query and document modes, image settings, normalize, dtype, device class<name>.manifest.json onlykeying the build's embedding cache

Together the three form the key of the shared embedding cache, so a rebuild reuses vectors only when the model, the input and the recipe are all unchanged. See Reproducible builds.

At query time only model_hash is checked. The recipe is not carried to the query side: ask and retrieve build the query embedder from the preset's defaults, not from overrides in the build spec. Where a recipe setting also enters model_hash (the dtype policy of a sentence-transformers model, for example), the gate catches the difference; where it does not, nothing does.

What model_hash covers

model_hash is sha256: followed by 64 hex digits, computed by the embedder over a canonical JSON fingerprint. What goes into the fingerprint depends on the embedder:

EmbedderFingerprint inputs
potion (the default, minishlab/potion-base-8M/v1)the bytes of config.json, tokenizer.json and model.safetensors, the tokenizer settings, mean pooling, normalize, the dimension (256) and the float32 policy
sentence-transformers presets (jina, wemm)the weight, tokenizer and processor files, the model-repo code files, pooling, normalize and the dtype policy
open_clip presets (clip, siglip2)the model id and pretrained tag, the preprocessing transform, a digest of every weight tensor, and L2 normalization
a sentence-transformers model fingerprinted with python/model_fingerprint.py (checkout)up to ten config, tokenizer, pooling and weight files, the dimension and normalize

The vendored potion table gives sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98. A different table gives a different hash, and corpora built with the old one stop answering. That is the gate doing its job.

For sentence-transformers presets the dtype policy follows the device: bfloat16 on CUDA, float16 on MPS, float32 on CPU. The same weights built on one device class and queried on another report different hashes unless URNA_ST_DTYPE (or the spec's dtype) pins the same dtype on both sides.

What the CLI checks at query time

urna ask, urna retrieve, urna search-text and the ask tab of urna tui run the query embedder, parse its JSON and check, in order:

  1. The reported embedding_model equals the manifest embedding_model.
  2. The reported embedding_dim and the vector length equal the manifest embedding_dim.
  3. The manifest model_hash is not the placeholder, sha256: followed by 64 zeros.
  4. The reported model_hash equals the manifest model_hash.

Each failure exits 1 before the search runs. A real mismatch on the quickstart corpus:

Error: model_hash mismatch: corpus was built with sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98, embedder reports sha256:abababababababababababababababababababababababababababababababab
fingerprint reported by embedder: {"fake":true}
hint: --model-path PATH to point at the exact snapshot, or rebuild the corpus with the model you intend to use.

ask and retrieve offer no way around the gate. search-text --skip-model-hash-check turns off checks 3 and 4 and keeps 1 and 2. The protocol and the exact messages are in Query embedder protocol.

Verbs that take a raw vector (search, search-ann, search-graph) never see a model, so they have no gate. search-space checks a named space's own model_hash only when you pass --expect-model-hash.

Opt-in in Python

In the Python API the gate is off unless you ask for it.

CallChecks model_hash
UrnaFile.retrieve(query, k, ..., expected_model_hash=None)only when expected_model_hash is given
UrnaFile.search_space(name, query, k, expected_model_hash=None)only when expected_model_hash is given; it compares against that space's hash
UrnaFile.search, search_ann, search_hybrid, search_graphnever

Pass the hash your query embedder reports. With the wheel installed as urna[embed]:

import urna
from urna.embed_potion import potion_embedder

db = urna.open("corpus.urna")
emb = potion_embedder()
query = emb.embed_texts(["can I use this offline"])[0]
hits = db.retrieve(query, 3, expected_model_hash=emb.model_hash())

On a mismatch, retrieve raises before it parses the query:

ValueError: model_hash mismatch: the query was embedded with ..., but the corpus was built with .... Results would be cosine-valid but semantically wrong. Pass expected_model_hash=None to bypass this check.

The comparison is a string compare: there is no name or dim layer in Python, although the runtime still rejects a vector of the wrong length. The FastAPI, Flask and notebook examples in the repository call retrieve without the hash, so the gate is off there. See Use urna from Python.

Placeholders are accepted on write

urna.build checks that model_hash has the sha256: plus 64 hex shape. It does not refuse the all-zero placeholder, so a file can be written, validated and opened with it. The refusal happens only at the CLI's query gate, with this message:

manifest carries the legacy placeholder model_hash (sha256:0000000000000000000000000000000000000000000000000000000000000000). Rebuild this corpus with a real fingerprint, or pass --skip-model-hash-check to proceed at your own risk.

A placeholder corpus cannot be asked

A corpus written with the placeholder model_hash passes urna validate and opens in Python, but urna ask and urna retrieve always refuse it, and in Python retrieve with expected_model_hash refuses it too. Only search-text --skip-model-hash-check, the raw-vector verbs and Python calls without expected_model_hash reach it. Pass a real model_hash when you build. See Known limits.

What the gate does not prove

The gate compares two claims: the hash the builder recorded and the hash the query embedder reports. It does not check that the stored vectors were produced by the model the manifest names; urna.build trusts the vectors and the hash it is given. And a hash identifies a model version, it does not vouch for it. For how far the file's own hashes go, see Citations and hashes and Security.

On this page