Reproducible builds
What makes two urna builds byte-identical, the L1, L2 and L3 reproduction levels, the build lock, the embed cache keyed by three hashes, and --rebuild-only.
A .urna build is deterministic: the same rows, the same vectors and the same settings produce the same bytes, and therefore the same file_hash, content_hash and citations. This page covers what the writer pins, what the forge (urna build --spec) records so you can repeat a build, and what you still have to check yourself.
What the writer pins
The Rust writer has no clock, no randomness and no dependency on the CPU it runs on:
- The manifest's
createdfield is the only timestamp. The writer never fills it by itself.reproducible(true)sets it to1970-01-01T00:00:00Z, so a build that passes acreatedvalue still comes out identical. - The HNSW index is built from a fixed seed (
hnsw_seed = 42inurna.build) with a distance function that does not use SIMD, so the index bytes do not depend on the machine. - Sections are sorted by id, aligned to 64 bytes and padded with zeros.
Byte identity still needs identical inputs: the same chunks in the same order, the same provenance JSON with the same key order, the same text encoding, dtype and mrl_dim, and the same HNSW parameters and seed.
| Entry point | reproducible default |
|---|---|
urna.build | False |
builder.BuildConfig (checkout only) | True |
build spec, [corpus] reproducible | true |
With reproducible off and no created value, nothing is written into created and the build is deterministic anyway.
Reproduction levels
The forge documentation declares three levels:
| Level | Claim | Checked by code |
|---|---|---|
| L1 | the same top-k on any machine | no |
| L2 | each vector's cosine within 1e-5 of the original, on the same device class | no |
| L3 | a byte-identical file_hash | partly: the build lock is compared; the byte comparison is yours |
L1 and L2 are statements about the embedding model and hardware, and no tool in 0.5.1 measures them. L3 is the one the forge supports directly, with the build lock and --rebuild-only.
The build lock
Every urna build --spec run writes <name>.build.lock.json next to the outputs:
| Field | Content |
|---|---|
lock_schema_version | 1 |
platform | os, machine, release |
packages | the Python version and the installed versions of numpy, torch, transformers, sentence-transformers, open-clip-torch, tokenizers, pillow |
tools | ffmpeg, ffprobe, cjxl, djxl, ssimulacra2: path, version and SHA-256 of the binary, or null when absent |
models | model_hash per preset |
device | URNA_ST_DEVICE, or auto |
resolved_spec | the full spec after profiles and ${VAR} expansion, without output.cache_dir |
The comparison ignores where things live: resolved_spec.spec_path, output.dir, output.cache_dir, source.path, source.db and source.image.path_template. Row content is covered per item by its input hash, so a rebuild under another data root still compares clean.
Everything else counts. The lock records all five media tools and all seven packages for every build, text-only builds included, so a different ffmpeg or torch version shows up as a divergence even when the build never used it. source.input_dir and source.labels are not in the ignore list, so an image_dir corpus rebuilt from a moved directory diverges on location alone.
The version field of cjxl, djxl and ssimulacra2 holds their usage text instead of a version, because the lock runs <tool> -version, which only ffmpeg and ffprobe accept. The SHA-256 still pins the binary.
The embed cache and its key
Computing embeddings is the slow part of a build, so the forge caches them outside the output directory. The cache key is three hashes, called the triad:
| Hash | Identifies | Changes when |
|---|---|---|
model_hash | the model | weights, tokenizer, processor, remote code, normalize or dtype policy change |
embedding_recipe_hash | how the model is used | query and document modes, image prompt, preprocess version, image_max_side, encode_kwargs, dtype, device class, image input mode, and for image spaces over decoded media the decoder settings |
corpus_input_hash | the rows | any item's text, source image bytes, label or chunker_version changes, or rows are added, removed or reordered |
Each entry is <root>/embed/<preset>/<16 hex>.npz, where the name is a hash of the triad plus which arrays the spec needs (text, image, deduplicated). A checksum sidecar and a lock file sit beside it. Two specs with the same rows and model read the same entry wherever their outputs go, a changed knob adds a new entry instead of overwriting one, and a torn file is recomputed.
The cache root is, in order: --cache-dir, [output] cache_dir, URNA_CACHE_DIR, then ${XDG_CACHE_HOME:-~/.cache}/urna. The location is not part of the lock.
model_hash probes are cached beside the entries, under <root>/models/, so a warm build does not load a model only to learn its hash. A probe is trusted while the model directory's file listing (names and sizes) is unchanged. potion and the open_clip presets have no model directory to list, so swapping their weights is noticed only when an embed actually runs.
Rebuilding with --rebuild-only
--rebuild-only re-emits the outputs from cached vectors and compares the new lock with the one on disk:
urna build --spec corpus.toml --rebuild-onlyWhat it does:
- Loads the rows again and recomputes
corpus_input_hash. - Reuses the encoded media when the saved media state matches. When it does not match, it re-encodes silently.
- Requires every embed cache entry to hit. A miss stops the build:
--rebuild-only: cache for '<preset>' is missing or stale (triad mismatch); run a full build. - Writes the
.urnafiles, overwriting the old ones. - Compares the new lock with the old one and prints
[forge] warning: build.lock divergence (L3 not claimable): ...when they differ. - Writes the new lock and the manifest, overwriting the old ones.
It does not compare file_hash. To claim L3, keep the old hash and compare it yourself. The reported file_hash is the SHA-256 of the whole file, the same value sha256sum prints (shasum -a 256 on macOS):
cp out/corpus.build.lock.json /tmp/corpus.lock.before
sha256sum out/corpus.urna > /tmp/corpus.sha256
urna build --spec corpus.toml --rebuild-only
sha256sum -c /tmp/corpus.sha256Keep a copy of the lock too: after a divergence warning the new lock replaces the old one.
--strict-env is not an urna build flag
Making a lock divergence an error needs --strict-env, which exists only on the Python tool: python python/tools/urna_forge.py --spec corpus.toml --rebuild-only --strict-env. urna build --strict-env is a usage error (exit 2). The same goes for --seed and --json.
--resume is narrower than the name suggests: only the media stage reads saved state. Embed caches are read on every build, with or without it, and rows and outputs are always recomputed.
What moves citations anyway
A reproducible build reproduces its citations, but several ordinary edits change them:
chunker_versionis part of everychunk_id: changing it changes every citation.- In a forge build, a row's span is its position after sorting by
order_by. Adding or removing a row shifts every later row'schunk_id. order_bysorts by string value, so"10"sorts before"2".--sample Nrenumbers rows, so a pilot build never shares citations with the full build.- The preset's dtype,
mrl_dimand index type are part ofcontent_hash.
Media encoders are pinned to fixed parallelism (lp=2 for SVT-AV1, -j 8 for avifenc) so the encoded bytes do not depend on the core count. For st_multimodal models, the dtype policy follows the device unless you pin it, so building on cuda and on mps gives different model_hash values: see Model registry.
See Citations and hashes for what each hash covers, and Build artifacts for the full lock and manifest schemas.
Presets and stored precision
How urna stores vectors as float32, float16, int8 or int4, what that costs in recall, and how every score discloses its rerank precision.
Media and named spaces
How a .urna file carries media blobs and extra named vector spaces, each space with its own model_hash and dim, and how to query one.