docsv0.5.1

Reproducible builds

What makes two urna builds byte-identical, the L1, L2 and L3 reproduction levels, the build lock, the embed cache keyed by three hashes, and --rebuild-only.

A .urna build is deterministic: the same rows, the same vectors and the same settings produce the same bytes, and therefore the same file_hash, content_hash and citations. This page covers what the writer pins, what the forge (urna build --spec) records so you can repeat a build, and what you still have to check yourself.

What the writer pins

The Rust writer has no clock, no randomness and no dependency on the CPU it runs on:

  • The manifest's created field is the only timestamp. The writer never fills it by itself. reproducible(true) sets it to 1970-01-01T00:00:00Z, so a build that passes a created value still comes out identical.
  • The HNSW index is built from a fixed seed (hnsw_seed = 42 in urna.build) with a distance function that does not use SIMD, so the index bytes do not depend on the machine.
  • Sections are sorted by id, aligned to 64 bytes and padded with zeros.

Byte identity still needs identical inputs: the same chunks in the same order, the same provenance JSON with the same key order, the same text encoding, dtype and mrl_dim, and the same HNSW parameters and seed.

Entry pointreproducible default
urna.buildFalse
builder.BuildConfig (checkout only)True
build spec, [corpus] reproducibletrue

With reproducible off and no created value, nothing is written into created and the build is deterministic anyway.

Reproduction levels

The forge documentation declares three levels:

LevelClaimChecked by code
L1the same top-k on any machineno
L2each vector's cosine within 1e-5 of the original, on the same device classno
L3a byte-identical file_hashpartly: the build lock is compared; the byte comparison is yours

L1 and L2 are statements about the embedding model and hardware, and no tool in 0.5.1 measures them. L3 is the one the forge supports directly, with the build lock and --rebuild-only.

The build lock

Every urna build --spec run writes <name>.build.lock.json next to the outputs:

FieldContent
lock_schema_version1
platformos, machine, release
packagesthe Python version and the installed versions of numpy, torch, transformers, sentence-transformers, open-clip-torch, tokenizers, pillow
toolsffmpeg, ffprobe, cjxl, djxl, ssimulacra2: path, version and SHA-256 of the binary, or null when absent
modelsmodel_hash per preset
deviceURNA_ST_DEVICE, or auto
resolved_specthe full spec after profiles and ${VAR} expansion, without output.cache_dir

The comparison ignores where things live: resolved_spec.spec_path, output.dir, output.cache_dir, source.path, source.db and source.image.path_template. Row content is covered per item by its input hash, so a rebuild under another data root still compares clean.

Everything else counts. The lock records all five media tools and all seven packages for every build, text-only builds included, so a different ffmpeg or torch version shows up as a divergence even when the build never used it. source.input_dir and source.labels are not in the ignore list, so an image_dir corpus rebuilt from a moved directory diverges on location alone.

The version field of cjxl, djxl and ssimulacra2 holds their usage text instead of a version, because the lock runs <tool> -version, which only ffmpeg and ffprobe accept. The SHA-256 still pins the binary.

The embed cache and its key

Computing embeddings is the slow part of a build, so the forge caches them outside the output directory. The cache key is three hashes, called the triad:

HashIdentifiesChanges when
model_hashthe modelweights, tokenizer, processor, remote code, normalize or dtype policy change
embedding_recipe_hashhow the model is usedquery and document modes, image prompt, preprocess version, image_max_side, encode_kwargs, dtype, device class, image input mode, and for image spaces over decoded media the decoder settings
corpus_input_hashthe rowsany item's text, source image bytes, label or chunker_version changes, or rows are added, removed or reordered

Each entry is <root>/embed/<preset>/<16 hex>.npz, where the name is a hash of the triad plus which arrays the spec needs (text, image, deduplicated). A checksum sidecar and a lock file sit beside it. Two specs with the same rows and model read the same entry wherever their outputs go, a changed knob adds a new entry instead of overwriting one, and a torn file is recomputed.

The cache root is, in order: --cache-dir, [output] cache_dir, URNA_CACHE_DIR, then ${XDG_CACHE_HOME:-~/.cache}/urna. The location is not part of the lock.

model_hash probes are cached beside the entries, under <root>/models/, so a warm build does not load a model only to learn its hash. A probe is trusted while the model directory's file listing (names and sizes) is unchanged. potion and the open_clip presets have no model directory to list, so swapping their weights is noticed only when an embed actually runs.

Rebuilding with --rebuild-only

--rebuild-only re-emits the outputs from cached vectors and compares the new lock with the one on disk:

urna build --spec corpus.toml --rebuild-only

What it does:

  1. Loads the rows again and recomputes corpus_input_hash.
  2. Reuses the encoded media when the saved media state matches. When it does not match, it re-encodes silently.
  3. Requires every embed cache entry to hit. A miss stops the build: --rebuild-only: cache for '<preset>' is missing or stale (triad mismatch); run a full build.
  4. Writes the .urna files, overwriting the old ones.
  5. Compares the new lock with the old one and prints [forge] warning: build.lock divergence (L3 not claimable): ... when they differ.
  6. Writes the new lock and the manifest, overwriting the old ones.

It does not compare file_hash. To claim L3, keep the old hash and compare it yourself. The reported file_hash is the SHA-256 of the whole file, the same value sha256sum prints (shasum -a 256 on macOS):

cp out/corpus.build.lock.json /tmp/corpus.lock.before
sha256sum out/corpus.urna > /tmp/corpus.sha256
urna build --spec corpus.toml --rebuild-only
sha256sum -c /tmp/corpus.sha256

Keep a copy of the lock too: after a divergence warning the new lock replaces the old one.

--strict-env is not an urna build flag

Making a lock divergence an error needs --strict-env, which exists only on the Python tool: python python/tools/urna_forge.py --spec corpus.toml --rebuild-only --strict-env. urna build --strict-env is a usage error (exit 2). The same goes for --seed and --json.

--resume is narrower than the name suggests: only the media stage reads saved state. Embed caches are read on every build, with or without it, and rows and outputs are always recomputed.

What moves citations anyway

A reproducible build reproduces its citations, but several ordinary edits change them:

  • chunker_version is part of every chunk_id: changing it changes every citation.
  • In a forge build, a row's span is its position after sorting by order_by. Adding or removing a row shifts every later row's chunk_id.
  • order_by sorts by string value, so "10" sorts before "2".
  • --sample N renumbers rows, so a pilot build never shares citations with the full build.
  • The preset's dtype, mrl_dim and index type are part of content_hash.

Media encoders are pinned to fixed parallelism (lp=2 for SVT-AV1, -j 8 for avifenc) so the encoded bytes do not depend on the core count. For st_multimodal models, the dtype policy follows the device unless you pin it, so building on cuda and on mps gives different model_hash values: see Model registry.

See Citations and hashes for what each hash covers, and Build artifacts for the full lock and manifest schemas.

On this page