Build artifacts
Reference for what urna build writes: the .urna files, manifest.json, build.lock.json, the shared embed cache, reproduction levels and every build error.
This page describes every file urna build writes, the embed cache it shares across builds, the reproduction levels the build lock supports, and the catalog of build errors.
Files
A build writes into [output] dir (or --out-dir). For a corpus named handbook with media on the av1 backend:
| Path | Written when | Content |
|---|---|---|
<name>.urna | mode is single or both | The corpus: every chunk, space 0, every named space, the indexes and the media references |
<name>-<preset>.urna | mode is per-model or both | Space 0 plus that preset's named spaces, with the same chunks and media tables |
<name>.manifest.json | Every build | The build record, see below |
<name>.build.lock.json | Every build, except a --models run when a lock already exists | The environment record, see below |
<name>.media/ | The media stage ran | The encoded media. Without embed_media this is what the corpus serves; with it, a build cache. Layouts per backend are on the [media] page |
.forge-state/media.json | The media stage ran | Media state for --resume: a hash of the rows and the [media] settings, the media record and the frame URIs |
.tmp/ | Every build | Staging for .urna files; empty after a successful build |
Each .urna is written into .tmp/, moved into place with an atomic rename (same filesystem), reopened and validated before the build continues. The per-model files have no manifest or media directory of their own: they share <name>.manifest.json and <name>.media/, so a tool that derives those names from a per-model file's name finds nothing.
The .urna file
Each chunk is one row: its stored text, its source_uri, and the span [ordinal, ordinal + 1). Space 0 holds the text = "default" model's vectors at the [build] preset. The file's provenance records {dataset: <name>, corpus_input_hash}, and its manifest takes title and version from [corpus].
Depending on the spec the file also has the chunk graph (with_graph, on by default), the space table and one band per named space, and, with media, the blob references and a span overlay that points each chunk at its frame or file. With embed_media = true the media bytes are stored in the file too. The graph, the space table and bands, and the media tables are outside the content_hash: two files with the same chunks and the same space 0 have the same content_hash, whatever named spaces or media tables they add. The sections are described on layout.
<name>.manifest.json
The build record, manifest_schema_version 1. It is written atomically as JSON with sorted keys and one-space indentation, and the writer refuses to save it without its required fields. It is separate from the manifest stored inside the .urna.
| Field | Content |
|---|---|
manifest_schema_version | 1. Readers check it first |
name | The corpus name |
n_items | Rows in the build |
n_unique_frames | Unique frames after deduplication |
corpus_input_hash | sha256: hash over the rows' input hashes |
text_quality | template when [source.text] template is set, else identity-only |
models | Per preset: model_hash, embedding_recipe_hash, the full recipe, and items_per_s when the model embedded in this run |
spaces | Per named space: name, preset, modality, dim (null for native) and space_dtype |
media | null, or the backend record (below) plus dedup, the ordering, crf_auto and embedded |
outputs | Per .urna: file (redacted by provenance), bytes, file_hash, build_s |
timings | Seconds for rows, media and embed.<preset> |
sql | The source query under full, "" under standard, absent under minimal |
items | Per row under standard and full: ordinal, key, label, media_uri, image_path. Under minimal: key and ordinal only |
items_compact | true under minimal, else false |
frame_of_row | Under minimal only, when deduplication merged rows: the frame index of each row |
provenance_mode | The [output] provenance value |
An excerpt of the quickstart manifest:
{
"corpus_input_hash": "sha256:00ebc2f0e96cbbf0effd0ce5b38a80f8b8eb13c972fdcde4968affa21caaa8bb",
"items": [
{
"image_path": null,
"key": "01-what-is-urna#1",
"label": null,
"media_uri": null,
"ordinal": 0
}
],
"items_compact": false,
"manifest_schema_version": 1,
"media": null,
"models": {
"potion": {
"embedding_recipe_hash": "sha256:f5c4ef99c235a02f1185b0830a16af090c5e3fdb28e6b20f2863786841539d45",
"items_per_s": 6270.29,
"model_hash": "sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98",
"recipe": {
"adapter_version": 1,
"device_class": "auto",
"encode_kwargs": {},
"image_doc_format": "dict",
"image_input_mode": "source",
"image_max_side": 0,
"image_mode": "image_document",
"image_prompt": "Represent this image.",
"model_dtype": "",
"normalize": true,
"preprocess_version": "processor-native-v1",
"preset": "potion",
"text_corpus_mode": "plain",
"text_query_mode": "plain"
}
}
},
"n_items": 12,
"n_unique_frames": 12,
"name": "quickstart",
"outputs": {
"quickstart.urna": {
"build_s": 0.056,
"bytes": 17942,
"file": "out/quickstart.urna",
"file_hash": "sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832"
}
},
"provenance_mode": "standard",
"spaces": [],
"sql": "",
"text_quality": "template",
"timings": {
"embed.potion": 0.002,
"rows": 0.001
}
}The quickstart has twelve items; the excerpt keeps the first. The file path is relative to the spec's directory because provenance is standard.
The media record
The record depends on the backend. Every backend records source_bytes as the size of the original files, so compression_ratio compares across backends.
| Backend | Keys |
|---|---|
av1, one stream | backend, codec (libsvtav1), crf, preset (the speed), fps, tune, keyint, pix_fmt (probed from the output), canvas, frame_count, source_bytes, output_bytes, compression_ratio, media_sha256, toolchain (ffmpeg and encoder versions and every encoder parameter, including tune_resolved), provenance_sha256, segments[] (uri, start_frame, n_frames, output_bytes, media_sha256), gop |
av1, sharded | As above, without the top-level fps, tune and media_sha256; with shard_size, a keyint per segment (top-level keyint is null when segments differ) and gop with per_segment: true, the overall decision (intra, inter or mixed) and each segment's probe |
avif | backend, codec (avifenc), quality, speed, yuv, frame_count, source_bytes, letterboxed_input_bytes, output_bytes, compression_ratio, toolchain, provenance_sha256, canvas |
jxl, jxl-transcode | backend, canvas (null), frame_count, files[], decisions[], source_bytes, output_bytes, compression_ratio, toolchain, provenance_sha256 |
control | backend (png-lossless), canvas, frame_count, output_bytes |
The build adds dedup (n_items, n_unique_frames), order and order_permutation when av1 applied an ordering, crf_auto (the crf gate report) when crf = "auto", and embedded (the embed_media value).
Reading items
Under minimal, items holds only key and ordinal. The forge's reader forge_manifest.manifest_items(manifest, need=(...)) raises an error naming provenance = "standard" when a caller asks for a dropped field, and forge_manifest.frame_resolver(manifest) maps an item to its frame in the media stream in every mode. Both live in the checkout.
<name>.build.lock.json
The environment record behind a byte-identical rebuild, lock_schema_version 1.
| Field | Content |
|---|---|
lock_schema_version | 1 |
platform | os, machine, release |
packages | The python version and the installed versions of numpy, torch, transformers, sentence-transformers, open-clip-torch, tokenizers and pillow (packages that are not installed are left out) |
tools | For ffmpeg, ffprobe, cjxl, djxl and ssimulacra2: path, version and sha256 of the binary, or null when not on PATH |
models | model_hash per preset |
device | URNA_ST_DEVICE, or auto when unset. The per-model device of the spec is not recorded here |
resolved_spec | The whole spec after profiles and ${VAR} expansion, without output.cache_dir |
models_subset | Only for a --models run that found no existing lock: the presets it built |
The tool version comes from running the tool with -version. Only ffmpeg and ffprobe accept that flag; for cjxl, djxl and ssimulacra2 the field holds their usage or error text, such as Unknown argument: -version. The sha256 still identifies the binary.
The lock records every tool and package on every build, text-only builds included.
What --rebuild-only compares
--rebuild-only builds a new lock and compares it with the one on disk, key by key. It ignores exactly these keys, because they only say where things live: resolved_spec.spec_path, resolved_spec.output.dir, resolved_spec.output.cache_dir, resolved_spec.source.path, resolved_spec.source.db and resolved_spec.source.image.path_template. Everything else counts: a new OS release, a changed tool binary or package version, a different URNA_ST_DEVICE. source.input_dir and source.labels are not ignored, so an image_dir corpus rebuilt under a moved data root diverges on location alone.
Each difference prints [forge] warning: build.lock divergence (L3 not claimable): ... (up to eight listed). With the Python tool's --strict-env it is an error instead (exit 1).
Shared embed cache
Vectors are cached per model under the cache root (by default ~/.cache/urna), shared by every spec and output directory:
~/.cache/urna/
embed/potion/76370c3d6d796e47.npz
embed/potion/76370c3d6d796e47.npz.sha256
embed/potion/76370c3d6d796e47.npz.lock
models/model_hash.potion.efe05bc4cb560530.jsonembed/<preset>/<16 hex>.npzholds one model's vectors for one set of rows:text(one vector per row) and/orimage_unique(one per unique frame), plusmetawith the three hashes below. The file name is the first 16 hex digits of a hash overmodel_hash,embedding_recipe_hash,corpus_input_hashand which arrays the spec needs (text, image, dedup). A changed knob adds a new entry beside the old one.- A cache entry is used only when the three hashes in
metamatch and the.sha256sidecar matches the file. Writes take an exclusive lock on the.lockfile and land through a temporary file and a rename; a torn or mismatched entry is recomputed, never reused. models/model_hash.<preset>.<16 hex>.jsoncaches a model'smodel_hashso a cache hit does not need to load the model. It is keyed by the preset plusnormalize,dtype,deviceandmodel_path, and holdsmodel_hash, a fingerprint of the model directory's file listing, and those knobs. It is trusted while the directory listing matches.potion,open_clippresets andfake-testhave no model directory to fingerprint, so replacing their weights does not invalidate the cached hash until the model is loaded again. When a loaded model's hash disagrees, the build rewrites the file.
The recipe hashed into embedding_recipe_hash holds adapter_version, preset, image_input_mode, text_corpus_mode, text_query_mode, image_mode, image_prompt, normalize, preprocess_version, image_max_side, image_doc_format, encode_kwargs, model_dtype and device_class. For image-space models embedding decoded media it also holds the decoder settings (backend, canvas, crf, pixel format, toolchain hash), so a text-only model's entry survives a change of media crf.
Deleting the cache is safe: the next build recomputes it.
Reproduction levels
The build declares three levels of reproduction:
| Level | Claim | What supports it |
|---|---|---|
| L1 | The same top-k results on any machine | Declared; no code checks it |
| L2 | Per-vector cosine within 1e-5 on the same device class | Declared; no code checks it |
| L3 | A byte-identical file_hash | Claimable only under a matching build lock |
The build supports L3 with reproducible = true (the default, which pins the file's creation time), fixed encoder parallelism, and --rebuild-only, which re-emits from the cached vectors and compares the lock. The byte comparison is left to you: record file_hash from the first build's result or manifest and compare it after the rebuild. See reproducible builds.
Error catalog
How a failure surfaces depends on when it happens:
| Kind | Printed as | Exit |
|---|---|---|
| Spec error | spec error: <message> on stderr | 2 |
| Registry error | registry error: <message> on stderr | 4 |
| Any other failure | A Python traceback | 1 |
At parse time
All exit 2, before anything else runs:
unsupported spec extension '<ext>' (use .toml or .json)yaml specs need pyyaml: pip install pyyaml (or use .toml)<key>: ${NAME} is not set; export NAME=/path (...), oris emptyunknown top-level table(s) [...]: valid tables are [...]models must be an array of tables; write [[models]], not [models]<table>: unknown key '<k>' (valid: ...)media.profile: unknown '<p>' (valid: [...])
At validation
Exit 2, also under --dry-run: every rule on the [corpus], [source], [[models]], [media] and [output] pages. An unknown preset, or fake-test without URNA_ENABLE_FAKE_PRESET=1, is a registry error (exit 4) even here.
While the build runs
Spec errors raised after work may already be done (exit 2):
| Message | Cause |
|---|---|
source.derive: unsupported expression ..., unknown helper ... | A [source.derive] entry that is not basename_stem(column) or lower(column) |
source.order_by [...] is not a total order ... | Two rows share the order_by key |
source.image: file missing for key ... | A rendered path_template that is not a file |
source.joins: join query has no column ..., 0 of N rows matched ... | A join key missing from the join query, or matching no row |
source.kind=pdf_dir: build via forge_pipeline (pages are temporary) | Every pdf_dir build |
output.cache_dir: cache root ... | A cache root that cannot be created or is not a directory |
media.quality.utility_floor_hit1: needs one label per item | The utility check without a label for every image |
media.quality.utility_floor_hit1: gate model has no text tower ..., cannot embed text ..., text tower dim X != image dim Y | A gate model that cannot run the utility check |
media.quality.utility_query_template: only {label} is a placeholder ... | Another placeholder in the query template |
Registry errors (exit 4):
unknown model preset '<p>'. valid presets: ...preset 'fake-test' requires URNA_ENABLE_FAKE_PRESET=1 (test-only)preset '<p>' needs the '<module>' package. install with: <pip line>- A pinned model-repository file that is missing (
pinned code file missing: <file>) or altered (... does not match the pinned allowlist (...). refusing to execute unreviewed remote code.) - A model worker that fails or dies, including a bad
[[models]] dtype
Build errors, each ending a traceback (exit 1):
source produced no rowsmedia enabled but N rows have no image (first: <key>)--rebuild-only: cache for '<p>' is missing or stale (triad mismatch); run a full buildbuild.lock divergence (L3 not claimable): ..., only with--strict-envon the Python toolmedia gate/cluster model '<p>' is not a spec model
Other failures with a traceback (exit 1): missing media tools (media.crf="auto" needs ssimulacra2 on PATH: brew install jpeg-xl, jxl backend needs 'cjxl' on PATH), encoder failures (encoder wrote yuv420p for a requested yuv444p ..., encoded N frames for M images ...), SQLite errors, a missing CSV or JSONL file, values that urna.build rejects ([build] preset, dtype, mrl_dim), value types the spec does not check, the crf = "auto" crash on rows without a label, and a sentence-transformers preset with no weights on disk.
[build] and [output]
Reference for [build] and [output] in a urna build spec: the engine preset, dtype, graph and mrl_dim of space 0, output modes, provenance and the cache.
Model registry
Every embedding model preset in the urna forge registry: model id, dims, MRL ladder, dependencies, device rules, model_hash and query contract.