docsv0.5.1

urna build

Reference for urna build, the launcher of the declarative corpus build: flags, how it finds the forge and Python, stages, output and exit codes.

urna build builds .urna files from a spec file (see the spec file). It is a launcher: it finds the Python build tool python/tools/urna_forge.py in a repository checkout, runs it with the flags you pass, streams its output and exits with its exit code.

Usage

urna build [OPTIONS] --spec <SPEC>

Arguments

None. The spec file is passed with --spec.

Options

OptionDefaultDescription
--spec <SPEC>requiredBuild spec path (.toml or .json)
--sample <SAMPLE>all rowsEvenly-spaced row subset for pilots
--models <MODELS>all modelsComma-separated preset subset (e.g. "potion,wemm-2b")
--out-dir <OUT_DIR>[output] dirOverride the spec's [output].dir
--cache-dir <CACHE_DIR>see descriptionOverride the spec's [output].cache_dir (the shared embed cache root; else URNA_CACHE_DIR, else ${XDG_CACHE_HOME:-~/.cache}/urna)
--resumeoffResume from per-stage state after an interrupted build
--rebuild-onlyoffRe-emit byte-identically from cached vectors (L3 check)
--dry-runoffResolve the plan + dependency status without loading models
--allow-heavyoffAllow presets flagged too heavy for this machine (wemm-4b/9b)
-h, --helpPrint help

What each option does

  • --spec: the forge also reads .yaml and .yml specs when pyyaml is installed. Any other extension is a spec error. A spec path that does not exist fails with a Python traceback and exit 1.
  • --sample N: keeps rows at evenly spaced positions (int(i * rows / N) for i in 0 to N - 1) and renumbers them. It is deterministic. 0, or a value at least the row count, keeps every row. A row's position is part of its citation, so a sampled build never shares citations with the full build. A dry run ignores it.
  • --models a,b: keeps only the named presets, then validates the spec again, so the text = "default" model must stay in the list. Names that match no model are ignored. The run rewrites <name>.urna and <name>.manifest.json in the output directory; an existing <name>.build.lock.json is left untouched. A dry run ignores it.
  • --out-dir: replaces [output] dir before validation. Point pilots at their own directory.
  • --cache-dir: replaces [output] cache_dir. The location never enters the build lock.
  • --resume: reuses the media encoded by an earlier run when .forge-state/media.json matches the current rows and [media] settings and every media file verifies (stream segments by SHA-256, per-image files non-empty). Rows are always reloaded and the files always written again.
  • --rebuild-only: reuses media like --resume (and encodes again when the state does not match), requires a cache hit for every model, writes the outputs again, then compares the new build lock with the one on disk and prints [forge] warning: build.lock divergence (L3 not claimable): ... on stdout for differences. It overwrites the outputs, the manifest and the lock. It does not compare file_hash; keep the old value and compare it yourself.
  • --dry-run: parses and validates the spec, then prints the plan. It loads no model, never opens the source and creates no directory.
  • --allow-heavy: lets validation pass wemm-4b and wemm-9b.

--strict-env, --seed and --json exist only on the Python tool. urna build --strict-env is a usage error (error: unexpected argument '--strict-env' found, exit 2). To turn lock divergence into an error, run the tool directly from the repository root with an interpreter that has the build's packages:

python python/tools/urna_forge.py --spec corpus.toml --rebuild-only --strict-env

On the Python tool, --json prints the dry-run plan as JSON and --seed (default 42) has no effect.

urna build needs a repository checkout

The final screen of urna setup and its --yes output suggest urna build --spec corpus.toml, and the not-found error says "install the forge payload". No release artifact ships urna_forge.py and there is no forge payload to install: urna build works only from a checkout. Outside one it exits 1:

Error: urna_forge.py not found (python/tools/urna_forge.py); run from the repo or install the forge payload

See Known limits.

How it works

Finding the forge

The launcher looks for urna_forge.py in this order and uses the first that exists:

  1. python/tools/urna_forge.py under the current directory.
  2. The same path under the parent of the current directory.
  3. The same path under the checkout of a binary you built yourself at <repo>/target/<profile>/urna.
  4. urna/tools/urna_forge.py under each data root: URNA_DATA_DIR, XDG_DATA_HOME, ~/.local/share, %LOCALAPPDATA%, then <exe>/../share. No installer writes a tools/ directory there.

Choosing Python

The interpreter is URNA_PYTHON when set, else the venv created by urna setup (<data root>/urna/venv/bin/python, or Scripts\python.exe on Windows, in the first data root that has one), else the first .venv/bin/python found in the current directory or its three nearest parents, else python3. The choice is printed on stderr:

[urna] embedder interpreter: /home/you/.local/share/urna/venv/bin/python

The forge needs Python 3.11 or later for tomllib, and the build's final stage imports the urna extension, which needs 3.12. See paths and resolution order.

Stages

The forge validates the spec, then runs:

  1. Filter models with --models and validate again.
  2. Create the output directory with .forge-state/ and .tmp/, and the cache root. A cache root that cannot be created is a spec error naming output.cache_dir.
  3. Load the rows (never cached) and compute corpus_input_hash.
  4. Deduplicate: rows whose source images are byte-identical share one frame.
  5. Media, only when [media] is set and rows have images: the crf gate, frame ordering, encoding, and the state file.
  6. Embed, per model: load cached vectors or compute and cache them.
  7. Emit: build each output file into .tmp/, move it into place, reopen it and validate it.
  8. Write the build lock (compared under --rebuild-only) and the manifest.

Importing the forge sets HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1 and HF_DATASETS_OFFLINE=1 when they are unset, so a build never downloads model weights.

Output

The launcher prints nothing of its own except the interpreter line on stderr. Everything else is the forge's stdout and stderr, streamed unchanged.

Dry run

urna build --spec examples/quickstart/corpus.toml --dry-run
[urna] embedder interpreter: /home/you/.local/share/urna/venv/bin/python
corpus: quickstart  (source=jsonl, image_input=source, output=single)
model:  potion               text=default image=none  dims=- deps=ok
spaces: (none)
files:  quickstart.urna
out:    /home/you/urna/examples/quickstart/out
cache:  /home/you/.cache/urna  (embed/<preset>/<triad>.npz, shared)

A spec with [media] adds a media: line with the backend, crf, tune, speed, order and dedup. A model whose packages are missing shows deps=MISSING and the fix on the next line:

model:  clip-vit-b32         text=none    image=space dims=- deps=MISSING
        -> preset 'clip-vit-b32' needs the 'torch' package. install with: pip install torch

A preset that runs model-repository code and is listed in allow_remote_code shows remote-code:ok. The dry run still exits 0 when a dependency is missing.

Build result

On success the forge prints one JSON object on stdout. For the quickstart spec:

{
 "outputs": {
  "quickstart.urna": {
   "file": "examples/quickstart/out/quickstart.urna",
   "bytes": 17942,
   "file_hash": "sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832",
   "build_s": 0.056
  }
 },
 "n_items": 12,
 "n_unique_frames": 12,
 "corpus_input_hash": "sha256:00ebc2f0e96cbbf0effd0ce5b38a80f8b8eb13c972fdcde4968affa21caaa8bb",
 "timings": {
  "rows": 0.001,
  "embed.potion": 0.002
 },
 "manifest": "examples/quickstart/out/quickstart.manifest.json",
 "build_lock": "examples/quickstart/out/quickstart.build.lock.json"
}
FieldContent
outputsOne entry per written .urna: path as written (not redacted), size in bytes, file_hash, seconds spent writing it
n_itemsRows in the build
n_unique_framesUnique frames after deduplication (equal to n_items without images)
corpus_input_hashHash over every row's input hash, in row order
timingsSeconds for rows, media (when it ran) and embed.<preset> (0.0 on a cache hit)
manifest, build_lockPaths of <name>.manifest.json and <name>.build.lock.json

Lines starting with [forge] can come before the JSON on stdout: partial join counts, the crf gate fallback warning, the lock divergence warning and the --models note. Encoder warnings go to stderr. A script should parse the last JSON object on stdout.

Exit codes

CodeMeaning
0Build succeeded, or the dry run succeeded (also when dependencies are missing)
1Uncaught error, with a Python traceback: build errors, missing media tools, encoder failures, SQLite errors, rejected [build] values, a wrong value type in the spec. From the launcher: urna_forge.py not found, Python could not be started, or the tool was killed by a signal
2Spec error, printed as spec error: <message> on stderr; also a usage error from urna build or the Python tool
4Model registry error, printed as registry error: <message>: unknown preset, missing Python package, a pinned model file that does not match, a model worker failure

The messages behind each code are listed in the error catalog. All exit codes of the CLI are on exit codes.

Environment

VariableEffect on urna build
URNA_PYTHONInterpreter that runs the forge
URNA_DATA_DIRThe first data root searched, both for urna/tools/urna_forge.py and for the setup venv
URNA_CACHE_DIREmbed cache root when neither --cache-dir nor [output] cache_dir is set
XDG_CACHE_HOMEBase of the default cache root, ${XDG_CACHE_HOME:-~/.cache}/urna
URNA_MODEL_DIR_<NAME>Model directory for a preset, for example URNA_MODEL_DIR_WEMM_2B
URNA_ST_DEVICE, URNA_ST_DTYPEDevice and dtype of sentence-transformers presets when the spec does not set them

Every variable the CLI reads is on environment variables.

Examples

Check a spec and the model dependencies without loading anything:

urna build --spec corpus.toml --dry-run

Build a 500-row pilot into its own directory:

urna build --spec corpus.toml --sample 500 --out-dir out/pilot

Build with a cache on another disk:

urna build --spec corpus.toml --cache-dir /mnt/cache/urna

Re-emit from cached vectors and compare the build lock:

urna build --spec corpus.toml --rebuild-only

A task-oriented walk-through is in Build from your own rows.

On this page