docsv0.5.1

Build from your own rows

Build a .urna from SQLite, CSV, JSONL or an image folder with urna build --spec: sources, models, dry runs, pilots, resume and common errors.

urna build --spec turns a declarative spec file into one or more .urna files. This guide covers the tasks around it: preparing the machine, describing your source, choosing models, checking the plan, running pilots, resuming and fixing the common errors. For a first walk-through with four JSONL rows, start with Your first corpus.

Prepare the machine

The build tool (the forge) is Python code in the repository, and urna build is a launcher that runs it. No release artifact ships the forge, so every build runs from a checkout:

  • The repository's python/ tree, with the Python extension built to python/_urna.so (see Your first corpus). The extension is imported only when the file is written, so a missing one fails after the rows are embedded.
  • Python 3.12 or later, with numpy, plus tokenizers for potion and pillow for anything with images.
  • The Python packages of each model you use. A dry run lists what is missing, with the pip install line to fix it.
  • For media: ffmpeg and ffprobe with libsvtav1 (the av1 backend), avifenc and avifdec (avif), cjxl and djxl (jxl, jxl-transcode), and ssimulacra2 for crf = "auto".

Run urna build from the repository root, or from a directory whose parent is the root; the launcher looks for python/tools/urna_forge.py there. A binary you built yourself at target/release/urna also finds its own checkout from any directory. It picks the Python interpreter in this order: URNA_PYTHON, the venv created by urna setup, the nearest .venv, then python3, and prints the choice on stderr. See paths and resolution order.

The forge never downloads a model. It sets HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1 and HF_DATASETS_OFFLINE=1 when they are unset, so model weights must already be on disk.

Write the spec

A spec is TOML (or JSON) with up to seven tables. The smallest working one needs [corpus], [source] and one [[models]]:

[corpus]
name = "support"
chunker_version = "support/1"

[source]
kind = "csv"
path = "${DATA}/tickets.csv"
order_by = ["ticket_id"]

[source.text]
template = "{subject}\n{body}"

[[models]]
preset = "potion"
text = "default"

Three rules catch most first specs:

  • Relative paths resolve against the directory you run the build from, not the spec's directory. Write data roots as ${VAR} and export the variable, so the same spec works on every machine. An unset or empty variable stops the build with an error naming the key.
  • order_by must be unique per row. It fixes the row order and the row key, and the row's position is part of its citation.
  • Unknown keys are errors. A typo such as dirr under [output] fails with spec error: output: unknown key 'dirr' and the list of valid keys.

The full key list is in the spec reference.

Describe the source

SQLite

[source]
kind = "sqlite"
db = "${DATA}/shop.sqlite"
query = "SELECT sku, name, description, photo_url FROM products"
order_by = ["sku"]

[[source.joins]]
query = "SELECT sku, category FROM categories"
on = "sku"

[source.derive]
photo_stem = "basename_stem(photo_url)"

[source.image]
path_template = "${DATA}/photos/{photo_stem}.jpg"
label_template = "{name}"

The database opens read-only. The query needs no ORDER BY, because the forge sorts on order_by itself. Each [[source.joins]] query runs on the same connection and adds its columns to the main rows that share the on value; the main query wins a name clash. A join that matches no row stops the build, and a partial match prints the count. [source.derive] adds computed columns with two helpers, basename_stem(column) and lower(column), before sorting and templating.

CSV and JSONL

kind = "csv" reads a file with a header row; every value is a string. kind = "jsonl" reads one JSON object per non-blank line. Both take path, order_by, [source.text], [source.derive] and [source.image].

Images in a folder

[source]
kind = "image_dir"
input_dir = "${DATA}/scans"
labels = "${DATA}/labels.csv"   # optional: image_id,label

Every .jpg, .jpeg, .png, .bmp, .webp, .tiff and .tif file under input_dir becomes one row, sorted by relative path. The stored text is the relative path, followed by the label in brackets when there is one. order_by, templates, joins and derive do not apply.

kind = "pdf_dir" passes validation but always fails when rows load. See Images and PDFs for the PDF path that works today.

The text that gets stored

[source.text] template is a Python format string over the row. A template line whose placeholders are all empty is dropped, so optional fields do not leave blank lines. Without a template, a row's stored text is its label, or its key. That stored text is what ask, retrieve and cite return.

Store the images

Rows with images can also carry the encoded images. Add a [media] table and the build encodes each unique image once, writes it to <name>.media/ beside the .urna, and links every chunk to its frame; [output] embed_media = true also stores the bytes inside the .urna, so the file serves them alone (the media directory then stays only as a build cache). Pick a starting point with profile (stills, retrieval, archive and others). Tune media compression explains the choices and the media reference lists every knob.

Choose models

Each [[models]] entry names a preset from the model registry and gives it roles:

  • text = "default": this model embeds every row into the file's main vectors (space 0). Exactly one model has this role.
  • text = "space" or image = "space": this model adds a named vector space (clip-vit-b32, wemm-2b-text@256) that you query with urna search-space.
[[models]]
preset = "potion"
text = "default"

[[models]]
preset = "clip-vit-b32"
image = "space"

potion is the default choice: it runs on the CPU with numpy and tokenizers, the weights ship in the repository, and the installed binary can query the result. Presets that run model-repository code (the jina-v5-omni-* and wemm-* presets) must be listed in [output] allow_remote_code. wemm-4b and wemm-9b also need --allow-heavy.

Put the weights on disk before the build: set model_path in the spec, export URNA_MODEL_DIR_<NAME> (for example URNA_MODEL_DIR_WEMM_2B), or populate the Hugging Face cache. With no local weights, a jina or wemm preset fails with a Python TypeError before any worker starts. Choose and bring embedding models covers each preset.

The installed binary answers potion corpora only

The query embedder payload that the release channels and urna setup install covers potion only. A corpus whose text = "default" model is any other preset needs a checkout and that preset's packages to run urna ask or urna retrieve. See Known limits.

Check the plan

urna build --spec corpus.toml --dry-run

A dry run parses and validates the spec, then prints the source kind, each model with its roles and dependency status, the named spaces, the output files and the resolved output and cache directories. It loads no model and never opens the source, and it exits 0 even when a dependency is missing. Missing files, bad SQL, a non-unique order_by and join mismatches only show up in a real build.

Run pilots

urna build --spec corpus.toml --sample 500 --out-dir out/pilot
urna build --spec corpus.toml --models potion --out-dir out/pilot

--sample N keeps N evenly spaced rows and renumbers them, so a pilot's citations never match the full build's. --models a,b keeps only the named presets; the text = "default" model must stay in the list, and names that match no model are ignored. Both flags write <name>.urna and <name>.manifest.json into the output directory, so point pilots at their own --out-dir. A --models run leaves an existing build lock untouched.

Build, resume and rebuild

urna build --spec corpus.toml

The stages run in order: load rows, deduplicate identical images, encode media, embed with each model, write each .urna into .tmp/ and move it into place, reopen and validate it, then write the manifest and the build lock. On success the forge prints one JSON result on stdout; Build artifacts describes every file it writes.

Embeddings go to a shared cache (by default ~/.cache/urna) keyed by the model, the embedding recipe and the rows. A second build over the same rows with the same model reuses them, whatever its output directory. Rows are always reloaded and files always rewritten.

  • --resume reuses the encoded media from an interrupted build when the media settings and rows are unchanged and every media file still verifies. The embed cache is used with or without it.
  • --rebuild-only re-emits from cached vectors only. It stops with an error if any model's vectors are not cached, then compares the new build lock with the previous one and prints a warning for each difference. It overwrites the outputs, the manifest and the lock, and it does not compare file_hash for you. See Reproducible builds.

Common errors

A spec error exits 2, a model registry error exits 4, and anything else, including a wrong value type in the spec, exits 1 with a Python traceback.

Message starts withCauseFix
spec error: source.db: ${DATA} is not setThe spec uses ${DATA} and the variable is unset or emptyexport DATA=/path/to/data
spec error: source.order_by [...] is not a total orderTwo rows share the order_by keyAppend a unique column to order_by
spec error: source.joins on '<column>': 0 of N rows matchedThe join key name or its SQLite type differs between the two queriesCast both sides to the same type, or fix the column name
spec error: source.image: file missing for keypath_template renders a path that does not existCheck the template and the data root
spec error: models.<preset>: executes model-repo codeThe preset runs repository code and is not opted inAdd it to [output] allow_remote_code
spec error: models.<preset>: flagged too heavywemm-4b or wemm-9bPass --allow-heavy
registry error: preset '<preset>' needs the '<module>' packageA model's Python package is missingRun the pip install line printed with the error
registry error: unknown model presetA typo in presetUse a name from the model registry
urna_forge.py not foundurna build ran outside a checkoutRun it from the repository root

The complete list of messages, grouped by when they fire, is in Build artifacts. Every flag is in the urna build reference.

On this page