Build from your own rows
Build a .urna from SQLite, CSV, JSONL or an image folder with urna build --spec: sources, models, dry runs, pilots, resume and common errors.
urna build --spec turns a declarative spec file into one or more .urna files. This guide covers the tasks around it: preparing the machine, describing your source, choosing models, checking the plan, running pilots, resuming and fixing the common errors. For a first walk-through with four JSONL rows, start with Your first corpus.
Prepare the machine
The build tool (the forge) is Python code in the repository, and urna build is a launcher that runs it. No release artifact ships the forge, so every build runs from a checkout:
- The repository's
python/tree, with the Python extension built topython/_urna.so(see Your first corpus). The extension is imported only when the file is written, so a missing one fails after the rows are embedded. - Python 3.12 or later, with
numpy, plustokenizersforpotionandpillowfor anything with images. - The Python packages of each model you use. A dry run lists what is missing, with the
pip installline to fix it. - For media:
ffmpegandffprobewithlibsvtav1(theav1backend),avifencandavifdec(avif),cjxlanddjxl(jxl,jxl-transcode), andssimulacra2forcrf = "auto".
Run urna build from the repository root, or from a directory whose parent is the root; the launcher looks for python/tools/urna_forge.py there. A binary you built yourself at target/release/urna also finds its own checkout from any directory. It picks the Python interpreter in this order: URNA_PYTHON, the venv created by urna setup, the nearest .venv, then python3, and prints the choice on stderr. See paths and resolution order.
The forge never downloads a model. It sets HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1 and HF_DATASETS_OFFLINE=1 when they are unset, so model weights must already be on disk.
Write the spec
A spec is TOML (or JSON) with up to seven tables. The smallest working one needs [corpus], [source] and one [[models]]:
[corpus]
name = "support"
chunker_version = "support/1"
[source]
kind = "csv"
path = "${DATA}/tickets.csv"
order_by = ["ticket_id"]
[source.text]
template = "{subject}\n{body}"
[[models]]
preset = "potion"
text = "default"Three rules catch most first specs:
- Relative paths resolve against the directory you run the build from, not the spec's directory. Write data roots as
${VAR}and export the variable, so the same spec works on every machine. An unset or empty variable stops the build with an error naming the key. order_bymust be unique per row. It fixes the row order and the row key, and the row's position is part of its citation.- Unknown keys are errors. A typo such as
dirrunder[output]fails withspec error: output: unknown key 'dirr'and the list of valid keys.
The full key list is in the spec reference.
Describe the source
SQLite
[source]
kind = "sqlite"
db = "${DATA}/shop.sqlite"
query = "SELECT sku, name, description, photo_url FROM products"
order_by = ["sku"]
[[source.joins]]
query = "SELECT sku, category FROM categories"
on = "sku"
[source.derive]
photo_stem = "basename_stem(photo_url)"
[source.image]
path_template = "${DATA}/photos/{photo_stem}.jpg"
label_template = "{name}"The database opens read-only. The query needs no ORDER BY, because the forge sorts on order_by itself. Each [[source.joins]] query runs on the same connection and adds its columns to the main rows that share the on value; the main query wins a name clash. A join that matches no row stops the build, and a partial match prints the count. [source.derive] adds computed columns with two helpers, basename_stem(column) and lower(column), before sorting and templating.
CSV and JSONL
kind = "csv" reads a file with a header row; every value is a string. kind = "jsonl" reads one JSON object per non-blank line. Both take path, order_by, [source.text], [source.derive] and [source.image].
Images in a folder
[source]
kind = "image_dir"
input_dir = "${DATA}/scans"
labels = "${DATA}/labels.csv" # optional: image_id,labelEvery .jpg, .jpeg, .png, .bmp, .webp, .tiff and .tif file under input_dir becomes one row, sorted by relative path. The stored text is the relative path, followed by the label in brackets when there is one. order_by, templates, joins and derive do not apply.
kind = "pdf_dir" passes validation but always fails when rows load. See Images and PDFs for the PDF path that works today.
The text that gets stored
[source.text] template is a Python format string over the row. A template line whose placeholders are all empty is dropped, so optional fields do not leave blank lines. Without a template, a row's stored text is its label, or its key. That stored text is what ask, retrieve and cite return.
Store the images
Rows with images can also carry the encoded images. Add a [media] table and the build encodes each unique image once, writes it to <name>.media/ beside the .urna, and links every chunk to its frame; [output] embed_media = true also stores the bytes inside the .urna, so the file serves them alone (the media directory then stays only as a build cache). Pick a starting point with profile (stills, retrieval, archive and others). Tune media compression explains the choices and the media reference lists every knob.
Choose models
Each [[models]] entry names a preset from the model registry and gives it roles:
text = "default": this model embeds every row into the file's main vectors (space 0). Exactly one model has this role.text = "space"orimage = "space": this model adds a named vector space (clip-vit-b32,wemm-2b-text@256) that you query withurna search-space.
[[models]]
preset = "potion"
text = "default"
[[models]]
preset = "clip-vit-b32"
image = "space"potion is the default choice: it runs on the CPU with numpy and tokenizers, the weights ship in the repository, and the installed binary can query the result. Presets that run model-repository code (the jina-v5-omni-* and wemm-* presets) must be listed in [output] allow_remote_code. wemm-4b and wemm-9b also need --allow-heavy.
Put the weights on disk before the build: set model_path in the spec, export URNA_MODEL_DIR_<NAME> (for example URNA_MODEL_DIR_WEMM_2B), or populate the Hugging Face cache. With no local weights, a jina or wemm preset fails with a Python TypeError before any worker starts. Choose and bring embedding models covers each preset.
The installed binary answers potion corpora only
The query embedder payload that the release channels and urna setup install covers potion only. A corpus whose text = "default" model is any other preset needs a checkout and that preset's packages to run urna ask or urna retrieve. See Known limits.
Check the plan
urna build --spec corpus.toml --dry-runA dry run parses and validates the spec, then prints the source kind, each model with its roles and dependency status, the named spaces, the output files and the resolved output and cache directories. It loads no model and never opens the source, and it exits 0 even when a dependency is missing. Missing files, bad SQL, a non-unique order_by and join mismatches only show up in a real build.
Run pilots
urna build --spec corpus.toml --sample 500 --out-dir out/pilot
urna build --spec corpus.toml --models potion --out-dir out/pilot--sample N keeps N evenly spaced rows and renumbers them, so a pilot's citations never match the full build's. --models a,b keeps only the named presets; the text = "default" model must stay in the list, and names that match no model are ignored. Both flags write <name>.urna and <name>.manifest.json into the output directory, so point pilots at their own --out-dir. A --models run leaves an existing build lock untouched.
Build, resume and rebuild
urna build --spec corpus.tomlThe stages run in order: load rows, deduplicate identical images, encode media, embed with each model, write each .urna into .tmp/ and move it into place, reopen and validate it, then write the manifest and the build lock. On success the forge prints one JSON result on stdout; Build artifacts describes every file it writes.
Embeddings go to a shared cache (by default ~/.cache/urna) keyed by the model, the embedding recipe and the rows. A second build over the same rows with the same model reuses them, whatever its output directory. Rows are always reloaded and files always rewritten.
--resumereuses the encoded media from an interrupted build when the media settings and rows are unchanged and every media file still verifies. The embed cache is used with or without it.--rebuild-onlyre-emits from cached vectors only. It stops with an error if any model's vectors are not cached, then compares the new build lock with the previous one and prints a warning for each difference. It overwrites the outputs, the manifest and the lock, and it does not comparefile_hashfor you. See Reproducible builds.
Common errors
A spec error exits 2, a model registry error exits 4, and anything else, including a wrong value type in the spec, exits 1 with a Python traceback.
| Message starts with | Cause | Fix |
|---|---|---|
spec error: source.db: ${DATA} is not set | The spec uses ${DATA} and the variable is unset or empty | export DATA=/path/to/data |
spec error: source.order_by [...] is not a total order | Two rows share the order_by key | Append a unique column to order_by |
spec error: source.joins on '<column>': 0 of N rows matched | The join key name or its SQLite type differs between the two queries | Cast both sides to the same type, or fix the column name |
spec error: source.image: file missing for key | path_template renders a path that does not exist | Check the template and the data root |
spec error: models.<preset>: executes model-repo code | The preset runs repository code and is not opted in | Add it to [output] allow_remote_code |
spec error: models.<preset>: flagged too heavy | wemm-4b or wemm-9b | Pass --allow-heavy |
registry error: preset '<preset>' needs the '<module>' package | A model's Python package is missing | Run the pip install line printed with the error |
registry error: unknown model preset | A typo in preset | Use a name from the model registry |
urna_forge.py not found | urna build ran outside a checkout | Run it from the repository root |
The complete list of messages, grouped by when they fire, is in Build artifacts. Every flag is in the urna build reference.
Media and named spaces
How a .urna file carries media blobs and extra named vector spaces, each space with its own model_hash and dim, and how to query one.
Ask and retrieve from the terminal
Ask a .urna corpus questions from the shell, tune k and candidates, read explain mode, script retrieve with jq, and verify each citation with cite.