The spec file
Reference for the urna build spec file: TOML, JSON and YAML formats, the seven tables, ${VAR} expansion, path resolution, validation and [corpus].
A spec file (by convention corpus.toml) declares one corpus build for urna build --spec: where the rows come from, which models embed them, how media is encoded and what files come out. This page covers the file as a whole and the [corpus] table. Each other table has its own page.
Formats
The format follows the file extension:
| Extension | Parser | Notes |
|---|---|---|
.toml | Python tomllib | The documented format |
.json | Python json | Same structure, tables as objects, [[models]] as an array |
.yaml, .yml | pyyaml | Works when pyyaml is installed; otherwise spec error: yaml specs need pyyaml: pip install pyyaml (or use .toml) |
Any other extension fails with spec error: unsupported spec extension '<ext>' (use .toml or .json).
The same spec in TOML and in JSON:
[corpus]
name = "quickstart"
chunker_version = "quickstart/1"
[source]
kind = "jsonl"
path = "examples/quickstart/docs.jsonl"
order_by = ["id"]
[source.text]
template = "{text}"
[[models]]
preset = "potion"
text = "default"{
"corpus": { "name": "quickstart", "chunker_version": "quickstart/1" },
"source": {
"kind": "jsonl",
"path": "examples/quickstart/docs.jsonl",
"order_by": ["id"],
"text": { "template": "{text}" }
},
"models": [{ "preset": "potion", "text": "default" }]
}Tables
A spec has at most seven top-level tables. Any other top-level name fails with spec error: unknown top-level table(s) [...]: valid tables are [...].
| Table | Required | Reference |
|---|---|---|
[corpus] | yes | below |
[source] | yes | [source], with [source.text], [source.image], [[source.joins]] and [source.derive] |
[[models]] | yes, at least one | [[models]] |
[embedding] | no | only [embedding.image_input], on the [source] page |
[media] | no | [media], with [media.quality], [media.cluster] and [media.jxl_transcode] |
[build] | no | [build] and [output] |
[output] | no | [build] and [output] |
models must be an array of tables. Writing [models] instead of [[models]] fails with spec error: models must be an array of tables; write [[models]], not [models].
Environment variables
Before parsing, every string value in the spec is expanded, including SQL queries and templates:
${NAME}is replaced with the value of the environment variableNAME. Only the braced form expands, so a bare$in a query and the{column}placeholders of a template are left alone.- Then a leading
~/becomes your home directory. A variable that holds~/dataworks, because this step runs after substitution. A~anywhere else in the string is left alone.
An unset or empty variable stops the build and names the key where it appears, list positions included:
spec error: source.db: ${DATA} is not set; export DATA=/path (spec paths stay portable, see docs/USAGE.md section 13)
spec error: models[1].model_path: ${WEMM_DIR} is empty; export WEMM_DIR=/path (spec paths stay portable, see docs/USAGE.md section 13)Write data roots as variables (db = "${DATA}/shop.sqlite") so the same spec builds on any machine. The build lock records the expanded values.
Paths
Relative paths resolve against the current directory of the build, not against the directory of the spec file. This applies to source.path, source.db, source.input_dir, source.labels, the rendered source.image.path_template, output.dir and model_path. Run the build from the directory the paths are written for, or use ${VAR} and absolute paths.
output.cache_dir, --cache-dir and URNA_CACHE_DIR also expand a leading ~.
Unknown keys
A key that is not part of a table fails with a message that lists the valid keys:
spec error: output: unknown key 'dirr' (valid: allow_remote_code, cache_dir, dir, embed_media, mode, provenance)Three places behave differently:
[embedding]reads onlyimage_input. Any other key under[embedding]is ignored without a message.[corpus]also accepts the keyssource,image_input,media,models,build,outputandspec_path, and then discards their values.[source.derive]is a free map: its keys are the names of the columns you add.
A [source] key that does not apply to the declared kind (for example query on a jsonl source) is ignored without a message.
When a spec is checked
A spec is checked in three stages. The earlier a problem is caught, the less work is lost.
| Stage | What it checks | Failure |
|---|---|---|
| Parse | Extension, ${VAR} values, table and key names, the [media] profile name | spec error: ..., exit 2 |
| Validate | Required keys, allowed values, model and media rules | spec error: ..., exit 2; an unknown or test-only preset is registry error: ..., exit 4 |
| Run | Row loading, joins, the crf gate, encoding, writing the file | spec error: ... (exit 2) for the run-time checks listed in the error catalog, otherwise a Python traceback and exit 1 |
A dry run (urna build --dry-run) stops after validation. It never opens the source, so a missing file, bad SQL, an order_by key that is not unique and a join that matches nothing all pass a dry run.
Value types and table shapes are not checked. A value of the wrong type is accepted and fails when the build first uses it, as a Python exception with a traceback and exit 1, not as a spec error. Examples: [[source.joins]] without on raises TypeError; text = "..." directly under [source] instead of [source.text] raises ValueError; media.width = "wide", an unknown media.pix_fmt, an unknown bucket in media.quality.buckets, an unknown build.preset or build.dtype, and an out-of-range build.mrl_dim all pass validation and fail later.
[corpus]
Names the corpus and pins the chunker identity.
| Key | Type | Default | Required | Meaning |
|---|---|---|---|---|
name | string | "" | yes | Names every output (<name>.urna, <name>-<preset>.urna, <name>.manifest.json, <name>.build.lock.json, <name>.media/). Also the default source_uri prefix item://<name>/, the dataset in the file's provenance, and the title when title is empty. Missing: spec error: corpus.name: required |
chunker_version | string | "" | yes | Written to the file's manifest and hashed into every row's input hash and every chunk_id. Changing it changes every citation. Missing: spec error: corpus.chunker_version: required |
title | string | "" | no | Title in the .urna manifest. Empty means name |
version | string | "0.1.0" | no | Version in the .urna manifest |
reproducible | boolean | true | no | When true, the file's creation time is pinned, so the same inputs give the same bytes |
Because name enters the provenance and the default source_uri, renaming a corpus changes its content_hash and, for rows without a source_uri column, their chunk_ids. See citations and hashes.
The next table to write is [source].
urna tui
Reference for urna tui, the full-screen terminal explorer that opens a .urna file, shows its sections, asks it questions and runs the install checks.
[source]
Reference for [source] in a urna build spec: sqlite, csv, jsonl and image_dir sources, text and image templates, joins, derive, and row identity.