Your first corpus
Turn your own JSONL rows into a .urna file with urna build, ask it a question, and resolve the citation back to the stored text.
This tutorial builds four rows of JSON Lines into a .urna file, asks the file a question, and resolves the citation it returns. It uses the bundled potion model, so nothing is downloaded and no GPU is needed.
Before you start
urna build is a launcher for the Python build tool (the forge) that lives in the repository under python/tools/urna_forge.py. No release artifact ships that tool, so the build runs from a checkout of the repository. The installed urna binary works as the launcher as long as you run it from the repository root.
You need:
gitwithgit-lfs, and a Rust toolchain to build the Python extension.- Python 3.12 or later with
numpyandtokenizers. If you ranurna setup, its venv already has both and the launcher picks it up. Otherwise create a.venvat the repository root; see paths and resolution order for how the interpreter is chosen.
Get the checkout
Clone the repository, fetch the real bytes of the potion table (it is stored with git-lfs), and build the Python extension that writes .urna files:
git clone https://github.com/hoffresearch/urna
cd urna
git lfs pull # or: sh scripts/fetch_potion.sh
cargo build --release -p urna-python --features pyo3/extension-module
cp target/release/lib_urna.dylib python/_urna.so # on Linux: lib_urna.soThe extension is imported only at the last stage of a build, so a missing or stale python/_urna.so fails after the rows are loaded and embedded. urna doctor checks the interpreter, its packages and the potion table; it exits 5 when the table is still an lfs pointer.
Write your rows
Every command in this tutorial runs from the repository root. Create my-corpus/handbook.jsonl with one JSON object per line:
{"id": "002", "section": "Leave", "text": "Request leave in the HR portal at least two weeks ahead.", "source_uri": "handbook/leave.md"}
{"id": "001", "section": "Onboarding", "text": "New hires get a laptop and an account on their first day.", "source_uri": "handbook/onboarding.md"}
{"id": "010", "section": "Expenses", "text": "Submit receipts within 30 days; the finance team pays them monthly.", "source_uri": "handbook/expenses.md"}
{"id": "003", "section": "Security", "text": "Lock your screen when you leave your desk and report lost devices at once."}One row becomes one chunk. The field names are yours; the spec says which ones to use. A field named source_uri is special: when it is present and non-empty it becomes the chunk's source_uri. The last row has none, so its source_uri becomes item://handbook/003.
Write the spec
Create my-corpus/corpus.toml:
[corpus]
name = "handbook"
chunker_version = "handbook/1" # part of every chunk_id
[source]
kind = "jsonl"
path = "my-corpus/handbook.jsonl"
order_by = ["id"] # must be unique per row
[source.text]
template = "{section}: {text}" # the stored, citable text of each row
[[models]]
preset = "potion" # offline static table, no torch, no download
text = "default"
[output]
dir = "my-corpus/out"namenames every output file (handbook.urna,handbook.manifest.json,handbook.build.lock.json).order_bysets the row order and the row key. The forge sorts rows by the string value of these columns and refuses the build if two rows share a key.templateis a Python format string over the row. The first row becomesOnboarding: New hires get a laptop and an account on their first day.text = "default"makespotionthe model that embeds every row into the file's main vectors.- Relative paths resolve against the directory you run
urna buildfrom, not the spec file's directory.
Every key is listed in the spec reference.
Check the plan with a dry run
urna build --spec my-corpus/corpus.toml --dry-run[urna] embedder interpreter: /home/you/.local/share/urna/venv/bin/python
corpus: handbook (source=jsonl, image_input=source, output=single)
model: potion text=default image=none dims=- deps=ok
spaces: (none)
files: handbook.urna
out: /home/you/urna/my-corpus/out
cache: /home/you/.cache/urna (embed/<preset>/<triad>.npz, shared)A dry run parses and validates the spec and reports whether each model's Python packages are installed. It loads no model and never opens the source, so a wrong path or a duplicate key passes here and fails in the real build. A spec error exits 2 and names the key:
spec error: source.order_by: required for total ordering (RFC-0 N1)Run a pilot on a sample
On a large source, build a few evenly spaced rows first. Point the pilot at its own output directory so it does not overwrite the real build:
urna build --spec my-corpus/corpus.toml --sample 2 --out-dir my-corpus/pilot--sample 2 keeps the rows at sorted positions 0 and 2 (001 and 003) and renumbers them. Row positions are part of every citation, so a pilot's citations never match the full build's.
Build the corpus
urna build --spec my-corpus/corpus.tomlThe build loads the rows, embeds them with potion, writes my-corpus/out/handbook.urna, reopens it and validates it. On success it prints one JSON object on stdout: each output file with its size and file_hash, the row count, the corpus_input_hash, the stage timings, and the paths of the manifest and the build lock. The build artifacts page describes each file.
Embeddings are cached outside the output directory, keyed by the model, the embedding recipe and the rows. Running the same build again reuses the vectors.
Ask a question
urna ask my-corpus/out/handbook.urna "how do I report a lost laptop" -k 1ask embeds the query offline with the corpus's model, checks its model_hash against the file, and prints the best chunk's stored text followed by its citation and source_uri:
Security: Lock your screen when you leave your desk and report lost devices at once.
-- urna://sha256:<content_hash>/sha256:<chunk_id> (item://handbook/003)The hashes are specific to your build. For machine-readable output use urna retrieve, which prints one JSON object per hit with its exact cosine score.
Resolve the citation
Paste the citation from the previous step:
urna cite my-corpus/out/handbook.urna 'urna://sha256:<content_hash>/sha256:<chunk_id>'cite prints the file and content hashes, the chunk_id, the source_uri, the span and the stored text. For a forge-built corpus the span is the row's position: byte_start is the ordinal and byte_end is the ordinal plus one, so this row, third in sorted order, shows 2 and 3.
Validate the file
urna validate my-corpus/out/handbook.urnavalidate checks the header, every section checksum, the footer hash, the manifest contract, the required sections and the embedding values, then prints the file_hash and content_hash. It exits 0 when every check passes. See urna validate for the output.
What changes a citation
A citation is urna://content_hash/chunk_id. The content_hash covers every chunk, every vector and the provenance, so any change to the rows, the model or the build options gives the file a new content_hash. urna cite then refuses old citations with content_hash mismatch. Rebuilding the same rows with the same spec gives the same file, because reproducible = true is the default.
The chunk_id identifies one chunk. For a forge-built corpus it is a hash of the row's stored text, its source_uri, its position and the chunker_version:
| Change | Effect on chunk_id |
|---|---|
Edit a row's text, or the template | That row's id changes (every row's, for a template edit) |
Change a row's source_uri, or corpus.name for rows without one | That row's id changes |
| Insert or delete a row | Every row sorted after it shifts position, so its id changes |
Change chunker_version | Every id changes |
Build with --sample | Positions are renumbered, so ids differ from the full build |
Change the model or [build] options | Ids stay; content_hash changes |
Two habits keep ids stable. Pick an order_by key that does not change when rows are added, and zero-pad numeric ids: rows sort by string value, so 10 sorts before 2. Bump chunker_version only when you mean to invalidate every citation.
To go further (SQLite sources with joins, CSV, images, several models), read Build from your own rows.
Quickstart
Build the twelve-paragraph example corpus from a checkout, ask it a question, read the route it took, then cite, validate and open it in the terminal.
The .urna file
What a .urna file holds, how its header, section table and footer fit together, and what the runtime checks before it answers a query.