# Architecture (https://docs.urna.dev/architecture)
urna is a Rust workspace of four crates plus a Python layer. Python builds files; Rust owns the format, serves the queries and runs the CLI. This page maps the pieces and follows a build and a query from end to end.
## The pieces [#the-pieces]
| Piece | Language | What it owns |
| -------------- | ------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `urna-format` | Rust library | The frozen v1 container: layout, manifest, reader, writer, section codecs, wire encodings, `chunk_id` and the hashes |
| `urna-runtime` | Rust library | Opening a file through mmap, query validation, SIMD dispatch, exact, HNSW, graph, hybrid and space search, the exact rerank, and the HNSW and BM25 index builders |
| `urna-cli` | Rust binary | The `urna` command: a clap surface over the two libraries, plus `setup` and `tui` |
| `urna-python` | Rust, pyo3 | The `_urna` extension: `urna.build`, `UrnaFile`, `urna.chunk_id` and the build presets |
| `python/` | Python | The `urna` module, the builder pipeline, model fingerprints, the query embedders and the forge |
`urna-runtime` depends on `urna-format`; the CLI and the Python bridge depend on both. The crates are published as `urna-format`, `urna-runtime` and `urna` (the CLI). `urna-python` is not published on crates.io: it reaches users as the PyPI wheel `urna`. Their public API is summarized in [Rust crates](/reference/rust) and [module urna](/reference/python).
Two more Cargo workspaces sit outside `crates/`. `forge-core/` holds the frozen schema of the forge's canonical intermediate format, kept apart so its dependencies never enter the format and runtime crates; the release gate does not build it. `fuzz/` holds the `cargo-fuzz` targets.
### urna-format [#urna-format]
The format crate knows bytes and nothing else. It defines the 128-byte header, the 32-byte section table entries, the 40-byte footer and the section id map. `UrnaFileBuilder` writes a file: it computes each `chunk_id`, and under zstd text it picks the smallest of several codecs per text section (zstd, intpack, dictionary, FSST, dedup; all decode to the same bytes, so `content_hash` does not move), stores embeddings as float32, float16, int8 or int4, aligns every payload to 64 bytes and writes the checksums and the footer hash. `UrnaView::from_bytes` reads a file and checks it in a fixed order. The crate has no `unsafe` block. See [the .urna file](/concepts/file) and [layout](/reference/format).
### urna-runtime [#urna-runtime]
The runtime maps a file with `MmapUrnaFile::open`, which runs the reader's checks, walks the embeddings for NaN and infinite values, and decodes the HNSW, BM25, graph, media and space tables. It answers five search calls: `search` (exact), `search_ann`, `search_graph`, `search_hybrid` and `search_space`. Every path that generates candidates ends in the same exact cosine rerank, read from one rerank source: a full-precision slab when the file has one, else the stored embeddings. Embeddings are never zstd-compressed, because the SIMD kernels (AVX2, NEON, scalar) score them straight from the map. The runtime never opens a socket. See [search paths and the exact rerank](/concepts/search).
### urna-cli [#urna-cli]
The binary groups its verbs in three sets:
* Engine verbs take a file and, for search, a vector: `inspect`, `validate`, `stats`, `media`, `search`, `search-ann`, `search-graph`, `search-space`, `benchmark`, `cite`. Two more sit in this group but start Python: `search-text` embeds the query with `python/embed_query.py`, and `doctor` runs one real embed as its last check.
* Agent verbs take text or a spec: `ask` and `retrieve` embed the query through a Python process and gate it, and `build` launches the forge.
* Setup verbs: `setup` installs the embedder payload and a Python env, and `tui` is the terminal explorer. Both sit behind the default `tui` feature; `--no-default-features` builds the engine-only CLI.
### urna-python and python/ [#urna-python-and-python]
The pyo3 bridge exposes `urna.build`, which resolves a preset, builds the HNSW, BM25 and graph payloads with the runtime crate, and hands everything to `UrnaFileBuilder`. `UrnaFile` wraps `MmapUrnaFile` with the same search calls plus `retrieve`. `python/urna.py` loads the extension; in the wheel it is the `urna` module.
The rest of `python/` exists only in a checkout of the repository:
* `builder.py`: chunking with real byte spans, an embedding cache and a `Pipeline` that calls `urna.build`.
* `model_fingerprint.py`: computes `model_hash` from a sentence-transformers snapshot.
* `forge/`: the declarative build (`build_spec`, `corpus_sources`, `forge_pipeline`, `forge_emit`), the model registry and its adapters, the media backends and the quality gate, and the two query embedders `embed_query_potion.py` and `embed_query_model.py`.
* `tools/urna_forge.py`: the program `urna build` launches, plus measurement and benchmark tools.
## Where each piece ships [#where-each-piece-ships]
| Piece | Release binary (archives, Homebrew, npm, cargo) | Embedder payload (`urna setup`, one-liners) | PyPI wheel `urna` | Repository checkout |
| ------------------------------------------------ | ----------------------------------------------- | ------------------------------------------- | ------------------------- | ------------------- |
| `urna` binary | yes | | | yes (`cargo build`) |
| `_urna` extension and the `urna` module | | | yes | yes (built by hand) |
| Potion query embedder and table | | yes | yes (`urna.embed_potion`) | yes (Git LFS) |
| Registry query embedder (`embed_query_model.py`) | | no | no | yes |
| `search-text` embedder (`embed_query.py`) | | no | no | yes |
| Forge and `urna_forge.py` | | no | no | yes |
An installed binary with the payload answers `ask` and `retrieve` on potion corpora. `urna build`, `search-text` and queries against a corpus built with a registry model (wemm, clip, jina) need a checkout and the model's Python dependencies. See [known limits](/limits) and [installation](/installation).
## A build, from rows to file [#a-build-from-rows-to-file]
```text
corpus.toml + rows your own chunks and vectors
| |
v |
urna build --spec (Rust launcher, checkout only) |
| finds python/tools/urna_forge.py |
| picks the interpreter, forwards the flags |
v |
forge (Python) |
validate the spec |
load rows, hash each item, dedup identical images |
media stage (encode, optional crf gate) |
embed once per model, through the shared cache |
| |
v v
urna.build(...) (urna-python bridge) <--------------+
resolve the preset, optional matryoshka truncation
HNSW from float32 rows, BM25 over the canonical text, graph
|
v
UrnaFileBuilder (urna-format)
chunk ids, per-section encoding, quantization
header, section table, manifest, 64-byte aligned payloads, footer
|
v
corpus.urna (+ corpus.manifest.json and corpus.build.lock.json from the forge)
```
1. `urna build --spec` is a thin Rust launcher. It finds `urna_forge.py` through the checkout layout, resolves the Python interpreter, forwards its flags and passes the child's exit code through.
2. The forge validates the spec, loads the rows (always recomputed), hashes each item and the whole corpus, and shares one frame between rows with byte-identical images. With a `[media]` table it encodes the media.
3. For each model in the spec it looks up the shared embed cache by model, recipe and corpus hash, and embeds only on a miss. Sentence-transformers models run in their own worker process.
4. For each output file it calls `urna.build` into a staging directory, renames the result into place, reopens it and validates it. Then it writes the sidecar manifest and the build lock.
5. `urna.build` resolves the preset (`exact`, `compressed`, `tiny`, `nano`, `hybrid`), truncates and renormalizes the vectors when `mrl_dim` is set, builds the HNSW graph from float32 rows before quantization, builds BM25 over the canonical text and the chunk graph when asked, and passes everything to the writer.
6. The writer computes the chunk ids, encodes each section, quantizes the embeddings to the preset's dtype, and writes the file with its checksums and footer hash.
Your own code can skip the forge and call `urna.build` directly, from the wheel or from a checkout. See [build from your own rows](/guides/build-spec), [presets and stored precision](/concepts/presets) and [reproducible builds](/concepts/reproducibility).
## A query, from text to citation [#a-query-from-text-to-citation]
```text
urna ask corpus.urna "question"
|
v
MmapUrnaFile::open (urna-runtime)
mmap, header and section checksums, footer hash,
manifest and contract, NaN walk, decode the indexes
|
| manifest: embedding_model, embedding_dim, model_hash
v
embedder process (Python, local)
potion script for minishlab/potion corpora, registry script otherwise
prints one JSON line: model name, dim, model_hash, vector
|
v
model gate (urna-cli)
name equal, dim equal, placeholder hash refused, model_hash equal
|
v
route by the declared index_type
exact -> search
hnsw -> search_ann (beam = max(candidates, k, the file's ef_construction))
hybrid -> search_hybrid
|
v
exact cosine rerank over the candidates
|
v
stored canonical text + urna://content_hash/chunk_id
```
1. The CLI opens the file with `MmapUrnaFile::open`, which checks every hash and decodes the index sections.
2. It reads `embedding_model`, `embedding_dim` and `model_hash` from the manifest and picks the embedder: `embed_query_potion.py` when the model name starts with `minishlab/potion`, `embed_query_model.py` otherwise. It resolves the interpreter and prints its choice on stderr.
3. The embedder runs as a separate Python process and prints one JSON line with the vector and the fingerprint of the model it used.
4. The gate compares the model name, the dimension and `model_hash` with the manifest. The all-zero placeholder hash is refused. Any mismatch stops the query with an error. See [the model gate](/concepts/model-gate).
5. The CLI routes by the manifest's declared `index_type`. The default candidate count is `max(4k, 64)`, set with `--candidates`.
6. Every candidate is rescored by exact cosine, so the score is real cosine, at stored precision when the embeddings are not float32. `--disclose explain` prints the route, the candidate counts and which precision the score has.
7. The CLI decodes the stored canonical text for each hit and prints it with its `urna://content_hash/chunk_id` citation, which `urna cite` resolves back. See [citations and hashes](/concepts/citations).
`preset="hybrid"` in `urna.build`, which is also the forge's default, builds HNSW and BM25 but writes `index_type = "hnsw"`. `ask`, `retrieve` and `search-text` route such a file to `search_ann`, so its BM25 section is never used from the CLI. When `search_hybrid` does run, BM25 only adds candidates: the final order is pure cosine. See [known limits](/limits).
The engine verbs skip the embedder and the gate: `search`, `search-ann`, `search-graph` and `search-space` take a JSON vector and call the runtime directly. In Python, `UrnaFile.retrieve` takes a vector you embedded yourself and checks `model_hash` only when you pass `expected_model_hash`. The script protocol is in [query embedder protocol](/reference/embedder-protocol), and the offline design is in [offline by construction](/concepts/offline).
# Benchmarks (https://docs.urna.dev/benchmarks)
This page collects the measurements the project has published for urna 0.5.1: the engine against four other vector stores, the size and rank stability of each build preset, and a 38,627-image corpus. Each table states the data, the machine and the ruler it was measured with. A number here holds for those conditions only; measure your own corpus before you rely on one. The same tables, drawn as charts, are on [urna.dev](https://urna.dev/#benchmarks).
All latencies below time the search after the query vector exists. `ask`, `retrieve` and `search-text` also start a Python process per call to embed the question, and none of these tables include that time.
## The engine against other stores [#the-engine-against-other-stores]
### Conditions [#conditions]
* Measured on 2026-09-10 on an arm64 Mac (Darwin 25.6.0) with Python 3.12.14, one thread for every store.
* Data: 100,000 synthetic rows of 384 dimensions, L2-normalized and clustered around 2,000 centers (random rows have no neighborhoods, and every HNSW implementation scores badly on them).
* Queries: 200, each a stored row plus a small amount of noise, re-normalized. `k = 10`, seed 7.
* Recall\@10 is measured against a brute-force top 10 over the same rows.
* Every store is driven from Python, so Python call overhead is inside every latency.
* Versions: usearch 2.26.2, hnswlib 0.8.0, sqlite-vec 0.1.9, lancedb 0.38.0, numpy 2.5.2.
### Results [#results]
| Store | Search path | Build (s) | Size (MB) | Cold open + first query (ms) | p50 (ms) | p99 (ms) | Recall\@10 | Rebuild byte-identical | Checks its own bytes |
| --------------------- | ----------- | --------: | --------: | ---------------------------: | -------: | -------: | ---------: | ---------------------- | --------------------------------------------- |
| urna, `exact` preset | exact | 2.14 | 165.3 | 292.3 | 7.803 | 8.315 | 1.0 | yes | yes: per section, whole file, decoded content |
| urna, `hybrid` preset | HNSW | 172.82 | 163.2 | 356.1 | 0.721 | 1.021 | 1.0 | yes | yes: per section, whole file, decoded content |
| usearch | HNSW | 109.26 | 168.5 | 58.1 | 0.672 | 61.422 | 0.995 | yes | no |
| hnswlib | HNSW | 83.07 | 168.4 | 181.5 | 0.324 | 0.525 | 1.0 | yes | no |
| sqlite-vec | exact | 0.81 | 156.6 | 50.0 | 19.762 | 24.836 | 1.0 | yes | structure only (`pragma integrity_check`) |
| lancedb | exact | 0.25 | 153.8 | 612.4 | 16.728 | 19.401 | 1.0 | no | no |
Size is bytes on disk divided by 10^6.
### How to read it [#how-to-read-it]
* The urna `hybrid` row is a file built with the `hybrid` preset (float32 vectors, HNSW with `m = 16` and `ef_construction = 200`, plus a BM25 index) and queried through the HNSW path with `ef = 100`. BM25 plays no part in it. urna raises the search beam to the file's `ef_construction`, so this row effectively searched with a beam of 200, while hnswlib searched with `ef = 100`. See [Known limits](/limits#low-ef-values-have-no-effect).
* Cold open is the time a fresh Python process takes to open the store and answer one query, minus the time of a process that does nothing, best of 3 runs. urna's number is dominated by checking every section checksum and the footer hash before the first answer; the other stores do not check their bytes.
* Build time is single-threaded everywhere, including hnswlib and usearch. urna's HNSW build is the slowest row.
* Rebuild byte-identical means two builds from the same rows produced the same SHA-256.
* The same rows written with raw text and with zstd text share one `content_hash`, so re-encoding a file never moves its citations.
The table says nothing about workloads urna does not serve: updates or deletes in place, metadata filtering, concurrent writers or a query language. A `.urna` file is built once and queried many times.
### Reproduce [#reproduce]
From a checkout, with the four other stores installed in the same Python environment (a store whose package is missing is skipped, not faked):
```sh
.venv/bin/python python/tools/bench_competitors.py --n 100000 --dim 384 --queries 200
```
`--n`, `--dim` and `--queries` change the corpus and the query count. The script regenerates [docs/BENCH.md](https://github.com/hoffresearch/urna/blob/main/docs/BENCH.md).
## Presets: size against rank stability [#presets-size-against-rank-stability]
A [preset](/concepts/presets) chooses how a file stores its text and vectors and which indices it carries. This table measures what each one costs in size and gives up in ranking.
### Conditions [#conditions-1]
* Baseline: a 30,725-chunk Brazilian Portuguese corpus embedded with a 384-dimension MiniLM model, stored as float32 with no index (119.94 MB) and searched exactly.
* 100 queries, `k = 10`, seed 0, NEON SIMD, hot cache, one query at a time from Python. The ladder file records the SIMD backend but not the machine or the date.
* `tiny`, `micro` and `nano` were queried through HNSW with `ef = 100` on files built with `ef_construction = 400`, so the effective beam was 400. `hybrid` was queried with an explicit hybrid search call (vector and BM25 candidates, 200 per path, with the first 200 characters of the chunk as the query text). `ask` and `retrieve` do not take that path on a `hybrid` preset file; they take HNSW.
* The ruler is self-perturbation: each query is a stored chunk's own vector plus up to 8e-5 of noise per dimension, and the ground truth is the float32 baseline's own top 10 for that query. It measures how stable the ranking stays under quantization. It does not measure retrieval quality on real questions, and it is easier than real retrieval, so these recall figures are likely inflated.
### Results [#results-1]
| Preset | Embeddings | Index | Size (MB) | Size ratio | Recall\@10 | p50 (ms) | p99 (ms) |
| ------------ | ---------------------- | ------------- | --------: | ---------: | ---------: | -------: | -------: |
| `exact` | float32 | none | 119.94 | 1.000 | 1.000 | 3.103 | 3.142 |
| `compressed` | float16 | none | 40.67 | 0.339 | 1.000 | 3.166 | 3.225 |
| `tiny` | int8 | HNSW | 30.69 | 0.256 | 0.992 | 1.157 | 1.529 |
| `micro` | int8 at 256 dimensions | HNSW | 26.76 | 0.223 | 0.810 | 0.771 | 1.023 |
| `nano` | int4 | HNSW | 25.04 | 0.209 | 0.913 | 2.060 | 2.721 |
| `hybrid` | float32 | HNSW and BM25 | 73.03 | 0.609 | 1.000 | 4.028 | 4.834 |
* `micro` is a recipe, not a preset value: `urna.build(text_encoding="zstd", dtype="int8", mrl_dim=256, with_hnsw=True)`.
* The int8 and int4 files keep no full-precision copy of the vectors, so their scores and recall are real cosine at the stored precision, not float32 cosine.
* `compressed` scoring 1.000 here does not make float16 lossless; it means the float16 ranking matched float32 under this ruler.
### Truncating dimensions [#truncating-dimensions]
`mrl_dim` keeps the first dimensions of each vector and re-normalizes them. It pays off on models trained for it. The baseline model was not, so on this corpus truncation costs recall:
| Variant | Size (MB) | Size ratio | Recall\@10 |
| ------------------------------ | --------: | ---------: | ---------: |
| int8, 256 dimensions (`micro`) | 26.76 | 0.223 | 0.810 |
| int8, 192 dimensions | 24.80 | 0.207 | 0.733 |
| int8, 128 dimensions | 22.83 | 0.190 | 0.659 |
| int8, 96 dimensions | 21.85 | 0.182 | 0.574 |
| int4, 256 dimensions | 22.95 | 0.191 | 0.777 |
| int4, 192 dimensions | 21.91 | 0.183 | 0.713 |
| int4, 128 dimensions | 20.86 | 0.174 | 0.627 |
Same conditions and ruler as the preset table. On this corpus `nano` (int4 at the full 384 dimensions) keeps more recall than every truncated variant.
## Image corpora: 38,627 card images [#image-corpora-38627-card-images]
These numbers come from a corpus of 38,627 Magic: The Gathering card images built into `.urna` files with the image [media profiles](/guides/media-compression). The code and data of that benchmark are private for now. The machine is not recorded in the published results.
### Media profiles [#media-profiles]
| Profile | Media | File | Against the JPEG source |
| --------------------- | --------------------------------------- | ------: | ----------------------: |
| `archive` | JPEG XL, byte-reversible JPEG transcode | 3.61 GB | 1.10x smaller |
| AV1 all-intra, crf 35 | one AV1 stream, every frame a keyframe | 1.37 GB | 2.89x smaller |
| `retrieval` | AV1 all-intra, crf 50 | 533 MB | 7.46x smaller |
* The AV1 crf 35 row is what the forge now calls the `stills-av1` profile. The forge's `stills` profile encodes one AVIF per image instead (quality 48, speed 8). Measured on 2026-09-12: 1,195,973,116 bytes against 1,374,431,484 bytes for the all-intra AV1 stream, 13% smaller, at a matched SSIMULACRA 2 mean of 61.96 on a 2,048-card sample. The cost is a CLIP embed at build time 4 to 10 times slower, because frames decode one AVIF at a time.
* The `retrieval` row was measured on 2026-09-03: 532,671,548 bytes in one self-contained file, with no measurable loss of text-to-image hit\@1 on 100 queries.
### Text-to-image search by model [#text-to-image-search-by-model]
Text-to-image hit\@1 over every card:
| Model preset | hit\@1 |
| ----------------------------------------- | -----: |
| `siglip2` | 0.750 |
| `wemm-2b` | 0.744 |
| a jina v5 omni preset (size not recorded) | 0.336 |
| `clip-vit-b32` | 0.098 |
hit\@1 counts a query as a hit when the card it describes is the top result.
### Measurements behind the media defaults [#measurements-behind-the-media-defaults]
* The still-picture tune: on 2,048 cards at crf 35, SSIMULACRA 2 p50 of 62.7 against 51.8 with the encoder's default tune, for 10% more bytes.
* The quality gate floors (`crf = "auto"`): cards at 488x680, yuv420, a 2,048-card sample, AV1 with the still tune at speed 6. crf 30 measured SSIMULACRA 2 p10 65.3, minimum 58.6 and embedding drift p10 0.967, and passes. crf 35 measured p50 62.7, p10 55.7, minimum 45.3 and drift p10 0.965, and fails on p10. The default floors therefore pick crf 30.
* Drift against task utility, measured on 2026-09-12: CLIP cosine drift p10 fell from 0.932 at crf 40 to 0.829 at crf 60, so the default drift floor refuses every rung, while text-to-image hit\@1 on 100 queries did not move up to crf 50. This is why the `retrieval` profiles gate on hit\@1 or pin the crf.
* Cluster ordering on a corpus of same-artwork reprints (2026-08-31): 29% smaller than all-intra, using 16-frame groups with scene-change detection off.
## Measure your own corpus [#measure-your-own-corpus]
`urna benchmark` times exact search on your file and, with `--ann`, the HNSW path and its recall against exact search on the same random queries. That recall measures agreement with exact search, not answer quality. See [urna benchmark](/reference/cli/benchmark).
```sh
urna benchmark my_corpus.urna -q 100 -k 10 --ann 100
```
# Changelog (https://docs.urna.dev/changelog)
The user-visible changes of each urna release, newest first. The complete history, with measurements, test counts and the reasoning behind each change, is [`docs/CHANGELOG`](https://github.com/hoffresearch/urna/blob/main/docs/CHANGELOG) in the repository, and the release artifacts are on the [GitHub releases page](https://github.com/hoffresearch/urna/releases).
The tagged releases are `v0.1.0`, `v0.2.0`, `v0.3.0`, `v0.5.0` and `v0.5.1`. Version 0.4.0 exists in the changelog but was never tagged. Until 0.5.0 the project was called nest (files `.nest`, magic `NEST`); the entries below use the current names.
## Unreleased [#unreleased]
Changes on `main` since `v0.5.1`. None of them changes the behavior of the binary, the Python package or the file format, so the 0.5.1 documentation applies to `main` as well.
* The file-size guard in `scripts/release_check.sh` and CI allows 639 lines per Rust source file.
* `install-test.yml` runs inside every release, after the announce step. 0.5.0 and 0.5.1 were tested by hand. It now covers bun, pnpm and yarn, and its wheel job installs the tag's version.
* Documentation: `AGENTS.md` rewritten as a contract, a YAML header on every repository doc, `llms.txt` in the llms.txt format, an afterwork checklist in the pull request template, Homebrew documented as `brew tap hoffresearch/urna` then `brew install urna`, and bun, pnpm and yarn documented as install channels for `@urna/cli`.
## 0.5.1 (2026-09-26) [#051-2026-09-26]
The same binaries as 0.5.0, republished so the registry pages show install commands that work: the 0.5.0 pages on crates.io and npm said `npm install -g urna`, and the PyPI page said `cargo install urna-cli`. Neither exists.
* Fixed the Windows one-liner: the release zip holds `urna.exe` at its root, and `install.ps1` looked for it in a subfolder. The script now takes `urna.exe` wherever the zip puts it. `install.ps1` is served from `main`, so the fix also reaches 0.5.0 installs.
* The npm package is `@urna/cli`, and this is the first tag whose npm package CI publishes.
* Repository docs renamed to `docs/BENCH.md`, `docs/USAGE.md` and `docs/arc/ARC.toml`.
* `pillow` joins the `forge` dependency group in `pyproject.toml`.
## 0.5.0 (2026-09-26) [#050-2026-09-26]
The first tag served by the release pipeline. No format change.
* The project is urna: repository `hoffresearch/urna`, binary `urna`, Python module and wheel `urna`, environment variables `URNA_*`, citation scheme `urna://`, file extension `.urna`, magic `URNA` and chunk-id domain `urna:chunk_id:v1`. The reader still opens files with the legacy `NEST` magic. Chunk ids in a legacy file were derived under the old domain, so `urna.chunk_id` does not reproduce them. The layout, section ids, encodings and manifest schema are unchanged (format v1, schema v1).
* Release channels: archives for five targets on the GitHub release, the Homebrew tap `hoffresearch/homebrew-urna`, npm `@urna/cli`, crates.io (`urna-format`, `urna-runtime`, and the CLI as `urna`, so `cargo install urna`), and the PyPI wheel `urna` with the potion table bundled.
* New `urna setup`: downloads the embedder payload through a system `curl` child, checks its SHA-256 while streaming, unpacks it into a staging directory, then builds a Python env at `/urna/venv` with uv, or `python3 -m venv` and pip, and runs the doctor checks. New exit codes `10` to `14`. Flags `--version`, `--force`, `--no-payload`, `--no-python` and `--uninstall`. See [urna setup](/reference/cli/setup).
* New `urna tui`, a terminal explorer with home, corpus, ask and health tabs. A bare `urna` on a terminal opens it; in a pipe it prints the help and exits `2`. See [the terminal explorer](/guides/tui).
* `cargo install urna --no-default-features` builds the engine-only CLI, without `setup` and `tui`.
* The CLI crate needs Rust 1.88. `urna-format`, `urna-runtime` and `urna-python` stay at 1.85.
* One data-root ladder for the payload: `URNA_DATA_DIR`, `XDG_DATA_HOME`, `~/.local/share`, `%LOCALAPPDATA%`, then `/../share`. This fixed the Windows one-liner, which put the payload under `%LOCALAPPDATA%\urna`, a directory the CLI never searched.
* The interpreter ladder checks the setup venv right after `URNA_PYTHON`, so an installed urna run from inside another project's `.venv` still uses its own env.
* `urna doctor` failures now name `urna setup`, and a missing embedder reports where it looked.
* `examples/quickstart/` builds from a clean checkout, and `urna --help` opens with the five verbs of the loop (`build`, `ask`, `retrieve`, `cite`, `validate`).
## 0.4.0 (2026-09-17, never tagged) [#040-2026-09-17-never-tagged]
There is no `v0.4.0` tag: everything below first shipped in a tagged release with 0.5.0.
* Declarative builds: `urna build --spec corpus.toml` describes sources (SQLite, CSV, JSONL, an image directory), media, models and outputs in one file. It launches the forge in `python/tools/urna_forge.py`, so it runs from a checkout of the repository. See [build from your own rows](/guides/build-spec).
* A model registry with the presets `potion`, `clip-vit-b32`, `siglip2`, `jina-v5-omni-nano`, `jina-v5-omni-small`, `wemm-2b`, `wemm-4b` and `wemm-9b`. Presets that run model-repository code need `[output] allow_remote_code` at build time and `URNA_ALLOW_REMOTE_CODE` at query time; a manifest can no longer grant it. See [the model registry](/reference/models).
* `ask` and `retrieve` pick the query embedder from the manifest's `embedding_model`: the potion script for potion corpora, the registry script for any other model. The name, dimension and `model_hash` checks they share with `search-text` live in one module.
* Media inside the file: the optional sections `blob_refs` (0x14), `blob_span_overlay` (0x16) and `blob_data` (0x17, from `[output] embed_media`), `urna media [--export DIR]`, media profiles and the `crf = "auto"` quality gate. None of these sections enters `content_hash`. See [media and named spaces](/concepts/multimodal).
* Named multimodal spaces: the `space_table` section (0x15) and one vector band per space (0x20 to 0x2F), `urna search-space`, spaces in `urna stats` and `urna inspect --json`, and `urna benchmark --space`.
* A shared embed cache under `${XDG_CACHE_HOME:-~/.cache}/urna`, overridable with `URNA_CACHE_DIR`, and `${VAR}` expansion in spec paths.
* New `urna doctor`, with typed exit codes `0` and `2` to `6`.
* `urna --help` groups the verbs as engine and agent.
* The HNSW build is 1.73x faster on 20k vectors of dimension 384, still single-threaded and byte-deterministic.
* Hardening: deterministic mutation-fuzz harnesses under `cargo test`, four `cargo-fuzz` targets, miri on `urna-format`, `cargo deny` and a semver check in CI. The first fuzz runs found and fixed five classes of malformed-file bug.
* A 0.3.0 reader validates a file with media and named spaces and searches its text embeddings.
## 0.3.0 (2026-06-10) [#030-2026-06-10]
Additive within format v1: 0.2 files load unchanged.
* int4 embeddings (encoding 7): 4-bit codes in blocks of 64 dimensions, one f16 scale per block, so `embedding_dim` must be divisible by 64. Scores are real cosine at the stored int4 precision. New `nano` preset: zstd text, int4 embeddings and HNSW.
* Matryoshka truncation: `urna.build(mrl_dim=K)` keeps the first K components of each vector and renormalizes before quantization. The manifest records `mrl_dim` and `full_dim`. `content_hash` covers the truncated vectors, so a citation is tied to its `mrl_dim`.
* The chunk graph: the optional `graph_adjacency` section (0x0C), `urna search-graph`, and `urna.build(with_graph=True, graph_top_m=...)`. The graph does not enter `content_hash`.
* New encodings for text sections: intpack (4), zstd\_dict (5), fsst (9) and txt\_streams (10), with the `dictionary` (0x0A) and `dedup_map` (0x0B) sections. They decode to the same bytes, so `content_hash` does not change. A 0.2 reader rejects a file that uses them, and a build with zstd text (every preset except `exact`) stores its chunk ids with intpack.
## 0.2.0 (2026-04-28) [#020-2026-04-28]
Extends format v1: 0.1 files load unchanged.
* Encodings: zstd (1) for text sections, float16 (2) and int8 (3) for embeddings. Embeddings are never zstd-compressed, so they can be scored straight from the memory map.
* Optional `hnsw_index` (0x07) and `bm25_index` (0x08) sections. HNSW candidates are always reranked by exact cosine, and the build is deterministic for a given seed.
* Build presets `exact`, `compressed`, `tiny` and `hybrid`. See [build presets](/reference/presets).
* New verbs and flags: `urna search-ann`, `urna search-text` (embeds the query in Python and checks `embedding_model`, the dimension and `model_hash` against the manifest), `urna benchmark --madvise-cold` and `urna inspect --json`.
* SIMD dispatch detected at runtime: AVX2 on x86\_64, NEON on aarch64, a scalar fallback, and `URNA_FORCE_SCALAR` to force the fallback.
* `python/model_fingerprint.py` computes `model_hash` from the model files, and `--model-path` points `search-text` at a local snapshot.
## 0.1.0 (2026-04-27) [#010-2026-04-27]
The first public release. It froze format v1: a change to the container, the hashes, the citation URI or the manifest contract bumps `URNA_FORMAT_VERSION` or `URNA_SCHEMA_VERSION`.
* The container: a 128-byte header, 32-byte section table entries, a 40-byte footer, section payloads aligned to 64 bytes, little-endian integers.
* Six required sections, all canonical: `chunk_ids`, `chunks_canonical`, `chunks_original_spans`, `embeddings`, `provenance` and `search_contract`.
* The hashes: 8-byte header and section checksums, a 32-byte footer hash, `content_hash` over the canonical sections and a domain-separated `chunk_id`. See [hashes](/reference/format/hashes).
* The citation URI `urna:///`. `urna cite` rejects a citation whose `content_hash` does not match the file.
* Reproducible builds: the reproducible flag pins `created` to `1970-01-01T00:00:00Z`.
* Version skew: a reader rejects a higher format or schema version and accepts an equal or lower one. See [compatibility](/reference/format/compatibility).
* The CLI verbs `inspect`, `validate`, `stats`, `search`, `benchmark` and `cite`, with exact search over raw float32 embeddings.
* The Python module: `urna.open`, `UrnaFile.search`, `inspect` and `validate`, `urna.build` and `urna.chunk_id`.
# Citations and hashes (https://docs.urna.dev/concepts/citations)
Every hit urna returns carries a citation of the form `urna://content_hash/chunk_id`. The citation names the corpus content and one chunk inside it, both by SHA-256, so it resolves to the same stored text on any machine that has the same corpus. This page covers the hashes behind it, what changes them, and what the span offsets in a hit mean.
## Anatomy of a citation [#anatomy-of-a-citation]
This is a real citation from the quickstart corpus:
```text
urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
```
* The first part is the file's `content_hash`, a SHA-256 over the six canonical sections.
* The second part is the `chunk_id`, a SHA-256 over one chunk's text and origin.
Both are written as `sha256:` followed by 64 lowercase hex digits. Resolve a citation with [urna cite](/reference/cli/cite):
```sh
urna cite examples/quickstart/out/quickstart.urna 'urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be'
```
```text
citation_id: urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
file: examples/quickstart/out/quickstart.urna
file_hash: sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
content_hash: sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df
chunk_id: sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
source_uri: demo/03-citations.md
byte_start: 7
byte_end: 8
text:
because the citation points at content, two people who build the same logical corpus on two machines get the same citation, and a stored corpus and a compressed one cite identically. resolving a citation returns the exact canonical text and the original byte span it came from, which is what lets an agent quote a source it can prove.
```
`cite` first recomputes the file's `content_hash` and refuses the citation if it does not match, so a citation never resolves against a different corpus by accident. The text it prints is the stored canonical text of the chunk, not a reopen of the original source file.
## chunk\_id [#chunk_id]
The `chunk_id` is the SHA-256 of this preimage, with every string length-prefixed:
```text
"urna:chunk_id:v1\n"
u32 LE len(canonical_text) canonical_text bytes
u32 LE len(source_uri) source_uri bytes
u64 LE byte_start
u64 LE byte_end
u32 LE len(chunker_version) chunker_version bytes
```
The embedding is not part of it. The same text from the same place, cut by the same chunker, gets the same `chunk_id` whatever model embedded it. Changing any of the five inputs (text, uri, either offset, or the manifest's `chunker_version`) gives a new id.
## content\_hash [#content_hash]
The `content_hash` is the SHA-256 of the six canonical sections, taken in this fixed order: `chunk_ids`, `chunks_canonical`, `chunks_original_spans`, `embeddings`, `provenance`, `search_contract`. For each section the hash takes the name length (`u32 LE`), the name, the decoded length (`u64 LE`) and the decoded bytes.
"Decoded" means after the wire codec: a section stored as `zstd` or `intpack` hashes the same as its `raw` form. Quantized embeddings are the exception, because they are hashed as stored: a `float16` corpus and a `float32` corpus of the same chunks have different `content_hash` values. The manifest is not part of the preimage.
## file\_hash has two meanings [#file_hash-has-two-meanings]
| Value | Covers | Where you see it |
| -------------------- | ----------------------------------------------- | ----------------------------------------------------------------------- |
| Footer hash | SHA-256 of every byte before the 40-byte footer | Stored as 32 raw bytes in the footer and checked on open; never printed |
| Reported `file_hash` | SHA-256 of the whole file, footer included | `validate`, `stats`, `inspect`, `cite`, every hit, `retrieve` JSON |
The reported `file_hash` is the same number `sha256sum` or `shasum -a 256` prints for the file:
```sh
shasum -a 256 examples/quickstart/out/quickstart.urna
```
```text
e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832 examples/quickstart/out/quickstart.urna
```
The header checksum and the per-section checksums are shorter: the first 8 bytes of a SHA-256, stored raw. `urna inspect` prints a section checksum as 16 hex digits with no prefix. They detect a corrupted header or payload; they are not identifiers.
Use `content_hash` to ask "is this the same corpus content?" and `file_hash` to ask "is this the same file, byte for byte?". Two files can share a `content_hash` and differ in `file_hash` (for example, a different `title` in the manifest, or a BM25 index added).
## What moves content\_hash [#what-moves-content_hash]
Anything that moves `content_hash` changes every citation into the file, because the hash is the first half of each citation.
| Change | Moves `content_hash` | Why |
| -------------------------------------------------------- | -------------------- | ----------------------------------------------------------------------------- |
| Raw vs `zstd` text, or which text codec the writer picks | No | Codecs decode to identical bytes |
| Embedding `dtype` (`float32`, `float16`, `int8`, `int4`) | Yes | Quantized bytes are hashed as stored |
| `mrl_dim` truncation | Yes | The truncated vectors are hashed |
| Chunk text, `source_uri`, span, `chunker_version` | Yes | They are in `chunk_ids`, `chunks_canonical` and `chunks_original_spans` |
| Provenance JSON, including key order | Yes | `provenance` is canonical |
| Attaching an HNSW index | Yes | It rewrites `index_type` and `rerank_policy`, which live in `search_contract` |
| Declaring `index_type = "hybrid"` | Yes | Same mechanism |
| Adding a BM25 index alone | No | Only a manifest capability changes |
| Graph, blob tables, blob data, space table and bands | No | Not canonical; only manifest flags change |
| Manifest-only fields (`title`, `created`, capabilities) | No | The manifest is covered by the file hash only |
In practice: rebuilding with a different [preset](/concepts/presets) or `dtype` gives new citations. Changing only the text compression does not.
Attaching an HNSW index sets `index_type = "hnsw"` and `rerank_policy = "exact"`, and the writer copies both into the canonical `search_contract` section. An `exact` build and an HNSW build of the same chunks therefore have different `content_hash` values and different citations. This also separates `exact` from the `tiny`, `nano` and `hybrid` presets, and any build with `with_hnsw=True`. A comment in the format crate says adding an optional section never invalidates citations; that holds for the graph, blobs and spaces, not for HNSW. See [Known limits](/limits).
## What the offsets mean [#what-the-offsets-mean]
`byte_start` and `byte_end` come from `chunks_original_spans` and mean whatever the builder wrote there. urna checks only that `byte_end` is not below `byte_start`.
| Built with | Offsets are |
| ------------------------------------------ | -------------------------------------------------------------------------------- |
| `builder.chunk_text` (repo checkout) | Real UTF-8 byte offsets into the source text |
| The Python examples in the repo | `0` to the UTF-8 byte length of the chunk text |
| `urna build --spec` (the forge) | The row ordinal: `[ordinal, ordinal + 1)` |
| A corpus with a blob span overlay (`0x16`) | In hits, a byte range inside the media blob, with the blob's uri as `source_uri` |
The quickstart corpus is built by the forge, which is why the citation above shows `byte_start: 7` and `byte_end: 8`: the chunk is ordinal 7, the eighth row after sorting by `order_by`.
Because the forge's ordinal enters the `chunk_id`, three things follow for spec builds:
* Adding or removing a row shifts the ordinal, and so the citation, of every later row.
* A `--sample` build renumbers its rows and never shares citations with the full build.
* Rows sort by the string value of their `order_by` columns, so numeric ids sort as text (`"10"` before `"2"`).
For a blob overlay, the runtime rewrites the spans at open time, so hits and `retrieve` report the blob range. `urna cite` reads `chunks_original_spans` directly and prints the stored span, which for a forge media corpus is the row ordinal. The `chunk_id` is always computed from the stored span.
## What the hashes prove [#what-the-hashes-prove]
A matching `content_hash` and `chunk_id` prove that the text you are reading is the text the citation was issued for, and a passing `urna validate` proves the file is internally consistent. All of these are unkeyed SHA-256: they detect corruption, truncation and mismatched corpora. They do not prove who built the file. See [Security](/security).
The hashes also make builds comparable. Two builds from the same rows, chunker, model, `dtype`, provenance and index parameters produce the same `content_hash`, and the forge pins `created` by default, so the files can be byte-identical. See [Reproducible builds](/concepts/reproducibility).
The byte-level preimages and golden values are in [Hashes and citation ids](/reference/format/hashes).
# The .urna file (https://docs.urna.dev/concepts/file)
A urna corpus is one `.urna` file. The chunk text, the source spans, the embeddings, the indices and the search contract all live inside it, and the runtime answers queries by memory-mapping that file and reading it in place. This page describes the file at the level you need to use it; the byte-level layout is in [File format](/reference/format).
## One file, memory-mapped [#one-file-memory-mapped]
There is no server process, no side directory and no unpack step. The Rust runtime opens the file read-only with `mmap`, checks it, and serves searches from the mapped bytes. Embeddings stay in their on-disk precision and are read row by row, so opening a quantized corpus does not expand it to float32 in memory.
Two consequences follow:
* Copying the file copies the whole corpus. Moving it between machines needs nothing else, as long as the query side has the matching embedding model (see [the model gate](/concepts/model-gate)).
* Do not rewrite a file while a process has it open. The runtime never writes to a `.urna`, and its checksums detect corruption after the fact; they do not protect a mapping from a live writer.
## What is inside [#what-is-inside]
A file is a set of numbered sections plus a JSON manifest. Six sections are required in every file. They are also the canonical sections: their decoded bytes are what `content_hash` covers, so they define the identity of the corpus and of every citation into it (see [citations and hashes](/concepts/citations)).
| Id | Section | Holds |
| ------ | ----------------------- | ------------------------------------------------------------------------------------ |
| `0x01` | `chunk_ids` | One `sha256:` id per chunk |
| `0x02` | `chunks_canonical` | The stored text of each chunk, the text `cite` returns |
| `0x03` | `chunks_original_spans` | `source_uri`, `byte_start` and `byte_end` per chunk |
| `0x04` | `embeddings` | One vector per chunk, in the file's `dtype` (`float32`, `float16`, `int8` or `int4`) |
| `0x05` | `provenance` | Free-form JSON written by the builder |
| `0x06` | `search_contract` | `metric`, `score_type`, `normalize`, `index_type`, `rerank_policy` |
Optional sections add capabilities without touching the canonical six:
| Id | Section | Adds |
| ---------------- | ------------------- | ---------------------------------------------------------- |
| `0x07` | `hnsw_index` | Approximate nearest-neighbour candidates |
| `0x08` | `bm25_index` | Lexical candidates for the hybrid path |
| `0x0C` | `graph_adjacency` | A chunk-to-chunk graph for the graph path |
| `0x14` | `blob_refs` | A table of media blobs (hash, uri, length, inlined or not) |
| `0x16` | `blob_span_overlay` | Per-chunk byte ranges inside a blob |
| `0x17` | `blob_data` | The media bytes themselves, when inlined |
| `0x15` | `space_table` | Named embedding spaces, one per extra model or tower |
| `0x20` to `0x2F` | `space_embeddings` | One vector band per named space |
A compressed text section can also bring two unnamed helper sections (`0x0A` dictionary, `0x0B` dedup map); tools list them as `unknown`. The full id map, including reserved ids, is in [Sections](/reference/format/sections).
The quickstart corpus has the six required sections plus HNSW, BM25 and the graph:
```sh
urna stats examples/quickstart/out/quickstart.urna
```
```text
sections: 9
0x01 chunk_ids encoding=intpack 389 bytes
0x02 chunks_canonical encoding=zstd 1484 bytes
0x03 chunks_original_spans encoding=zstd 154 bytes
0x04 embeddings encoding=raw 12288 bytes
0x05 provenance encoding=zstd 118 bytes
0x06 search_contract encoding=zstd 95 bytes
0x07 hnsw_index encoding=raw 182 bytes
0x08 bm25_index encoding=zstd 1695 bytes
0x0c graph_adjacency encoding=raw 174 bytes
```
## Layout [#layout]
```text
[0, 128) header
[128, 128 + 32 * count) section table, one 32-byte entry per section
manifest JSON, directly after the table
payloads each section at a 64-byte aligned offset
[file_size - 40, file_size) footer
```
* The header starts with the magic `URNA` and the version `1.0`, then the embedding dim, the chunk and embedding counts, the file size, the offsets of the table and the manifest, and an 8-byte header checksum. Files written by 0.4.0 and earlier carry the magic `NEST`; the reader still opens them.
* Each section table entry holds the section id, its wire encoding, its offset, its size and an 8-byte checksum of its payload bytes.
* The footer holds its own size and a SHA-256 of every byte before it.
* All integers are little-endian. Padding between payloads is zero and is not part of any checksum.
The section encoding is how the bytes are stored (`raw`, `zstd`, `intpack`, `float16`, `int8`, `int4` and a few text codecs). Text codecs decode to the same bytes as `raw`, which is why compressing the text does not change `content_hash`. See [Encodings](/reference/format/encodings).
## The manifest [#the-manifest]
The manifest is the JSON that describes the corpus: `embedding_model`, `embedding_dim`, `n_chunks`, `dtype`, `metric` (`ip`), `score_type`, `normalize` (`l2`), `index_type`, `rerank_policy`, `model_hash`, `chunker_version` and a `capabilities` object, plus optional fields such as `title`, `created` and `mrl_dim`. `urna inspect` prints it.
Two points matter when you reason about hashes:
* The manifest is covered by the footer hash only. It has no section checksum and is not part of `content_hash`, so a manifest-only change (a title, a capability flag) changes `file_hash` and leaves every citation intact.
* The `search_contract` section repeats five manifest fields. Because that section is canonical, those five fields do move `content_hash`. The reader rejects a file whose contract and manifest disagree.
The forge (`urna build`) also writes a `.manifest.json` next to the `.urna`. That is a separate build report, not the manifest inside the file. See [Build artifacts](/reference/spec/artifacts).
## What the reader checks on open [#what-the-reader-checks-on-open]
Every tool that opens a file runs the same parser first. It stops at the first failure with a typed error, and nothing is served from a file that fails. In order:
1. The file is at least 168 bytes (header plus footer).
2. The magic is `URNA` or `NEST`, and the version is `1.0`.
3. The header checksum matches.
4. The `file_size` in the header equals the real length.
5. The section table fits inside the file.
6. For each section: the encoding is legal for that section, the offset is 64-byte aligned, the payload fits inside the file, and the payload checksum matches.
7. The manifest parses and passes validation (allowed `dtype`, `metric`, `index_type` and so on; `model_hash` shaped as `sha256:` plus 64 hex digits).
8. The footer hash matches.
9. The manifest's `embedding_dim` and `n_chunks` equal the header.
10. All six required sections are present.
11. The embeddings encoding matches `dtype` and the section has the exact expected size; each named space band has its exact size too.
12. The `search_contract` section matches the manifest field by field.
When the runtime opens a file for search (`MmapUrnaFile::open`, which every search verb, `ask`, `retrieve`, `media` and `inspect --json` use), it then:
* walks the embeddings for NaN and Inf;
* decodes the chunk ids and spans;
* decodes the HNSW and BM25 indices when their sections are present;
* decodes the graph when the section is present and the manifest sets `capabilities_ext.graph_present`;
* decodes the blob tables and applies the span overlay when the manifest sets `blobs_present`;
* decodes the space table and checks each band for NaN and Inf;
* computes `file_hash` and `content_hash`.
`urna validate` runs the parser, the NaN walk on the embeddings, the contract check, and a SHA-256 proof of every inlined blob. It does not decode the index payloads or check space bands for NaN, so a file with a malformed index and a valid checksum passes `validate` and fails on open. See [urna validate](/reference/cli/validate).
The checksums and hashes are unkeyed SHA-256. They prove the bytes are consistent with themselves, so they catch corruption and truncation. They do not prove who produced the file. See [Security](/security).
## Look at a file [#look-at-a-file]
```sh
urna validate corpus.urna # every checksum, the footer hash, the contract
urna stats corpus.urna # model, dtype, index_type, sections, hashes
urna inspect corpus.urna # header, section table with checksums, manifest
urna inspect --json corpus.urna
```
Every hit, and every citation, is tied to this file through its hashes: see [citations and hashes](/concepts/citations).
# The model gate (https://docs.urna.dev/concepts/model-gate)
A vector search with the wrong query model still returns numbers. The cosine arithmetic is valid, the dimensions may even match, and the ranking is noise. urna records a fingerprint of the embedding model in every corpus, `model_hash`, and the CLI compares it with the fingerprint the query embedder reports before it searches. A mismatch is an error, never a silently bad answer.
## Three hashes, one gate [#three-hashes-one-gate]
A build through `urna build` keys the embeddings of each model by three hashes. Only one of them takes part in the query-time gate.
| Hash | What it identifies | Where it lives | Used for |
| ----------------------- | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | ---------------------------------- |
| `model_hash` | the model: the files and settings that decide what vector a text becomes | the `.urna` manifest, and the build's `.manifest.json` | the query-time gate |
| `corpus_input_hash` | the input rows: each row's text, source image, label and `chunker_version`, in row order | the `.urna` provenance section, and `.manifest.json` | keying the build's embedding cache |
| `embedding_recipe_hash` | how the model was driven: query and document modes, image settings, `normalize`, dtype, device class | `.manifest.json` only | keying the build's embedding cache |
Together the three form the key of the shared embedding cache, so a rebuild reuses vectors only when the model, the input and the recipe are all unchanged. See [Reproducible builds](/concepts/reproducibility).
At query time only `model_hash` is checked. The recipe is not carried to the query side: `ask` and `retrieve` build the query embedder from the preset's defaults, not from overrides in the build spec. Where a recipe setting also enters `model_hash` (the dtype policy of a sentence-transformers model, for example), the gate catches the difference; where it does not, nothing does.
## What `model_hash` covers [#what-model_hash-covers]
`model_hash` is `sha256:` followed by 64 hex digits, computed by the embedder over a canonical JSON fingerprint. What goes into the fingerprint depends on the embedder:
| Embedder | Fingerprint inputs |
| ----------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| potion (the default, `minishlab/potion-base-8M/v1`) | the bytes of `config.json`, `tokenizer.json` and `model.safetensors`, the tokenizer settings, mean pooling, `normalize`, the dimension (256) and the float32 policy |
| sentence-transformers presets (jina, wemm) | the weight, tokenizer and processor files, the model-repo code files, pooling, `normalize` and the dtype policy |
| open\_clip presets (clip, siglip2) | the model id and pretrained tag, the preprocessing transform, a digest of every weight tensor, and L2 normalization |
| a sentence-transformers model fingerprinted with `python/model_fingerprint.py` (checkout) | up to ten config, tokenizer, pooling and weight files, the dimension and `normalize` |
The vendored potion table gives `sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98`. A different table gives a different hash, and corpora built with the old one stop answering. That is the gate doing its job.
For sentence-transformers presets the dtype policy follows the device: bfloat16 on CUDA, float16 on MPS, float32 on CPU. The same weights built on one device class and queried on another report different hashes unless `URNA_ST_DTYPE` (or the spec's `dtype`) pins the same dtype on both sides.
## What the CLI checks at query time [#what-the-cli-checks-at-query-time]
`urna ask`, `urna retrieve`, `urna search-text` and the ask tab of `urna tui` run the query embedder, parse its JSON and check, in order:
1. The reported `embedding_model` equals the manifest `embedding_model`.
2. The reported `embedding_dim` and the vector length equal the manifest `embedding_dim`.
3. The manifest `model_hash` is not the placeholder, `sha256:` followed by 64 zeros.
4. The reported `model_hash` equals the manifest `model_hash`.
Each failure exits 1 before the search runs. A real mismatch on the quickstart corpus:
```text
Error: model_hash mismatch: corpus was built with sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98, embedder reports sha256:abababababababababababababababababababababababababababababababab
fingerprint reported by embedder: {"fake":true}
hint: --model-path PATH to point at the exact snapshot, or rebuild the corpus with the model you intend to use.
```
`ask` and `retrieve` offer no way around the gate. `search-text --skip-model-hash-check` turns off checks 3 and 4 and keeps 1 and 2. The protocol and the exact messages are in [Query embedder protocol](/reference/embedder-protocol).
Verbs that take a raw vector (`search`, `search-ann`, `search-graph`) never see a model, so they have no gate. `search-space` checks a named space's own `model_hash` only when you pass `--expect-model-hash`.
## Opt-in in Python [#opt-in-in-python]
In the Python API the gate is off unless you ask for it.
| Call | Checks `model_hash` |
| ----------------------------------------------------------------- | ------------------------------------------------------------------------------- |
| `UrnaFile.retrieve(query, k, ..., expected_model_hash=None)` | only when `expected_model_hash` is given |
| `UrnaFile.search_space(name, query, k, expected_model_hash=None)` | only when `expected_model_hash` is given; it compares against that space's hash |
| `UrnaFile.search`, `search_ann`, `search_hybrid`, `search_graph` | never |
Pass the hash your query embedder reports. With the wheel installed as `urna[embed]`:
```python
import urna
from urna.embed_potion import potion_embedder
db = urna.open("corpus.urna")
emb = potion_embedder()
query = emb.embed_texts(["can I use this offline"])[0]
hits = db.retrieve(query, 3, expected_model_hash=emb.model_hash())
```
On a mismatch, `retrieve` raises before it parses the query:
```text
ValueError: model_hash mismatch: the query was embedded with ..., but the corpus was built with .... Results would be cosine-valid but semantically wrong. Pass expected_model_hash=None to bypass this check.
```
The comparison is a string compare: there is no name or dim layer in Python, although the runtime still rejects a vector of the wrong length. The FastAPI, Flask and notebook examples in the repository call `retrieve` without the hash, so the gate is off there. See [Use urna from Python](/guides/python).
## Placeholders are accepted on write [#placeholders-are-accepted-on-write]
`urna.build` checks that `model_hash` has the `sha256:` plus 64 hex shape. It does not refuse the all-zero placeholder, so a file can be written, validated and opened with it. The refusal happens only at the CLI's query gate, with this message:
```text
manifest carries the legacy placeholder model_hash (sha256:0000000000000000000000000000000000000000000000000000000000000000). Rebuild this corpus with a real fingerprint, or pass --skip-model-hash-check to proceed at your own risk.
```
A corpus written with the placeholder `model_hash` passes `urna validate` and opens in Python, but `urna ask` and `urna retrieve` always refuse it, and in Python `retrieve` with `expected_model_hash` refuses it too. Only `search-text --skip-model-hash-check`, the raw-vector verbs and Python calls without `expected_model_hash` reach it. Pass a real `model_hash` when you build. See [Known limits](/limits).
## What the gate does not prove [#what-the-gate-does-not-prove]
The gate compares two claims: the hash the builder recorded and the hash the query embedder reports. It does not check that the stored vectors were produced by the model the manifest names; `urna.build` trusts the vectors and the hash it is given. And a hash identifies a model version, it does not vouch for it. For how far the file's own hashes go, see [Citations and hashes](/concepts/citations) and [Security](/security).
# Media and named spaces (https://docs.urna.dev/concepts/multimodal)
Besides its text chunks and their default vectors, a `.urna` file can carry two optional layers: media blobs (the encoded images or video segments the chunks came from) and named spaces (extra vector sets, such as image embeddings from a second model). Both sit in optional sections outside `content_hash`, so adding them never changes a citation.
## The default space and named spaces [#the-default-space-and-named-spaces]
Every file has one default space: the vectors in section 0x04, one per chunk, embedded with the model the manifest names in `embedding_model`. `urna ask`, `urna retrieve`, `urna search-text` and every `search*` verb except `search-space` read this space and only this space.
A named space is a second set of vectors over the same chunks, with its own name, model, dim and dtype. A file can hold up to 15. `search-space` reads one named space and nothing else, so a text query can never be scored against image vectors by accident, and the reverse.
## Media blobs [#media-blobs]
Three sections describe media. The manifest sets `capabilities_ext.blobs_present` when any of them is written.
| Section | Name | Content |
| ------- | ------------------- | -------------------------------------------------------------------------------------------------------------------- |
| 0x14 | `blob_refs` | one record per blob: the SHA-256 of the blob's bytes, its `original_uri`, its byte length, and whether it is inlined |
| 0x16 | `blob_span_overlay` | one entry per chunk, in chunk order: a blob index and a range inside that blob, or "none" |
| 0x17 | `blob_data` | the inlined blob bytes, behind an offset table parallel to 0x14 |
The records keep their build order, and the overlay points into them by position. An overlay entry of "none" keeps the chunk's own span from section 0x03, so a file can mix media chunks and plain text chunks.
When a file with an overlay is opened, the runtime replaces each pointing chunk's `source_uri` and span with the blob's `original_uri` and the range in the overlay. Search hits, `urna retrieve` and `UrnaFile.retrieve` report those values. `urna cite` reads section 0x03 directly and prints the chunk's original span instead: for a forge corpus that is `item:///` and the row ordinal.
### What a range means [#what-a-range-means]
The forge writes the overlay in one of two shapes, depending on the media backend (see [Tune media compression](/guides/media-compression)):
| Backend | One blob per | Range for a chunk |
| ----------------------------------------- | -------------------------- | ------------------------------------- |
| `av1` | stream segment (an `.mp4`) | the frame index, `[frame, frame + 1)` |
| `avif`, `jxl`, `jxl-transcode`, `control` | image file | the whole file, `[0, byte_len)` |
The `original_uri` is `media://` followed by the blob's path inside the corpus media directory.
### Sidecar or inlined [#sidecar-or-inlined]
By default the forge writes the encoded media to a `.media/` directory next to the `.urna`, and the 0x14 records say `inlined = false`. With `[output] embed_media = true` the bytes also go into section 0x17 and the file is self-contained. The `.media/` directory stays on disk as the build cache, and peak build memory is about twice the media size.
`urna media` lists the records, and `--export` writes the inlined ones back to files:
```sh
urna media corpus.urna
urna media corpus.urna --export exported/
```
Each blob is checked against its 0x14 SHA-256 before it is written, and the export stops at the first mismatch. `urna validate` runs the same check over every inlined blob without writing anything. In Python, `UrnaFile.blob_bytes(i)` returns one inlined blob.
## Named spaces [#named-spaces]
A named space is one entry in the space table (section 0x15) plus one band section holding its vectors. The manifest sets `capabilities_ext.supports_multimodal`.
| Space table field | Meaning |
| ----------------- | --------------------------------------------------------------------------- |
| `space_index` | 1 to 15; the band lives in section `0x20 + space_index` (0x21 to 0x2F) |
| `name` | unique within the file |
| `dim` | the vector length |
| `dtype` | `float32`, `float16`, `int8` or `int4` (`int4` needs `dim` divisible by 64) |
| `model_hash` | the identity of the model that produced the vectors, `sha256:` prefixed |
| `n_vectors` | must equal the number of chunks |
Bands are parallel to the chunks: row `i` of every band belongs to chunk `i`, so a hit in a named space maps to a chunk, its text and its citation the same way a default-space hit does. At open, each listed band must be present at exactly the size its dim, dtype and count imply, and its values must be finite.
### How the forge names spaces [#how-the-forge-names-spaces]
In a build spec, each `[[models]]` block with `image = "space"` or `text = "space"` emits named spaces (see [`[[models]]`](/reference/spec/models)):
| Role | No `dims` | With `dims = [256, 512]` |
| ----------------- | --------------- | ---------------------------------------- |
| `image = "space"` | `` | `@256`, `@512` |
| `text = "space"` | `-text` | `-text@256`, `-text@512` |
The `text = "default"` model is the default space, not a named one. For each dim the forge keeps the first `dim` components of every vector and L2-normalizes them again, and stores the result at the model's `space_dtype` (default `int8`). Every space of one preset shares that preset's `model_hash`, image and text alike.
### Listing spaces [#listing-spaces]
`urna stats` prints a `spaces:` block with each space's name, dim, dtype, vector count and `model_hash`. `urna inspect --json` has a `spaces` array, and Python has `UrnaFile.space_names`.
## Querying a named space [#querying-a-named-space]
`urna search-space` runs an exact scan over one band. It takes a query vector, not text, as a JSON array of floats at the space's dim. With the vector in `query.json`:
```sh
urna search-space corpus.urna "$(cat query.json)" --space clip-vit-b32 -k 10 \
--expect-model-hash "sha256:"
```
* An unknown name fails with `embedding space not found: `. There is no fallback to the default space.
* A query of the wrong length fails with a dimension mismatch.
* `--expect-model-hash` fails the query when the space's `model_hash` differs. Without it, the hash is not checked: `search-space` has no embedder of its own and cannot compute one. In Python the same check is `expected_model_hash=`.
* Scores are exact cosine over the whole band, `recall = 1.0`, and hits report `index_type = "space"`. The rerank source follows the band's dtype, so an `int8` space scores at stored precision.
* A hit's `embedding_model` field still names the file's default model, not the space's model.
No verb embeds text for a named space: `urna ask` and `urna retrieve` always query the default space. Embed the query with the same preset yourself. From a repository checkout, with the preset's packages installed:
```python
import sys
sys.path.insert(0, "python")
import urna
from forge import model_registry
emb = model_registry.create_embedder("clip-vit-b32")
query = emb.embed_texts(["a red dragon over a castle"], role="query")[0].tolist()
db = urna.open("out/cards/cards.urna")
hits = db.search_space("clip-vit-b32", query, 10, expected_model_hash=emb.model_hash)
for hit in hits:
print(round(hit.score, 4), hit.source_uri, hit.citation_id)
```
For a space with a dim suffix, such as `clip-vit-b32@256` on a model with a ladder, slice the query with `model_registry.slice_renorm(vectors, 256)` before searching. `urna benchmark --space ` measures the latency of one space.
## Outside content\_hash [#outside-content_hash]
`content_hash` covers the six required sections only. Blob records, overlay, inlined bytes, space table and bands are excluded, so:
* Adding media or a named space to a corpus keeps every `chunk_id`, `content_hash` and citation.
* A self-contained file and its sidecar twin cite identically.
* These sections are still covered by their section checksums and by the file hash, which `urna validate` checks.
To build a corpus with images, see [Images and PDFs](/guides/images).
# Offline by construction (https://docs.urna.dev/concepts/offline)
Answering a query never needs the network. The Rust binary links no network stack, the file is read from a memory map, and the query embedders the CLI runs are offline by default. The network appears only when something is installed: the binary, the embedder payload, the Python environment, or a model you explicitly allow to download. This page lists each of those places. For the step-by-step air-gapped procedure, see [Air-gapped install and queries](/guides/offline).
## What runs at query time [#what-runs-at-query-time]
| Piece | Network |
| --------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The `urna` binary (open, validate, every search path, `cite`, `inspect`, `stats`) | None. The Rust crates link no network library; queries are answered from `mmap`. |
| `embed_query_potion.py`, the embedder for potion corpora | None. It loads the vendored potion table with numpy and tokenizers, no torch, and never opens a socket. |
| `embed_query_model.py`, the embedder for registry-model corpora (checkout only) | It sets the Hugging Face offline variables before loading a model. For the sentence-transformers presets (jina, wemm) a download needs `URNA_ALLOW_DOWNLOAD=1` and no local model directory. |
| `embed_query.py`, the `search-text` default (checkout only) | None by default. Same offline variables, same `URNA_ALLOW_DOWNLOAD=1` opt-in. |
| `urna doctor` | None. Its last check is one real potion embed, offline. |
Because the file carries its own vectors, text and indices, nothing else is fetched: no index server, no remote store, no telemetry.
## Where sockets exist [#where-sockets-exist]
Network access in urna happens during installation or on explicit opt-in.
| Where | What it downloads | Through |
| --------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `install.sh` | the release archive for the platform, the embedder payload, and a `.sha256` for each | `curl`, from GitHub releases or `URNA_RELEASE_BASE` |
| `install.ps1` | the same four files for Windows | `Invoke-WebRequest`, from GitHub releases or `URNA_RELEASE_BASE` |
| Homebrew, npm (and bun, pnpm, yarn), `cargo install`, `cargo binstall`, pip | the package or release archive | the package manager. The npm wrapper fetches the release archive at install time, or on the first run when the install step was skipped. |
| `urna setup`, payload step | `urna-embedder-payload.tar.gz` and its `.sha256` | a `curl` child process: the binary itself opens no socket. Only `https://` and `file://` sources, redirects only to `https://`, two retries, a 20 second connect timeout. |
| `urna setup`, Python step | `numpy` and `tokenizers` | `uv pip install` or `pip install`, against the package index they are configured for |
| Python embedders and the build forge | model weights from the Hugging Face hub | only with `URNA_ALLOW_DOWNLOAD=1` |
The setup screen's "network: curl, only while downloading" covers the payload step. The Python step also reaches a package index, unless you skip it with `--no-python` and point `URNA_PYTHON` at an interpreter that already has numpy and tokenizers.
## The potion payload [#the-potion-payload]
The default embedder is model2vec potion-base-8M, a static token table of about 30 MB (`model.safetensors` is 30,236,760 bytes). A static table needs no GPU and no model runtime: a query becomes the mean of its token rows, L2-normalized, with numpy. It is deterministic, so the same text gives the same vector on every machine.
The table travels with urna in two ways:
* The embedder payload, `urna-embedder-payload.tar.gz` in each release, holds the potion scripts and the table. The one-line installers unpack it; `urna setup` downloads and verifies it for every other channel. It lands in `/urna/forge/`.
* The Python wheel bundles the same table under `urna/models/potion-base-8M/`, so `urna.embed_potion` works right after `pip install "urna[embed]"`.
Once the payload and a Python with numpy and tokenizers are on the machine, `urna doctor` proves the chain with no network: interpreter, dependencies, embedder script, table, and one real embed that must return `dim=256` and a `model_hash`.
The payload carries only the potion embedder. A corpus built with a registry model needs the repository checkout and that model's own dependencies and weights on disk; see [Open a corpus you downloaded](/guides/use-a-corpus).
## How model downloads are gated [#how-model-downloads-are-gated]
Every Python entry point that can reach the Hugging Face hub (the potion module, the registry embedder, the `search-text` embedder, the fingerprint helper, the build forge) sets these variables to `1` before it imports a Hugging Face library:
* `HF_HUB_OFFLINE`
* `TRANSFORMERS_OFFLINE`
* `HF_DATASETS_OFFLINE`
They are set with `setdefault`, so a value already present in your environment wins. If your shell exports `HF_HUB_OFFLINE=0`, the scripts keep it.
`URNA_ALLOW_DOWNLOAD=1` is the opt-in, and only the exact value `1` counts. With it, `embed_query.py`, the fingerprint helper and the registry's sentence-transformers presets may fetch a model that is not on disk. The registry presets allow the download only when no local model directory resolved (no `--model-path`, no `URNA_MODEL_DIR_`, nothing in the cache). The potion embedder never downloads anything: a missing table is an error (exit 3 from the script, a failed check 5 in `urna doctor`), fixed by `urna setup` or, in a checkout, `git lfs pull`.
A separate switch guards model code rather than bytes. Presets that need `trust_remote_code` (jina, wemm) run only with an explicit opt-in: `allow_remote_code` in the build spec, `URNA_ALLOW_REMOTE_CODE=""` on the query side, and a pinned sha256 for each code file where the registry has pins. A downloaded `.urna` cannot turn that on by itself.
## What offline does not mean [#what-offline-does-not-mean]
Offline is about sockets, not trust. The query embedder runs Python code from the payload or the checkout, under an interpreter chosen by a fixed ladder that can pick up a nearby `.venv`; pin `URNA_PYTHON` in directory trees you do not control. And a `.urna` file from elsewhere is untrusted input even with no network involved. See [Security](/security).
# Presets and stored precision (https://docs.urna.dev/concepts/presets)
A `.urna` file stores one embedding per chunk, and you choose the precision it is stored at. Lower precision makes the file smaller and moves scores slightly away from the `float32` values. A build preset bundles that choice with the text encoding and the index sections; this page explains the trade, and [Build presets](/reference/presets) lists the exact settings.
## The four stored precisions [#the-four-stored-precisions]
| dtype | Bytes per vector at dim `d` | Layout | Constraint |
| --------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------ | ---------------------------- |
| `float32` | `4d` | IEEE float32, row-major | none |
| `float16` | `2d` | IEEE float16 | none |
| `int8` | `d + 4` | one `float32` scale per vector (max absolute value / 127) and `d` codes from -127 to 127 | none |
| `int4` | `d/2 + 2 x (d/64)` | one `float16` scale per block of 64 components (max absolute value / 7) and two 4-bit codes per byte, from -7 to 7 | `d` must be a multiple of 64 |
`int8` and `int4` sections also carry an 8-byte header. At dim 384 a vector takes 1,536 bytes as `float32`, 768 as `float16`, 388 as `int8` and 204 as `int4`.
The dtype applies to the default space (section 0x04). Named spaces have their own dtype, set per model with `space_dtype` in the build spec (see [Media and named spaces](/concepts/multimodal)).
## Truncation is a second axis [#truncation-is-a-second-axis]
`mrl_dim` keeps only the first `K` components of every vector and L2-normalizes the prefix again, before quantization and before the HNSW index is built. It multiplies with the dtype: `int8` at `mrl_dim = 256` is the `micro` point.
Truncation only makes sense for a model trained for it (matryoshka representation learning), where the first components carry most of the meaning. The file records `mrl_dim` and `full_dim`, `urna stats` prints them, and `urna ask` slices the query to the same length. A query at the full dim against a truncated file is a dimension mismatch. With `int4`, `K` must be a multiple of 64.
## Size against recall [#size-against-recall]
The repository publishes one measured run of the whole curve in [`data/measure/ladder.json`](https://github.com/hoffresearch/urna/blob/main/data/measure/ladder.json):
| Point | dtype | Dim | Size ratio | recall\@10 |
| ----------------------- | --------- | --: | ---------: | ---------: |
| `exact` | `float32` | 384 | 1.000 | 1.000 |
| `compressed` | `float16` | 384 | 0.339 | 1.000 |
| `tiny` | `int8` | 384 | 0.256 | 0.992 |
| `nano` | `int4` | 384 | 0.209 | 0.913 |
| `micro` (`mrl256-int8`) | `int8` | 256 | 0.223 | 0.810 |
| `mrl192-int8` | `int8` | 192 | 0.207 | 0.733 |
| `mrl128-int8` | `int8` | 128 | 0.190 | 0.659 |
| `mrl96-int8` | `int8` | 96 | 0.182 | 0.574 |
| `mrl256-int4` | `int4` | 256 | 0.191 | 0.777 |
| `mrl192-int4` | `int4` | 192 | 0.183 | 0.713 |
| `mrl128-int4` | `int4` | 128 | 0.174 | 0.627 |
The size ratio is the whole file against the `exact` build, text and indexes included. The run used `data/corpus_next.v1.urna` (30,725 chunks, dim 384, a MiniLM model not trained for truncation), 100 queries, k = 10, on the NEON backend.
Each query is a corpus chunk's own embedding plus tiny noise, and recall\@10 compares against the `float32` exact top 10 for that query. It tells you how stable the ranking stays when the vectors are quantized. It does not measure retrieval quality on real questions, and it is likely inflated.
On this corpus, `nano` (full dim, `int4`) keeps more recall than any truncated point, and truncation costs recall at every step, because MiniLM was not trained for it. Reach for `mrl_dim` when raw size matters more than the last recall points, or when the model is trained for truncation.
## How a score is computed [#how-a-score-is-computed]
Every score urna returns is a cosine recomputed by an exact rerank. HNSW, the chunk graph and BM25 only choose candidates; each candidate is then scored against the query with the same kernel the flat scan uses. The candidate list can miss a chunk, but a score is never an index-side estimate.
The rerank reads one source: a full-precision slab (section 0x09) when the file has one, otherwise the stored embeddings (0x04). The reader accepts a 0x09 slab, but no 0.5.1 writer emits one. So in practice:
* A `float32` file reranks at full precision.
* A `float16`, `int8` or `int4` file reranks at stored precision: the query stays `float32` and is scored against the stored codes and scales, with `float32` accumulation.
"Real cosine at stored precision" is still a cosine between the query and the vector the file holds, not an approximation from the index. It differs from the `float32` cosine by the quantization error of the stored vector.
## The rerank source is disclosed [#the-rerank-source-is-disclosed]
The rerank source is reported, so a reader never has to guess which kind of score they got.
`urna ask --disclose explain` prints it before the answer:
```text
route: hnsw
candidates: exact=0 ann=12 bm25=0 graph=0 fusion=none
rerank_source: real cosine
recall: (not computed; rerank guarantees real cosine)
```
On a `float16`, `int8` or `int4` file the line reads `real cosine at stored precision`. `urna retrieve` writes the same fact on every hit as `"rerank_source": "full_precision"` or `"stored_precision"`, and so does `RetrieveHit.rerank_source` in Python. `urna stats` prints the `dtype`, and `mrl_dim` and `full_dim` when the file is truncated.
## Choosing a preset [#choosing-a-preset]
* `exact` when file size is not the constraint, or when you need the `float32` ground truth to compare other builds against.
* `compressed` for a file about a third the size with scores at `float16`. There is no index, so every query scans all vectors.
* `tiny` for about a quarter of the size, with an HNSW index for large corpora and `int8` scores.
* `nano` for the smallest full-dim file, when the dim is a multiple of 64 and the recall above is acceptable for your data.
* `micro` (a recipe, built with `dtype="int8", mrl_dim=256`) only with a model trained for truncation, or when size wins over recall.
* `hybrid` writes `float32` vectors, an HNSW index and a BM25 index. It is the default of `urna build --spec`.
A `hybrid` file declares `index_type = "hnsw"`, so `urna ask`, `urna retrieve` and `urna search-text` never read its BM25 section. See [known limits](/limits).
Changing the preset changes `content_hash`, and with it every citation: the dtype, the truncation and the index type are all part of the hashed content. Pick the preset before you publish citations. The search routes themselves are described in [Search paths and the exact rerank](/concepts/search).
# Reproducible builds (https://docs.urna.dev/concepts/reproducibility)
A `.urna` build is deterministic: the same rows, the same vectors and the same settings produce the same bytes, and therefore the same `file_hash`, `content_hash` and citations. This page covers what the writer pins, what the forge (`urna build --spec`) records so you can repeat a build, and what you still have to check yourself.
## What the writer pins [#what-the-writer-pins]
The Rust writer has no clock, no randomness and no dependency on the CPU it runs on:
* The manifest's `created` field is the only timestamp. The writer never fills it by itself. `reproducible(true)` sets it to `1970-01-01T00:00:00Z`, so a build that passes a `created` value still comes out identical.
* The HNSW index is built from a fixed seed (`hnsw_seed = 42` in `urna.build`) with a distance function that does not use SIMD, so the index bytes do not depend on the machine.
* Sections are sorted by id, aligned to 64 bytes and padded with zeros.
Byte identity still needs identical inputs: the same chunks in the same order, the same provenance JSON with the same key order, the same text encoding, dtype and `mrl_dim`, and the same HNSW parameters and seed.
| Entry point | `reproducible` default |
| ------------------------------------- | ---------------------- |
| `urna.build` | `False` |
| `builder.BuildConfig` (checkout only) | `True` |
| build spec, `[corpus] reproducible` | `true` |
With `reproducible` off and no `created` value, nothing is written into `created` and the build is deterministic anyway.
## Reproduction levels [#reproduction-levels]
The forge documentation declares three levels:
| Level | Claim | Checked by code |
| ----- | -------------------------------------------------------------------------- | ---------------------------------------------------------------- |
| L1 | the same top-k on any machine | no |
| L2 | each vector's cosine within 1e-5 of the original, on the same device class | no |
| L3 | a byte-identical `file_hash` | partly: the build lock is compared; the byte comparison is yours |
L1 and L2 are statements about the embedding model and hardware, and no tool in 0.5.1 measures them. L3 is the one the forge supports directly, with the build lock and `--rebuild-only`.
## The build lock [#the-build-lock]
Every `urna build --spec` run writes `.build.lock.json` next to the outputs:
| Field | Content |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| `lock_schema_version` | `1` |
| `platform` | `os`, `machine`, `release` |
| `packages` | the Python version and the installed versions of `numpy`, `torch`, `transformers`, `sentence-transformers`, `open-clip-torch`, `tokenizers`, `pillow` |
| `tools` | `ffmpeg`, `ffprobe`, `cjxl`, `djxl`, `ssimulacra2`: path, version and SHA-256 of the binary, or `null` when absent |
| `models` | `model_hash` per preset |
| `device` | `URNA_ST_DEVICE`, or `auto` |
| `resolved_spec` | the full spec after profiles and `${VAR}` expansion, without `output.cache_dir` |
The comparison ignores where things live: `resolved_spec.spec_path`, `output.dir`, `output.cache_dir`, `source.path`, `source.db` and `source.image.path_template`. Row content is covered per item by its input hash, so a rebuild under another data root still compares clean.
Everything else counts. The lock records all five media tools and all seven packages for every build, text-only builds included, so a different `ffmpeg` or `torch` version shows up as a divergence even when the build never used it. `source.input_dir` and `source.labels` are not in the ignore list, so an `image_dir` corpus rebuilt from a moved directory diverges on location alone.
The version field of `cjxl`, `djxl` and `ssimulacra2` holds their usage text instead of a version, because the lock runs ` -version`, which only `ffmpeg` and `ffprobe` accept. The SHA-256 still pins the binary.
## The embed cache and its key [#the-embed-cache-and-its-key]
Computing embeddings is the slow part of a build, so the forge caches them outside the output directory. The cache key is three hashes, called the triad:
| Hash | Identifies | Changes when |
| ----------------------- | --------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `model_hash` | the model | weights, tokenizer, processor, remote code, normalize or dtype policy change |
| `embedding_recipe_hash` | how the model is used | query and document modes, image prompt, preprocess version, `image_max_side`, `encode_kwargs`, dtype, device class, image input mode, and for image spaces over decoded media the decoder settings |
| `corpus_input_hash` | the rows | any item's text, source image bytes, label or `chunker_version` changes, or rows are added, removed or reordered |
Each entry is `/embed//<16 hex>.npz`, where the name is a hash of the triad plus which arrays the spec needs (text, image, deduplicated). A checksum sidecar and a lock file sit beside it. Two specs with the same rows and model read the same entry wherever their outputs go, a changed knob adds a new entry instead of overwriting one, and a torn file is recomputed.
The cache root is, in order: `--cache-dir`, `[output] cache_dir`, `URNA_CACHE_DIR`, then `${XDG_CACHE_HOME:-~/.cache}/urna`. The location is not part of the lock.
`model_hash` probes are cached beside the entries, under `/models/`, so a warm build does not load a model only to learn its hash. A probe is trusted while the model directory's file listing (names and sizes) is unchanged. `potion` and the open\_clip presets have no model directory to list, so swapping their weights is noticed only when an embed actually runs.
## Rebuilding with --rebuild-only [#rebuilding-with---rebuild-only]
`--rebuild-only` re-emits the outputs from cached vectors and compares the new lock with the one on disk:
```sh
urna build --spec corpus.toml --rebuild-only
```
What it does:
1. Loads the rows again and recomputes `corpus_input_hash`.
2. Reuses the encoded media when the saved media state matches. When it does not match, it re-encodes silently.
3. Requires every embed cache entry to hit. A miss stops the build: `--rebuild-only: cache for '' is missing or stale (triad mismatch); run a full build`.
4. Writes the `.urna` files, overwriting the old ones.
5. Compares the new lock with the old one and prints `[forge] warning: build.lock divergence (L3 not claimable): ...` when they differ.
6. Writes the new lock and the manifest, overwriting the old ones.
It does not compare `file_hash`. To claim L3, keep the old hash and compare it yourself. The reported `file_hash` is the SHA-256 of the whole file, the same value `sha256sum` prints (`shasum -a 256` on macOS):
```sh
cp out/corpus.build.lock.json /tmp/corpus.lock.before
sha256sum out/corpus.urna > /tmp/corpus.sha256
urna build --spec corpus.toml --rebuild-only
sha256sum -c /tmp/corpus.sha256
```
Keep a copy of the lock too: after a divergence warning the new lock replaces the old one.
Making a lock divergence an error needs `--strict-env`, which exists only on the Python tool: `python python/tools/urna_forge.py --spec corpus.toml --rebuild-only --strict-env`. `urna build --strict-env` is a usage error (exit 2). The same goes for `--seed` and `--json`.
`--resume` is narrower than the name suggests: only the media stage reads saved state. Embed caches are read on every build, with or without it, and rows and outputs are always recomputed.
## What moves citations anyway [#what-moves-citations-anyway]
A reproducible build reproduces its citations, but several ordinary edits change them:
* `chunker_version` is part of every `chunk_id`: changing it changes every citation.
* In a forge build, a row's span is its position after sorting by `order_by`. Adding or removing a row shifts every later row's `chunk_id`.
* `order_by` sorts by string value, so `"10"` sorts before `"2"`.
* `--sample N` renumbers rows, so a pilot build never shares citations with the full build.
* The preset's dtype, `mrl_dim` and index type are part of `content_hash`.
Media encoders are pinned to fixed parallelism (`lp=2` for SVT-AV1, `-j 8` for `avifenc`) so the encoded bytes do not depend on the core count. For st\_multimodal models, the dtype policy follows the device unless you pin it, so building on cuda and on mps gives different `model_hash` values: see [Model registry](/reference/models#identity-model_hash).
See [Citations and hashes](/concepts/citations) for what each hash covers, and [Build artifacts](/reference/spec/artifacts) for the full lock and manifest schemas.
# Search paths and the exact rerank (https://docs.urna.dev/concepts/search)
urna has several ways to find candidate chunks and one way to score them. Every path, however it generates candidates, ends by recomputing the cosine between the query and each candidate from the stored vectors. The score you get back is that recomputed cosine, never an index distance or a fusion score.
## The scoring contract [#the-scoring-contract]
Every file declares `metric = "ip"` and `normalize = "l2"`: vectors are stored L2-normalized and scored by inner product, which on unit vectors is cosine similarity. Hits always report `score_type: "cosine"`. The writer does not check the norm of the vectors it is given, so this holds when the embedder that built the file normalized its output, as the bundled embedders do.
Before any path runs, the runtime checks the query and stops with a typed error on the first problem:
1. `k` is greater than 0.
2. The query is not empty.
3. Its length equals the file's `embedding_dim`.
4. Every value is finite (no NaN or Inf).
5. Its norm is not zero.
Then it L2-normalizes the query, so you do not have to.
## The paths [#the-paths]
| Path | Candidates come from | Reached by | `recall` in the result |
| ----------- | -------------------------------------------------------- | ---------------------------------------------------------------------------------- | ---------------------- |
| Exact | Every chunk | [search](/reference/cli/search), `ask`/`retrieve` on `index_type = "exact"` | `1` |
| HNSW | The HNSW graph (`0x07`) | [search-ann](/reference/cli/search-ann), `ask`/`retrieve` on `index_type = "hnsw"` | not computed |
| Hybrid | HNSW (or exact) shortlist plus a BM25 shortlist (`0x08`) | `ask`/`retrieve`/`search-text` on `index_type = "hybrid"`, Python `search_hybrid` | not computed |
| Graph | Exact seeds expanded over the chunk graph (`0x0C`) | [search-graph](/reference/cli/search-graph), Python `search_graph` | not computed |
| Named space | Every vector in one space band | [search-space](/reference/cli/search-space), Python `search_space` | `1` |
The CLI prints a `recall` that was not measured as `(not computed; rerank guarantees real cosine)`. The rerank guarantees that each returned score is a real cosine. It does not guarantee that an approximate path found the true top `k`; measure that with [urna benchmark](/reference/cli/benchmark) `--ann`.
### Exact [#exact]
The exact path scores every chunk and sorts, with ties kept in file order. It is the ground truth the other paths are compared against.
### HNSW [#hnsw]
The HNSW path walks the graph stored in `hnsw_index` to collect a candidate list, then reranks it exactly. If the file has no HNSW section, the path falls back to exact and the result says `index_type: exact`.
The beam width is `max(ef, k, ef_construction)`, where `ef_construction` is the value the file was built with. Python and the forge build with `ef_construction = 400`.
The runtime never searches with a beam narrower than the file's `ef_construction`. On a file built with the default 400, any `--ef` or `--candidates` value below 400 runs at 400, so an `ef` sweep below that value is flat. Only values above `ef_construction` change the result. See [Known limits](/limits).
### Hybrid [#hybrid]
The hybrid path builds two shortlists. The vector shortlist comes from HNSW when the file has it (with the same beam floor as above), else it is the top `candidates_per_path` of an exact scan. The BM25 shortlist is the top `candidates_per_path` chunks by BM25 over the stored text. There is no BM25-only path: BM25 runs only inside hybrid. The path merges the two lists with reciprocal-rank fusion (RRF, `k = 60`), rescores every member of the merged set by exact cosine, and returns the top `k` by cosine.
That last step defines what hybrid does in urna: the RRF order is discarded and the final order is pure cosine. BM25 can bring a chunk into the candidate set that the vector shortlist missed; it never lifts a chunk above one with a higher cosine. A rare term or proper noun only reaches the top `k` if its chunk's cosine earns it.
Two edge cases:
* The path never falls back. Without a BM25 section the lexical list is empty and the route is still `hybrid`.
* Without HNSW, the vector shortlist holds at most `candidates_per_path` chunks, with no minimum of `k`. With no HNSW, no BM25 and `candidates_per_path` below `k`, fewer than `k` hits come back.
BM25 uses `k1 = 1.5` and `b = 0.75`. Its tokenizer lowercases runs of letters and digits, splits on anything else, and drops tokens shorter than 2 characters. There is no stemming and no stop-word list. A run with no spaces or punctuation stays one token, so an unspaced CJK clause only matches an identical whole run.
`preset="hybrid"` in `urna.build`, and the forge's default `[build] preset = "hybrid"`, write an HNSW index and a BM25 index but declare `index_type = "hnsw"`. `ask`, `retrieve` and `search-text` route on the declared value, so on these files they take the HNSW path and never read the BM25 index. The BM25 index is reached only by calling `search_hybrid` yourself, from Python (`UrnaFile.search_hybrid`) or the Rust runtime. A file declares `index_type = "hybrid"` only when built with the Rust builder's `.hybrid()`. See [Known limits](/limits).
### Graph [#graph]
The graph path seeds from the exact top `max(ef, k)` chunks, expands them over the chunk-to-chunk graph for `hops` steps (all edge types, capped at 8 times the seed count), and reranks the union exactly. If the file has no `graph_adjacency` section, or its manifest does not set `capabilities_ext.graph_present`, it falls back to exact.
Because the seeds already are the exact top `max(ef, k)` and the rerank uses the same scores and ordering, the hits equal the exact path's hits. The graph adds neighbours to the candidate set, but a neighbour cannot outrank a seed it would need to displace. The path costs an exact scan plus the traversal.
### Named spaces [#named-spaces]
A file can carry extra embedding spaces, for example an image tower next to the text embeddings (see [Media and named spaces](/concepts/multimodal)). A named-space search is an exact scan over one space's band. The query must come from that space's model and have that space's dim. Hits map back to the same chunks, so they carry the same `chunk_id` and citation as a text hit on that chunk. The text paths never read a space band, and a space search never falls back to the text embeddings.
## How ask and retrieve choose a path [#how-ask-and-retrieve-choose-a-path]
[urna ask](/reference/cli/ask), [urna retrieve](/reference/cli/retrieve) and [urna search-text](/reference/cli/search-text) embed the query text, pass it through [the model gate](/concepts/model-gate), and then route on the manifest's declared `index_type`:
| Declared `index_type` | Path | Candidates |
| --------------------- | ------ | ------------------------------------------------------------------- |
| `exact` | Exact | Every chunk |
| `hnsw` | HNSW | `--candidates`, default `max(4 * k, 64)`, then the beam floor above |
| `hybrid` | Hybrid | `--candidates` per path, same default |
The manifest admits only `exact`, `hnsw` and `hybrid`, so none of these verbs takes the graph path. Routing does not look at which sections exist or at the `capabilities` flags; `urna stats` shows the `index_type` that decides it.
The quickstart corpus declares `hnsw` and also carries BM25 and a graph. `ask --disclose explain` shows the route it took:
```sh
urna ask examples/quickstart/out/quickstart.urna "can I use this offline" -k 1 --disclose explain
```
```text
route: hnsw
candidates: exact=0 ann=12 bm25=0 graph=0 fusion=none
rerank_source: real cosine
recall: (not computed; rerank guarantees real cosine)
```
The file has 12 chunks, so the HNSW candidate list is the whole corpus. BM25 and the graph are present and unused.
## The rerank source [#the-rerank-source]
The rerank reads vectors from one slab. When a full-precision slab (`embeddings_fp`, `0x09`) is present, it reads that. Otherwise it reads the stored `embeddings` section at its stored `dtype`. No writer in 0.5.1 emits `0x09`, so in practice the rerank reads the stored embeddings.
| Effective dtype | `rerank_source` in `retrieve` JSON | `--disclose explain` line |
| ------------------------- | ---------------------------------- | --------------------------------- |
| `float32` | `full_precision` | `real cosine` |
| `float16`, `int8`, `int4` | `stored_precision` | `real cosine at stored precision` |
A stored-precision score is still a cosine recomputed from the vectors, not a proxy from the index. It is the cosine between the query and the quantized vector, so it can differ from the `float32` score of the same pair. See [Presets and stored precision](/concepts/presets).
Scoring runs on AVX2 (with FMA) on x86\_64, NEON on aarch64, or a scalar fallback, detected once per process; `float16` has no AVX2 kernel and scores with the scalar kernel on x86\_64. Set `URNA_FORCE_SCALAR` to any value other than `0` to force the scalar kernels (see [Environment variables](/reference/environment)). The `int4` kernel gives bit-identical scores on all three.
The per-verb flags are in the reference, starting with [urna search](/reference/cli/search).
# Contributing (https://docs.urna.dev/contributing)
This page covers how to set up a checkout of [hoffresearch/urna](https://github.com/hoffresearch/urna), what the release gate and CI run, how to fuzz a decoder change, and the rules a pull request has to follow.
## How a change lands [#how-a-change-lands]
1. Fork the repository and branch from `main`, for example `git checkout -b feature/short-description origin/main`.
2. Keep each pull request to one concern.
3. Add or update tests. New behavior needs a new test, written against real artifacts (built `.urna` files, the golden fixtures, real corpora) rather than mocks, covering the happy path, the error path and one edge case.
4. Run `scripts/release_check.sh` locally before you push.
5. Open the pull request against `main` and fill in the template. Its description becomes the squash commit.
The rules on `main`: squash merge only, one approving review, signed commits and linear history, and the branch is deleted after the merge. CI results are not a required status check, so run the gate yourself. Sign your commits with an SSH or GPG key registered on GitHub (`git config commit.gpgsign true`). Write commit messages in plain English, without a conventional-commits prefix; the message explains why, the diff shows what.
## Set up a checkout [#set-up-a-checkout]
You need:
* Rust 1.88 or newer, edition 2024. The CLI crate (`urna`) sets `rust-version = "1.88"`; `urna-format`, `urna-runtime` and `urna-python` keep the workspace's 1.85, but `cargo build --workspace` builds the CLI, so 1.88 is the floor for a checkout.
* A C compiler, because `zstd-sys` compiles C.
* Python 3.12 or newer.
* Git LFS, for the potion table and the measurement corpus.
* macOS or Linux for the Python side. The gate copies the extension only on those two, and the checkout loader in `python/urna.py` has no Windows path. The Windows CI job covers the CLI.
```sh
git clone https://github.com/hoffresearch/urna.git
cd urna
git lfs pull
python3 -m venv .venv
. .venv/bin/activate
pip install numpy tokenizers pillow ruff
cargo build --release --workspace
PYO3_PYTHON="$PWD/.venv/bin/python" cargo build --release -p urna-python --features pyo3/extension-module
cp target/release/lib_urna.dylib python/_urna.so # macOS
cp target/release/lib_urna.so python/_urna.so # Linux
cp scripts/pre-commit .git/hooks/pre-commit && chmod +x .git/hooks/pre-commit
```
Build the extension with `--features pyo3/extension-module`. Without it the library links `libpython` directly, and under a statically linked interpreter such as uv's standalone Python it loads a second runtime and segfaults at import. `PYO3_PYTHON` pins the build to the interpreter that runs the tests. Rebuild `python/_urna.so` after any change to `urna-format`, `urna-runtime` or `urna-python`.
`numpy`, `tokenizers` and `pillow` are the `forge` dependency group in the root `pyproject.toml`. Registry models and media backends need more; each preset prints its own install line (see [the model registry](/reference/models)).
If the LFS budget or a missing `git-lfs` gets in the way, `sh scripts/fetch_potion.sh` fetches the potion table alone from Hugging Face at a pinned revision and accepts it only when its SHA-256 matches the LFS pointer. Without the real table, embedding fails and `urna doctor` exits `5`. `urna setup` cannot fix that inside a checkout, because the checkout's `python/forge` wins the embedder lookup.
The demo datasets under `data/demo/` are local only and gitignored; `data/demo/Instructions.md` has the commands to fetch them. The test suites do not need them.
The pre-commit hook blocks any staged `.urna`, `*.sidecar.jsonl`, `cohort_*.urna` or file under `pairs/`, except three sanctioned files: `data/corpus_next.v1.urna`, `data/measure/fakerecogna_exact.urna` and `crates/urna-format/tests/fixtures/golden_v1_minimal.urna`. Copy it into `.git/hooks` rather than pointing `core.hooksPath` at `scripts/`, which would disable the LFS hooks. Corpus data, personal data above all, stays outside the repository.
## The release gate [#the-release-gate]
`scripts/release_check.sh` is the local gate. It stops at the first failure.
| Step | What runs |
| ---- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 | `cargo build --release --workspace` |
| 2 | `cargo test --release --workspace`, with a passed, failed and ignored summary |
| 3 | `cargo clippy --workspace --all-targets -- -D warnings` (output hidden; rerun it by hand to see a lint) |
| 4 | `cargo fmt --all --check` |
| 5 | The 639-line guard on every non-test `.rs` file under `crates/` |
| 6 | The extension rebuild with `pyo3/extension-module`, copied to `python/_urna.so` |
| 7 | Eight Python suites: `test_e2e`, `test_builder`, `test_search_text_model_hash`, `test_image_corpus`, `test_forge_spec`, `test_quality_gate`, `test_cli_space`, `test_query_embedder_routing` |
| 8 | Ruff through `scripts/ruff_check.sh`, skipped when ruff is not importable |
| 9 | `python/tools/measure_presets.py` on the LFS corpus `data/corpus_next.v1.urna` |
| 10 | `python/tools/compare_measure.py`, the regression gates against `data/measure/baseline.json` |
The Python suites that need `ffmpeg`, `ssimulacra2` or `cjxl` skip those legs when the tools are missing.
| Variable | Default | Effect |
| --------------- | ------------------------------------ | ------------------------------------------------------------- |
| `URNA_PYTHON` | `./.venv/bin/python`, else `python3` | The interpreter for the extension build and every Python step |
| `URNA_BASELINE` | `data/measure/baseline.json` | The baseline for step 10 |
| `URNA_QUERIES` | `100` | Query count for step 9 |
| `URNA_K` | `10` | Top-k for step 9 |
| `URNA_OUT` | `/tmp/release_check_post.json` | Where step 9 writes its JSON |
The gate does not run `cargo deny`, the semver check, the Windows job, the engine-only clippy, `forge-core`, `cargo-fuzz`, `tests/test_offline_guard.py`, `tests/test_blob_bridge.py`, `tests/test_space_bridge.py` or the self-tests under `python/forge/`. Its last line suggests a tag command; releases are tagged by the maintainer with a signed tag, so a contributor can ignore it.
## CI [#ci]
`.github/workflows/ci.yml` runs on every pull request, on push to `main`, nightly at 03:17 UTC and on manual dispatch, with `RUSTFLAGS=-D warnings`.
| Job | When | What it runs |
| ------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `rust` | Not on the nightly schedule; Ubuntu and macOS | fmt, clippy on the workspace, clippy on `urna` with `--no-default-features`, build and test (debug), the runtime benches compiled, the mutation harnesses at 6000 (format) and 800 (runtime) iterations, the 639-line guard on `crates/**/src`, and `forge-core` fmt, clippy and test |
| `cli-windows` | Not on the schedule; Windows | clippy on `urna`, its unit tests and `setup_e2e` |
| `fuzz` | Not on the schedule | Every `cargo-fuzz` target for 90 seconds |
| `deny` | Every event | `cargo deny` on the workspace, `forge-core` and `fuzz` |
| `semver` | Pull requests | `cargo semver-checks` on `urna-format` against the base commit |
| `fuzz-soak` | Schedule or dispatch | Each target for 1800 seconds by default, corpus carried between nights |
| `miri` | Schedule or dispatch | `cargo miri test -p urna-format` |
| `python` | Not on the schedule | Ruff through `scripts/ruff_check.sh` |
CI and the local gate overlap but are not the same. CI builds no Python extension and runs none of the eight Python suites, only ruff. It adds `deny`, `semver`, the Windows job, fuzzing, the benches build and `forge-core`, which the local gate skips. A pull request needs both.
## Fuzzing and miri [#fuzzing-and-miri]
A change to a section decoder or a search path should run the deterministic mutation harnesses, which also run under `cargo test`:
```sh
cargo test --release -p urna-format --test mutation_fuzz
cargo test --release -p urna-runtime --test mutation_fuzz
URNA_MUTATION_ITERS=25000 cargo test --release -p urna-format --test mutation_fuzz
```
The default is 1500 iterations for the format harness and 250 for the runtime harness. Half of the mutations reseal the checksums so the corruption reaches the decoders.
The coverage-guided targets live in the separate `fuzz/` workspace and need a nightly toolchain and `cargo install cargo-fuzz`. The targets are `urna-view`, `section-decoders`, `runtime-indexes` and `mmap-open-search`:
```sh
cd fuzz
cargo +nightly fuzz run urna-view -- -max_total_time=600
```
`sh scripts/fuzz_soak.sh [seconds]` runs every target for an hour each by default, keeping the corpus under `fuzz/corpus/`; `URNA_FUZZ_TARGETS` narrows the list. A new codec gets an arm in `fuzz/fuzz_targets/section_decoders.rs`. A finding becomes a `crates/*/tests/negative_*.rs` test and a `fuzz/seeds/regress-*.bin` seed before the fix.
To run miri the way CI does:
```sh
MIRIFLAGS="-Zmiri-disable-isolation -Zmiri-tree-borrows" cargo +nightly miri test -p urna-format
```
Tests that call zstd, the `half` f16 conversions on aarch64 or the FSST table build are marked to skip under miri, with the reason in the source.
## Code rules [#code-rules]
Rust:
* `cargo fmt --all` with the settings pinned in `rustfmt.toml` (width 100).
* `cargo clippy --workspace --all-targets -- -D warnings` is a hard gate. Suppress a lint on one item with `#[allow(clippy::name)]` and a one-line justification, never globally. `clippy.toml` pins cognitive complexity at 15, type complexity at 250 and at most 7 arguments.
* `clippy::unwrap_used` and `clippy::undocumented_unsafe_blocks` are denied workspace-wide. Tests may unwrap. Parse paths read fixed-width fields through `urna_format::bytes`, and every `unsafe` block carries a `// SAFETY:` comment that names its invariant.
* Public items get a doc comment that explains why, not what.
Python:
* Ruff with `target-version = "py312"`, line length 100 and the lint set `E F W I B UP SIM`, configured in the root `pyproject.toml`.
* `scripts/ruff_check.sh` holds the one file list that the gate and CI both check. When you touch a Python module, add it to the list and make it clean.
* Private helpers in `python/tools/` start with `_`, for example `_baseline_decoder.py`.
File size: 639 lines per code file. Above that, split along single-responsibility lines in the same pull request. Tests, data, generated files, lockfiles, JSON, YAML, TOML and vendored files are exempt. The guards enforce it for Rust under `crates/` only; Python, `forge-core` and `fuzz` follow it by convention.
The format:
* The container is frozen at v1. A byte-level change either fits inside v1, with a new section id or encoding id that does not collide with one already emitted, or bumps `URNA_FORMAT_VERSION` and ships as v2. The ids in use are listed in [sections](/reference/format/sections) and [encodings](/reference/format/encodings).
* New manifest fields are `Option` with `skip_serializing_if`, so an unset field writes nothing and old files stay byte-identical. New capability flags go in `capabilities_ext` or the flattened `extra` map, never as a new required bool. `crates/urna-format/tests/manifest_additivity.rs` guards this.
Repository docs start with a YAML header (`project`, `audience`, `status`, `last-updated`, `domain`); the README, the license, the pull request template and `llms.txt` are exempt. No emoji and no em dash in text. A change to architecture, module boundaries, data flow or doc locations updates `docs/arc/ARC.toml` in the same pull request, and a user-visible change gets a line under `[Unreleased]` in `docs/CHANGELOG`. The pull request template lists the rest as checkboxes. A coding agent reads `.contracts/.agents/AGENTS.md`; the root `CLAUDE.md` is a link to it.
## Report issues [#report-issues]
* Bugs and feature requests: [GitHub issues](https://github.com/hoffresearch/urna/issues).
* Security problems: never a public issue. Follow [report a vulnerability](/security#report-a-vulnerability).
* Questions about the format: a GitHub discussion, or `docs/arc/ARC.toml`.
A bug report should carry the `file_hash` and `content_hash` of the `.urna` involved and the `simd_backend` (all three from `urna stats `), the exact CLI or Python invocation, and the error output.
## Conduct and license [#conduct-and-license]
The project follows [`docs/CODE_OF_CONDUCT.md`](https://github.com/hoffresearch/urna/blob/main/docs/CODE_OF_CONDUCT.md). Contributions are licensed under the [MIT license](https://github.com/hoffresearch/urna/blob/main/docs/LICENSE); copyright vests in Hoff Research as the maintainer.
How the pieces fit together is on the [architecture](/architecture) page.
# Data governance (https://docs.urna.dev/data-governance)
A `.urna` file is meant to be copied: to laptops, edge nodes and machines with no network. When its content includes personal or sensitive data, that design has consequences. This page lists what the file stores, what it does not, how removal works, where processing happens and what a build leaves on disk besides the file. It describes the software; it does not assess your legal obligations.
## A .urna file is a copy of the data [#a-urna-file-is-a-copy-of-the-data]
A `.urna` is a datastore, not a cache. It holds the text of every chunk in cleartext, next to embeddings derived from that text. The format has no encryption at rest: zstd is compression, not confidentiality. A file built over personal data is a copy of that data and needs the same controls as its source, including full-disk encryption (FileVault, LUKS or equivalent) on every volume that holds it.
## What the file contains [#what-the-file-contains]
Six sections are required in every file. The others appear only when the build asked for them.
| Part | What it stores |
| -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `chunks_canonical` (0x02) | The canonical text of every chunk. This is the text that `ask`, `retrieve` and `cite` return. |
| `chunks_original_spans` (0x03) | A `source_uri` and a byte span per chunk. The URI is whatever the builder wrote: a `source_uri` column or `item:///` in a forge build, `media://...` for media rows, any string in a direct `urna.build` call, local paths included. |
| `chunk_ids` (0x01) | One `sha256:` id per chunk, computed over the text, the URI, the span and `chunker_version`. Anyone holding a candidate text and its source can recompute the id and confirm a match. |
| `embeddings` (0x04) | One vector per chunk, a derived representation of the text. |
| `provenance` (0x05) | Any JSON the builder passed. `urna build --spec` writes `{"dataset": , "corpus_input_hash": }`; `urna.build` writes `{}` unless you pass a dict. |
| `search_contract` (0x06) | The metric, score type, normalization, index type and rerank policy. |
| Manifest | `embedding_model`, `model_hash`, `chunker_version`, dimensions and dtype, plus the optional `title`, `version`, `created`, `description`, `license`, `authors` and any extra keys the builder added. |
| `bm25_index` (0x08, optional) | The lowercased token vocabulary of the canonical text with term frequencies per chunk: a second copy of the words, without their order. |
| `hnsw_index` (0x07), `graph_adjacency` (0x0C) (optional) | Neighbor lists between chunk ordinals. No text. |
| `blob_refs` (0x14, optional) | Per media record: the SHA-256 of the original bytes, the original URI, the byte length and whether the bytes are inlined. |
| `blob_data` (0x17, optional) | The media bytes themselves, as encoded by the build, when the spec sets `[output] embed_media = true`. |
| `space_table` (0x15) and bands (optional) | Per named space: name, dimension, dtype, `model_hash` and one vector per chunk (image vectors, for example). |
The layout of every section is in [sections](/reference/format/sections), and the manifest fields in [manifest](/reference/format/manifest).
## What the file does not contain [#what-the-file-does-not-contain]
* Encryption, access control or per-chunk permissions. Anyone who can read the file can read all of it.
* A signature or a verified author. The hashes prove integrity, not origin, and `authors` is free text. See [security](/security#what-the-hashes-prove).
* Model weights. The manifest records the model name and its `model_hash`; the query embedder lives outside the file.
* A record of queries. The runtime maps the file read-only, so querying never changes it.
* A delete or tombstone mechanism. Chunks cannot be marked as removed; the `edit_journal` section id is reserved and nothing writes it.
* The original source files, with one exception: inlined media in `blob_data`. For text rows, only the canonical text and the URI are stored.
## Remove or correct content [#remove-or-correct-content]
A distributed file cannot be edited in place. To erase or correct a chunk, fix the source rows and rebuild. The rebuild changes the file:
* `content_hash` covers the six canonical sections, so any changed chunk gives the new file a new `content_hash` and a new `file_hash`.
* A citation is `urna:///`. `urna cite` rejects a citation whose `content_hash` does not match the file, so every citation issued from the old build stops resolving against the new one.
* In a forge build the byte span of a row is its ordinal after sorting, and the span enters the `chunk_id`. Removing a row shifts the `chunk_id` of every row after it.
Copies already shipped are not recalled. The runtime never opens a socket, so there is no channel to reach them. Before you distribute a file that holds personal data, plan the removal process:
* Version the corpus. Each build has its own `content_hash` and `file_hash`, which identify it exactly.
* Publish the hashes of superseded builds, and require operators to pull the current build and delete the old copies.
* Treat the embeddings and the BM25 vocabulary as derived personal data, in scope for the same requests as the text.
* Record consent and provenance for third-party content before it goes into a build.
Data-subject rights such as erasure and rectification (GDPR articles 16 and 17, LGPD article 18) cannot be served by editing a copy that has left your hands. For special-category data, such as health records, confirm with counsel that distributing immutable copies is compatible with the rights that apply before you ship.
Attaching an HNSW index writes `index_type = "hnsw"` and `rerank_policy = "exact"` into the canonical `search_contract` section. An `exact` build and a `tiny`, `nano` or `hybrid` build of the same rows therefore have different `content_hash` values and cite differently, even with identical text. Keep the preset fixed across rebuilds when you need old citations to line up with new ones. See [known limits](/limits).
Some changes leave `content_hash` and every citation unchanged: the text encoding (raw or zstd), the graph, media sections, named spaces, and manifest-only fields such as `title` or `created`. The full table is in [citations and hashes](/concepts/citations).
## Where processing happens [#where-processing-happens]
Building and querying run on your machine. The Rust runtime never opens a socket and answers from the memory-mapped file. The query embedders run as local Python processes, and the Python embedders and builders set the Hugging Face offline variables unless you opt into downloads with `URNA_ALLOW_DOWNLOAD=1`. Media encoders (`ffmpeg`, `avifenc`, `cjxl`) run as local subprocesses.
The exceptions are the installers, `urna setup` and explicit download opt-ins. They are listed in [where urna opens a network connection](/security#where-urna-opens-a-network-connection), and the design is explained in [offline by construction](/concepts/offline).
## What a build leaves on disk [#what-a-build-leaves-on-disk]
Personal data can outlive a `.urna` in these places. Remove them together with the file.
| Location | Written by | Contents |
| ------------------------------------------------- | ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `