Architecture
How urna is built: the four Rust crates, the Python forge, how a build becomes a .urna file, and how a query flows through the runtime to a citation.
urna is a Rust workspace of four crates plus a Python layer. Python builds files; Rust owns the format, serves the queries and runs the CLI. This page maps the pieces and follows a build and a query from end to end.
The pieces
| Piece | Language | What it owns |
|---|---|---|
urna-format | Rust library | The frozen v1 container: layout, manifest, reader, writer, section codecs, wire encodings, chunk_id and the hashes |
urna-runtime | Rust library | Opening a file through mmap, query validation, SIMD dispatch, exact, HNSW, graph, hybrid and space search, the exact rerank, and the HNSW and BM25 index builders |
urna-cli | Rust binary | The urna command: a clap surface over the two libraries, plus setup and tui |
urna-python | Rust, pyo3 | The _urna extension: urna.build, UrnaFile, urna.chunk_id and the build presets |
python/ | Python | The urna module, the builder pipeline, model fingerprints, the query embedders and the forge |
urna-runtime depends on urna-format; the CLI and the Python bridge depend on both. The crates are published as urna-format, urna-runtime and urna (the CLI). urna-python is not published on crates.io: it reaches users as the PyPI wheel urna. Their public API is summarized in Rust crates and module urna.
Two more Cargo workspaces sit outside crates/. forge-core/ holds the frozen schema of the forge's canonical intermediate format, kept apart so its dependencies never enter the format and runtime crates; the release gate does not build it. fuzz/ holds the cargo-fuzz targets.
urna-format
The format crate knows bytes and nothing else. It defines the 128-byte header, the 32-byte section table entries, the 40-byte footer and the section id map. UrnaFileBuilder writes a file: it computes each chunk_id, and under zstd text it picks the smallest of several codecs per text section (zstd, intpack, dictionary, FSST, dedup; all decode to the same bytes, so content_hash does not move), stores embeddings as float32, float16, int8 or int4, aligns every payload to 64 bytes and writes the checksums and the footer hash. UrnaView::from_bytes reads a file and checks it in a fixed order. The crate has no unsafe block. See the .urna file and layout.
urna-runtime
The runtime maps a file with MmapUrnaFile::open, which runs the reader's checks, walks the embeddings for NaN and infinite values, and decodes the HNSW, BM25, graph, media and space tables. It answers five search calls: search (exact), search_ann, search_graph, search_hybrid and search_space. Every path that generates candidates ends in the same exact cosine rerank, read from one rerank source: a full-precision slab when the file has one, else the stored embeddings. Embeddings are never zstd-compressed, because the SIMD kernels (AVX2, NEON, scalar) score them straight from the map. The runtime never opens a socket. See search paths and the exact rerank.
urna-cli
The binary groups its verbs in three sets:
- Engine verbs take a file and, for search, a vector:
inspect,validate,stats,media,search,search-ann,search-graph,search-space,benchmark,cite. Two more sit in this group but start Python:search-textembeds the query withpython/embed_query.py, anddoctorruns one real embed as its last check. - Agent verbs take text or a spec:
askandretrieveembed the query through a Python process and gate it, andbuildlaunches the forge. - Setup verbs:
setupinstalls the embedder payload and a Python env, andtuiis the terminal explorer. Both sit behind the defaulttuifeature;--no-default-featuresbuilds the engine-only CLI.
urna-python and python/
The pyo3 bridge exposes urna.build, which resolves a preset, builds the HNSW, BM25 and graph payloads with the runtime crate, and hands everything to UrnaFileBuilder. UrnaFile wraps MmapUrnaFile with the same search calls plus retrieve. python/urna.py loads the extension; in the wheel it is the urna module.
The rest of python/ exists only in a checkout of the repository:
builder.py: chunking with real byte spans, an embedding cache and aPipelinethat callsurna.build.model_fingerprint.py: computesmodel_hashfrom a sentence-transformers snapshot.forge/: the declarative build (build_spec,corpus_sources,forge_pipeline,forge_emit), the model registry and its adapters, the media backends and the quality gate, and the two query embeddersembed_query_potion.pyandembed_query_model.py.tools/urna_forge.py: the programurna buildlaunches, plus measurement and benchmark tools.
Where each piece ships
| Piece | Release binary (archives, Homebrew, npm, cargo) | Embedder payload (urna setup, one-liners) | PyPI wheel urna | Repository checkout |
|---|---|---|---|---|
urna binary | yes | yes (cargo build) | ||
_urna extension and the urna module | yes | yes (built by hand) | ||
| Potion query embedder and table | yes | yes (urna.embed_potion) | yes (Git LFS) | |
Registry query embedder (embed_query_model.py) | no | no | yes | |
search-text embedder (embed_query.py) | no | no | yes | |
Forge and urna_forge.py | no | no | yes |
An installed binary with the payload answers ask and retrieve on potion corpora. urna build, search-text and queries against a corpus built with a registry model (wemm, clip, jina) need a checkout and the model's Python dependencies. See known limits and installation.
A build, from rows to file
corpus.toml + rows your own chunks and vectors
| |
v |
urna build --spec (Rust launcher, checkout only) |
| finds python/tools/urna_forge.py |
| picks the interpreter, forwards the flags |
v |
forge (Python) |
validate the spec |
load rows, hash each item, dedup identical images |
media stage (encode, optional crf gate) |
embed once per model, through the shared cache |
| |
v v
urna.build(...) (urna-python bridge) <--------------+
resolve the preset, optional matryoshka truncation
HNSW from float32 rows, BM25 over the canonical text, graph
|
v
UrnaFileBuilder (urna-format)
chunk ids, per-section encoding, quantization
header, section table, manifest, 64-byte aligned payloads, footer
|
v
corpus.urna (+ corpus.manifest.json and corpus.build.lock.json from the forge)urna build --specis a thin Rust launcher. It findsurna_forge.pythrough the checkout layout, resolves the Python interpreter, forwards its flags and passes the child's exit code through.- The forge validates the spec, loads the rows (always recomputed), hashes each item and the whole corpus, and shares one frame between rows with byte-identical images. With a
[media]table it encodes the media. - For each model in the spec it looks up the shared embed cache by model, recipe and corpus hash, and embeds only on a miss. Sentence-transformers models run in their own worker process.
- For each output file it calls
urna.buildinto a staging directory, renames the result into place, reopens it and validates it. Then it writes the sidecar manifest and the build lock. urna.buildresolves the preset (exact,compressed,tiny,nano,hybrid), truncates and renormalizes the vectors whenmrl_dimis set, builds the HNSW graph from float32 rows before quantization, builds BM25 over the canonical text and the chunk graph when asked, and passes everything to the writer.- The writer computes the chunk ids, encodes each section, quantizes the embeddings to the preset's dtype, and writes the file with its checksums and footer hash.
Your own code can skip the forge and call urna.build directly, from the wheel or from a checkout. See build from your own rows, presets and stored precision and reproducible builds.
A query, from text to citation
urna ask corpus.urna "question"
|
v
MmapUrnaFile::open (urna-runtime)
mmap, header and section checksums, footer hash,
manifest and contract, NaN walk, decode the indexes
|
| manifest: embedding_model, embedding_dim, model_hash
v
embedder process (Python, local)
potion script for minishlab/potion corpora, registry script otherwise
prints one JSON line: model name, dim, model_hash, vector
|
v
model gate (urna-cli)
name equal, dim equal, placeholder hash refused, model_hash equal
|
v
route by the declared index_type
exact -> search
hnsw -> search_ann (beam = max(candidates, k, the file's ef_construction))
hybrid -> search_hybrid
|
v
exact cosine rerank over the candidates
|
v
stored canonical text + urna://content_hash/chunk_id- The CLI opens the file with
MmapUrnaFile::open, which checks every hash and decodes the index sections. - It reads
embedding_model,embedding_dimandmodel_hashfrom the manifest and picks the embedder:embed_query_potion.pywhen the model name starts withminishlab/potion,embed_query_model.pyotherwise. It resolves the interpreter and prints its choice on stderr. - The embedder runs as a separate Python process and prints one JSON line with the vector and the fingerprint of the model it used.
- The gate compares the model name, the dimension and
model_hashwith the manifest. The all-zero placeholder hash is refused. Any mismatch stops the query with an error. See the model gate. - The CLI routes by the manifest's declared
index_type. The default candidate count ismax(4k, 64), set with--candidates. - Every candidate is rescored by exact cosine, so the score is real cosine, at stored precision when the embeddings are not float32.
--disclose explainprints the route, the candidate counts and which precision the score has. - The CLI decodes the stored canonical text for each hit and prints it with its
urna://content_hash/chunk_idcitation, whichurna citeresolves back. See citations and hashes.
The hybrid preset never takes the hybrid route
preset="hybrid" in urna.build, which is also the forge's default, builds HNSW and BM25 but writes index_type = "hnsw". ask, retrieve and search-text route such a file to search_ann, so its BM25 section is never used from the CLI. When search_hybrid does run, BM25 only adds candidates: the final order is pure cosine. See known limits.
The engine verbs skip the embedder and the gate: search, search-ann, search-graph and search-space take a JSON vector and call the runtime directly. In Python, UrnaFile.retrieve takes a vector you embedded yourself and checks model_hash only when you pass expected_model_hash. The script protocol is in query embedder protocol, and the offline design is in offline by construction.
Data governance
What a .urna file stores, what it leaves out, how to remove or correct content, where processing runs, and which files a build leaves on disk.
Contributing
Set up a urna checkout, run the release gate, pass CI, fuzz the decoders and open a pull request that follows the repository rules.