docsv0.5.1

Architecture

How urna is built: the four Rust crates, the Python forge, how a build becomes a .urna file, and how a query flows through the runtime to a citation.

urna is a Rust workspace of four crates plus a Python layer. Python builds files; Rust owns the format, serves the queries and runs the CLI. This page maps the pieces and follows a build and a query from end to end.

The pieces

PieceLanguageWhat it owns
urna-formatRust libraryThe frozen v1 container: layout, manifest, reader, writer, section codecs, wire encodings, chunk_id and the hashes
urna-runtimeRust libraryOpening a file through mmap, query validation, SIMD dispatch, exact, HNSW, graph, hybrid and space search, the exact rerank, and the HNSW and BM25 index builders
urna-cliRust binaryThe urna command: a clap surface over the two libraries, plus setup and tui
urna-pythonRust, pyo3The _urna extension: urna.build, UrnaFile, urna.chunk_id and the build presets
python/PythonThe urna module, the builder pipeline, model fingerprints, the query embedders and the forge

urna-runtime depends on urna-format; the CLI and the Python bridge depend on both. The crates are published as urna-format, urna-runtime and urna (the CLI). urna-python is not published on crates.io: it reaches users as the PyPI wheel urna. Their public API is summarized in Rust crates and module urna.

Two more Cargo workspaces sit outside crates/. forge-core/ holds the frozen schema of the forge's canonical intermediate format, kept apart so its dependencies never enter the format and runtime crates; the release gate does not build it. fuzz/ holds the cargo-fuzz targets.

urna-format

The format crate knows bytes and nothing else. It defines the 128-byte header, the 32-byte section table entries, the 40-byte footer and the section id map. UrnaFileBuilder writes a file: it computes each chunk_id, and under zstd text it picks the smallest of several codecs per text section (zstd, intpack, dictionary, FSST, dedup; all decode to the same bytes, so content_hash does not move), stores embeddings as float32, float16, int8 or int4, aligns every payload to 64 bytes and writes the checksums and the footer hash. UrnaView::from_bytes reads a file and checks it in a fixed order. The crate has no unsafe block. See the .urna file and layout.

urna-runtime

The runtime maps a file with MmapUrnaFile::open, which runs the reader's checks, walks the embeddings for NaN and infinite values, and decodes the HNSW, BM25, graph, media and space tables. It answers five search calls: search (exact), search_ann, search_graph, search_hybrid and search_space. Every path that generates candidates ends in the same exact cosine rerank, read from one rerank source: a full-precision slab when the file has one, else the stored embeddings. Embeddings are never zstd-compressed, because the SIMD kernels (AVX2, NEON, scalar) score them straight from the map. The runtime never opens a socket. See search paths and the exact rerank.

urna-cli

The binary groups its verbs in three sets:

  • Engine verbs take a file and, for search, a vector: inspect, validate, stats, media, search, search-ann, search-graph, search-space, benchmark, cite. Two more sit in this group but start Python: search-text embeds the query with python/embed_query.py, and doctor runs one real embed as its last check.
  • Agent verbs take text or a spec: ask and retrieve embed the query through a Python process and gate it, and build launches the forge.
  • Setup verbs: setup installs the embedder payload and a Python env, and tui is the terminal explorer. Both sit behind the default tui feature; --no-default-features builds the engine-only CLI.

urna-python and python/

The pyo3 bridge exposes urna.build, which resolves a preset, builds the HNSW, BM25 and graph payloads with the runtime crate, and hands everything to UrnaFileBuilder. UrnaFile wraps MmapUrnaFile with the same search calls plus retrieve. python/urna.py loads the extension; in the wheel it is the urna module.

The rest of python/ exists only in a checkout of the repository:

  • builder.py: chunking with real byte spans, an embedding cache and a Pipeline that calls urna.build.
  • model_fingerprint.py: computes model_hash from a sentence-transformers snapshot.
  • forge/: the declarative build (build_spec, corpus_sources, forge_pipeline, forge_emit), the model registry and its adapters, the media backends and the quality gate, and the two query embedders embed_query_potion.py and embed_query_model.py.
  • tools/urna_forge.py: the program urna build launches, plus measurement and benchmark tools.

Where each piece ships

PieceRelease binary (archives, Homebrew, npm, cargo)Embedder payload (urna setup, one-liners)PyPI wheel urnaRepository checkout
urna binaryyesyes (cargo build)
_urna extension and the urna moduleyesyes (built by hand)
Potion query embedder and tableyesyes (urna.embed_potion)yes (Git LFS)
Registry query embedder (embed_query_model.py)nonoyes
search-text embedder (embed_query.py)nonoyes
Forge and urna_forge.pynonoyes

An installed binary with the payload answers ask and retrieve on potion corpora. urna build, search-text and queries against a corpus built with a registry model (wemm, clip, jina) need a checkout and the model's Python dependencies. See known limits and installation.

A build, from rows to file

  corpus.toml + rows                          your own chunks and vectors
        |                                                |
        v                                                |
  urna build --spec   (Rust launcher, checkout only)     |
        |  finds python/tools/urna_forge.py              |
        |  picks the interpreter, forwards the flags     |
        v                                                |
  forge (Python)                                         |
    validate the spec                                    |
    load rows, hash each item, dedup identical images    |
    media stage (encode, optional crf gate)              |
    embed once per model, through the shared cache      |
        |                                                |
        v                                                v
  urna.build(...)   (urna-python bridge)  <--------------+
    resolve the preset, optional matryoshka truncation
    HNSW from float32 rows, BM25 over the canonical text, graph
        |
        v
  UrnaFileBuilder   (urna-format)
    chunk ids, per-section encoding, quantization
    header, section table, manifest, 64-byte aligned payloads, footer
        |
        v
  corpus.urna   (+ corpus.manifest.json and corpus.build.lock.json from the forge)
  1. urna build --spec is a thin Rust launcher. It finds urna_forge.py through the checkout layout, resolves the Python interpreter, forwards its flags and passes the child's exit code through.
  2. The forge validates the spec, loads the rows (always recomputed), hashes each item and the whole corpus, and shares one frame between rows with byte-identical images. With a [media] table it encodes the media.
  3. For each model in the spec it looks up the shared embed cache by model, recipe and corpus hash, and embeds only on a miss. Sentence-transformers models run in their own worker process.
  4. For each output file it calls urna.build into a staging directory, renames the result into place, reopens it and validates it. Then it writes the sidecar manifest and the build lock.
  5. urna.build resolves the preset (exact, compressed, tiny, nano, hybrid), truncates and renormalizes the vectors when mrl_dim is set, builds the HNSW graph from float32 rows before quantization, builds BM25 over the canonical text and the chunk graph when asked, and passes everything to the writer.
  6. The writer computes the chunk ids, encodes each section, quantizes the embeddings to the preset's dtype, and writes the file with its checksums and footer hash.

Your own code can skip the forge and call urna.build directly, from the wheel or from a checkout. See build from your own rows, presets and stored precision and reproducible builds.

A query, from text to citation

  urna ask corpus.urna "question"
        |
        v
  MmapUrnaFile::open         (urna-runtime)
    mmap, header and section checksums, footer hash,
    manifest and contract, NaN walk, decode the indexes
        |
        |  manifest: embedding_model, embedding_dim, model_hash
        v
  embedder process           (Python, local)
    potion script for minishlab/potion corpora, registry script otherwise
    prints one JSON line: model name, dim, model_hash, vector
        |
        v
  model gate                 (urna-cli)
    name equal, dim equal, placeholder hash refused, model_hash equal
        |
        v
  route by the declared index_type
    exact  -> search
    hnsw   -> search_ann     (beam = max(candidates, k, the file's ef_construction))
    hybrid -> search_hybrid
        |
        v
  exact cosine rerank over the candidates
        |
        v
  stored canonical text + urna://content_hash/chunk_id
  1. The CLI opens the file with MmapUrnaFile::open, which checks every hash and decodes the index sections.
  2. It reads embedding_model, embedding_dim and model_hash from the manifest and picks the embedder: embed_query_potion.py when the model name starts with minishlab/potion, embed_query_model.py otherwise. It resolves the interpreter and prints its choice on stderr.
  3. The embedder runs as a separate Python process and prints one JSON line with the vector and the fingerprint of the model it used.
  4. The gate compares the model name, the dimension and model_hash with the manifest. The all-zero placeholder hash is refused. Any mismatch stops the query with an error. See the model gate.
  5. The CLI routes by the manifest's declared index_type. The default candidate count is max(4k, 64), set with --candidates.
  6. Every candidate is rescored by exact cosine, so the score is real cosine, at stored precision when the embeddings are not float32. --disclose explain prints the route, the candidate counts and which precision the score has.
  7. The CLI decodes the stored canonical text for each hit and prints it with its urna://content_hash/chunk_id citation, which urna cite resolves back. See citations and hashes.

The hybrid preset never takes the hybrid route

preset="hybrid" in urna.build, which is also the forge's default, builds HNSW and BM25 but writes index_type = "hnsw". ask, retrieve and search-text route such a file to search_ann, so its BM25 section is never used from the CLI. When search_hybrid does run, BM25 only adds candidates: the final order is pure cosine. See known limits.

The engine verbs skip the embedder and the gate: search, search-ann, search-graph and search-space take a JSON vector and call the runtime directly. In Python, UrnaFile.retrieve takes a vector you embedded yourself and checks model_hash only when you pass expected_model_hash. The script protocol is in query embedder protocol, and the offline design is in offline by construction.

On this page