The .urna file
What a .urna file holds, how its header, section table and footer fit together, and what the runtime checks before it answers a query.
A urna corpus is one .urna file. The chunk text, the source spans, the embeddings, the indices and the search contract all live inside it, and the runtime answers queries by memory-mapping that file and reading it in place. This page describes the file at the level you need to use it; the byte-level layout is in File format.
One file, memory-mapped
There is no server process, no side directory and no unpack step. The Rust runtime opens the file read-only with mmap, checks it, and serves searches from the mapped bytes. Embeddings stay in their on-disk precision and are read row by row, so opening a quantized corpus does not expand it to float32 in memory.
Two consequences follow:
- Copying the file copies the whole corpus. Moving it between machines needs nothing else, as long as the query side has the matching embedding model (see the model gate).
- Do not rewrite a file while a process has it open. The runtime never writes to a
.urna, and its checksums detect corruption after the fact; they do not protect a mapping from a live writer.
What is inside
A file is a set of numbered sections plus a JSON manifest. Six sections are required in every file. They are also the canonical sections: their decoded bytes are what content_hash covers, so they define the identity of the corpus and of every citation into it (see citations and hashes).
| Id | Section | Holds |
|---|---|---|
0x01 | chunk_ids | One sha256: id per chunk |
0x02 | chunks_canonical | The stored text of each chunk, the text cite returns |
0x03 | chunks_original_spans | source_uri, byte_start and byte_end per chunk |
0x04 | embeddings | One vector per chunk, in the file's dtype (float32, float16, int8 or int4) |
0x05 | provenance | Free-form JSON written by the builder |
0x06 | search_contract | metric, score_type, normalize, index_type, rerank_policy |
Optional sections add capabilities without touching the canonical six:
| Id | Section | Adds |
|---|---|---|
0x07 | hnsw_index | Approximate nearest-neighbour candidates |
0x08 | bm25_index | Lexical candidates for the hybrid path |
0x0C | graph_adjacency | A chunk-to-chunk graph for the graph path |
0x14 | blob_refs | A table of media blobs (hash, uri, length, inlined or not) |
0x16 | blob_span_overlay | Per-chunk byte ranges inside a blob |
0x17 | blob_data | The media bytes themselves, when inlined |
0x15 | space_table | Named embedding spaces, one per extra model or tower |
0x20 to 0x2F | space_embeddings | One vector band per named space |
A compressed text section can also bring two unnamed helper sections (0x0A dictionary, 0x0B dedup map); tools list them as unknown. The full id map, including reserved ids, is in Sections.
The quickstart corpus has the six required sections plus HNSW, BM25 and the graph:
urna stats examples/quickstart/out/quickstart.urnasections: 9
0x01 chunk_ids encoding=intpack 389 bytes
0x02 chunks_canonical encoding=zstd 1484 bytes
0x03 chunks_original_spans encoding=zstd 154 bytes
0x04 embeddings encoding=raw 12288 bytes
0x05 provenance encoding=zstd 118 bytes
0x06 search_contract encoding=zstd 95 bytes
0x07 hnsw_index encoding=raw 182 bytes
0x08 bm25_index encoding=zstd 1695 bytes
0x0c graph_adjacency encoding=raw 174 bytesLayout
[0, 128) header
[128, 128 + 32 * count) section table, one 32-byte entry per section
manifest JSON, directly after the table
payloads each section at a 64-byte aligned offset
[file_size - 40, file_size) footer- The header starts with the magic
URNAand the version1.0, then the embedding dim, the chunk and embedding counts, the file size, the offsets of the table and the manifest, and an 8-byte header checksum. Files written by 0.4.0 and earlier carry the magicNEST; the reader still opens them. - Each section table entry holds the section id, its wire encoding, its offset, its size and an 8-byte checksum of its payload bytes.
- The footer holds its own size and a SHA-256 of every byte before it.
- All integers are little-endian. Padding between payloads is zero and is not part of any checksum.
The section encoding is how the bytes are stored (raw, zstd, intpack, float16, int8, int4 and a few text codecs). Text codecs decode to the same bytes as raw, which is why compressing the text does not change content_hash. See Encodings.
The manifest
The manifest is the JSON that describes the corpus: embedding_model, embedding_dim, n_chunks, dtype, metric (ip), score_type, normalize (l2), index_type, rerank_policy, model_hash, chunker_version and a capabilities object, plus optional fields such as title, created and mrl_dim. urna inspect prints it.
Two points matter when you reason about hashes:
- The manifest is covered by the footer hash only. It has no section checksum and is not part of
content_hash, so a manifest-only change (a title, a capability flag) changesfile_hashand leaves every citation intact. - The
search_contractsection repeats five manifest fields. Because that section is canonical, those five fields do movecontent_hash. The reader rejects a file whose contract and manifest disagree.
The forge (urna build) also writes a <name>.manifest.json next to the .urna. That is a separate build report, not the manifest inside the file. See Build artifacts.
What the reader checks on open
Every tool that opens a file runs the same parser first. It stops at the first failure with a typed error, and nothing is served from a file that fails. In order:
- The file is at least 168 bytes (header plus footer).
- The magic is
URNAorNEST, and the version is1.0. - The header checksum matches.
- The
file_sizein the header equals the real length. - The section table fits inside the file.
- For each section: the encoding is legal for that section, the offset is 64-byte aligned, the payload fits inside the file, and the payload checksum matches.
- The manifest parses and passes validation (allowed
dtype,metric,index_typeand so on;model_hashshaped assha256:plus 64 hex digits). - The footer hash matches.
- The manifest's
embedding_dimandn_chunksequal the header. - All six required sections are present.
- The embeddings encoding matches
dtypeand the section has the exact expected size; each named space band has its exact size too. - The
search_contractsection matches the manifest field by field.
When the runtime opens a file for search (MmapUrnaFile::open, which every search verb, ask, retrieve, media and inspect --json use), it then:
- walks the embeddings for NaN and Inf;
- decodes the chunk ids and spans;
- decodes the HNSW and BM25 indices when their sections are present;
- decodes the graph when the section is present and the manifest sets
capabilities_ext.graph_present; - decodes the blob tables and applies the span overlay when the manifest sets
blobs_present; - decodes the space table and checks each band for NaN and Inf;
- computes
file_hashandcontent_hash.
urna validate runs the parser, the NaN walk on the embeddings, the contract check, and a SHA-256 proof of every inlined blob. It does not decode the index payloads or check space bands for NaN, so a file with a malformed index and a valid checksum passes validate and fails on open. See urna validate.
Integrity, not authenticity
The checksums and hashes are unkeyed SHA-256. They prove the bytes are consistent with themselves, so they catch corruption and truncation. They do not prove who produced the file. See Security.
Look at a file
urna validate corpus.urna # every checksum, the footer hash, the contract
urna stats corpus.urna # model, dtype, index_type, sections, hashes
urna inspect corpus.urna # header, section table with checksums, manifest
urna inspect --json corpus.urnaEvery hit, and every citation, is tied to this file through its hashes: see citations and hashes.
Your first corpus
Turn your own JSONL rows into a .urna file with urna build, ask it a question, and resolve the citation back to the stored text.
Citations and hashes
How a urna:// citation is built from content_hash and chunk_id, which changes move it, what the offsets mean, and what file_hash covers.