docsv0.5.1

The .urna file

What a .urna file holds, how its header, section table and footer fit together, and what the runtime checks before it answers a query.

A urna corpus is one .urna file. The chunk text, the source spans, the embeddings, the indices and the search contract all live inside it, and the runtime answers queries by memory-mapping that file and reading it in place. This page describes the file at the level you need to use it; the byte-level layout is in File format.

One file, memory-mapped

There is no server process, no side directory and no unpack step. The Rust runtime opens the file read-only with mmap, checks it, and serves searches from the mapped bytes. Embeddings stay in their on-disk precision and are read row by row, so opening a quantized corpus does not expand it to float32 in memory.

Two consequences follow:

  • Copying the file copies the whole corpus. Moving it between machines needs nothing else, as long as the query side has the matching embedding model (see the model gate).
  • Do not rewrite a file while a process has it open. The runtime never writes to a .urna, and its checksums detect corruption after the fact; they do not protect a mapping from a live writer.

What is inside

A file is a set of numbered sections plus a JSON manifest. Six sections are required in every file. They are also the canonical sections: their decoded bytes are what content_hash covers, so they define the identity of the corpus and of every citation into it (see citations and hashes).

IdSectionHolds
0x01chunk_idsOne sha256: id per chunk
0x02chunks_canonicalThe stored text of each chunk, the text cite returns
0x03chunks_original_spanssource_uri, byte_start and byte_end per chunk
0x04embeddingsOne vector per chunk, in the file's dtype (float32, float16, int8 or int4)
0x05provenanceFree-form JSON written by the builder
0x06search_contractmetric, score_type, normalize, index_type, rerank_policy

Optional sections add capabilities without touching the canonical six:

IdSectionAdds
0x07hnsw_indexApproximate nearest-neighbour candidates
0x08bm25_indexLexical candidates for the hybrid path
0x0Cgraph_adjacencyA chunk-to-chunk graph for the graph path
0x14blob_refsA table of media blobs (hash, uri, length, inlined or not)
0x16blob_span_overlayPer-chunk byte ranges inside a blob
0x17blob_dataThe media bytes themselves, when inlined
0x15space_tableNamed embedding spaces, one per extra model or tower
0x20 to 0x2Fspace_embeddingsOne vector band per named space

A compressed text section can also bring two unnamed helper sections (0x0A dictionary, 0x0B dedup map); tools list them as unknown. The full id map, including reserved ids, is in Sections.

The quickstart corpus has the six required sections plus HNSW, BM25 and the graph:

urna stats examples/quickstart/out/quickstart.urna
sections:     9
  0x01 chunk_ids                encoding=intpack  389 bytes
  0x02 chunks_canonical         encoding=zstd     1484 bytes
  0x03 chunks_original_spans    encoding=zstd     154 bytes
  0x04 embeddings               encoding=raw      12288 bytes
  0x05 provenance               encoding=zstd     118 bytes
  0x06 search_contract          encoding=zstd     95 bytes
  0x07 hnsw_index               encoding=raw      182 bytes
  0x08 bm25_index               encoding=zstd     1695 bytes
  0x0c graph_adjacency          encoding=raw      174 bytes

Layout

[0, 128)                     header
[128, 128 + 32 * count)      section table, one 32-byte entry per section
manifest                     JSON, directly after the table
payloads                     each section at a 64-byte aligned offset
[file_size - 40, file_size)  footer
  • The header starts with the magic URNA and the version 1.0, then the embedding dim, the chunk and embedding counts, the file size, the offsets of the table and the manifest, and an 8-byte header checksum. Files written by 0.4.0 and earlier carry the magic NEST; the reader still opens them.
  • Each section table entry holds the section id, its wire encoding, its offset, its size and an 8-byte checksum of its payload bytes.
  • The footer holds its own size and a SHA-256 of every byte before it.
  • All integers are little-endian. Padding between payloads is zero and is not part of any checksum.

The section encoding is how the bytes are stored (raw, zstd, intpack, float16, int8, int4 and a few text codecs). Text codecs decode to the same bytes as raw, which is why compressing the text does not change content_hash. See Encodings.

The manifest

The manifest is the JSON that describes the corpus: embedding_model, embedding_dim, n_chunks, dtype, metric (ip), score_type, normalize (l2), index_type, rerank_policy, model_hash, chunker_version and a capabilities object, plus optional fields such as title, created and mrl_dim. urna inspect prints it.

Two points matter when you reason about hashes:

  • The manifest is covered by the footer hash only. It has no section checksum and is not part of content_hash, so a manifest-only change (a title, a capability flag) changes file_hash and leaves every citation intact.
  • The search_contract section repeats five manifest fields. Because that section is canonical, those five fields do move content_hash. The reader rejects a file whose contract and manifest disagree.

The forge (urna build) also writes a <name>.manifest.json next to the .urna. That is a separate build report, not the manifest inside the file. See Build artifacts.

What the reader checks on open

Every tool that opens a file runs the same parser first. It stops at the first failure with a typed error, and nothing is served from a file that fails. In order:

  1. The file is at least 168 bytes (header plus footer).
  2. The magic is URNA or NEST, and the version is 1.0.
  3. The header checksum matches.
  4. The file_size in the header equals the real length.
  5. The section table fits inside the file.
  6. For each section: the encoding is legal for that section, the offset is 64-byte aligned, the payload fits inside the file, and the payload checksum matches.
  7. The manifest parses and passes validation (allowed dtype, metric, index_type and so on; model_hash shaped as sha256: plus 64 hex digits).
  8. The footer hash matches.
  9. The manifest's embedding_dim and n_chunks equal the header.
  10. All six required sections are present.
  11. The embeddings encoding matches dtype and the section has the exact expected size; each named space band has its exact size too.
  12. The search_contract section matches the manifest field by field.

When the runtime opens a file for search (MmapUrnaFile::open, which every search verb, ask, retrieve, media and inspect --json use), it then:

  • walks the embeddings for NaN and Inf;
  • decodes the chunk ids and spans;
  • decodes the HNSW and BM25 indices when their sections are present;
  • decodes the graph when the section is present and the manifest sets capabilities_ext.graph_present;
  • decodes the blob tables and applies the span overlay when the manifest sets blobs_present;
  • decodes the space table and checks each band for NaN and Inf;
  • computes file_hash and content_hash.

urna validate runs the parser, the NaN walk on the embeddings, the contract check, and a SHA-256 proof of every inlined blob. It does not decode the index payloads or check space bands for NaN, so a file with a malformed index and a valid checksum passes validate and fails on open. See urna validate.

Integrity, not authenticity

The checksums and hashes are unkeyed SHA-256. They prove the bytes are consistent with themselves, so they catch corruption and truncation. They do not prove who produced the file. See Security.

Look at a file

urna validate corpus.urna      # every checksum, the footer hash, the contract
urna stats corpus.urna         # model, dtype, index_type, sections, hashes
urna inspect corpus.urna       # header, section table with checksums, manifest
urna inspect --json corpus.urna

Every hit, and every citation, is tied to this file through its hashes: see citations and hashes.

On this page