Sections
Every section id of the .urna v1 format, which are required, who writes each one, the encodings a reader accepts and the byte layout of each payload.
A section is one payload named by a u32 id in the section table. This page lists every id that format v1 defines, what writes it, what reads it, and the byte layout of each payload.
Required and optional
Six sections are required. They are also the canonical sections: content_hash covers exactly these six, in this order, and a reader rejects a file missing any of them with MissingRequiredSection.
| Id | Name | Content |
|---|---|---|
0x01 | chunk_ids | one sha256:<hex> id per chunk |
0x02 | chunks_canonical | the stored text of each chunk |
0x03 | chunks_original_spans | source URI and byte span of each chunk |
0x04 | embeddings | one vector per chunk |
0x05 | provenance | free-form JSON |
0x06 | search_contract | metric, score type, normalization, index type, rerank policy |
The order is alphabetical by name and fixed by spec, so a new section id can never reshuffle the hash. Every other section is optional and outside content_hash. Adding one of them changes file_hash but not the citations, with one exception on Hashes and citations: attaching HNSW also rewrites search_contract.
Section id map
"Text encodings" means the set {0 raw, 1 zstd, 4 intpack, 5 zstd_dict, 9 fsst, 10 txt_streams}. "Vector encodings" means {0 raw float32, 2 float16, 3 int8, 7 int4}. The reader applies the vector set to 0x04 and to both band ranges, and the text set to every other id. Encoding ids are on Encodings.
| Id | Name | Status | Written by | Read by the runtime | Gate at open | Encodings accepted |
|---|---|---|---|---|---|---|
0x01 | chunk_ids | required | always | at open | none | text |
0x02 | chunks_canonical | required | always | on demand (texts, content_hash) | none | text |
0x03 | chunks_original_spans | required | always | at open | none | text |
0x04 | embeddings | required | always | per query, from the map | none | vector, must match dtype |
0x05 | provenance | required | always (default {}) | only for content_hash | none | text |
0x06 | search_contract | required | always | cross-checked at parse | none | text |
0x07 | hnsw_index | optional | hnsw_index(), always raw | at open | section presence | text |
0x08 | bm25_index | optional | bm25_index(), raw or zstd | at open | section presence | text |
0x09 | embeddings_fp | reserved, read only | never | yes, as the rerank source when present | presence | text (raw in practice) |
0x0A | dictionary | auxiliary | the text codec, when the dictionary candidate wins | through 0x02 | presence | text, written raw |
0x0B | dedup_map | auxiliary | the text codec, when the dedup candidate wins | through 0x02 | presence | text, written raw |
0x0C | graph_adjacency | optional | graph_adjacency(), raw | at open | capabilities_ext.graph_present | text |
0x0D | chunk_scalars | reserved name | never | never | n/a | text |
0x0E | tokenizer_model | reserved name | never | never | n/a | text |
0x0F | edit_journal | reserved name | never | never | n/a | text |
0x10 | repro_manifest | reserved name | never | never | n/a | text |
0x11 | graph_nodes | reserved name | never | never | n/a | text |
0x12 | graph_edge_props | reserved name | never | never | n/a | text |
0x13 | graph_entity_map | reserved name | never | never | n/a | text |
0x14 | blob_refs | optional | blob_refs(), raw | at open | capabilities_ext.blobs_present | text |
0x15 | space_table | optional | space_table(), raw | at open | capabilities_ext.supports_multimodal | text |
0x16 | blob_span_overlay | optional | blob_span_overlay(), raw | at open, rewrites spans | blobs_present | text |
0x17 | blob_data | optional | blob_data(), raw | offset table at open, bytes on demand | blobs_present and 0x14 present; must be raw | text (the runtime requires raw) |
0x20 to 0x2F | space_embeddings | optional band | space_band(index, encoding, payload) | per search_space | supports_multimodal and listed in 0x15 | vector |
0x30 to 0x3F | space_embeddings_fp | reserved band | never | never | n/a | vector |
Names in the table for 0x09 to 0x13 come from the Rust constants (SECTION_EMBEDDINGS_FP, SECTION_DICTIONARY, SECTION_CHUNK_SCALARS and so on); they are not stored in the file. section_name() does not resolve those ids, so urna inspect --json labels them "unknown". 0x0A and 0x0B stay unnamed on purpose: they are private companions of 0x02.
"Written by" names the UrnaFileBuilder method. The HNSW and BM25 index sections only open on section presence; the manifest flags supports_ann and supports_bm25 are not consulted by the runtime. A 0x0C section without graph_present is ignored, and search_graph then falls back to exact search.
Who writes what
| Sections | Rust builder | urna.build (Python) | urna build --spec (forge) |
|---|---|---|---|
0x01 to 0x06 | always | always | always |
0x07 hnsw_index | hnsw_index() | presets tiny, nano, hybrid, or with_hnsw=True | from the [build] preset (default hybrid) |
0x08 bm25_index | bm25_index() | preset hybrid, or with_bm25=True | from the [build] preset |
0x0A, 0x0B | under SectionEncoding::Zstd, when their candidate wins | zstd-text presets | zstd-text presets |
0x0C graph_adjacency | graph_adjacency() | with_graph=True | on by default |
0x14, 0x16 | blob_refs(), blob_span_overlay() | blob_refs=, chunk_blob_spans= | with a [media] section |
0x17 blob_data | blob_data() | blob_data_paths= | [output] embed_media = true |
0x15 and bands | space_table(), space_band() | spaces= | named spaces from [[models]] |
The builder takes optional payloads as opaque bytes. It does not decode or validate them, so a bad HNSW, BM25, graph, blob or space payload builds and fails later, when MmapUrnaFile::open decodes it. See urna.build and the spec file for the Python and forge options.
Space bands
- Band
0x20 + iholds the vectors of spacei. Space 0 is the text space in0x04and is never listed in0x15, so usable band ids are0x21to0x2F: at most 15 named spaces. - A
space_tableentry withspace_index0 or 16 and above is rejected. The builder'sspace_band(index, ...)does not range-check its index. - Every band listed in
0x15must be present, with the exact size its(n_vectors, dim, dtype)implies, andn_vectorsmust equaln_chunks: bands are parallel to the chunks.
Payload formats
Most sections open with a common 12-byte prefix: u32 version = 1 followed by u64 count. Strings are length-prefixed: u32 len then UTF-8 bytes, no NUL. Decoders reject trailing bytes and bound every count against the remaining bytes before allocating.
0x01 chunk_ids
| Encoding | Layout |
|---|---|
| raw | prefix + count strings, each sha256:<64 hex> (71 ASCII bytes) |
| intpack, kind 0 | u8 0 + u32 count + count x 32 raw digest bytes |
The writer uses intpack under zstd text and falls back to raw when any id is not lowercase sha256:<64 hex>.
0x02 chunks_canonical
Raw: prefix + count strings, the canonical text of each chunk. Under zstd text the writer picks the smallest of five codecs; all of them decode to these exact raw bytes. See the text-codec chooser.
0x03 chunks_original_spans
| Encoding | Layout |
|---|---|
| raw | prefix + count x (string source_uri, u64 byte_start, u64 byte_end) |
| intpack, kind 1 | u8 1 + u32 count + u32 n_uris + n_uris strings (first-seen URI pool) + three columns, each u32 len + intpack: URI index, byte_start, byte_end - byte_start |
Under zstd text the writer keeps the smaller of intpack and zstd.
0x04 embeddings
No prefix. A fixed-stride slab of n_embeddings rows at embedding_dim, in the layout of the manifest dtype: see embedding dtypes.
0x05 provenance
u32 1 + u64 json_len (the count slot carries the JSON byte length) + compact JSON of any value. The reader only requires it to parse. It is canonical, so its bytes, key order included, move content_hash.
0x06 search_contract
u32 1 + u64 json_len + JSON {metric, score_type, normalize, index_type, rerank_policy}. The writer fills it from the manifest, and the reader rejects a file where the two disagree, with the matching Unsupported* error and the text section says X but manifest says Y.
0x07 hnsw_index
Payload version 2: seven u32 (payload version 2, m, m_max0, ef_construction, entry_point, max_level, n_nodes), then three columns of u32 len + intpack: the level of each node, the neighbour count per layer, and the neighbour ids in build order. Version 1 payloads still decode.
0x08 bm25_index
Payload version 2: u32 2, f32 k1, f32 b, f32 avgdl, u32 n_docs, u32 n_terms, intpack of document lengths, then per term (sorted by bytes) a string token and u32 df, then intpack of delta-gapped document ids and intpack of term frequencies. Version 1 payloads still decode.
0x09 embeddings_fp
A fixed-stride slab of n x dim float32 or float16 values. The runtime infers the dtype from size / (n * dim) (4 or 2 bytes). No writer emits this section in 0.5.1.
0x0A dictionary
The raw trained zstd dictionary that the zstd_dict codec of 0x02 decodes against.
0x0B dedup_map
u8 0 + intpack of one u32 per chunk: the index of its text in the unique pool stored in 0x02.
0x0C graph_adjacency
u32 1 + u64 n_nodes + (u32 len + intpack of n_nodes + 1 offsets) + (u32 len + intpack of destination gaps, delta-coded within each (source, edge type) run) + an edge-type column: u8 0 + u8 type when every edge has one type, or u8 1 + u32 len + intpack of types.
Edge types: EDGE_TYPE_NEXT_CHUNK = 0, EDGE_TYPE_SEMANTIC = 1, EDGE_TYPE_CITATION = 2. Edges are sorted by (source, type, destination), and a node has at most GRAPH_MAX_DEGREE = 2^20 edges. The runtime requires n_nodes == n_embeddings.
0x14 blob_refs
u32 1 + u64 n + n x (32-byte raw SHA-256 of the original bytes, string original_uri, u64 byte_len, u8 inlined 0 or 1). Entry order is the contract: the overlay and blob_data address records by position.
0x15 space_table
u32 1 + u64 n + n x (u8 space_index, string name, u32 dim, u8 dtype, string model_hash, u64 n_vectors). The dtype codes are 0 float32, 1 float16, 2 int8, 3 int4. Duplicate indices or names are rejected, and model_hash must start with sha256:.
0x16 blob_span_overlay
u32 1 + u64 n + n x (u32 blob_ref_index, u64 byte_start, u64 byte_end), one entry per chunk in chunk order. BLOB_REF_NONE = 0xFFFFFFFF keeps the chunk's 0x03 span. At open the runtime replaces the span of every other chunk with the blob-relative range, so hits and cite report offsets inside the blob. A blob_ref_index past the end of 0x14 fails the open.
0x17 blob_data
u32 1 + u64 n + n x (u64 offset, u64 len), offsets relative to the first data byte, then the concatenated blob bytes. A record kept outside the file has (0, 0). Media is already compressed by its own codec, so this section is always raw.
0x20 to 0x2F bands
A fixed-stride slab of n_vectors rows at the space's dim, in the layout of its dtype, the same layouts as 0x04.
Accepted, never written
The reader has no allow-list of ids. Any id, the reserved names 0x0D to 0x13 and the 0x30 to 0x3F band included, loads when its encoding is legal for its class and its checksum matches. The runtime ignores what it does not know. A future optional section therefore opens in a 0.5.1 reader, as long as it uses an encoding that reader knows.
Encodings and dtype layouts are next: Encodings.
Layout
Byte layout of a .urna v1 file: the 128-byte header, 32-byte section table entries, the manifest, 64-byte aligned payloads and the 40-byte footer.
Encodings
The wire encoding registry of .urna v1: ids 0 to 10, zstd limits, the float16, int8 and int4 embedding layouts, the text-codec chooser and intpack.