docsv0.5.1

Sections

Every section id of the .urna v1 format, which are required, who writes each one, the encodings a reader accepts and the byte layout of each payload.

A section is one payload named by a u32 id in the section table. This page lists every id that format v1 defines, what writes it, what reads it, and the byte layout of each payload.

Required and optional

Six sections are required. They are also the canonical sections: content_hash covers exactly these six, in this order, and a reader rejects a file missing any of them with MissingRequiredSection.

IdNameContent
0x01chunk_idsone sha256:<hex> id per chunk
0x02chunks_canonicalthe stored text of each chunk
0x03chunks_original_spanssource URI and byte span of each chunk
0x04embeddingsone vector per chunk
0x05provenancefree-form JSON
0x06search_contractmetric, score type, normalization, index type, rerank policy

The order is alphabetical by name and fixed by spec, so a new section id can never reshuffle the hash. Every other section is optional and outside content_hash. Adding one of them changes file_hash but not the citations, with one exception on Hashes and citations: attaching HNSW also rewrites search_contract.

Section id map

"Text encodings" means the set {0 raw, 1 zstd, 4 intpack, 5 zstd_dict, 9 fsst, 10 txt_streams}. "Vector encodings" means {0 raw float32, 2 float16, 3 int8, 7 int4}. The reader applies the vector set to 0x04 and to both band ranges, and the text set to every other id. Encoding ids are on Encodings.

IdNameStatusWritten byRead by the runtimeGate at openEncodings accepted
0x01chunk_idsrequiredalwaysat opennonetext
0x02chunks_canonicalrequiredalwayson demand (texts, content_hash)nonetext
0x03chunks_original_spansrequiredalwaysat opennonetext
0x04embeddingsrequiredalwaysper query, from the mapnonevector, must match dtype
0x05provenancerequiredalways (default {})only for content_hashnonetext
0x06search_contractrequiredalwayscross-checked at parsenonetext
0x07hnsw_indexoptionalhnsw_index(), always rawat opensection presencetext
0x08bm25_indexoptionalbm25_index(), raw or zstdat opensection presencetext
0x09embeddings_fpreserved, read onlyneveryes, as the rerank source when presentpresencetext (raw in practice)
0x0Adictionaryauxiliarythe text codec, when the dictionary candidate winsthrough 0x02presencetext, written raw
0x0Bdedup_mapauxiliarythe text codec, when the dedup candidate winsthrough 0x02presencetext, written raw
0x0Cgraph_adjacencyoptionalgraph_adjacency(), rawat opencapabilities_ext.graph_presenttext
0x0Dchunk_scalarsreserved namenevernevern/atext
0x0Etokenizer_modelreserved namenevernevern/atext
0x0Fedit_journalreserved namenevernevern/atext
0x10repro_manifestreserved namenevernevern/atext
0x11graph_nodesreserved namenevernevern/atext
0x12graph_edge_propsreserved namenevernevern/atext
0x13graph_entity_mapreserved namenevernevern/atext
0x14blob_refsoptionalblob_refs(), rawat opencapabilities_ext.blobs_presenttext
0x15space_tableoptionalspace_table(), rawat opencapabilities_ext.supports_multimodaltext
0x16blob_span_overlayoptionalblob_span_overlay(), rawat open, rewrites spansblobs_presenttext
0x17blob_dataoptionalblob_data(), rawoffset table at open, bytes on demandblobs_present and 0x14 present; must be rawtext (the runtime requires raw)
0x20 to 0x2Fspace_embeddingsoptional bandspace_band(index, encoding, payload)per search_spacesupports_multimodal and listed in 0x15vector
0x30 to 0x3Fspace_embeddings_fpreserved bandnevernevern/avector

Names in the table for 0x09 to 0x13 come from the Rust constants (SECTION_EMBEDDINGS_FP, SECTION_DICTIONARY, SECTION_CHUNK_SCALARS and so on); they are not stored in the file. section_name() does not resolve those ids, so urna inspect --json labels them "unknown". 0x0A and 0x0B stay unnamed on purpose: they are private companions of 0x02.

"Written by" names the UrnaFileBuilder method. The HNSW and BM25 index sections only open on section presence; the manifest flags supports_ann and supports_bm25 are not consulted by the runtime. A 0x0C section without graph_present is ignored, and search_graph then falls back to exact search.

Who writes what

SectionsRust builderurna.build (Python)urna build --spec (forge)
0x01 to 0x06alwaysalwaysalways
0x07 hnsw_indexhnsw_index()presets tiny, nano, hybrid, or with_hnsw=Truefrom the [build] preset (default hybrid)
0x08 bm25_indexbm25_index()preset hybrid, or with_bm25=Truefrom the [build] preset
0x0A, 0x0Bunder SectionEncoding::Zstd, when their candidate winszstd-text presetszstd-text presets
0x0C graph_adjacencygraph_adjacency()with_graph=Trueon by default
0x14, 0x16blob_refs(), blob_span_overlay()blob_refs=, chunk_blob_spans=with a [media] section
0x17 blob_datablob_data()blob_data_paths=[output] embed_media = true
0x15 and bandsspace_table(), space_band()spaces=named spaces from [[models]]

The builder takes optional payloads as opaque bytes. It does not decode or validate them, so a bad HNSW, BM25, graph, blob or space payload builds and fails later, when MmapUrnaFile::open decodes it. See urna.build and the spec file for the Python and forge options.

Space bands

  • Band 0x20 + i holds the vectors of space i. Space 0 is the text space in 0x04 and is never listed in 0x15, so usable band ids are 0x21 to 0x2F: at most 15 named spaces.
  • A space_table entry with space_index 0 or 16 and above is rejected. The builder's space_band(index, ...) does not range-check its index.
  • Every band listed in 0x15 must be present, with the exact size its (n_vectors, dim, dtype) implies, and n_vectors must equal n_chunks: bands are parallel to the chunks.

Payload formats

Most sections open with a common 12-byte prefix: u32 version = 1 followed by u64 count. Strings are length-prefixed: u32 len then UTF-8 bytes, no NUL. Decoders reject trailing bytes and bound every count against the remaining bytes before allocating.

0x01 chunk_ids

EncodingLayout
rawprefix + count strings, each sha256:<64 hex> (71 ASCII bytes)
intpack, kind 0u8 0 + u32 count + count x 32 raw digest bytes

The writer uses intpack under zstd text and falls back to raw when any id is not lowercase sha256:<64 hex>.

0x02 chunks_canonical

Raw: prefix + count strings, the canonical text of each chunk. Under zstd text the writer picks the smallest of five codecs; all of them decode to these exact raw bytes. See the text-codec chooser.

0x03 chunks_original_spans

EncodingLayout
rawprefix + count x (string source_uri, u64 byte_start, u64 byte_end)
intpack, kind 1u8 1 + u32 count + u32 n_uris + n_uris strings (first-seen URI pool) + three columns, each u32 len + intpack: URI index, byte_start, byte_end - byte_start

Under zstd text the writer keeps the smaller of intpack and zstd.

0x04 embeddings

No prefix. A fixed-stride slab of n_embeddings rows at embedding_dim, in the layout of the manifest dtype: see embedding dtypes.

0x05 provenance

u32 1 + u64 json_len (the count slot carries the JSON byte length) + compact JSON of any value. The reader only requires it to parse. It is canonical, so its bytes, key order included, move content_hash.

0x06 search_contract

u32 1 + u64 json_len + JSON {metric, score_type, normalize, index_type, rerank_policy}. The writer fills it from the manifest, and the reader rejects a file where the two disagree, with the matching Unsupported* error and the text section says X but manifest says Y.

0x07 hnsw_index

Payload version 2: seven u32 (payload version 2, m, m_max0, ef_construction, entry_point, max_level, n_nodes), then three columns of u32 len + intpack: the level of each node, the neighbour count per layer, and the neighbour ids in build order. Version 1 payloads still decode.

0x08 bm25_index

Payload version 2: u32 2, f32 k1, f32 b, f32 avgdl, u32 n_docs, u32 n_terms, intpack of document lengths, then per term (sorted by bytes) a string token and u32 df, then intpack of delta-gapped document ids and intpack of term frequencies. Version 1 payloads still decode.

0x09 embeddings_fp

A fixed-stride slab of n x dim float32 or float16 values. The runtime infers the dtype from size / (n * dim) (4 or 2 bytes). No writer emits this section in 0.5.1.

0x0A dictionary

The raw trained zstd dictionary that the zstd_dict codec of 0x02 decodes against.

0x0B dedup_map

u8 0 + intpack of one u32 per chunk: the index of its text in the unique pool stored in 0x02.

0x0C graph_adjacency

u32 1 + u64 n_nodes + (u32 len + intpack of n_nodes + 1 offsets) + (u32 len + intpack of destination gaps, delta-coded within each (source, edge type) run) + an edge-type column: u8 0 + u8 type when every edge has one type, or u8 1 + u32 len + intpack of types.

Edge types: EDGE_TYPE_NEXT_CHUNK = 0, EDGE_TYPE_SEMANTIC = 1, EDGE_TYPE_CITATION = 2. Edges are sorted by (source, type, destination), and a node has at most GRAPH_MAX_DEGREE = 2^20 edges. The runtime requires n_nodes == n_embeddings.

0x14 blob_refs

u32 1 + u64 n + n x (32-byte raw SHA-256 of the original bytes, string original_uri, u64 byte_len, u8 inlined 0 or 1). Entry order is the contract: the overlay and blob_data address records by position.

0x15 space_table

u32 1 + u64 n + n x (u8 space_index, string name, u32 dim, u8 dtype, string model_hash, u64 n_vectors). The dtype codes are 0 float32, 1 float16, 2 int8, 3 int4. Duplicate indices or names are rejected, and model_hash must start with sha256:.

0x16 blob_span_overlay

u32 1 + u64 n + n x (u32 blob_ref_index, u64 byte_start, u64 byte_end), one entry per chunk in chunk order. BLOB_REF_NONE = 0xFFFFFFFF keeps the chunk's 0x03 span. At open the runtime replaces the span of every other chunk with the blob-relative range, so hits and cite report offsets inside the blob. A blob_ref_index past the end of 0x14 fails the open.

0x17 blob_data

u32 1 + u64 n + n x (u64 offset, u64 len), offsets relative to the first data byte, then the concatenated blob bytes. A record kept outside the file has (0, 0). Media is already compressed by its own codec, so this section is always raw.

0x20 to 0x2F bands

A fixed-stride slab of n_vectors rows at the space's dim, in the layout of its dtype, the same layouts as 0x04.

Accepted, never written

The reader has no allow-list of ids. Any id, the reserved names 0x0D to 0x13 and the 0x30 to 0x3F band included, loads when its encoding is legal for its class and its checksum matches. The runtime ignores what it does not know. A future optional section therefore opens in a 0.5.1 reader, as long as it uses an encoding that reader knows.

Encodings and dtype layouts are next: Encodings.

On this page