Encodings
The wire encoding registry of .urna v1: ids 0 to 10, zstd limits, the float16, int8 and int4 embedding layouts, the text-codec chooser and intpack.
Every section table entry carries an encoding id that says how the payload bytes are stored. This page lists every id, where each one is legal, and the byte layout of the embedding dtypes and the text codecs.
Registry
| Id | Constant | Name | Status | Legal on | The writer emits it when |
|---|---|---|---|---|---|
| 0 | SECTION_ENCODING_RAW | raw | implemented | every section; float32 on 0x04 and bands | default; float32 embeddings; HNSW, graph, blobs, space table, dictionary and dedup map always |
| 1 | SECTION_ENCODING_ZSTD | zstd | implemented | non-vector sections | zstd text: 0x05, 0x06, 0x08 always; 0x02 and 0x03 when they win |
| 2 | SECTION_ENCODING_FLOAT16 | float16 | implemented | 0x04, bands | dtype = "float16" |
| 3 | SECTION_ENCODING_INT8 | int8 | implemented | 0x04, bands | dtype = "int8" |
| 4 | SECTION_ENCODING_INTPACK | intpack repack | implemented | non-vector sections | zstd text: 0x01 always (unless an id is not canonical), 0x03 when smaller than zstd |
| 5 | SECTION_ENCODING_ZSTD_DICT | txt_streams V2 with a shared dictionary | implemented, needs 0x0A | non-vector sections | 0x02, chooser candidate 3 |
| 6 | SECTION_ENCODING_FRONTCODE | frontcode | reserved, rejected | none | never |
| 7 | SECTION_ENCODING_INT4 | int4 block-64 | implemented | 0x04, bands | dtype = "int4" |
| 8 | SECTION_ENCODING_RABITQ | rabitq | reserved, rejected | none | never |
| 9 | SECTION_ENCODING_FSST | txt_streams V3 (FSST) | implemented | non-vector sections | 0x02, chooser candidate 4 |
| 10 | SECTION_ENCODING_TXT_STREAMS | txt_streams V1 | implemented | non-vector sections | 0x02, chooser candidate 2 |
"Zstd text" means SectionEncoding::Zstd in Rust, which every preset except exact selects (see Build presets). Under SectionEncoding::Raw every non-vector section is written raw.
Rules the reader enforces:
- Vector encodings (0, 2, 3, 7) are the only ones legal on
0x04and the band ranges. Zstd on embeddings isUnsupportedSectionEncoding: vectors stay uncompressed so the runtime can score them straight from the memory map. - Non-vector sections accept 0, 1, 4, 5, 9 and 10.
- The embeddings encoding must match the manifest
dtype: float32 with 0, float16 with 2, int8 with 3, int4 with 7. Any other pair isManifestInvalid. - Ids 6 and 8, and any id above 10, are
UnsupportedSectionEncoding.
Every non-vector codec decodes byte-identically to the raw payload. That is why switching between raw and compressed text never changes content_hash or a citation.
Zstd
- Compression level 19 (
DEFAULT_ZSTD_LEVEL). - Decompression is capped at the larger of 64 MiB and 128 times the compressed length. A frame that declares more is refused before allocating, as
MalformedSectionPayload. - The dictionary codec caps each stream at 64 MiB.
Under zstd text the writer compresses 0x05 and 0x06 even when the result is larger than raw.
Embedding dtypes
The expected size of 0x04 (and of a band) follows from the dtype, the row count n and the dimension dim. The reader rejects any other size with EmbeddingSizeMismatch.
| Dtype | Encoding | Layout | Size in bytes |
|---|---|---|---|
| float32 | 0 | n x dim f32, little-endian, row-major | 4 * n * dim |
| float16 | 2 | n x dim IEEE f16, little-endian | 2 * n * dim |
| int8 | 3 | u32 1 (version) + u32 0 (per-vector scale) + n f32 scales + n x dim i8 | 8 + 4 * n + n * dim |
| int4 | 7 | see below | 8 + 2 * n * (dim / 64) + n * dim / 2 |
Quantization:
- float16: each value converted with
half::f16::from_f32. - int8: per vector,
scale = max(abs(v)) / 127, codes rounded and clamped to -127 to 127. A zero vector gets scale 1. - int4: per 64-value block, described next.
The writer expects L2-normalized vectors. It does not normalize or check the norm; it rejects a wrong length, NaN and Inf.
int4 block-64 layout
u32 LE payload_version = 1 INT4_PAYLOAD_VERSION
u32 LE scale_kind = 1 INT4_SCALE_KIND_PER_GROUP
f16 LE x (n * dim/64) block scales, row-major
u8 x (n * dim/2) packed codesdimmust be a nonzero multiple of 64 (INT4_BLOCK). All scales of all rows come first, then all codes.- The scale of row
i, blockgis at byte8 + (i * dim/64 + g) * 2. The codes of rowisit at[i * dim/2, (i + 1) * dim/2)within the code region. - Per block,
scale = f16(max(abs(v)) / 7), or 1.0 for an all-zero block. Codes areround(v / scale)clamped to -7 to 7; -8 is never written. - Element
jof a row lives in bytej / 2: low nibble for evenj, high nibble for oddj, as a two's-complement 4-bit value. - The value is reconstructed as
code * scaleof its block. The runtime sums each block in f32 and multiplies by the block scale once, with identical results on AVX2, NEON and the scalar kernel.
No writer emits the full-precision source 0x09, so int4, int8 and float16 files rerank at stored precision: every score is a real cosine computed on the stored values, and --disclose explain says real cosine at stored precision. See Presets and stored precision.
The chunks_canonical text-codec chooser
Under zstd text the writer builds up to five candidate encodings of 0x02 and keeps the smallest. The comparison counts the auxiliary sections a candidate needs, and a tie keeps the earlier candidate.
| Order | Candidate | Encoding | Auxiliary section | Offered when |
|---|---|---|---|---|
| 1 | one zstd frame over the raw payload | 1 | none | always, so a build never gets larger |
| 2 | txt_streams V1: one zstd frame per chunk behind an intpack offset table | 10 | none | always |
| 3 | txt_streams V2 with a trained zstd dictionary | 5 | 0x0A dictionary | the trainer succeeds: at least 8 distinct texts |
| 4 | txt_streams V3 with an FSST symbol table | 9 | none, the table is embedded | always |
| 5 | deduplicated texts in one zstd frame | 1 | 0x0B dedup map | some text repeats |
The trained dictionary is sized at a tenth of the total text, clamped between 4 KiB and 112 KiB.
The three txt_streams variants share one container: u8 kind (V1 = 0, V2 = 1, V3 = 2) + u64 count + an intpack table of count + 1 offsets + the streams. One chunk decodes on its own, without the rest. The FSST region is u32 table_len + the table (u16 count, then per symbol a u8 len and its bytes; up to 255 symbols of 1 to 8 bytes; 0xFF escapes one raw byte) + the frames.
Every candidate decodes to the same raw payload, so which one wins never changes content_hash. The quickstart corpus, 12 short chunks, keeps candidate 1:
0x02 chunks_canonical encoding=zstd offset=1536 size=1484 checksum=37c32466fed5b2aaIntpack
Intpack stores a list of unsigned integers in blocks of 128 (INTPACK_BLOCK):
u32 count
u32 n_blocks
u32 x n_blocks absolute byte offset of each block
per block:
u64 min
u8 width
values minus min, packed LSB-first at `width` bits eachIt keeps order and reads any single value in constant time. It is the building block of encoding 4 (the chunk_ids and spans repacks, dispatched by their kind byte 0 or 1) and of the txt_streams offset tables, the dedup map, the graph columns, HNSW v2 and BM25 v2.
Rust entry points
| Item | Use |
|---|---|
SectionEncoding::{Raw, Zstd} | text encoding of a build, id() gives 0 or 1 |
EmbeddingDType::{Float32, Float16, Int8, Int4} | dtype of a build, with manifest_str() and encoding() |
expected_embeddings_size(dtype, n, dim) | expected slab size, None on overflow or unknown dtype |
f32_to_f16_bytes, f16_bytes_to_f32 | float16 conversion |
encode_int8_embeddings, Int8EmbeddingsView, quantize_f32_to_i8 | int8 encode and read |
encode_int4_embeddings, Int4EmbeddingsView, quantize_f32_to_i4, pack_nibbles, nibble_to_i4, int4_blocks_per_row, INT4_BLOCK | int4 encode and read |
UrnaView::decoded_section(id) | a section after its wire codec, dictionary and dedup expansion |
The same encoders produce the payloads for space_band. The crates are described on Rust crates.
How these bytes feed the hashes is on Hashes and citations.
Sections
Every section id of the .urna v1 format, which are required, who writes each one, the encodings a reader accepts and the byte layout of each payload.
Hashes and citations
Exact preimages of every integrity value in a .urna file: header and section checksums, the footer hash, file_hash, content_hash, chunk_id and the citation id.