docsv0.5.1

Encodings

The wire encoding registry of .urna v1: ids 0 to 10, zstd limits, the float16, int8 and int4 embedding layouts, the text-codec chooser and intpack.

Every section table entry carries an encoding id that says how the payload bytes are stored. This page lists every id, where each one is legal, and the byte layout of the embedding dtypes and the text codecs.

Registry

IdConstantNameStatusLegal onThe writer emits it when
0SECTION_ENCODING_RAWrawimplementedevery section; float32 on 0x04 and bandsdefault; float32 embeddings; HNSW, graph, blobs, space table, dictionary and dedup map always
1SECTION_ENCODING_ZSTDzstdimplementednon-vector sectionszstd text: 0x05, 0x06, 0x08 always; 0x02 and 0x03 when they win
2SECTION_ENCODING_FLOAT16float16implemented0x04, bandsdtype = "float16"
3SECTION_ENCODING_INT8int8implemented0x04, bandsdtype = "int8"
4SECTION_ENCODING_INTPACKintpack repackimplementednon-vector sectionszstd text: 0x01 always (unless an id is not canonical), 0x03 when smaller than zstd
5SECTION_ENCODING_ZSTD_DICTtxt_streams V2 with a shared dictionaryimplemented, needs 0x0Anon-vector sections0x02, chooser candidate 3
6SECTION_ENCODING_FRONTCODEfrontcodereserved, rejectednonenever
7SECTION_ENCODING_INT4int4 block-64implemented0x04, bandsdtype = "int4"
8SECTION_ENCODING_RABITQrabitqreserved, rejectednonenever
9SECTION_ENCODING_FSSTtxt_streams V3 (FSST)implementednon-vector sections0x02, chooser candidate 4
10SECTION_ENCODING_TXT_STREAMStxt_streams V1implementednon-vector sections0x02, chooser candidate 2

"Zstd text" means SectionEncoding::Zstd in Rust, which every preset except exact selects (see Build presets). Under SectionEncoding::Raw every non-vector section is written raw.

Rules the reader enforces:

  • Vector encodings (0, 2, 3, 7) are the only ones legal on 0x04 and the band ranges. Zstd on embeddings is UnsupportedSectionEncoding: vectors stay uncompressed so the runtime can score them straight from the memory map.
  • Non-vector sections accept 0, 1, 4, 5, 9 and 10.
  • The embeddings encoding must match the manifest dtype: float32 with 0, float16 with 2, int8 with 3, int4 with 7. Any other pair is ManifestInvalid.
  • Ids 6 and 8, and any id above 10, are UnsupportedSectionEncoding.

Every non-vector codec decodes byte-identically to the raw payload. That is why switching between raw and compressed text never changes content_hash or a citation.

Zstd

  • Compression level 19 (DEFAULT_ZSTD_LEVEL).
  • Decompression is capped at the larger of 64 MiB and 128 times the compressed length. A frame that declares more is refused before allocating, as MalformedSectionPayload.
  • The dictionary codec caps each stream at 64 MiB.

Under zstd text the writer compresses 0x05 and 0x06 even when the result is larger than raw.

Embedding dtypes

The expected size of 0x04 (and of a band) follows from the dtype, the row count n and the dimension dim. The reader rejects any other size with EmbeddingSizeMismatch.

DtypeEncodingLayoutSize in bytes
float320n x dim f32, little-endian, row-major4 * n * dim
float162n x dim IEEE f16, little-endian2 * n * dim
int83u32 1 (version) + u32 0 (per-vector scale) + n f32 scales + n x dim i88 + 4 * n + n * dim
int47see below8 + 2 * n * (dim / 64) + n * dim / 2

Quantization:

  • float16: each value converted with half::f16::from_f32.
  • int8: per vector, scale = max(abs(v)) / 127, codes rounded and clamped to -127 to 127. A zero vector gets scale 1.
  • int4: per 64-value block, described next.

The writer expects L2-normalized vectors. It does not normalize or check the norm; it rejects a wrong length, NaN and Inf.

int4 block-64 layout

u32 LE   payload_version = 1          INT4_PAYLOAD_VERSION
u32 LE   scale_kind      = 1          INT4_SCALE_KIND_PER_GROUP
f16 LE   x (n * dim/64)               block scales, row-major
u8       x (n * dim/2)                packed codes
  • dim must be a nonzero multiple of 64 (INT4_BLOCK). All scales of all rows come first, then all codes.
  • The scale of row i, block g is at byte 8 + (i * dim/64 + g) * 2. The codes of row i sit at [i * dim/2, (i + 1) * dim/2) within the code region.
  • Per block, scale = f16(max(abs(v)) / 7), or 1.0 for an all-zero block. Codes are round(v / scale) clamped to -7 to 7; -8 is never written.
  • Element j of a row lives in byte j / 2: low nibble for even j, high nibble for odd j, as a two's-complement 4-bit value.
  • The value is reconstructed as code * scale of its block. The runtime sums each block in f32 and multiplies by the block scale once, with identical results on AVX2, NEON and the scalar kernel.

No writer emits the full-precision source 0x09, so int4, int8 and float16 files rerank at stored precision: every score is a real cosine computed on the stored values, and --disclose explain says real cosine at stored precision. See Presets and stored precision.

The chunks_canonical text-codec chooser

Under zstd text the writer builds up to five candidate encodings of 0x02 and keeps the smallest. The comparison counts the auxiliary sections a candidate needs, and a tie keeps the earlier candidate.

OrderCandidateEncodingAuxiliary sectionOffered when
1one zstd frame over the raw payload1nonealways, so a build never gets larger
2txt_streams V1: one zstd frame per chunk behind an intpack offset table10nonealways
3txt_streams V2 with a trained zstd dictionary50x0A dictionarythe trainer succeeds: at least 8 distinct texts
4txt_streams V3 with an FSST symbol table9none, the table is embeddedalways
5deduplicated texts in one zstd frame10x0B dedup mapsome text repeats

The trained dictionary is sized at a tenth of the total text, clamped between 4 KiB and 112 KiB.

The three txt_streams variants share one container: u8 kind (V1 = 0, V2 = 1, V3 = 2) + u64 count + an intpack table of count + 1 offsets + the streams. One chunk decodes on its own, without the rest. The FSST region is u32 table_len + the table (u16 count, then per symbol a u8 len and its bytes; up to 255 symbols of 1 to 8 bytes; 0xFF escapes one raw byte) + the frames.

Every candidate decodes to the same raw payload, so which one wins never changes content_hash. The quickstart corpus, 12 short chunks, keeps candidate 1:

  0x02 chunks_canonical         encoding=zstd offset=1536 size=1484 checksum=37c32466fed5b2aa

Intpack

Intpack stores a list of unsigned integers in blocks of 128 (INTPACK_BLOCK):

u32 count
u32 n_blocks
u32 x n_blocks          absolute byte offset of each block
per block:
  u64 min
  u8  width
  values minus min, packed LSB-first at `width` bits each

It keeps order and reads any single value in constant time. It is the building block of encoding 4 (the chunk_ids and spans repacks, dispatched by their kind byte 0 or 1) and of the txt_streams offset tables, the dedup map, the graph columns, HNSW v2 and BM25 v2.

Rust entry points

ItemUse
SectionEncoding::{Raw, Zstd}text encoding of a build, id() gives 0 or 1
EmbeddingDType::{Float32, Float16, Int8, Int4}dtype of a build, with manifest_str() and encoding()
expected_embeddings_size(dtype, n, dim)expected slab size, None on overflow or unknown dtype
f32_to_f16_bytes, f16_bytes_to_f32float16 conversion
encode_int8_embeddings, Int8EmbeddingsView, quantize_f32_to_i8int8 encode and read
encode_int4_embeddings, Int4EmbeddingsView, quantize_f32_to_i4, pack_nibbles, nibble_to_i4, int4_blocks_per_row, INT4_BLOCKint4 encode and read
UrnaView::decoded_section(id)a section after its wire codec, dictionary and dedup expansion

The same encoders produce the payloads for space_band. The crates are described on Rust crates.

How these bytes feed the hashes is on Hashes and citations.

On this page