docsv0.5.1

Build presets

The storage presets of urna.build (exact, compressed, tiny, nano, hybrid) and the micro recipe: dtype, text encoding, indexes and measured size.

A build preset picks four storage settings at once: the text encoding, the embedding dtype, whether an HNSW index is written, and whether a BM25 index is written. Presets are resolved by the Python bridge behind urna.build; the Rust writer has no preset concept.

Presets

PresetText encodingEmbedding dtypeHNSW (0x07)BM25 (0x08)index_typererankRerank source
exactrawfloat32nonoexactnonefull precision
compressedzstdfloat16nonoexactnonestored precision
tinyzstdint8yesnohnswexactstored precision
nanozstdint4 (block 64)yesnohnswexactstored precision
hybridzstdfloat32yesyeshnswexactfull precision

index_type and rerank are the values urna stats prints (the manifest's index_type and rerank_policy). The rerank source is what urna ask --disclose explain reports as real cosine (full precision) or real cosine at stored precision, and what urna retrieve writes as rerank_source. Only float32 counts as full precision: no 0.5.1 writer emits the full-precision slab (section 0x09), so float16, int8 and int4 files rerank at the precision they store. See Presets and stored precision.

Preset names are case-sensitive. Any other value fails:

ValueError: unknown preset: micro (expected exact|compressed|tiny|nano|hybrid)

The hybrid preset declares index_type hnsw

The hybrid preset writes both indexes but records index_type = "hnsw" and score_type = "cosine". urna ask, urna retrieve, urna search-text and UrnaFile.retrieve route by the declared index_type, so on a hybrid file they take the HNSW route and never read the BM25 section. BM25 runs only when you call UrnaFile.search_hybrid yourself, and even then it only adds candidates: the final order is by cosine. See known limits.

What each preset writes

Every file has the six required sections (0x01 chunk ids, 0x02 canonical text, 0x03 spans, 0x04 embeddings, 0x05 provenance, 0x06 search contract). The preset decides how they are encoded and which index sections follow.

exact

Text sections raw, embeddings float32 raw (4 bytes per component). No index. Every search is a flat scan with recall = 1.0.

compressed

The zstd text path: chunk ids as intpack, canonical text through the text-codec chooser (the smallest of plain zstd, per-chunk zstd streams, zstd with a trained dictionary in 0x0A, fsst, or deduplication with a back-reference map in 0x0B), spans as the smaller of intpack and zstd, provenance and search contract as zstd. Embeddings float16 (2 bytes per component). No index.

tiny

The zstd text path, embeddings int8 (1 byte per component plus one float32 scale per vector), and an HNSW index in 0x07. Searches through ask and retrieve take the HNSW route, then rerank the shortlist exactly.

nano

The zstd text path, embeddings int4 in blocks of 64 components (half a byte per component plus one float16 scale per block), and an HNSW index. The effective dim must be a multiple of 64:

ValueError: int4 requires (effective) embedding_dim divisible by 64, got 96

hybrid

The zstd text path, embeddings float32, an HNSW index in 0x07 and a BM25 index over the canonical text in 0x08 (zstd). The manifest sets supports_bm25 = true and, as the callout above says, index_type = "hnsw". The quickstart corpus is built with this preset:

index_type:   hnsw
rerank:       exact
supports_ann: true
supports_bm25:true
sections:     9
  0x01 chunk_ids                encoding=intpack  389 bytes
  0x02 chunks_canonical         encoding=zstd     1484 bytes
  0x03 chunks_original_spans    encoding=zstd     154 bytes
  0x04 embeddings               encoding=raw      12288 bytes
  0x05 provenance               encoding=zstd     118 bytes
  0x06 search_contract          encoding=zstd     95 bytes
  0x07 hnsw_index               encoding=raw      182 bytes
  0x08 bm25_index               encoding=zstd     1695 bytes
  0x0c graph_adjacency          encoding=raw      174 bytes

The 0x0C graph comes from the forge's [build] with_graph = true, not from the preset.

HNSW parameters

The indexed presets build HNSW from the float32 vectors (after any mrl_dim truncation, before quantization) with hnsw_m = 16, hnsw_ef_construction = 400 and hnsw_seed = 42 unless you pass other values to urna.build. The seed makes the index bytes deterministic.

micro is a recipe, not a preset

micro is the published name of one point on the size curve: int8 embeddings truncated to their first 256 components. preset="micro" raises ValueError. Build it with explicit settings:

urna.build(
    "out/corpus-micro.urna",
    embedding_model, embedding_dim, chunker_version, model_hash, chunks,
    text_encoding="zstd", dtype="int8", mrl_dim=256, with_hnsw=True,
)

The file records mrl_dim = 256 and full_dim (the source dim), and urna stats prints both. micro is int8, so the multiple-of-64 rule does not apply to it; it applies to any int4 build, truncated or not.

Overrides

The preset only sets defaults. Explicit arguments win:

ArgumentValuesOverrides
text_encoding"raw", "zstd"the text encoding
dtype"float32", "float16", "int8", "int4"the embedding dtype
with_hnswTrue, Falsethe HNSW section
with_bm25True, Falsethe BM25 section; it never sets index_type = "hybrid"
mrl_dim0 < K <= embedding_dimnot set by any preset; truncates and renormalizes every vector

with_hnsw=True on any preset sets index_type = "hnsw" and rerank_policy = "exact".

Where presets are chosen

Entry pointSettingDefault
urna.buildpreset="exact"
builder.BuildConfig (checkout only)preset"exact"
urna build --spec[build] preset, [build] dtype, [build] mrl_dim"hybrid"
python/tools/urna_build_image_corpus.py--preset, --dtypecompressed

The build spec does not validate [build] preset or [build] dtype itself: a bad value fails inside urna.build at the emit stage, after the embeddings are computed. See [build] and [output].

In Rust, call the builder methods directly. The hybrid preset corresponds to:

UrnaFileBuilder::new(manifest)
    .text_encoding(SectionEncoding::Zstd)
    .embedding_dtype(EmbeddingDType::Float32)
    .hnsw_index(HnswIndex::build(vectors, n, dim, 16, 400, 42).to_bytes())
    .bm25_index(Bm25Index::build(&docs, DEFAULT_K1, DEFAULT_B).to_bytes())

Adding .hybrid() declares index_type = "hybrid" and score_type = "hybrid_rrf"; only then do ask, retrieve and search-text take the hybrid route.

Presets and citations

The embedding dtype and mrl_dim are part of content_hash, and so is the search contract. Two builds of the same rows with different presets have the same chunk_id values but different content_hash values, so a urna://content_hash/chunk_id citation from one file does not resolve against the other. The text encoding alone does not move content_hash: every text codec decodes to the same bytes. See Citations and hashes.

Measured sizes

The repository publishes one measured run of every preset in data/measure/ladder.json.

PresetSizeSize ratiorecall@10p50 latencyRoute measured
exact119.94 MB1.0001.0003.10 msexact
compressed40.67 MB0.3391.0003.17 msexact
tiny30.69 MB0.2560.9921.16 msHNSW
micro26.76 MB0.2230.8100.77 msHNSW
nano25.04 MB0.2090.9132.06 msHNSW
hybrid73.03 MB0.6091.0004.03 msexplicit search_hybrid

Conditions:

  • Corpus: data/corpus_next.v1.urna, 30,725 chunks at dim 384, built with a MiniLM model that is not trained for matryoshka truncation.
  • 100 queries, k = 10, seed 0, NEON SIMD backend, python/tools/measure_presets.py.
  • recall@10 is measured against the float32 exact top 10 on a self-perturbation ruler: each query is a corpus chunk's own embedding plus tiny noise. It measures how stable the ranking stays under quantization, not retrieval quality on real queries, and it is likely inflated.
  • The hybrid row was measured by calling search_hybrid directly. urna ask on a hybrid file takes the HNSW route instead.

More points of the truncation curve are on Presets and stored precision.

On this page