Presets and stored precision
How urna stores vectors as float32, float16, int8 or int4, what that costs in recall, and how every score discloses its rerank precision.
A .urna file stores one embedding per chunk, and you choose the precision it is stored at. Lower precision makes the file smaller and moves scores slightly away from the float32 values. A build preset bundles that choice with the text encoding and the index sections; this page explains the trade, and Build presets lists the exact settings.
The four stored precisions
| dtype | Bytes per vector at dim d | Layout | Constraint |
|---|---|---|---|
float32 | 4d | IEEE float32, row-major | none |
float16 | 2d | IEEE float16 | none |
int8 | d + 4 | one float32 scale per vector (max absolute value / 127) and d codes from -127 to 127 | none |
int4 | d/2 + 2 x (d/64) | one float16 scale per block of 64 components (max absolute value / 7) and two 4-bit codes per byte, from -7 to 7 | d must be a multiple of 64 |
int8 and int4 sections also carry an 8-byte header. At dim 384 a vector takes 1,536 bytes as float32, 768 as float16, 388 as int8 and 204 as int4.
The dtype applies to the default space (section 0x04). Named spaces have their own dtype, set per model with space_dtype in the build spec (see Media and named spaces).
Truncation is a second axis
mrl_dim keeps only the first K components of every vector and L2-normalizes the prefix again, before quantization and before the HNSW index is built. It multiplies with the dtype: int8 at mrl_dim = 256 is the micro point.
Truncation only makes sense for a model trained for it (matryoshka representation learning), where the first components carry most of the meaning. The file records mrl_dim and full_dim, urna stats prints them, and urna ask slices the query to the same length. A query at the full dim against a truncated file is a dimension mismatch. With int4, K must be a multiple of 64.
Size against recall
The repository publishes one measured run of the whole curve in data/measure/ladder.json:
| Point | dtype | Dim | Size ratio | recall@10 |
|---|---|---|---|---|
exact | float32 | 384 | 1.000 | 1.000 |
compressed | float16 | 384 | 0.339 | 1.000 |
tiny | int8 | 384 | 0.256 | 0.992 |
nano | int4 | 384 | 0.209 | 0.913 |
micro (mrl256-int8) | int8 | 256 | 0.223 | 0.810 |
mrl192-int8 | int8 | 192 | 0.207 | 0.733 |
mrl128-int8 | int8 | 128 | 0.190 | 0.659 |
mrl96-int8 | int8 | 96 | 0.182 | 0.574 |
mrl256-int4 | int4 | 256 | 0.191 | 0.777 |
mrl192-int4 | int4 | 192 | 0.183 | 0.713 |
mrl128-int4 | int4 | 128 | 0.174 | 0.627 |
The size ratio is the whole file against the exact build, text and indexes included. The run used data/corpus_next.v1.urna (30,725 chunks, dim 384, a MiniLM model not trained for truncation), 100 queries, k = 10, on the NEON backend.
What this recall measures
Each query is a corpus chunk's own embedding plus tiny noise, and recall@10 compares against the float32 exact top 10 for that query. It tells you how stable the ranking stays when the vectors are quantized. It does not measure retrieval quality on real questions, and it is likely inflated.
On this corpus, nano (full dim, int4) keeps more recall than any truncated point, and truncation costs recall at every step, because MiniLM was not trained for it. Reach for mrl_dim when raw size matters more than the last recall points, or when the model is trained for truncation.
How a score is computed
Every score urna returns is a cosine recomputed by an exact rerank. HNSW, the chunk graph and BM25 only choose candidates; each candidate is then scored against the query with the same kernel the flat scan uses. The candidate list can miss a chunk, but a score is never an index-side estimate.
The rerank reads one source: a full-precision slab (section 0x09) when the file has one, otherwise the stored embeddings (0x04). The reader accepts a 0x09 slab, but no 0.5.1 writer emits one. So in practice:
- A
float32file reranks at full precision. - A
float16,int8orint4file reranks at stored precision: the query staysfloat32and is scored against the stored codes and scales, withfloat32accumulation.
"Real cosine at stored precision" is still a cosine between the query and the vector the file holds, not an approximation from the index. It differs from the float32 cosine by the quantization error of the stored vector.
The rerank source is disclosed
The rerank source is reported, so a reader never has to guess which kind of score they got.
urna ask --disclose explain prints it before the answer:
route: hnsw
candidates: exact=0 ann=12 bm25=0 graph=0 fusion=none
rerank_source: real cosine
recall: (not computed; rerank guarantees real cosine)On a float16, int8 or int4 file the line reads real cosine at stored precision. urna retrieve writes the same fact on every hit as "rerank_source": "full_precision" or "stored_precision", and so does RetrieveHit.rerank_source in Python. urna stats prints the dtype, and mrl_dim and full_dim when the file is truncated.
Choosing a preset
exactwhen file size is not the constraint, or when you need thefloat32ground truth to compare other builds against.compressedfor a file about a third the size with scores atfloat16. There is no index, so every query scans all vectors.tinyfor about a quarter of the size, with an HNSW index for large corpora andint8scores.nanofor the smallest full-dim file, when the dim is a multiple of 64 and the recall above is acceptable for your data.micro(a recipe, built withdtype="int8", mrl_dim=256) only with a model trained for truncation, or when size wins over recall.hybridwritesfloat32vectors, an HNSW index and a BM25 index. It is the default ofurna build --spec.
hybrid files are searched as hnsw
A hybrid file declares index_type = "hnsw", so urna ask, urna retrieve and urna search-text never read its BM25 section. See known limits.
Changing the preset changes content_hash, and with it every citation: the dtype, the truncation and the index type are all part of the hashed content. Pick the preset before you publish citations. The search routes themselves are described in Search paths and the exact rerank.
Offline by construction
Where urna touches the network and where it cannot: the runtime has no network stack, queries embed offline, and only installers and setup download.
Reproducible builds
What makes two urna builds byte-identical, the L1, L2 and L3 reproduction levels, the build lock, the embed cache keyed by three hashes, and --rebuild-only.