Benchmarks
urna against hnswlib, usearch, sqlite-vec and lancedb, the size and recall of each build preset, and 38,627 card images in one file, with every condition.
This page collects the measurements the project has published for urna 0.5.1: the engine against four other vector stores, the size and rank stability of each build preset, and a 38,627-image corpus. Each table states the data, the machine and the ruler it was measured with. A number here holds for those conditions only; measure your own corpus before you rely on one. The same tables, drawn as charts, are on urna.dev.
All latencies below time the search after the query vector exists. ask, retrieve and search-text also start a Python process per call to embed the question, and none of these tables include that time.
The engine against other stores
Conditions
- Measured on 2026-09-10 on an arm64 Mac (Darwin 25.6.0) with Python 3.12.14, one thread for every store.
- Data: 100,000 synthetic rows of 384 dimensions, L2-normalized and clustered around 2,000 centers (random rows have no neighborhoods, and every HNSW implementation scores badly on them).
- Queries: 200, each a stored row plus a small amount of noise, re-normalized.
k = 10, seed 7. - Recall@10 is measured against a brute-force top 10 over the same rows.
- Every store is driven from Python, so Python call overhead is inside every latency.
- Versions: usearch 2.26.2, hnswlib 0.8.0, sqlite-vec 0.1.9, lancedb 0.38.0, numpy 2.5.2.
Results
| Store | Search path | Build (s) | Size (MB) | Cold open + first query (ms) | p50 (ms) | p99 (ms) | Recall@10 | Rebuild byte-identical | Checks its own bytes |
|---|---|---|---|---|---|---|---|---|---|
urna, exact preset | exact | 2.14 | 165.3 | 292.3 | 7.803 | 8.315 | 1.0 | yes | yes: per section, whole file, decoded content |
urna, hybrid preset | HNSW | 172.82 | 163.2 | 356.1 | 0.721 | 1.021 | 1.0 | yes | yes: per section, whole file, decoded content |
| usearch | HNSW | 109.26 | 168.5 | 58.1 | 0.672 | 61.422 | 0.995 | yes | no |
| hnswlib | HNSW | 83.07 | 168.4 | 181.5 | 0.324 | 0.525 | 1.0 | yes | no |
| sqlite-vec | exact | 0.81 | 156.6 | 50.0 | 19.762 | 24.836 | 1.0 | yes | structure only (pragma integrity_check) |
| lancedb | exact | 0.25 | 153.8 | 612.4 | 16.728 | 19.401 | 1.0 | no | no |
Size is bytes on disk divided by 10^6.
How to read it
- The urna
hybridrow is a file built with thehybridpreset (float32 vectors, HNSW withm = 16andef_construction = 200, plus a BM25 index) and queried through the HNSW path withef = 100. BM25 plays no part in it. urna raises the search beam to the file'sef_construction, so this row effectively searched with a beam of 200, while hnswlib searched withef = 100. See Known limits. - Cold open is the time a fresh Python process takes to open the store and answer one query, minus the time of a process that does nothing, best of 3 runs. urna's number is dominated by checking every section checksum and the footer hash before the first answer; the other stores do not check their bytes.
- Build time is single-threaded everywhere, including hnswlib and usearch. urna's HNSW build is the slowest row.
- Rebuild byte-identical means two builds from the same rows produced the same SHA-256.
- The same rows written with raw text and with zstd text share one
content_hash, so re-encoding a file never moves its citations.
The table says nothing about workloads urna does not serve: updates or deletes in place, metadata filtering, concurrent writers or a query language. A .urna file is built once and queried many times.
Reproduce
From a checkout, with the four other stores installed in the same Python environment (a store whose package is missing is skipped, not faked):
.venv/bin/python python/tools/bench_competitors.py --n 100000 --dim 384 --queries 200--n, --dim and --queries change the corpus and the query count. The script regenerates docs/BENCH.md.
Presets: size against rank stability
A preset chooses how a file stores its text and vectors and which indices it carries. This table measures what each one costs in size and gives up in ranking.
Conditions
- Baseline: a 30,725-chunk Brazilian Portuguese corpus embedded with a 384-dimension MiniLM model, stored as float32 with no index (119.94 MB) and searched exactly.
- 100 queries,
k = 10, seed 0, NEON SIMD, hot cache, one query at a time from Python. The ladder file records the SIMD backend but not the machine or the date. tiny,microandnanowere queried through HNSW withef = 100on files built withef_construction = 400, so the effective beam was 400.hybridwas queried with an explicit hybrid search call (vector and BM25 candidates, 200 per path, with the first 200 characters of the chunk as the query text).askandretrievedo not take that path on ahybridpreset file; they take HNSW.- The ruler is self-perturbation: each query is a stored chunk's own vector plus up to 8e-5 of noise per dimension, and the ground truth is the float32 baseline's own top 10 for that query. It measures how stable the ranking stays under quantization. It does not measure retrieval quality on real questions, and it is easier than real retrieval, so these recall figures are likely inflated.
Results
| Preset | Embeddings | Index | Size (MB) | Size ratio | Recall@10 | p50 (ms) | p99 (ms) |
|---|---|---|---|---|---|---|---|
exact | float32 | none | 119.94 | 1.000 | 1.000 | 3.103 | 3.142 |
compressed | float16 | none | 40.67 | 0.339 | 1.000 | 3.166 | 3.225 |
tiny | int8 | HNSW | 30.69 | 0.256 | 0.992 | 1.157 | 1.529 |
micro | int8 at 256 dimensions | HNSW | 26.76 | 0.223 | 0.810 | 0.771 | 1.023 |
nano | int4 | HNSW | 25.04 | 0.209 | 0.913 | 2.060 | 2.721 |
hybrid | float32 | HNSW and BM25 | 73.03 | 0.609 | 1.000 | 4.028 | 4.834 |
microis a recipe, not a preset value:urna.build(text_encoding="zstd", dtype="int8", mrl_dim=256, with_hnsw=True).- The int8 and int4 files keep no full-precision copy of the vectors, so their scores and recall are real cosine at the stored precision, not float32 cosine.
compressedscoring 1.000 here does not make float16 lossless; it means the float16 ranking matched float32 under this ruler.
Truncating dimensions
mrl_dim keeps the first dimensions of each vector and re-normalizes them. It pays off on models trained for it. The baseline model was not, so on this corpus truncation costs recall:
| Variant | Size (MB) | Size ratio | Recall@10 |
|---|---|---|---|
int8, 256 dimensions (micro) | 26.76 | 0.223 | 0.810 |
| int8, 192 dimensions | 24.80 | 0.207 | 0.733 |
| int8, 128 dimensions | 22.83 | 0.190 | 0.659 |
| int8, 96 dimensions | 21.85 | 0.182 | 0.574 |
| int4, 256 dimensions | 22.95 | 0.191 | 0.777 |
| int4, 192 dimensions | 21.91 | 0.183 | 0.713 |
| int4, 128 dimensions | 20.86 | 0.174 | 0.627 |
Same conditions and ruler as the preset table. On this corpus nano (int4 at the full 384 dimensions) keeps more recall than every truncated variant.
Image corpora: 38,627 card images
These numbers come from a corpus of 38,627 Magic: The Gathering card images built into .urna files with the image media profiles. The code and data of that benchmark are private for now. The machine is not recorded in the published results.
Media profiles
| Profile | Media | File | Against the JPEG source |
|---|---|---|---|
archive | JPEG XL, byte-reversible JPEG transcode | 3.61 GB | 1.10x smaller |
| AV1 all-intra, crf 35 | one AV1 stream, every frame a keyframe | 1.37 GB | 2.89x smaller |
retrieval | AV1 all-intra, crf 50 | 533 MB | 7.46x smaller |
- The AV1 crf 35 row is what the forge now calls the
stills-av1profile. The forge'sstillsprofile encodes one AVIF per image instead (quality 48, speed 8). Measured on 2026-09-12: 1,195,973,116 bytes against 1,374,431,484 bytes for the all-intra AV1 stream, 13% smaller, at a matched SSIMULACRA 2 mean of 61.96 on a 2,048-card sample. The cost is a CLIP embed at build time 4 to 10 times slower, because frames decode one AVIF at a time. - The
retrievalrow was measured on 2026-09-03: 532,671,548 bytes in one self-contained file, with no measurable loss of text-to-image hit@1 on 100 queries.
Text-to-image search by model
Text-to-image hit@1 over every card:
| Model preset | hit@1 |
|---|---|
siglip2 | 0.750 |
wemm-2b | 0.744 |
| a jina v5 omni preset (size not recorded) | 0.336 |
clip-vit-b32 | 0.098 |
hit@1 counts a query as a hit when the card it describes is the top result.
Measurements behind the media defaults
- The still-picture tune: on 2,048 cards at crf 35, SSIMULACRA 2 p50 of 62.7 against 51.8 with the encoder's default tune, for 10% more bytes.
- The quality gate floors (
crf = "auto"): cards at 488x680, yuv420, a 2,048-card sample, AV1 with the still tune at speed 6. crf 30 measured SSIMULACRA 2 p10 65.3, minimum 58.6 and embedding drift p10 0.967, and passes. crf 35 measured p50 62.7, p10 55.7, minimum 45.3 and drift p10 0.965, and fails on p10. The default floors therefore pick crf 30. - Drift against task utility, measured on 2026-09-12: CLIP cosine drift p10 fell from 0.932 at crf 40 to 0.829 at crf 60, so the default drift floor refuses every rung, while text-to-image hit@1 on 100 queries did not move up to crf 50. This is why the
retrievalprofiles gate on hit@1 or pin the crf. - Cluster ordering on a corpus of same-artwork reprints (2026-08-31): 29% smaller than all-intra, using 16-frame groups with scene-change detection off.
Measure your own corpus
urna benchmark times exact search on your file and, with --ann, the HNSW path and its recall against exact search on the same random queries. That recall measures agreement with exact search, not answer quality. See urna benchmark.
urna benchmark my_corpus.urna -q 100 -k 10 --ann 100Run in Docker
Build the urna Docker image from the repository's Dockerfile, a static binary in a scratch image, and run the engine verbs on a mounted corpus.
Known limits
Where urna 0.5.1 behaves differently from what you might expect, grouped by building, querying, file format and install, with a workaround for each.