docsv0.5.1

Benchmarks

urna against hnswlib, usearch, sqlite-vec and lancedb, the size and recall of each build preset, and 38,627 card images in one file, with every condition.

This page collects the measurements the project has published for urna 0.5.1: the engine against four other vector stores, the size and rank stability of each build preset, and a 38,627-image corpus. Each table states the data, the machine and the ruler it was measured with. A number here holds for those conditions only; measure your own corpus before you rely on one. The same tables, drawn as charts, are on urna.dev.

All latencies below time the search after the query vector exists. ask, retrieve and search-text also start a Python process per call to embed the question, and none of these tables include that time.

The engine against other stores

Conditions

  • Measured on 2026-09-10 on an arm64 Mac (Darwin 25.6.0) with Python 3.12.14, one thread for every store.
  • Data: 100,000 synthetic rows of 384 dimensions, L2-normalized and clustered around 2,000 centers (random rows have no neighborhoods, and every HNSW implementation scores badly on them).
  • Queries: 200, each a stored row plus a small amount of noise, re-normalized. k = 10, seed 7.
  • Recall@10 is measured against a brute-force top 10 over the same rows.
  • Every store is driven from Python, so Python call overhead is inside every latency.
  • Versions: usearch 2.26.2, hnswlib 0.8.0, sqlite-vec 0.1.9, lancedb 0.38.0, numpy 2.5.2.

Results

StoreSearch pathBuild (s)Size (MB)Cold open + first query (ms)p50 (ms)p99 (ms)Recall@10Rebuild byte-identicalChecks its own bytes
urna, exact presetexact2.14165.3292.37.8038.3151.0yesyes: per section, whole file, decoded content
urna, hybrid presetHNSW172.82163.2356.10.7211.0211.0yesyes: per section, whole file, decoded content
usearchHNSW109.26168.558.10.67261.4220.995yesno
hnswlibHNSW83.07168.4181.50.3240.5251.0yesno
sqlite-vecexact0.81156.650.019.76224.8361.0yesstructure only (pragma integrity_check)
lancedbexact0.25153.8612.416.72819.4011.0nono

Size is bytes on disk divided by 10^6.

How to read it

  • The urna hybrid row is a file built with the hybrid preset (float32 vectors, HNSW with m = 16 and ef_construction = 200, plus a BM25 index) and queried through the HNSW path with ef = 100. BM25 plays no part in it. urna raises the search beam to the file's ef_construction, so this row effectively searched with a beam of 200, while hnswlib searched with ef = 100. See Known limits.
  • Cold open is the time a fresh Python process takes to open the store and answer one query, minus the time of a process that does nothing, best of 3 runs. urna's number is dominated by checking every section checksum and the footer hash before the first answer; the other stores do not check their bytes.
  • Build time is single-threaded everywhere, including hnswlib and usearch. urna's HNSW build is the slowest row.
  • Rebuild byte-identical means two builds from the same rows produced the same SHA-256.
  • The same rows written with raw text and with zstd text share one content_hash, so re-encoding a file never moves its citations.

The table says nothing about workloads urna does not serve: updates or deletes in place, metadata filtering, concurrent writers or a query language. A .urna file is built once and queried many times.

Reproduce

From a checkout, with the four other stores installed in the same Python environment (a store whose package is missing is skipped, not faked):

.venv/bin/python python/tools/bench_competitors.py --n 100000 --dim 384 --queries 200

--n, --dim and --queries change the corpus and the query count. The script regenerates docs/BENCH.md.

Presets: size against rank stability

A preset chooses how a file stores its text and vectors and which indices it carries. This table measures what each one costs in size and gives up in ranking.

Conditions

  • Baseline: a 30,725-chunk Brazilian Portuguese corpus embedded with a 384-dimension MiniLM model, stored as float32 with no index (119.94 MB) and searched exactly.
  • 100 queries, k = 10, seed 0, NEON SIMD, hot cache, one query at a time from Python. The ladder file records the SIMD backend but not the machine or the date.
  • tiny, micro and nano were queried through HNSW with ef = 100 on files built with ef_construction = 400, so the effective beam was 400. hybrid was queried with an explicit hybrid search call (vector and BM25 candidates, 200 per path, with the first 200 characters of the chunk as the query text). ask and retrieve do not take that path on a hybrid preset file; they take HNSW.
  • The ruler is self-perturbation: each query is a stored chunk's own vector plus up to 8e-5 of noise per dimension, and the ground truth is the float32 baseline's own top 10 for that query. It measures how stable the ranking stays under quantization. It does not measure retrieval quality on real questions, and it is easier than real retrieval, so these recall figures are likely inflated.

Results

PresetEmbeddingsIndexSize (MB)Size ratioRecall@10p50 (ms)p99 (ms)
exactfloat32none119.941.0001.0003.1033.142
compressedfloat16none40.670.3391.0003.1663.225
tinyint8HNSW30.690.2560.9921.1571.529
microint8 at 256 dimensionsHNSW26.760.2230.8100.7711.023
nanoint4HNSW25.040.2090.9132.0602.721
hybridfloat32HNSW and BM2573.030.6091.0004.0284.834
  • micro is a recipe, not a preset value: urna.build(text_encoding="zstd", dtype="int8", mrl_dim=256, with_hnsw=True).
  • The int8 and int4 files keep no full-precision copy of the vectors, so their scores and recall are real cosine at the stored precision, not float32 cosine.
  • compressed scoring 1.000 here does not make float16 lossless; it means the float16 ranking matched float32 under this ruler.

Truncating dimensions

mrl_dim keeps the first dimensions of each vector and re-normalizes them. It pays off on models trained for it. The baseline model was not, so on this corpus truncation costs recall:

VariantSize (MB)Size ratioRecall@10
int8, 256 dimensions (micro)26.760.2230.810
int8, 192 dimensions24.800.2070.733
int8, 128 dimensions22.830.1900.659
int8, 96 dimensions21.850.1820.574
int4, 256 dimensions22.950.1910.777
int4, 192 dimensions21.910.1830.713
int4, 128 dimensions20.860.1740.627

Same conditions and ruler as the preset table. On this corpus nano (int4 at the full 384 dimensions) keeps more recall than every truncated variant.

Image corpora: 38,627 card images

These numbers come from a corpus of 38,627 Magic: The Gathering card images built into .urna files with the image media profiles. The code and data of that benchmark are private for now. The machine is not recorded in the published results.

Media profiles

ProfileMediaFileAgainst the JPEG source
archiveJPEG XL, byte-reversible JPEG transcode3.61 GB1.10x smaller
AV1 all-intra, crf 35one AV1 stream, every frame a keyframe1.37 GB2.89x smaller
retrievalAV1 all-intra, crf 50533 MB7.46x smaller
  • The AV1 crf 35 row is what the forge now calls the stills-av1 profile. The forge's stills profile encodes one AVIF per image instead (quality 48, speed 8). Measured on 2026-09-12: 1,195,973,116 bytes against 1,374,431,484 bytes for the all-intra AV1 stream, 13% smaller, at a matched SSIMULACRA 2 mean of 61.96 on a 2,048-card sample. The cost is a CLIP embed at build time 4 to 10 times slower, because frames decode one AVIF at a time.
  • The retrieval row was measured on 2026-09-03: 532,671,548 bytes in one self-contained file, with no measurable loss of text-to-image hit@1 on 100 queries.

Text-to-image search by model

Text-to-image hit@1 over every card:

Model presethit@1
siglip20.750
wemm-2b0.744
a jina v5 omni preset (size not recorded)0.336
clip-vit-b320.098

hit@1 counts a query as a hit when the card it describes is the top result.

Measurements behind the media defaults

  • The still-picture tune: on 2,048 cards at crf 35, SSIMULACRA 2 p50 of 62.7 against 51.8 with the encoder's default tune, for 10% more bytes.
  • The quality gate floors (crf = "auto"): cards at 488x680, yuv420, a 2,048-card sample, AV1 with the still tune at speed 6. crf 30 measured SSIMULACRA 2 p10 65.3, minimum 58.6 and embedding drift p10 0.967, and passes. crf 35 measured p50 62.7, p10 55.7, minimum 45.3 and drift p10 0.965, and fails on p10. The default floors therefore pick crf 30.
  • Drift against task utility, measured on 2026-09-12: CLIP cosine drift p10 fell from 0.932 at crf 40 to 0.829 at crf 60, so the default drift floor refuses every rung, while text-to-image hit@1 on 100 queries did not move up to crf 50. This is why the retrieval profiles gate on hit@1 or pin the crf.
  • Cluster ordering on a corpus of same-artwork reprints (2026-08-31): 29% smaller than all-intra, using 16-frame groups with scene-change detection off.

Measure your own corpus

urna benchmark times exact search on your file and, with --ann, the HNSW path and its recall against exact search on the same random queries. That recall measures agreement with exact search, not answer quality. See urna benchmark.

urna benchmark my_corpus.urna -q 100 -k 10 --ann 100

On this page