urna benchmark
Reference for urna benchmark, which times exact search on random query vectors, and optionally HNSW with recall@k, a cold-cache pass, or a named space.
urna benchmark measures search latency on a .urna file with random query vectors. It always times the exact path. It can also time the HNSW path and report its recall@k against exact, repeat each run with the page cache dropped, or time one named space instead of the text embeddings.
Usage
urna benchmark [OPTIONS] <FILE>Arguments
| Argument | Description |
|---|---|
<FILE> | Path to the .urna file |
Options
| Option | Default | Description |
|---|---|---|
-q, --queries <QUERIES> | 100 | Number of random queries |
-k, --k <K> | 10 | Hits per query. Must be greater than 0 |
--ann <ANN> | If set, also benchmark search_ann with the given ef | |
--madvise-cold | off | Force a "madvise-cold" cache between queries by calling posix_madvise(MADV_DONTNEED) on the mmap. Approximates the first hit after boot, but it is a hint, not a guarantee |
--space <SPACE> | Benchmark the named multimodal space instead of the default path | |
-h, --help | Print help |
Behavior
- Opens the file with the full runtime check.
- Builds
--queriesrandom query vectors at the file'sembedding_dim(or the space's dim with--space). Each component is drawn from 0 to 1 and the vector is L2-normalized. The generator is not seeded, so two runs use different queries and print slightly different numbers. - Runs each query once through the exact path and records its time. The time covers the whole search call: query validation, scoring and building the hits.
- With
--madvise-cold, runs the same queries again, asking the kernel to drop the file's pages before each one. This is a hint on Unix and does nothing on other systems, so treat the result as an estimate of cold-cache latency, not a measurement of a cold boot. - With
--ann EF: if the file has no HNSW section, prints(no HNSW section - ANN bench skipped)and stops. Otherwise times the HNSW path (and its cold pass with--madvise-cold), then runs every query through both paths and prints the mean recall@k: for each query, how many of the exact topkchunks the HNSW topkcontains, divided byk. - With
--space NAME: times the exact scan of that space's band instead, with an optional cold pass.--annis ignored and no recall is printed.
Things to keep in mind when reading the numbers:
- The random queries all point into the positive orthant. They exercise the kernels and the memory path, but their recall@k can differ from the recall on real queries. For recall on your own queries, compare
search-annandsearchhits yourself. - The HNSW beam is
max(ef, k, ef_construction). On a file built with the defaultef_construction = 400, any--annvalue of 400 or less measures the same beam. See Search paths and the exact rerank. - recall@k divides by
k. On a corpus with fewer thankchunks it cannot reach 1.
Output
To stdout, one block per run, with latencies in milliseconds:
Exact (<n> queries, dim=<dim>, dtype=<dtype>, simd=<backend>) [hot]:
mean: <ms> ms
p50: <ms> ms
p95: <ms> ms
p99: <ms> msFurther blocks are labelled [madvise-cold], ANN ef=<ef> (<n> queries) [hot] and ANN ef=<ef> (<n> queries) [madvise-cold], followed by recall@<k> (ANN vs exact): <value>. A space run is labelled Space '<name>' (<n> queries, dim=<dim>, dtype=<dtype>).
Exit codes
| Code | Meaning |
|---|---|
0 | The benchmark ran |
1 | File missing or unreadable, a failed check, k of 0 or less, or an unknown space. Printed to stderr as Error: <message> |
2 | Usage error: a missing or unparseable argument |
With --space on a file that has no space table, the error names the missing section by its decimal id: Error: section 21 not found (0x15). With a space table but an unknown name: Error: space '<name>' not found in the space_table.
Examples
Exact and HNSW on the quickstart corpus:
urna benchmark examples/quickstart/out/quickstart.urna -q 100 -k 10 --ann 100Exact (100 queries, dim=256, dtype=float32, simd=neon) [hot]:
mean: 0.005 ms
p50: 0.004 ms
p95: 0.006 ms
p99: 0.022 ms
ANN ef=100 (100 queries) [hot]:
mean: 0.007 ms
p50: 0.006 ms
p95: 0.008 ms
p99: 0.023 ms
recall@10 (ANN vs exact): 1.0000This corpus has 12 chunks, so the HNSW beam covers all of them and recall is 1. The numbers say little about a real corpus; run the benchmark on yours.
With a cold pass:
urna benchmark examples/quickstart/out/quickstart.urna -q 20 -k 5 --madvise-coldExact (20 queries, dim=256, dtype=float32, simd=neon) [hot]:
mean: 0.003 ms
p50: 0.003 ms
p95: 0.004 ms
p99: 0.015 ms
Exact (20 queries, dim=256, dtype=float32, simd=neon) [madvise-cold]:
mean: 0.003 ms
p50: 0.003 ms
p95: 0.003 ms
p99: 0.004 ms
(note: posix_madvise(MADV_DONTNEED) is a hint, not a guarantee. Treat as an upper bound on cold-cache latency, not absolute cold.)Published measurements on larger corpora, with their conditions, are on Benchmarks.
urna media
Reference for urna media, which lists the media blobs a .urna corpus references and exports the inlined ones to files after checking each SHA-256.
urna doctor
Reference for urna doctor, the offline install check that tests the Python env, the potion embedder and one real embed, and exits with a typed code.