docsv0.5.1

urna benchmark

Reference for urna benchmark, which times exact search on random query vectors, and optionally HNSW with recall@k, a cold-cache pass, or a named space.

urna benchmark measures search latency on a .urna file with random query vectors. It always times the exact path. It can also time the HNSW path and report its recall@k against exact, repeat each run with the page cache dropped, or time one named space instead of the text embeddings.

Usage

urna benchmark [OPTIONS] <FILE>

Arguments

ArgumentDescription
<FILE>Path to the .urna file

Options

OptionDefaultDescription
-q, --queries <QUERIES>100Number of random queries
-k, --k <K>10Hits per query. Must be greater than 0
--ann <ANN>If set, also benchmark search_ann with the given ef
--madvise-coldoffForce a "madvise-cold" cache between queries by calling posix_madvise(MADV_DONTNEED) on the mmap. Approximates the first hit after boot, but it is a hint, not a guarantee
--space <SPACE>Benchmark the named multimodal space instead of the default path
-h, --helpPrint help

Behavior

  1. Opens the file with the full runtime check.
  2. Builds --queries random query vectors at the file's embedding_dim (or the space's dim with --space). Each component is drawn from 0 to 1 and the vector is L2-normalized. The generator is not seeded, so two runs use different queries and print slightly different numbers.
  3. Runs each query once through the exact path and records its time. The time covers the whole search call: query validation, scoring and building the hits.
  4. With --madvise-cold, runs the same queries again, asking the kernel to drop the file's pages before each one. This is a hint on Unix and does nothing on other systems, so treat the result as an estimate of cold-cache latency, not a measurement of a cold boot.
  5. With --ann EF: if the file has no HNSW section, prints (no HNSW section - ANN bench skipped) and stops. Otherwise times the HNSW path (and its cold pass with --madvise-cold), then runs every query through both paths and prints the mean recall@k: for each query, how many of the exact top k chunks the HNSW top k contains, divided by k.
  6. With --space NAME: times the exact scan of that space's band instead, with an optional cold pass. --ann is ignored and no recall is printed.

Things to keep in mind when reading the numbers:

  • The random queries all point into the positive orthant. They exercise the kernels and the memory path, but their recall@k can differ from the recall on real queries. For recall on your own queries, compare search-ann and search hits yourself.
  • The HNSW beam is max(ef, k, ef_construction). On a file built with the default ef_construction = 400, any --ann value of 400 or less measures the same beam. See Search paths and the exact rerank.
  • recall@k divides by k. On a corpus with fewer than k chunks it cannot reach 1.

Output

To stdout, one block per run, with latencies in milliseconds:

Exact (<n> queries, dim=<dim>, dtype=<dtype>, simd=<backend>) [hot]:
  mean:   <ms> ms
  p50:    <ms> ms
  p95:    <ms> ms
  p99:    <ms> ms

Further blocks are labelled [madvise-cold], ANN ef=<ef> (<n> queries) [hot] and ANN ef=<ef> (<n> queries) [madvise-cold], followed by recall@<k> (ANN vs exact): <value>. A space run is labelled Space '<name>' (<n> queries, dim=<dim>, dtype=<dtype>).

Exit codes

CodeMeaning
0The benchmark ran
1File missing or unreadable, a failed check, k of 0 or less, or an unknown space. Printed to stderr as Error: <message>
2Usage error: a missing or unparseable argument

With --space on a file that has no space table, the error names the missing section by its decimal id: Error: section 21 not found (0x15). With a space table but an unknown name: Error: space '<name>' not found in the space_table.

Examples

Exact and HNSW on the quickstart corpus:

urna benchmark examples/quickstart/out/quickstart.urna -q 100 -k 10 --ann 100
Exact (100 queries, dim=256, dtype=float32, simd=neon) [hot]:
  mean:   0.005 ms
  p50:    0.004 ms
  p95:    0.006 ms
  p99:    0.022 ms
ANN ef=100 (100 queries) [hot]:
  mean:   0.007 ms
  p50:    0.006 ms
  p95:    0.008 ms
  p99:    0.023 ms
  recall@10 (ANN vs exact): 1.0000

This corpus has 12 chunks, so the HNSW beam covers all of them and recall is 1. The numbers say little about a real corpus; run the benchmark on yours.

With a cold pass:

urna benchmark examples/quickstart/out/quickstart.urna -q 20 -k 5 --madvise-cold
Exact (20 queries, dim=256, dtype=float32, simd=neon) [hot]:
  mean:   0.003 ms
  p50:    0.003 ms
  p95:    0.004 ms
  p99:    0.015 ms
Exact (20 queries, dim=256, dtype=float32, simd=neon) [madvise-cold]:
  mean:   0.003 ms
  p50:    0.003 ms
  p95:    0.003 ms
  p99:    0.004 ms
  (note: posix_madvise(MADV_DONTNEED) is a hint, not a guarantee. Treat as an upper bound on cold-cache latency, not absolute cold.)

Published measurements on larger corpora, with their conditions, are on Benchmarks.

On this page