Search paths and the exact rerank
The exact, HNSW, hybrid, graph and named-space search paths in urna, how ask and retrieve pick one, and why every returned score is a real cosine.
urna has several ways to find candidate chunks and one way to score them. Every path, however it generates candidates, ends by recomputing the cosine between the query and each candidate from the stored vectors. The score you get back is that recomputed cosine, never an index distance or a fusion score.
The scoring contract
Every file declares metric = "ip" and normalize = "l2": vectors are stored L2-normalized and scored by inner product, which on unit vectors is cosine similarity. Hits always report score_type: "cosine". The writer does not check the norm of the vectors it is given, so this holds when the embedder that built the file normalized its output, as the bundled embedders do.
Before any path runs, the runtime checks the query and stops with a typed error on the first problem:
kis greater than 0.- The query is not empty.
- Its length equals the file's
embedding_dim. - Every value is finite (no NaN or Inf).
- Its norm is not zero.
Then it L2-normalizes the query, so you do not have to.
The paths
| Path | Candidates come from | Reached by | recall in the result |
|---|---|---|---|
| Exact | Every chunk | search, ask/retrieve on index_type = "exact" | 1 |
| HNSW | The HNSW graph (0x07) | search-ann, ask/retrieve on index_type = "hnsw" | not computed |
| Hybrid | HNSW (or exact) shortlist plus a BM25 shortlist (0x08) | ask/retrieve/search-text on index_type = "hybrid", Python search_hybrid | not computed |
| Graph | Exact seeds expanded over the chunk graph (0x0C) | search-graph, Python search_graph | not computed |
| Named space | Every vector in one space band | search-space, Python search_space | 1 |
The CLI prints a recall that was not measured as (not computed; rerank guarantees real cosine). The rerank guarantees that each returned score is a real cosine. It does not guarantee that an approximate path found the true top k; measure that with urna benchmark --ann.
Exact
The exact path scores every chunk and sorts, with ties kept in file order. It is the ground truth the other paths are compared against.
HNSW
The HNSW path walks the graph stored in hnsw_index to collect a candidate list, then reranks it exactly. If the file has no HNSW section, the path falls back to exact and the result says index_type: exact.
The beam width is max(ef, k, ef_construction), where ef_construction is the value the file was built with. Python and the forge build with ef_construction = 400.
--ef below the build value does nothing
The runtime never searches with a beam narrower than the file's ef_construction. On a file built with the default 400, any --ef or --candidates value below 400 runs at 400, so an ef sweep below that value is flat. Only values above ef_construction change the result. See Known limits.
Hybrid
The hybrid path builds two shortlists. The vector shortlist comes from HNSW when the file has it (with the same beam floor as above), else it is the top candidates_per_path of an exact scan. The BM25 shortlist is the top candidates_per_path chunks by BM25 over the stored text. There is no BM25-only path: BM25 runs only inside hybrid. The path merges the two lists with reciprocal-rank fusion (RRF, k = 60), rescores every member of the merged set by exact cosine, and returns the top k by cosine.
That last step defines what hybrid does in urna: the RRF order is discarded and the final order is pure cosine. BM25 can bring a chunk into the candidate set that the vector shortlist missed; it never lifts a chunk above one with a higher cosine. A rare term or proper noun only reaches the top k if its chunk's cosine earns it.
Two edge cases:
- The path never falls back. Without a BM25 section the lexical list is empty and the route is still
hybrid. - Without HNSW, the vector shortlist holds at most
candidates_per_pathchunks, with no minimum ofk. With no HNSW, no BM25 andcandidates_per_pathbelowk, fewer thankhits come back.
BM25 uses k1 = 1.5 and b = 0.75. Its tokenizer lowercases runs of letters and digits, splits on anything else, and drops tokens shorter than 2 characters. There is no stemming and no stop-word list. A run with no spaces or punctuation stays one token, so an unspaced CJK clause only matches an identical whole run.
The hybrid preset does not route to hybrid
preset="hybrid" in urna.build, and the forge's default [build] preset = "hybrid", write an HNSW index and a BM25 index but declare index_type = "hnsw". ask, retrieve and search-text route on the declared value, so on these files they take the HNSW path and never read the BM25 index. The BM25 index is reached only by calling search_hybrid yourself, from Python (UrnaFile.search_hybrid) or the Rust runtime. A file declares index_type = "hybrid" only when built with the Rust builder's .hybrid(). See Known limits.
Graph
The graph path seeds from the exact top max(ef, k) chunks, expands them over the chunk-to-chunk graph for hops steps (all edge types, capped at 8 times the seed count), and reranks the union exactly. If the file has no graph_adjacency section, or its manifest does not set capabilities_ext.graph_present, it falls back to exact.
Because the seeds already are the exact top max(ef, k) and the rerank uses the same scores and ordering, the hits equal the exact path's hits. The graph adds neighbours to the candidate set, but a neighbour cannot outrank a seed it would need to displace. The path costs an exact scan plus the traversal.
Named spaces
A file can carry extra embedding spaces, for example an image tower next to the text embeddings (see Media and named spaces). A named-space search is an exact scan over one space's band. The query must come from that space's model and have that space's dim. Hits map back to the same chunks, so they carry the same chunk_id and citation as a text hit on that chunk. The text paths never read a space band, and a space search never falls back to the text embeddings.
How ask and retrieve choose a path
urna ask, urna retrieve and urna search-text embed the query text, pass it through the model gate, and then route on the manifest's declared index_type:
Declared index_type | Path | Candidates |
|---|---|---|
exact | Exact | Every chunk |
hnsw | HNSW | --candidates, default max(4 * k, 64), then the beam floor above |
hybrid | Hybrid | --candidates per path, same default |
The manifest admits only exact, hnsw and hybrid, so none of these verbs takes the graph path. Routing does not look at which sections exist or at the capabilities flags; urna stats shows the index_type that decides it.
The quickstart corpus declares hnsw and also carries BM25 and a graph. ask --disclose explain shows the route it took:
urna ask examples/quickstart/out/quickstart.urna "can I use this offline" -k 1 --disclose explainroute: hnsw
candidates: exact=0 ann=12 bm25=0 graph=0 fusion=none
rerank_source: real cosine
recall: (not computed; rerank guarantees real cosine)The file has 12 chunks, so the HNSW candidate list is the whole corpus. BM25 and the graph are present and unused.
The rerank source
The rerank reads vectors from one slab. When a full-precision slab (embeddings_fp, 0x09) is present, it reads that. Otherwise it reads the stored embeddings section at its stored dtype. No writer in 0.5.1 emits 0x09, so in practice the rerank reads the stored embeddings.
| Effective dtype | rerank_source in retrieve JSON | --disclose explain line |
|---|---|---|
float32 | full_precision | real cosine |
float16, int8, int4 | stored_precision | real cosine at stored precision |
A stored-precision score is still a cosine recomputed from the vectors, not a proxy from the index. It is the cosine between the query and the quantized vector, so it can differ from the float32 score of the same pair. See Presets and stored precision.
Scoring runs on AVX2 (with FMA) on x86_64, NEON on aarch64, or a scalar fallback, detected once per process; float16 has no AVX2 kernel and scores with the scalar kernel on x86_64. Set URNA_FORCE_SCALAR to any value other than 0 to force the scalar kernels (see Environment variables). The int4 kernel gives bit-identical scores on all three.
The per-verb flags are in the reference, starting with urna search.
Citations and hashes
How a urna:// citation is built from content_hash and chunk_id, which changes move it, what the offsets mean, and what file_hash covers.
The model gate
Why urna refuses a text query embedded by a different model than the corpus, what model_hash covers, and where the check runs or does not.