urna search
Reference for urna search, the exact search verb that takes a query vector as a JSON array and prints every hit with its cosine score and citation.
urna search runs an exact search over a .urna file with a query vector you supply as a JSON array. It scores every chunk, returns the top k by cosine, and prints each hit with its citation. It runs no Python and embeds nothing: to search with text, use urna ask, urna retrieve or urna search-text.
Usage
urna search [OPTIONS] <FILE> <QUERY>Arguments
| Argument | Description |
|---|---|
<FILE> | Path to the .urna file |
<QUERY> | The query vector as a JSON array of numbers, with exactly embedding_dim values (see urna stats) |
Options
| Option | Default | Description |
|---|---|---|
-k, --k <K> | 10 | Number of hits to return. Must be greater than 0 |
-h, --help | Print help |
Behavior
- Opens the file with the full runtime check (see what the reader checks on open).
- Parses
<QUERY>as JSON. - Validates the query:
kabove 0, not empty, length equal toembedding_dim, no NaN or Inf, norm not zero. Then L2-normalizes it, so the vector does not need unit length. - Scores every chunk by cosine and returns the top
k. Ties keep file order.
There is no model gate on this verb. urna search cannot know which model produced your vector, so a vector from a different model with the right length returns hits with meaningless scores. Embed the query with the model the file records in model and model_hash (shown by urna stats; see the model gate).
When k is larger than the number of chunks, every chunk is returned and k_returned is the chunk count.
Output
All output goes to stdout. The result header comes first, then one line per hit.
| Field | Meaning |
|---|---|
index_type | The path that ran: exact for this verb |
recall | 1 on the exact path. Paths that did not measure it print (not computed; rerank guarantees real cosine) |
truncated | true when k is smaller than the number of chunks |
k_requested, k_returned | The k you asked for and the number of hits printed |
query_time | Milliseconds for validation, scoring and building the hits |
Each hit line holds:
| Field | Meaning |
|---|---|
[ n] | Rank, starting at 1 |
chunk_id | The chunk's sha256: id |
score | Cosine similarity, 6 decimals |
score_type | Always cosine |
source_uri | Where the chunk came from |
offset | byte_start-byte_end from the stored span (see what the offsets mean) |
model | The manifest's embedding_model |
index_type | The path that produced the hit |
reranked | false on the exact path, true on paths that rerank candidates |
file_hash | SHA-256 of the whole file |
content_hash | SHA-256 of the canonical sections |
citation_id | urna://content_hash/chunk_id, resolvable with urna cite |
The same format is printed by search-ann, search-graph, search-space and search-text.
Exit codes
| Code | Meaning |
|---|---|
0 | Search ran and printed its result |
1 | Any error: file missing or unreadable, a failed integrity check, invalid query JSON, wrong dimension, NaN or Inf, zero-norm query, k of 0 or less. The message is printed to stderr as Error: <message> |
2 | Usage error: a missing or unparseable argument |
Examples
Search the quickstart corpus. Here query.json holds a 256-value array: the stored embedding of the corpus's first chunk, so that chunk comes back with a score of 1.000000.
urna search examples/quickstart/out/quickstart.urna "$(cat query.json)" -k 2index_type: exact
recall: 1
truncated: true
k_requested: 2
k_returned: 2
query_time: 1.503 ms
hits:
[ 1] chunk_id=sha256:19f36b3e072d553eb83626bf30db5e8f3b1f729a495f7826999ffeae53848e6e score=1.000000 score_type=cosine source_uri=demo/01-what-is-urna.md offset=0-1 model=minishlab/potion-base-8M/v1 index_type=exact reranked=false file_hash=sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832 content_hash=sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df citation_id=urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:19f36b3e072d553eb83626bf30db5e8f3b1f729a495f7826999ffeae53848e6e
[ 2] chunk_id=sha256:47ecc7e1141759d285f08545d74de807a0196a73d5db27b024e7bbe3b2f78bb4 score=0.664475 score_type=cosine source_uri=demo/03-citations.md offset=6-7 model=minishlab/potion-base-8M/v1 index_type=exact reranked=false file_hash=sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832 content_hash=sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df citation_id=urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:47ecc7e1141759d285f08545d74de807a0196a73d5db27b024e7bbe3b2f78bb4A vector of the wrong length:
urna search examples/quickstart/out/quickstart.urna "[0.1, 0.2]"Error: dimension mismatch: expected 256, got 2How the exact path relates to the others is in Search paths and the exact rerank.
urna build
Reference for urna build, the launcher of the declarative corpus build: flags, how it finds the forge and Python, stages, output and exit codes.
urna search-text
Search a .urna file by raw text through the sentence-transformers query embedder, with the model gate, and print every hit field for inspection.