docsv0.5.1

Ask and retrieve from the terminal

Ask a .urna corpus questions from the shell, tune k and candidates, read explain mode, script retrieve with jq, and verify each citation with cite.

Three verbs cover querying from the shell: urna ask prints an answer for you to read, urna retrieve prints JSON for a script, and urna cite turns a citation back into the text it points at. This guide uses the quickstart corpus; any potion corpus works the same way.

Before you start

Run urna doctor. Exit 0 means the interpreter, numpy and tokenizers, the embedder script and the potion table are in place, and one real embed worked. If it fails, urna setup lays down what is missing. See Installation.

The examples use examples/quickstart/out/quickstart.urna, which urna build --spec examples/quickstart/corpus.toml writes in a checkout (see Quickstart). Check which model a corpus needs before asking it:

urna stats corpus.urna

A model: line that starts with minishlab/potion is answered by the installed binary. Any other model needs a checkout of the repository; see Open a corpus you downloaded.

Ask a question

urna ask examples/quickstart/out/quickstart.urna "can I use this offline" -k 1
[urna] embedder interpreter: /Users/nn/.local/share/urna/venv/bin/python
to keep that promise for a brand-new user, the default embedder is a static, offline embedder that ships with the tool. it needs no model download and no network round-trip on first use, and it is deterministic, so a build is byte-identical and reproducible. a power user can bring a stronger embedding model instead, and the model's fingerprint is recorded so the corpus and the query embedder must agree or the search fails loudly.
  -- urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:eed9a60b68133464e91c831f8af5960491f6c444cf1645fc5e4864435ab4bd44 (demo/04-offline-sovereignty.md)

The first line is on stderr and names the Python that embedded the query. Then comes the stored text of the hit, and under it the citation and the source it came from. The text is exactly what the file stores, the same bytes urna cite returns.

Tune k and candidates

-k sets how many hits come back. The default is 10; hits print highest score first.

urna ask corpus.urna "how do citations work" -k 3

--candidates sets how wide the approximate stage looks before the exact rerank. It matters only on files whose manifest declares index_type = "hnsw" (or hybrid); on an exact file every chunk is scored and the flag is ignored. The default is 4*k, at least 64. Raise it when you suspect the HNSW shortlist misses a chunk that exact search would find:

urna ask corpus.urna "how do citations work" -k 5 --candidates 800

The scores do not change with --candidates: every hit is rescored with exact cosine. Only which chunks make the shortlist can change.

Small values have no effect

The HNSW beam is never smaller than the ef_construction the file was built with, which is 400 by default in urna.build and urna build. On those files any --candidates below 400 runs at 400. And a file built with the hybrid preset declares index_type = "hnsw", so its BM25 section is not used by ask. See Known limits.

See how an answer was found

--disclose explain adds four lines before the answer:

urna ask examples/quickstart/out/quickstart.urna "can I use this offline" -k 1 --disclose explain
route:         hnsw
candidates:    exact=0 ann=12 bm25=0 graph=0 fusion=none
rerank_source: real cosine
recall:        (not computed; rerank guarantees real cosine)
  • route is the path that ran, chosen by the manifest index_type.
  • candidates counts what each stage proposed. The quickstart corpus has 12 chunks, so the HNSW shortlist held all of them.
  • rerank_source says what precision the final score was computed at: real cosine for float32 vectors, real cosine at stored precision for a file stored as float16, int8 or int4.
  • recall is 1 on the exact path. On the other paths it is not estimated, because the rerank makes each returned score exact even when the shortlist is approximate.

Search paths and the exact rerank explains each route.

Retrieve for scripts

urna retrieve runs the same search and prints JSON on stdout. The interpreter line goes to stderr, so a pipe sees only JSON. The default format is JSONL, one hit per line:

urna retrieve examples/quickstart/out/quickstart.urna "how do citations work" -k 3 2>/dev/null \
  | jq -r '[.score, .citation_id] | @tsv'
0.5004974007606506	urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
0.27243560552597046	urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:2be5a0f62d1bb7a69556b1a59d15e3a7d7cf9991baa579e83c1ae67cdd657748
0.25516077876091003	urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:19f36b3e072d553eb83626bf30db5e8f3b1f729a495f7826999ffeae53848e6e

More filters:

# the text of the best hit
urna retrieve corpus.urna "how do citations work" -k 1 2>/dev/null | jq -r .text

# hits above a score threshold
urna retrieve corpus.urna "how do citations work" -k 10 2>/dev/null | jq -c 'select(.score >= 0.3)'

# one array instead of lines
urna retrieve corpus.urna "how do citations work" -k 3 --format json 2>/dev/null | jq 'map({score, source_uri})'

A score is a cosine similarity between the query and the chunk under the corpus's model. It ranks hits within one corpus; a threshold that works on one corpus and model does not carry over to another. Every field is described in urna retrieve.

Each call is a new process that opens and verifies the whole file and starts Python. For many queries in a row, a long-running process that keeps the file open avoids that cost: see Use urna from Python and Serve a corpus over HTTP.

Verify with cite

A citation is urna://<content_hash>/<chunk_id>. urna cite checks that the file's content_hash matches the citation, finds the chunk and prints its stored text and span:

urna cite examples/quickstart/out/quickstart.urna urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
citation_id:  urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
file:         examples/quickstart/out/quickstart.urna
file_hash:    sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
content_hash: sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df
chunk_id:     sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
source_uri:   demo/03-citations.md
byte_start:   7
byte_end:     8
text:
because the citation points at content, two people who build the same logical corpus on two machines get the same citation, and a stored corpus and a compressed one cite identically. resolving a citation returns the exact canonical text and the original byte span it came from, which is what lets an agent quote a source it can prove.

A citation from a different build of the corpus fails with content_hash mismatch: citation says ... but file is ..., and a chunk the file does not hold fails with chunk_id ... not found in file. Both exit 1. In a urna build corpus the span is the row ordinal (7 to 8 above), not a byte range. See Citations and hashes.

When the model gate fails

ask and retrieve refuse to search when the query embedder is missing or does not match the corpus. Every error exits 1 with a message on stderr:

Message starts withCauseWhat to do
embedder script not found: python/forge/embed_query_model.pyThe corpus was built with a registry model (wemm, jina, clip, siglip2). The installed payload carries only the potion embedder.Run from a checkout of the repository with the model's dependencies. See Open a corpus you downloaded.
embedder script not found: <path>The path given to --embedder does not exist.Fix the path.
embedder failed (status=...)The script exited non-zero. Its stderr follows: a missing table (exit 3), missing dependencies, or a registry preset that needs URNA_ALLOW_REMOTE_CODE or URNA_ALLOW_HEAVY=1 (exit 4). A ModuleNotFoundError means the interpreter lacks numpy or tokenizers.Run urna doctor, then urna setup, or set URNA_PYTHON to an interpreter with the dependencies.
model name mismatchThe embedder reports a different model than the manifest names.Drop --embedder, or point it at the embedder for this corpus's model.
dim mismatchThe embedder produced vectors of another size.Same as above. From search-text on a corpus built with mrl_dim, use ask or retrieve instead: they pass --mrl-dim.
manifest carries the legacy placeholder model_hashThe corpus was built without a real fingerprint.Rebuild it with a real model_hash. ask and retrieve cannot skip this check.
model_hash mismatchSame model name and dim, different model bytes or settings: another potion table, another snapshot, or a different dtype policy for a sentence-transformers preset.Point --model-path at the exact snapshot the corpus was built with, set URNA_ST_DTYPE to the dtype used at build time, or rebuild the corpus with the model you have.

The model gate explains what each check protects. For a corpus built with a sentence-transformers model in a checkout, urna search-text is the verb, and the only one with --skip-model-hash-check.

To hand these results to an agent instead of a shell, continue with Give a corpus to an agent.

On this page