urna search-text
Search a .urna file by raw text through the sentence-transformers query embedder, with the model gate, and print every hit field for inspection.
urna search-text embeds a text query with a Python script, checks the result against the corpus manifest, runs the path the manifest declares and prints every field of every hit. Its default embedder is the sentence-transformers script python/embed_query.py from a checkout of the repository. It is listed as an engine verb, but it runs Python on every call.
For everyday questions use urna ask or urna retrieve: they route to the right embedder by themselves. search-text is for corpora built with a sentence-transformers model from a checkout, for debugging a query embedder, and for the legacy placeholder corpora that need --skip-model-hash-check.
Usage
urna search-text [OPTIONS] <FILE> <QUERY>Arguments
| Argument | Description |
|---|---|
<FILE> | Path to the .urna file. |
<QUERY> | The query text, as one argument. |
Options
| Option | Default | Description |
|---|---|---|
-k, --k <K> | 10 | Number of hits. |
--embedder <EMBEDDER> | python/embed_query.py | Path to the query-embedder script. The default is looked up only in a repository layout (see Behavior). |
--candidates <CANDIDATES> | 4*k, at least 64 | The HNSW beam on an hnsw file, the candidates per path on a hybrid file. Ignored on an exact file. |
--model-path <MODEL_PATH> | none | Local path to the model snapshot directory, passed to the embedder as --model-path. Without it, the default embedder loads the model named in the manifest from the Hugging Face cache. |
--skip-model-hash-check | off | Skip the model_hash layer of the gate. Meant for legacy corpora whose model_hash is the all-zero placeholder. |
-h, --help | Print help. |
Behavior
- Opens and fully validates the file, then reads
embedding_model,embedding_dimandmodel_hashfrom the manifest. - Resolves the embedder:
--embedderwhen given, elsepython/embed_query.pyunder the current directory, its parent, or the checkout of a dev-built binary. The installed embedder payload does not carry this script, so outside a checkout the default fails withembedder script not found: python/embed_query.py (override with --embedder). - Prints
[urna] embedding query with <model> via <script>on stderr, resolves the interpreter (printed as[urna] embedder interpreter: <path>) and runs<interpreter> <script> [--model-path P] <embedding_model> <query>. The query embedder protocol has both lookups and the JSON contract. - Applies the model gate: the reported model name equals the manifest name; the reported dim and the vector length equal the manifest dim; unless
--skip-model-hash-checkis set, the placeholdermodel_hashis refused and the reportedmodel_hashmust equal the manifest's. - Routes by the manifest
index_type:hnswruns the HNSW path with--candidatesas the beam,hybridruns the vector and BM25 legs with the query text, anything else runs the exact scan. Every score comes from the exact cosine rerank.
The default embedder and the network
python/embed_query.py loads the model with sentence-transformers and fingerprints the snapshot it loaded (python/model_fingerprint.py). Before it imports anything from Hugging Face it sets HF_HUB_OFFLINE, TRANSFORMERS_OFFLINE and HF_DATASETS_OFFLINE to 1, unless URNA_ALLOW_DOWNLOAD=1. So the model must already be on disk: in the Hugging Face cache, or at the directory --model-path names. With URNA_ALLOW_DOWNLOAD=1 it may download the model on first use. The interpreter needs sentence-transformers and its dependencies. See Offline by construction.
The reported embedding_model is the manifest name the CLI passed in, also when --model-path points elsewhere. The model_hash covers the snapshot's config, tokenizer, pooling and weight files, so --model-path must point at the same snapshot the corpus was built with.
What --skip-model-hash-check skips
The flag turns off the whole model_hash layer: the placeholder refusal and the equality check. The model name and dim checks still run. With the flag, a query embedded by a different model of the same name and dim returns cosine scores that are valid arithmetic and meaningless ranking. Use it only on a corpus that carries the placeholder and only when you know the query model is the one the corpus was built with. Rebuilding the corpus with a real fingerprint is the fix.
Other embedders
--embedder accepts any script that speaks the protocol. The potion script works on a potion corpus, and gives the same hits as ask:
urna search-text corpus.urna "how do citations work" -k 2 \
--embedder ~/.local/share/urna/forge/embed_query_potion.pysearch-text never passes --mrl-dim, so on a corpus built with mrl_dim (the manifest records full_dim) the embedder reports the full dimension and the dim check fails. Use ask or retrieve there.
Preset-built files take the HNSW route
Files built with the hybrid preset declare index_type = "hnsw", so search-text never runs their BM25 section, and a --candidates value below the file's ef_construction (400 by default) does not change the beam. See Known limits.
Output
stdout carries a summary block and one line per hit. stderr carries the two [urna] lines and, on failure, Error: <message>.
| Line | Meaning |
|---|---|
index_type: | The path that produced the hits: exact, hnsw or hybrid. |
recall: | 1 on the exact path; (not computed; rerank guarantees real cosine) on the others. |
truncated: | true when k is below the number of chunks in the file. |
k_requested:, k_returned: | k as passed and the number of hits returned. |
query_time: | Search time in milliseconds, excluding the embedder run. |
hits: | One line per hit: rank, chunk_id, score (6 decimals), score_type, source_uri, offset, model, index_type, reranked, file_hash, content_hash, citation_id. |
Real output with the potion script on the quickstart corpus:
[urna] embedding query with minishlab/potion-base-8M/v1 via /Users/nn/.local/share/urna/forge/embed_query_potion.py
[urna] embedder interpreter: /Users/nn/.local/share/urna/venv/bin/python
index_type: hnsw
recall: (not computed; rerank guarantees real cosine)
truncated: true
k_requested: 2
k_returned: 2
query_time: 0.037 ms
hits:
[ 1] chunk_id=sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be score=0.500497 score_type=cosine source_uri=demo/03-citations.md offset=7-8 model=minishlab/potion-base-8M/v1 index_type=hnsw reranked=true file_hash=sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832 content_hash=sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df citation_id=urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
[ 2] chunk_id=sha256:2be5a0f62d1bb7a69556b1a59d15e3a7d7cf9991baa579e83c1ae67cdd657748 score=0.272436 score_type=cosine source_uri=demo/03-citations.md offset=8-9 model=minishlab/potion-base-8M/v1 index_type=hnsw reranked=true file_hash=sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832 content_hash=sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df citation_id=urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:2be5a0f62d1bb7a69556b1a59d15e3a7d7cf9991baa579e83c1ae67cdd657748Exit codes
| Code | Meaning |
|---|---|
0 | The search ran. |
1 | Any error: the file failed to open or validate, the embedder was not found, exited non-zero or printed invalid JSON, the model gate refused, or the search rejected the query. |
2 | Usage error: a missing argument or an unknown flag. |
Examples
From the root of a checkout, on a corpus built with a sentence-transformers model that is in the local cache:
urna search-text my_corpus.urna "vector search on the edge" -k 5Fully offline, with the model snapshot copied next to the corpus:
urna search-text my_corpus.urna "vector search on the edge" --model-path ./models/all-MiniLM-L6-v2A legacy corpus with the placeholder model_hash:
urna search-text legacy.urna "vector search on the edge" --skip-model-hash-checkurna search
Reference for urna search, the exact search verb that takes a query vector as a JSON array and prints every hit with its cosine score and citation.
urna search-ann
Reference for urna search-ann, which forces the HNSW path with a JSON query vector, reranks the candidates by exact cosine, and falls back to exact.