Query embedder protocol
How the urna CLI finds, runs and checks the Python scripts that embed a text query: arguments, the JSON on stdout, exit status and lookups.
The Rust binary links no model runtime. When a verb needs a text query turned into a vector, it runs a Python script as a child process, reads one JSON document from its stdout and checks it against the corpus manifest before searching. This page is the contract between the two sides: how the script is chosen and invoked, what it must print, and what the CLI does with it.
Who calls an embedder
| Caller | Script | Arguments after the script path |
|---|---|---|
urna ask, urna retrieve, the ask tab of urna tui | routed by the manifest model, or --embedder | [--model-path P] [--mrl-dim D] <embedding_model> <query> |
urna search-text | python/embed_query.py, or --embedder | [--model-path P] <embedding_model> <query> |
urna doctor (check 7) | the resolved embed_query_potion.py | potion-base-8M "urna doctor probe" |
<embedding_model> is the manifest embedding_model string, verbatim. <query> is the query text as one argument. The CLI spawns the process directly, with no shell, so the query needs no escaping beyond what your own shell requires. The child inherits the working directory and the environment of urna.
--mrl-dim <embedding_dim> is added by ask and retrieve when the manifest records full_dim, which means the corpus was built with mrl_dim. search-text never adds it.
Routing
ask and retrieve pick the script from the manifest embedding_model:
Manifest embedding_model | Script |
|---|---|
starts with minishlab/potion | forge/embed_query_potion.py |
| anything else | forge/embed_query_model.py |
--embedder <path> replaces the routed script, and the path is used as given. A path that does not exist fails before any process starts: embedder script not found: <path> (override with --embedder).
Script lookup
The first existing path wins.
For forge/embed_query_potion.py and forge/embed_query_model.py:
python/forge/<name>under the current directory.python/forge/<name>under the parent of the current directory.python/forge/<name>under the checkout of a dev-built binary: for an executable at<repo>/target/<profile>/urna, that is<repo>, used when it containspython/.<root>/urna/forge/<name>for each data root, in this order:$URNA_DATA_DIR,$XDG_DATA_HOME,$HOME/.local/share,%LOCALAPPDATA%,<directory of the executable>/../share. Empty variables are skipped.- Otherwise the relative path
python/forge/<name>, which then fails as not found.
For search-text's default python/embed_query.py, only steps 1 to 3 apply; no data root is searched.
The installed embedder payload lays down urna/forge/ with __init__.py, embed_default.py, embed_potion.py, embed_query_potion.py and the potion table. It has neither embed_query_model.py nor embed_query.py, so outside a checkout only potion corpora can be queried by text.
Registry-model corpora need a checkout
An installed binary asked about a corpus built with a registry model (wemm, jina, clip, siglip2) stops with embedder script not found: python/forge/embed_query_model.py (override with --embedder). urna setup cannot fix this, because the release payload does not ship the script. Run from a checkout of the repository, or pass --embedder with the script's path in a checkout. See Known limits.
Interpreter lookup
The script runs as <interpreter> <script> .... The interpreter is the first match of:
$URNA_PYTHON, used verbatim when non-empty.- The venv
urna setupcreates: the first data root whoseurna/venv/bin/python(on Windowsurna\venv\Scripts\python.exe) is a file. .venv/bin/pythonin the current directory or one of its three nearest ancestors. This step uses the Unix layout on every platform, so a Windows.venv\Scripts\python.exeis never found.python3fromPATH.
The choice is printed on stderr as [urna] embedder interpreter: <path>, except inside urna tui. Step 3 executes whatever .venv/bin/python sits closest to the current directory. In a directory tree you do not control, set URNA_PYTHON to pin the interpreter.
Output contract
On success the script exits 0 and prints exactly one JSON document on stdout. Anything else on stdout breaks the parse, so logs belong on stderr.
| Field | Type | Required | Rule |
|---|---|---|---|
model_hash | string | yes | Must equal the manifest model_hash. The gate's source of truth. |
embedding_model | string | yes | Must equal the manifest embedding_model the script received. |
embedding_dim | integer | yes | Must equal the manifest embedding_dim. |
vector | array of numbers | yes | Its length must equal the manifest embedding_dim. |
fingerprint | any JSON value | no | Diagnostic only. Printed inside a model_hash mismatch error. |
Unknown fields are ignored. The runtime L2-normalizes the query before scoring, so the vector's scale does not change the scores; the shipped scripts normalize anyway. The vector must be finite and non-zero, or the search fails with NaN or Inf in query or zero-norm query.
What the CLI checks
The model gate runs on the parsed output, in this order, and each failure exits 1:
| Check | Error |
|---|---|
embedding_model equals the manifest name | model name mismatch: manifest=<m>, embedder reports=<r> |
embedding_dim and the vector length equal the manifest dim | dim mismatch: manifest=<d>, embedder dim=<e>, vector len=<n> |
the manifest model_hash is not the placeholder sha256: followed by 64 zeros | manifest carries the legacy placeholder model_hash (...) |
model_hash equals the manifest model_hash | model_hash mismatch: corpus was built with <a>, embedder reports <b>, followed by the fingerprint and a hint |
search-text --skip-model-hash-check skips the last two rows. ask and retrieve always run all four.
The gate trusts the script's report. It catches the accidents it exists for (a different model, a different snapshot, a different table), not a script written to lie.
urna doctor applies a lighter check to its probe: embedding_dim greater than 0, equal to the vector length, and a model_hash that starts with sha256:.
Exit status and errors
Any non-zero exit becomes embedder failed (status=<status>): <stderr of the script>, exit 1. The CLI does not branch on the script's own codes; they only help a person reading stderr. Output that is not valid JSON becomes invalid embedder output: <parse error> (stdout=<...>).
The shipped scripts
| Script | Used by | Arguments | Reports as embedding_model | Exit codes | Shipped in |
|---|---|---|---|---|---|
python/forge/embed_query_potion.py | ask, retrieve, doctor for potion corpora | [--model-path DIR] [--mrl-dim N] <model> [<query>] | its own name, minishlab/potion-base-8M/v1; the <model> argument is ignored for inference | 0; 2 query missing; 3 table missing or a git-lfs pointer; a traceback with 1 when numpy or tokenizers is missing | the checkout and the installed payload |
python/forge/embed_query_model.py | ask, retrieve for registry-model corpora | [--model-path P] [--preset NAME] [--mrl-dim N] <model> [<query>] | the preset's manifest name | 0; 2 usage; 3 model asset missing; 4 missing dependencies, unknown model, remote-code or heavy gate | the checkout only |
python/embed_query.py | search-text | [--embed-dim] [--model-path P] <model> [<query>] | the <model> argument as given | 0; 2 query missing; an uncaught exception (a model not in the cache while offline, for example) is a traceback with 1. --embed-dim prints the bare dim and exits | the checkout only |
Details that matter when you run them yourself:
embed_query_potion.pyneeds numpy and tokenizers, never opens a socket, and--model-pathpoints at a potion table directory. With--mrl-dim Nit slices the vector to N, renormalizes and reports N.embed_query_model.pyfinds the preset from--preset, else by matching<model>against the registry's manifest names. It embeds with the preset's query route (role="query"), so asymmetric models encode queries and documents differently. It readsURNA_ALLOW_REMOTE_CODE(a comma list of presets allowed to run model-repo code; jina and wemm need it),URNA_ALLOW_HEAVY=1(wemm-4b, wemm-9b),URNA_MODEL_DIR_<PRESET>andURNA_ALLOW_DOWNLOAD=1. Spec-level overrides from the build (query mode, dtype, device) are not read on the query side. See Model registry.embed_query.pyforces the Hugging Face offline variables unlessURNA_ALLOW_DOWNLOAD=1, and computesmodel_hashfrom the snapshot it loaded withpython/model_fingerprint.py.
Write your own embedder
A script that follows the contract can serve a corpus through --embedder. It must reproduce the corpus's model exactly, because the gate compares the model_hash your script reports with the one the builder recorded. A minimal shape:
import json
import sys
from my_model import load, model_hash # your model and the hash you built the corpus with
def main() -> int:
args = sys.argv[1:]
mrl_dim = 0
if "--mrl-dim" in args:
i = args.index("--mrl-dim")
mrl_dim = int(args[i + 1])
del args[i : i + 2]
if "--model-path" in args:
i = args.index("--model-path")
del args[i : i + 2]
model_name, query = args[-2], args[-1]
model = load()
vector = model.embed(query) # list of floats, L2-normalized
if mrl_dim:
vector = vector[:mrl_dim]
norm = sum(x * x for x in vector) ** 0.5
vector = [x / norm for x in vector]
json.dump(
{
"model_hash": model_hash(),
"embedding_model": model_name,
"embedding_dim": len(vector),
"vector": vector,
},
sys.stdout,
)
return 0
if __name__ == "__main__":
sys.exit(main())urna ask corpus.urna "your question" --embedder ./my_embedder.pyEchoing model_name back passes the name check only because your script claims that model; the model_hash check is the one that proves the match.