docsv0.5.1

Query embedder protocol

How the urna CLI finds, runs and checks the Python scripts that embed a text query: arguments, the JSON on stdout, exit status and lookups.

The Rust binary links no model runtime. When a verb needs a text query turned into a vector, it runs a Python script as a child process, reads one JSON document from its stdout and checks it against the corpus manifest before searching. This page is the contract between the two sides: how the script is chosen and invoked, what it must print, and what the CLI does with it.

Who calls an embedder

CallerScriptArguments after the script path
urna ask, urna retrieve, the ask tab of urna tuirouted by the manifest model, or --embedder[--model-path P] [--mrl-dim D] <embedding_model> <query>
urna search-textpython/embed_query.py, or --embedder[--model-path P] <embedding_model> <query>
urna doctor (check 7)the resolved embed_query_potion.pypotion-base-8M "urna doctor probe"

<embedding_model> is the manifest embedding_model string, verbatim. <query> is the query text as one argument. The CLI spawns the process directly, with no shell, so the query needs no escaping beyond what your own shell requires. The child inherits the working directory and the environment of urna.

--mrl-dim <embedding_dim> is added by ask and retrieve when the manifest records full_dim, which means the corpus was built with mrl_dim. search-text never adds it.

Routing

ask and retrieve pick the script from the manifest embedding_model:

Manifest embedding_modelScript
starts with minishlab/potionforge/embed_query_potion.py
anything elseforge/embed_query_model.py

--embedder <path> replaces the routed script, and the path is used as given. A path that does not exist fails before any process starts: embedder script not found: <path> (override with --embedder).

Script lookup

The first existing path wins.

For forge/embed_query_potion.py and forge/embed_query_model.py:

  1. python/forge/<name> under the current directory.
  2. python/forge/<name> under the parent of the current directory.
  3. python/forge/<name> under the checkout of a dev-built binary: for an executable at <repo>/target/<profile>/urna, that is <repo>, used when it contains python/.
  4. <root>/urna/forge/<name> for each data root, in this order: $URNA_DATA_DIR, $XDG_DATA_HOME, $HOME/.local/share, %LOCALAPPDATA%, <directory of the executable>/../share. Empty variables are skipped.
  5. Otherwise the relative path python/forge/<name>, which then fails as not found.

For search-text's default python/embed_query.py, only steps 1 to 3 apply; no data root is searched.

The installed embedder payload lays down urna/forge/ with __init__.py, embed_default.py, embed_potion.py, embed_query_potion.py and the potion table. It has neither embed_query_model.py nor embed_query.py, so outside a checkout only potion corpora can be queried by text.

Registry-model corpora need a checkout

An installed binary asked about a corpus built with a registry model (wemm, jina, clip, siglip2) stops with embedder script not found: python/forge/embed_query_model.py (override with --embedder). urna setup cannot fix this, because the release payload does not ship the script. Run from a checkout of the repository, or pass --embedder with the script's path in a checkout. See Known limits.

Interpreter lookup

The script runs as <interpreter> <script> .... The interpreter is the first match of:

  1. $URNA_PYTHON, used verbatim when non-empty.
  2. The venv urna setup creates: the first data root whose urna/venv/bin/python (on Windows urna\venv\Scripts\python.exe) is a file.
  3. .venv/bin/python in the current directory or one of its three nearest ancestors. This step uses the Unix layout on every platform, so a Windows .venv\Scripts\python.exe is never found.
  4. python3 from PATH.

The choice is printed on stderr as [urna] embedder interpreter: <path>, except inside urna tui. Step 3 executes whatever .venv/bin/python sits closest to the current directory. In a directory tree you do not control, set URNA_PYTHON to pin the interpreter.

Output contract

On success the script exits 0 and prints exactly one JSON document on stdout. Anything else on stdout breaks the parse, so logs belong on stderr.

FieldTypeRequiredRule
model_hashstringyesMust equal the manifest model_hash. The gate's source of truth.
embedding_modelstringyesMust equal the manifest embedding_model the script received.
embedding_dimintegeryesMust equal the manifest embedding_dim.
vectorarray of numbersyesIts length must equal the manifest embedding_dim.
fingerprintany JSON valuenoDiagnostic only. Printed inside a model_hash mismatch error.

Unknown fields are ignored. The runtime L2-normalizes the query before scoring, so the vector's scale does not change the scores; the shipped scripts normalize anyway. The vector must be finite and non-zero, or the search fails with NaN or Inf in query or zero-norm query.

What the CLI checks

The model gate runs on the parsed output, in this order, and each failure exits 1:

CheckError
embedding_model equals the manifest namemodel name mismatch: manifest=<m>, embedder reports=<r>
embedding_dim and the vector length equal the manifest dimdim mismatch: manifest=<d>, embedder dim=<e>, vector len=<n>
the manifest model_hash is not the placeholder sha256: followed by 64 zerosmanifest carries the legacy placeholder model_hash (...)
model_hash equals the manifest model_hashmodel_hash mismatch: corpus was built with <a>, embedder reports <b>, followed by the fingerprint and a hint

search-text --skip-model-hash-check skips the last two rows. ask and retrieve always run all four.

The gate trusts the script's report. It catches the accidents it exists for (a different model, a different snapshot, a different table), not a script written to lie.

urna doctor applies a lighter check to its probe: embedding_dim greater than 0, equal to the vector length, and a model_hash that starts with sha256:.

Exit status and errors

Any non-zero exit becomes embedder failed (status=<status>): <stderr of the script>, exit 1. The CLI does not branch on the script's own codes; they only help a person reading stderr. Output that is not valid JSON becomes invalid embedder output: <parse error> (stdout=<...>).

The shipped scripts

ScriptUsed byArgumentsReports as embedding_modelExit codesShipped in
python/forge/embed_query_potion.pyask, retrieve, doctor for potion corpora[--model-path DIR] [--mrl-dim N] <model> [<query>]its own name, minishlab/potion-base-8M/v1; the <model> argument is ignored for inference0; 2 query missing; 3 table missing or a git-lfs pointer; a traceback with 1 when numpy or tokenizers is missingthe checkout and the installed payload
python/forge/embed_query_model.pyask, retrieve for registry-model corpora[--model-path P] [--preset NAME] [--mrl-dim N] <model> [<query>]the preset's manifest name0; 2 usage; 3 model asset missing; 4 missing dependencies, unknown model, remote-code or heavy gatethe checkout only
python/embed_query.pysearch-text[--embed-dim] [--model-path P] <model> [<query>]the <model> argument as given0; 2 query missing; an uncaught exception (a model not in the cache while offline, for example) is a traceback with 1. --embed-dim prints the bare dim and exitsthe checkout only

Details that matter when you run them yourself:

  • embed_query_potion.py needs numpy and tokenizers, never opens a socket, and --model-path points at a potion table directory. With --mrl-dim N it slices the vector to N, renormalizes and reports N.
  • embed_query_model.py finds the preset from --preset, else by matching <model> against the registry's manifest names. It embeds with the preset's query route (role="query"), so asymmetric models encode queries and documents differently. It reads URNA_ALLOW_REMOTE_CODE (a comma list of presets allowed to run model-repo code; jina and wemm need it), URNA_ALLOW_HEAVY=1 (wemm-4b, wemm-9b), URNA_MODEL_DIR_<PRESET> and URNA_ALLOW_DOWNLOAD=1. Spec-level overrides from the build (query mode, dtype, device) are not read on the query side. See Model registry.
  • embed_query.py forces the Hugging Face offline variables unless URNA_ALLOW_DOWNLOAD=1, and computes model_hash from the snapshot it loaded with python/model_fingerprint.py.

Write your own embedder

A script that follows the contract can serve a corpus through --embedder. It must reproduce the corpus's model exactly, because the gate compares the model_hash your script reports with the one the builder recorded. A minimal shape:

import json
import sys

from my_model import load, model_hash  # your model and the hash you built the corpus with


def main() -> int:
    args = sys.argv[1:]
    mrl_dim = 0
    if "--mrl-dim" in args:
        i = args.index("--mrl-dim")
        mrl_dim = int(args[i + 1])
        del args[i : i + 2]
    if "--model-path" in args:
        i = args.index("--model-path")
        del args[i : i + 2]
    model_name, query = args[-2], args[-1]

    model = load()
    vector = model.embed(query)  # list of floats, L2-normalized
    if mrl_dim:
        vector = vector[:mrl_dim]
        norm = sum(x * x for x in vector) ** 0.5
        vector = [x / norm for x in vector]

    json.dump(
        {
            "model_hash": model_hash(),
            "embedding_model": model_name,
            "embedding_dim": len(vector),
            "vector": vector,
        },
        sys.stdout,
    )
    return 0


if __name__ == "__main__":
    sys.exit(main())
urna ask corpus.urna "your question" --embedder ./my_embedder.py

Echoing model_name back passes the name check only because your script claims that model; the model_hash check is the one that proves the match.

On this page