docsv0.5.1

Choose and bring embedding models

Pick an embedding model from the urna registry, keep the build and query sides on the same model, run fully offline with --model-path, and load heavy presets.

A corpus is embedded with one model at build time and must be queried with the same model: urna checks the model's name, dim and model_hash on every query and refuses a mismatch. This guide covers choosing a registry preset, what each machine needs, and keeping model files local.

What you need where

TaskInstalled urna binary (release)Repository checkout with the preset's Python packages
Build any corpus (urna build --spec)noyes
Query a potion corpus (urna ask, urna retrieve)yes, after urna setupyes
Query a corpus built with any other registry modelnoyes
Query a sentence-transformers corpus (urna search-text)noyes

The release artifacts ship the offline potion query embedder and its table, nothing else. The build tool and the registry query embedder are Python files in python/ of the repository.

The installed binary answers potion corpora only

On a corpus built with clip-vit-b32, siglip2, a jina or a wemm preset, the installed binary stops with embedder script not found: python/forge/embed_query_model.py (override with --embedder), because that script is not in the installer payload. Run the query from a repository checkout. See known limits.

Choose a preset

PresetUse it forDimNeeds
potionthe default: text, offline, CPU, no torch256numpy, tokenizers
clip-vit-b32images and short text in one space512torch, open_clip_torch, pillow
siglip2images and short text, a larger open_clip model768same as clip
jina-v5-omni-nano, jina-v5-omni-smalltext and images with an asymmetric query and document route, truncatableprobed at loadtorch, sentence-transformers>=5.7, transformers==5.2.0; runs repository code
wemm-2btext and images, the largest preset that runs by default2048the jina set plus qwen-vl-utils==0.0.14; runs pinned repository code
wemm-4b, wemm-9bthe same family, larger2560, 4096as wemm-2b, plus --allow-heavy

Only potion is answered by the installed binary. Pick another preset when you need image search or a stronger text model, and plan for the checkout on every machine that queries the corpus. Every field of every preset is on Model registry.

Declare the models in the spec

Each [[models]] block names a preset and its role. Exactly one model has text = "default": it embeds the chunk text into the default space, the one urna ask queries. Other models add named spaces.

[[models]]
preset = "wemm-2b"
text = "default"
image = "space"
dims = [256]

[output]
allow_remote_code = ["wemm-2b"]

This embeds the text with wemm-2b into the default space and adds an image space named wemm-2b@256. dims must be on the preset's ladder. See [[models]] for every key.

Check the plan and the dependencies before a real build. The dry run loads no model:

urna build --spec corpus.toml --dry-run
model:  clip-vit-b32         text=none    image=space dims=- deps=MISSING
        -> preset 'clip-vit-b32' needs the 'torch' package. install with: pip install torch

Install what it names into the interpreter urna build uses. The launcher prints its choice on stderr as [urna] embedder interpreter: <path>.

Keep build and query on the same model

The query side has to reproduce the build's model_hash exactly. Three things break that:

  • A different model snapshot. The hash covers the weights, tokenizer, processor and repository code files, so a re-downloaded or updated model is a different model. Keep the snapshot you built with.
  • A different dtype. For jina and wemm the hash includes the dtype policy, which follows the device: bfloat16 on cuda, float16 on mps, float32 on cpu. A corpus built on a cuda machine and queried on a Mac fails the gate. Set the same URNA_ST_DTYPE for the build and for every query, for example URNA_ST_DTYPE=float32.
  • Spec overrides. The query embedder uses the preset defaults and never reads the spec. A build that sets normalize = false, or a dtype or device that changes the dtype policy, produces a hash the query side does not reproduce. text_query_mode, encode_kwargs and image_prompt in the spec do not reach the query side either.

A mismatch fails before any search runs:

model_hash mismatch: corpus was built with sha256:..., embedder reports sha256:...

See The model gate for the three checks.

Query a registry corpus

From the repository root, with the preset's packages installed in a virtual environment:

export URNA_PYTHON="$PWD/.venv/bin/python"
export URNA_ALLOW_REMOTE_CODE="wemm-2b"
urna ask out/corpus/corpus.urna "which cards draw two" -k 5
  • URNA_PYTHON picks the interpreter. Without it, the CLI uses the urna setup environment when one exists, then the nearest .venv in the working directory or up to three parents, then python3. The urna setup environment has only numpy and tokenizers, so on a machine that ran setup, set URNA_PYTHON explicitly.
  • The CLI finds python/forge/embed_query_model.py relative to the working directory, so run it from the checkout.
  • Presets that run repository code (the jina and wemm presets) need their names in URNA_ALLOW_REMOTE_CODE, comma-separated. The file's manifest alone never authorizes code.

Run fully offline

The forge and the query embedders set HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1 and HF_DATASETS_OFFLINE=1 unless you already set them, so model files must be on disk before you build or query. The registry resolves a model directory in this order: model_path in the spec (or --model-path on urna ask and urna retrieve), then URNA_MODEL_DIR_<NAME>, then the preset's local_dir, then the Hugging Face cache ($HF_HOME/hub, default ~/.cache/huggingface/hub).

For jina and wemm, the build computes model_hash from the local files before loading the model. With no local directory it fails with a Python TypeError (exit 1) instead of downloading. Put the snapshot in place first.

To move a corpus to an air-gapped machine, copy the model directory with it and point every query at it:

urna ask corpus.urna "which cards draw two" --model-path /mnt/models/WeMM-Embedding-2B

or set the variable once:

export URNA_MODEL_DIR_WEMM_2B=/mnt/models/WeMM-Embedding-2B

The variable name is URNA_MODEL_DIR_ plus the preset name upper-cased, with - replaced by _. On a potion corpus, --model-path points at the potion table directory instead. The open_clip presets load by name and pretrained tag from their own cache, which must already hold the weights. See Air-gapped install and queries for the rest of the offline setup.

Heavy presets

wemm-4b and wemm-9b are registered but flagged too heavy for a typical machine. The build refuses them unless you pass --allow-heavy:

urna build --spec corpus.toml --allow-heavy

The query embedder loads them only with URNA_ALLOW_HEAVY=1 in the environment.

Bring a model the registry does not have

The registry is a Python table, PRESETS in python/forge/model_registry.py. A new preset there becomes available to urna build and, through its manifest name, to urna ask, once you run both from that checkout.

For a sentence-transformers text model without editing the registry, build with urna.build directly and stamp the file with the model's fingerprint from python/model_fingerprint.py (checkout only):

import sys
sys.path.insert(0, "python")

from model_fingerprint import compute_model_fingerprint, fingerprint_to_model_hash, resolve_model_dir

model_id = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
model_dir = resolve_model_dir(model_id)
model_hash = fingerprint_to_model_hash(compute_model_fingerprint(model_dir, model_id=model_id))

Pass embedding_model=model_id and that model_hash to urna.build. Query the file with urna search-text, which embeds through python/embed_query.py and accepts --model-path:

urna search-text corpus.urna "vacina contra covid" -k 5 \
  --model-path /mnt/models/paraphrase-multilingual-MiniLM-L12-v2

urna ask does not serve these files: it looks the manifest name up in the registry. See Use urna from Python for the build call.

On this page