Model registry
Every embedding model preset in the urna forge registry: model id, dims, MRL ladder, dependencies, device rules, model_hash and query contract.
The model registry is the list of named embedding models that the forge can build with and that urna ask can query with. It lives in python/forge/model_registry.py in the repository checkout: it is not in the urna wheel or in any release archive. A build spec picks models from it by name in [[models]] preset (see [[models]]).
Presets
| Preset | Kind | Manifest embedding_model | Model id | Default dim | MRL ladder | Modalities |
|---|---|---|---|---|---|---|
potion | potion | minishlab/potion-base-8M/v1 | minishlab/potion-base-8M | 256 | none | text |
clip-vit-b32 | open_clip | open_clip/ViT-B-32/openai | ViT-B-32, pretrained openai | 512 | none | text, image |
siglip2 | open_clip | open_clip/ViT-B-16-SigLIP2/webli | ViT-B-16-SigLIP2, pretrained webli | 768 | none | text, image |
jina-v5-omni-nano | st_multimodal | jinaai/jina-embeddings-v5-omni-nano | same | probed at load | 32, 64, 128, 256, 512, 768 | text, image, video |
jina-v5-omni-small | st_multimodal | jinaai/jina-embeddings-v5-omni-small | same | probed at load | 32, 64, 128, 256, 512, 768, 1024 | text, image, video |
wemm-2b | st_multimodal | tencent/WeMM-Embedding-2B | same | 2048 | 128, 256, 512, 1024, 2048 | text, image, video |
wemm-4b | st_multimodal | tencent/WeMM-Embedding-4B | same | 2560 | 128, 256, 512, 1024, 2560 | text, image, video |
wemm-9b | st_multimodal | tencent/WeMM-Embedding-9B | same | 4096 | 128, 256, 512, 1024, 4096 | text, image, video |
fake-test | fake | urna-forge-fake-test/v1 | none | 8 | 8, 4 | text, image |
The manifest name is what a built file records as embedding_model. It is unique per preset, so the query side can find the preset from a file's manifest alone.
The video modality is declared and not used by the forge. fake-test is a deterministic, model-free preset for tests: it loads only with URNA_ENABLE_FAKE_PRESET=1, and any other value makes get_preset raise RegistryError.
Flags per preset
| Preset | trust_remote_code | Pinned code files | Heavy (needs --allow-heavy) | local_dir |
|---|---|---|---|---|
potion | no | n/a | no | none (the table is vendored under python/forge/models/potion-base-8M/) |
clip-vit-b32, siglip2 | no | n/a | no | none |
jina-v5-omni-nano, jina-v5-omni-small | yes | none | no | none |
wemm-2b | yes | 5 files | no | ~/models/modelsdownload/WeMM-Embedding-2B |
wemm-4b, wemm-9b | yes | none | yes | none |
fake-test | no | n/a | no | none |
MRL ladders
dims in a [[models]] block emits one named space per dim, and each dim must be on the preset's ladder. The ladder is a fixed list of the prefix lengths the model is trained for, not a range: jina-v5-omni-small accepts 256, and refuses 300. Every ladder uses the method prefix_slice_l2: keep the first dim components, then L2-normalize again.
Presets without a ladder (potion, clip-vit-b32, siglip2) refuse dims:
spec error: models.potion.dims: preset is not MRL-trained; dims not allowedA dim off the ladder:
spec error: models.wemm-2b.dims: 300 not in the validated ladder [128, 256, 512, 1024, 2048] (method prefix_slice_l2)[build] mrl_dim skips the ladder
The ladder check applies to [[models]] dims only. [build] mrl_dim, which truncates the default space of the file, accepts any value from 1 to the model's dim, so potion at mrl_dim = 128 builds without an error. See known limits.
Dependencies
Each preset lists the Python packages it imports. The check runs before a model loads and names the exact install line for the first missing package.
| Kind | Packages checked | Fix lines printed |
|---|---|---|
| potion | numpy, tokenizers | pip install numpy, pip install tokenizers |
| open_clip | torch, open_clip, PIL | pip install torch, pip install open_clip_torch, pip install pillow |
| st_multimodal (jina) | torch, sentence_transformers, transformers | pip install torch, pip install "sentence-transformers>=5.7", pip install "transformers==5.2.0" |
| st_multimodal (wemm) | the jina set plus qwen_vl_utils | the jina lines plus pip install "qwen-vl-utils==0.0.14" |
| fake | none | none |
The message has this shape:
registry error: preset 'clip-vit-b32' needs the 'torch' package. install with: pip install torchThe check tests that a module can be imported, not its version. When the sentence-transformers stack fails to import inside the loader, its own error message also lists "accelerate>=1.1.0".
urna build --dry-run runs the same check without loading anything. A missing package shows as deps=MISSING with the fix line under it, and the dry run still exits 0. The real build stops with the registry error and exit code 4.
Device and dtype
| Kind | Device | Dtype |
|---|---|---|
| st_multimodal | [[models]] device, else URNA_ST_DEVICE, else cuda, else mps, else cpu | [[models]] dtype, else URNA_ST_DTYPE, else bfloat16 on cuda, float16 on mps, float32 on cpu |
| open_clip | [[models]] device, else cuda, else mps, else cpu (URNA_ST_DEVICE is not read) | the weights' loaded dtype |
| potion | CPU, numpy | float32 |
Every st_multimodal model runs in its own worker process, because two models that execute repository code in one process collide in the transformers dynamic module cache.
Heavy presets
wemm-4b and wemm-9b are registered as too heavy for a typical machine. A spec that names either one fails validation unless you pass --allow-heavy:
spec error: models.wemm-9b: flagged too heavy for this machine; pass --allow-heavyAt query time the registry query embedder loads them only when URNA_ALLOW_HEAVY=1 is set.
Remote code
The jina and wemm presets set trust_remote_code: loading them runs Python files shipped in the model repository. A spec must name each such preset in [output] allow_remote_code:
[output]
allow_remote_code = ["wemm-2b"]Without it, validation fails:
spec error: models.wemm-2b: executes model-repo code; opt in with output.allow_remote_code = ["wemm-2b"] (RFC-0 N11)The query side has the same rule. urna ask and urna retrieve on a corpus built with one of these presets refuse to load the model until you list the preset in URNA_ALLOW_REMOTE_CODE, a comma-separated list such as URNA_ALLOW_REMOTE_CODE="wemm-2b,jina-v5-omni-nano". A file's manifest is never enough to run code.
Pinned files
wemm-2b is the only preset with pinned code files. When its model directory resolves on disk, each file must match its SHA-256 before the model loads, in the parent process and again in the worker:
| File | SHA-256 |
|---|---|
modeling_st_wemm.py | 521d02c1c60ae727cc9dc6500cdb0b28c53b259e0ce3d37197920a33ba4dd333 |
modeling_wemm_embedding.py | ac255e1fad459cc3e68891d6c3327f4486922aed02fb3c5c13fb53277ba8e94f |
chat_template.jinja | 273d8e0e683b885071fb17e08d71e5f2a5ddfb5309756181681de4f5a1822d80 |
embedding_chat_template.jinja | 7c3df2aab83ab9096428ec27b6b99ad87c4790418b830d119957634f28c677ba |
processor_config.json | d601e2fe0de1bc11852de3aa843f01a1677ca84f4dc743916c9ce4b8d30fb384 |
A changed file is refused, not warned about:
registry error: preset 'wemm-2b': <file> sha256 <actual> does not match the pinned allowlist (<expected>). refusing to execute unreviewed remote code.The jina presets, wemm-4b and wemm-9b have no pins, so the opt-in alone loads them. Build with those in an environment you are willing to run their repository code in.
Model directory resolution
For each preset the registry resolves a model directory in this order:
[[models]] model_pathin the spec (or--model-pathonurna askandurna retrieve).URNA_MODEL_DIR_<NAME>, where<NAME>is the preset name upper-cased with-replaced by_:URNA_MODEL_DIR_WEMM_2B,URNA_MODEL_DIR_JINA_V5_OMNI_NANO.- The preset's
local_dir, when that directory exists. - For st_multimodal presets only: the Hugging Face cache snapshot that
refs/mainpoints to, under$HF_HOME/hub(default~/.cache/huggingface/hub). With several snapshots and no usablerefs/main, nothing resolves. - Nothing.
potion and the open_clip presets do not load from this directory: potion uses its vendored table and open_clip loads by model id and pretrained tag. urna build --dry-run prints the resolved model_dir, or null.
Importing the forge sets HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1 and HF_DATASETS_OFFLINE=1 when they are unset, so weights must already be on disk. For st_multimodal presets the build computes model_hash from the local files before any model loads: with no local directory it fails with a Python TypeError (exit 1) instead of downloading. Put the snapshot in the Hugging Face cache, or point model_path or URNA_MODEL_DIR_<NAME> at it.
Identity: model_hash
model_hash identifies the model, not the corpus. It is written into the file and checked by the model gate on every query.
| Kind | Preimage | Needs the model loaded |
|---|---|---|
| potion | SHA-256 of canonical JSON over the embedder name, version, the hash of config.json, tokenizer.json and model.safetensors, the tokenizer hash, dim, tokenizer name, add_special_tokens, mean pooling, normalize, and float32-stable | no |
| open_clip | SHA-256 over model id, pretrained tag, the preprocess transform, a digest of every weight tensor, and l2 | yes |
| st_multimodal | SHA-256 of canonical JSON over the model file fingerprint (weights, tokenizer, processor), the SHA-256 of each modeling_*.py, *chat_template*.jinja and processor_config.json, the pooling, normalize, and the dtype policy | no, computed from the files |
| fake | SHA-256 of urna-forge-fake-test/v1 | no |
With the vendored table, potion's hash is sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98.
The st_multimodal dtype policy comes from the device, so the same weights hash differently on cuda (bfloat16), mps (float16) and cpu (float32). A corpus built on one device class and queried on another fails the gate unless the dtype is pinned. Set URNA_ST_DTYPE to the same value for the build and for every query. A [[models]] dtype in the spec pins the build side only: the query side never reads the spec.
Query and document contract
Asymmetric models embed queries and documents through different routes. The preset sets the default, and the spec can override the document side.
| Kind | Corpus text | Query text | Images |
|---|---|---|---|
| potion | one route, role ignored | same | refused, text only |
| open_clip | text tower, role ignored | same | image tower, on source files or decoded frames, L2-normalized |
| st_multimodal | text_corpus_mode, default document, calls encode_document | text_query_mode, default query, calls encode_query | encode_document. wemm receives {"image": img, "text": image_prompt} and thumbnails to 768 px; jina receives the bare image |
jina also passes encode_kwargs = {task = "retrieval"}. A mode of plain, or any unknown value, calls plain encode(). The default image_prompt is "Represent this image.".
The jina presets take the bare image because the dict form collapses their image vectors onto the shared prompt text: the registry comment records pairwise similarity of 0.99 across different images with the dict, against 0.45 bare.
The query side always uses the preset defaults. Spec overrides of text_query_mode, encode_kwargs, image_prompt, normalize, dtype and device do not reach urna ask or urna retrieve. normalize and the dtype policy are part of the st_multimodal model_hash, so a corpus built with normalize = false, or with a dtype that differs from the one the query machine resolves, fails the query gate.
The query embedder
urna ask and urna retrieve pick a query embedder from the file's embedding_model:
- A name that starts with
minishlab/potionrunspython/forge/embed_query_potion.py. This script ships in the installer payload, so the installed binary answers potion corpora. - Any other name runs
python/forge/embed_query_model.py, which finds the preset by reverse lookup of the manifest name, applies the remote-code and heavy gates, resolves the model directory as above and embeds the query withrole="query". This script is not in any release artifact: it runs from a repository checkout, with the preset's packages installed in the interpreterurnauses.
When the file's default space was truncated (the manifest records full_dim), the CLI passes --mrl-dim so the query is sliced to the same length.
Exit code of embed_query_model.py | Meaning |
|---|---|
| 0 | the query vector was printed |
| 2 | usage error, or no query |
| 3 | a model asset is missing |
| 4 | a dependency, an unknown preset, the remote-code gate or the heavy gate |
The CLI does not branch on these codes: any non-zero exit becomes embedder failed (status=...): <stderr>. The protocol is documented in query embedder protocol.
Errors
Registry errors reach urna build as registry error: <message> on stderr with exit code 4:
| Message | Cause |
|---|---|
unknown model preset '<p>'. valid presets: ... | a name not in the table |
preset 'fake-test' requires URNA_ENABLE_FAKE_PRESET=1 (test-only) | fake-test without the variable |
preset '<p>' needs the '<module>' package. install with: <line> | a missing dependency |
preset '<p>': pinned code file missing: <file> | a wemm-2b pin with the file absent |
preset '<p>': <file> sha256 ... does not match the pinned allowlist ... | a wemm-2b pin mismatch |
The heavy and remote-code refusals come from spec validation first, as spec error: with exit code 2. For choosing a model and keeping build and query in step, see Choose and bring embedding models.
Build artifacts
Reference for what urna build writes: the .urna files, manifest.json, build.lock.json, the shared embed cache, reproduction levels and every build error.
Build presets
The storage presets of urna.build (exact, compressed, tiny, nano, hybrid) and the micro recipe: dtype, text encoding, indexes and measured size.