docsv0.5.1

Model registry

Every embedding model preset in the urna forge registry: model id, dims, MRL ladder, dependencies, device rules, model_hash and query contract.

The model registry is the list of named embedding models that the forge can build with and that urna ask can query with. It lives in python/forge/model_registry.py in the repository checkout: it is not in the urna wheel or in any release archive. A build spec picks models from it by name in [[models]] preset (see [[models]]).

Presets

PresetKindManifest embedding_modelModel idDefault dimMRL ladderModalities
potionpotionminishlab/potion-base-8M/v1minishlab/potion-base-8M256nonetext
clip-vit-b32open_clipopen_clip/ViT-B-32/openaiViT-B-32, pretrained openai512nonetext, image
siglip2open_clipopen_clip/ViT-B-16-SigLIP2/webliViT-B-16-SigLIP2, pretrained webli768nonetext, image
jina-v5-omni-nanost_multimodaljinaai/jina-embeddings-v5-omni-nanosameprobed at load32, 64, 128, 256, 512, 768text, image, video
jina-v5-omni-smallst_multimodaljinaai/jina-embeddings-v5-omni-smallsameprobed at load32, 64, 128, 256, 512, 768, 1024text, image, video
wemm-2bst_multimodaltencent/WeMM-Embedding-2Bsame2048128, 256, 512, 1024, 2048text, image, video
wemm-4bst_multimodaltencent/WeMM-Embedding-4Bsame2560128, 256, 512, 1024, 2560text, image, video
wemm-9bst_multimodaltencent/WeMM-Embedding-9Bsame4096128, 256, 512, 1024, 4096text, image, video
fake-testfakeurna-forge-fake-test/v1none88, 4text, image

The manifest name is what a built file records as embedding_model. It is unique per preset, so the query side can find the preset from a file's manifest alone.

The video modality is declared and not used by the forge. fake-test is a deterministic, model-free preset for tests: it loads only with URNA_ENABLE_FAKE_PRESET=1, and any other value makes get_preset raise RegistryError.

Flags per preset

Presettrust_remote_codePinned code filesHeavy (needs --allow-heavy)local_dir
potionnon/anonone (the table is vendored under python/forge/models/potion-base-8M/)
clip-vit-b32, siglip2non/anonone
jina-v5-omni-nano, jina-v5-omni-smallyesnonenonone
wemm-2byes5 filesno~/models/modelsdownload/WeMM-Embedding-2B
wemm-4b, wemm-9byesnoneyesnone
fake-testnon/anonone

MRL ladders

dims in a [[models]] block emits one named space per dim, and each dim must be on the preset's ladder. The ladder is a fixed list of the prefix lengths the model is trained for, not a range: jina-v5-omni-small accepts 256, and refuses 300. Every ladder uses the method prefix_slice_l2: keep the first dim components, then L2-normalize again.

Presets without a ladder (potion, clip-vit-b32, siglip2) refuse dims:

spec error: models.potion.dims: preset is not MRL-trained; dims not allowed

A dim off the ladder:

spec error: models.wemm-2b.dims: 300 not in the validated ladder [128, 256, 512, 1024, 2048] (method prefix_slice_l2)

[build] mrl_dim skips the ladder

The ladder check applies to [[models]] dims only. [build] mrl_dim, which truncates the default space of the file, accepts any value from 1 to the model's dim, so potion at mrl_dim = 128 builds without an error. See known limits.

Dependencies

Each preset lists the Python packages it imports. The check runs before a model loads and names the exact install line for the first missing package.

KindPackages checkedFix lines printed
potionnumpy, tokenizerspip install numpy, pip install tokenizers
open_cliptorch, open_clip, PILpip install torch, pip install open_clip_torch, pip install pillow
st_multimodal (jina)torch, sentence_transformers, transformerspip install torch, pip install "sentence-transformers>=5.7", pip install "transformers==5.2.0"
st_multimodal (wemm)the jina set plus qwen_vl_utilsthe jina lines plus pip install "qwen-vl-utils==0.0.14"
fakenonenone

The message has this shape:

registry error: preset 'clip-vit-b32' needs the 'torch' package. install with: pip install torch

The check tests that a module can be imported, not its version. When the sentence-transformers stack fails to import inside the loader, its own error message also lists "accelerate>=1.1.0".

urna build --dry-run runs the same check without loading anything. A missing package shows as deps=MISSING with the fix line under it, and the dry run still exits 0. The real build stops with the registry error and exit code 4.

Device and dtype

KindDeviceDtype
st_multimodal[[models]] device, else URNA_ST_DEVICE, else cuda, else mps, else cpu[[models]] dtype, else URNA_ST_DTYPE, else bfloat16 on cuda, float16 on mps, float32 on cpu
open_clip[[models]] device, else cuda, else mps, else cpu (URNA_ST_DEVICE is not read)the weights' loaded dtype
potionCPU, numpyfloat32

Every st_multimodal model runs in its own worker process, because two models that execute repository code in one process collide in the transformers dynamic module cache.

Heavy presets

wemm-4b and wemm-9b are registered as too heavy for a typical machine. A spec that names either one fails validation unless you pass --allow-heavy:

spec error: models.wemm-9b: flagged too heavy for this machine; pass --allow-heavy

At query time the registry query embedder loads them only when URNA_ALLOW_HEAVY=1 is set.

Remote code

The jina and wemm presets set trust_remote_code: loading them runs Python files shipped in the model repository. A spec must name each such preset in [output] allow_remote_code:

[output]
allow_remote_code = ["wemm-2b"]

Without it, validation fails:

spec error: models.wemm-2b: executes model-repo code; opt in with output.allow_remote_code = ["wemm-2b"] (RFC-0 N11)

The query side has the same rule. urna ask and urna retrieve on a corpus built with one of these presets refuse to load the model until you list the preset in URNA_ALLOW_REMOTE_CODE, a comma-separated list such as URNA_ALLOW_REMOTE_CODE="wemm-2b,jina-v5-omni-nano". A file's manifest is never enough to run code.

Pinned files

wemm-2b is the only preset with pinned code files. When its model directory resolves on disk, each file must match its SHA-256 before the model loads, in the parent process and again in the worker:

FileSHA-256
modeling_st_wemm.py521d02c1c60ae727cc9dc6500cdb0b28c53b259e0ce3d37197920a33ba4dd333
modeling_wemm_embedding.pyac255e1fad459cc3e68891d6c3327f4486922aed02fb3c5c13fb53277ba8e94f
chat_template.jinja273d8e0e683b885071fb17e08d71e5f2a5ddfb5309756181681de4f5a1822d80
embedding_chat_template.jinja7c3df2aab83ab9096428ec27b6b99ad87c4790418b830d119957634f28c677ba
processor_config.jsond601e2fe0de1bc11852de3aa843f01a1677ca84f4dc743916c9ce4b8d30fb384

A changed file is refused, not warned about:

registry error: preset 'wemm-2b': <file> sha256 <actual> does not match the pinned allowlist (<expected>). refusing to execute unreviewed remote code.

The jina presets, wemm-4b and wemm-9b have no pins, so the opt-in alone loads them. Build with those in an environment you are willing to run their repository code in.

Model directory resolution

For each preset the registry resolves a model directory in this order:

  1. [[models]] model_path in the spec (or --model-path on urna ask and urna retrieve).
  2. URNA_MODEL_DIR_<NAME>, where <NAME> is the preset name upper-cased with - replaced by _: URNA_MODEL_DIR_WEMM_2B, URNA_MODEL_DIR_JINA_V5_OMNI_NANO.
  3. The preset's local_dir, when that directory exists.
  4. For st_multimodal presets only: the Hugging Face cache snapshot that refs/main points to, under $HF_HOME/hub (default ~/.cache/huggingface/hub). With several snapshots and no usable refs/main, nothing resolves.
  5. Nothing.

potion and the open_clip presets do not load from this directory: potion uses its vendored table and open_clip loads by model id and pretrained tag. urna build --dry-run prints the resolved model_dir, or null.

Importing the forge sets HF_HUB_OFFLINE=1, TRANSFORMERS_OFFLINE=1 and HF_DATASETS_OFFLINE=1 when they are unset, so weights must already be on disk. For st_multimodal presets the build computes model_hash from the local files before any model loads: with no local directory it fails with a Python TypeError (exit 1) instead of downloading. Put the snapshot in the Hugging Face cache, or point model_path or URNA_MODEL_DIR_<NAME> at it.

Identity: model_hash

model_hash identifies the model, not the corpus. It is written into the file and checked by the model gate on every query.

KindPreimageNeeds the model loaded
potionSHA-256 of canonical JSON over the embedder name, version, the hash of config.json, tokenizer.json and model.safetensors, the tokenizer hash, dim, tokenizer name, add_special_tokens, mean pooling, normalize, and float32-stableno
open_clipSHA-256 over model id, pretrained tag, the preprocess transform, a digest of every weight tensor, and l2yes
st_multimodalSHA-256 of canonical JSON over the model file fingerprint (weights, tokenizer, processor), the SHA-256 of each modeling_*.py, *chat_template*.jinja and processor_config.json, the pooling, normalize, and the dtype policyno, computed from the files
fakeSHA-256 of urna-forge-fake-test/v1no

With the vendored table, potion's hash is sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98.

The st_multimodal dtype policy comes from the device, so the same weights hash differently on cuda (bfloat16), mps (float16) and cpu (float32). A corpus built on one device class and queried on another fails the gate unless the dtype is pinned. Set URNA_ST_DTYPE to the same value for the build and for every query. A [[models]] dtype in the spec pins the build side only: the query side never reads the spec.

Query and document contract

Asymmetric models embed queries and documents through different routes. The preset sets the default, and the spec can override the document side.

KindCorpus textQuery textImages
potionone route, role ignoredsamerefused, text only
open_cliptext tower, role ignoredsameimage tower, on source files or decoded frames, L2-normalized
st_multimodaltext_corpus_mode, default document, calls encode_documenttext_query_mode, default query, calls encode_queryencode_document. wemm receives {"image": img, "text": image_prompt} and thumbnails to 768 px; jina receives the bare image

jina also passes encode_kwargs = {task = "retrieval"}. A mode of plain, or any unknown value, calls plain encode(). The default image_prompt is "Represent this image.".

The jina presets take the bare image because the dict form collapses their image vectors onto the shared prompt text: the registry comment records pairwise similarity of 0.99 across different images with the dict, against 0.45 bare.

The query side always uses the preset defaults. Spec overrides of text_query_mode, encode_kwargs, image_prompt, normalize, dtype and device do not reach urna ask or urna retrieve. normalize and the dtype policy are part of the st_multimodal model_hash, so a corpus built with normalize = false, or with a dtype that differs from the one the query machine resolves, fails the query gate.

The query embedder

urna ask and urna retrieve pick a query embedder from the file's embedding_model:

  • A name that starts with minishlab/potion runs python/forge/embed_query_potion.py. This script ships in the installer payload, so the installed binary answers potion corpora.
  • Any other name runs python/forge/embed_query_model.py, which finds the preset by reverse lookup of the manifest name, applies the remote-code and heavy gates, resolves the model directory as above and embeds the query with role="query". This script is not in any release artifact: it runs from a repository checkout, with the preset's packages installed in the interpreter urna uses.

When the file's default space was truncated (the manifest records full_dim), the CLI passes --mrl-dim so the query is sliced to the same length.

Exit code of embed_query_model.pyMeaning
0the query vector was printed
2usage error, or no query
3a model asset is missing
4a dependency, an unknown preset, the remote-code gate or the heavy gate

The CLI does not branch on these codes: any non-zero exit becomes embedder failed (status=...): <stderr>. The protocol is documented in query embedder protocol.

Errors

Registry errors reach urna build as registry error: <message> on stderr with exit code 4:

MessageCause
unknown model preset '<p>'. valid presets: ...a name not in the table
preset 'fake-test' requires URNA_ENABLE_FAKE_PRESET=1 (test-only)fake-test without the variable
preset '<p>' needs the '<module>' package. install with: <line>a missing dependency
preset '<p>': pinned code file missing: <file>a wemm-2b pin with the file absent
preset '<p>': <file> sha256 ... does not match the pinned allowlist ...a wemm-2b pin mismatch

The heavy and remote-code refusals come from spec validation first, as spec error: with exit code 2. For choosing a model and keeping build and query in step, see Choose and bring embedding models.

On this page