Embedders
The offline potion embedder in the urna wheel, the lexical floor and registry adapters in the repo checkout, and the two embedder protocols they follow.
urna.build and the search methods take vectors, not text. An embedder turns text into those vectors and reports the identity (embedding_model, model_hash) that the file records. The wheel ships one embedder, urna.embed_potion. A repo checkout adds a stdlib-only lexical embedder and the adapters of the model registry.
urna.embed_potion
The potion-base-8M static table from model2vec, bundled in the wheel. It needs numpy and tokenizers (the embed extra), no torch, no network. It never opens a socket.
pip install "urna[embed]"from urna.embed_potion import potion_embedder
emb = potion_embedder()
vectors = emb.embed_texts(["first text", "second text"]) # two lists of 256 floatsIn a checkout the same module is forge.embed_potion, with the same contents.
Importing the module sets HF_HUB_OFFLINE, TRANSFORMERS_OFFLINE and HF_DATASETS_OFFLINE to 1 and TOKENIZERS_PARALLELISM to false, each only when it is not already set.
Module names
| Name | Value |
|---|---|
MODEL_ID | "minishlab/potion-base-8M" |
POTION_VERSION | "1" |
MODEL_DIR | the bundled table directory, next to the module |
RELEVANT_FILES | ("config.json", "tokenizer.json", "model.safetensors"), the files hashed into the fingerprint |
PotionEmbedder | the embedder class |
potion_embedder(model_dir=MODEL_DIR) | returns a PotionEmbedder |
default_embedder | alias of potion_embedder |
PotionEmbedder
PotionEmbedder(model_dir: Path | str = MODEL_DIR, normalize: bool = True)Loading is lazy: the table loads on the first use of embedding_dim, fingerprint(), model_hash() or embed_texts(), and is cached per directory.
| Member | Kind | Returns |
|---|---|---|
embedding_model | property | "minishlab/potion-base-8M/v1", without loading the table |
embedding_dim | property | 256 |
fingerprint() | method | dict with embedder, version, files_hash, tokenizer_hash, embedding_dim, tokenizer ("bert-wordpiece-lower"), add_special_tokens (False), pooling ("mean"), normalize ("l2" or "none"), dtype ("float32-stable") |
model_hash() | method | sha256: of the fingerprint as canonical JSON |
embed_texts(texts) | method | list[list[float]], one vector per text |
__call__(specs) | method | embed_texts over each item's canonical_text, or the item itself when it is a string. This lets it act as the embedder of the checkout's builder.Pipeline |
model_hash is a method, not a property: call emb.model_hash().
embed_texts tokenizes without special tokens, averages the token rows, L2-normalizes and casts through float32. An empty text maps to the [UNK] row, so it never returns a zero vector.
With the table in the 0.5.1 wheel, model_hash() is:
sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98The hash covers the bytes of the table files, so a different table gives a different hash, and corpora built with the old one fail the model gate against the new one.
Failures
| Exception | When |
|---|---|
FileNotFoundError | the table is missing, or model.safetensors is a git-lfs pointer (under 1024 bytes). The message suggests git lfs pull |
ModuleNotFoundError | numpy or tokenizers is missing (no embed extra). Raised on first table use, not at import |
ValueError | the table is not float32 |
Make a Python-built file answerable by the CLI
The installed urna ask and urna retrieve embed the query themselves. They choose the potion embedder when the manifest embedding_model starts with minishlab/potion, then check the name, the dimension and the model_hash. Pass the embedder's identity to urna.build exactly:
urna.build(
"corpus.urna",
emb.embedding_model, # "minishlab/potion-base-8M/v1"
emb.embedding_dim, # 256
"my-chunker/1",
emb.model_hash(),
chunks,
)Then urna ask corpus.urna "your question" works with the installed binary after urna setup. See urna ask.
Other models
Any model works with urna.build, as long as you pass its vectors, its name and a model_hash that identifies it. What changes is who can query the file afterwards.
| Built with | Query from Python | Query from the CLI |
|---|---|---|
| potion (wheel or checkout) | retrieve(..., expected_model_hash=emb.model_hash()) | urna ask and urna retrieve, installed binary |
a registry model (wemm-*, clip-vit-b32, siglip2, jina-*), through urna build --spec | embed with the registry adapter, checkout only | urna ask and urna retrieve from a checkout with the model's Python deps |
a sentence-transformers model you fingerprint with model_fingerprint | embed with the same model, pass its hash | urna search-text from a checkout, with sentence-transformers installed |
The installed binary answers potion corpora only
The query embedder payload that urna setup installs carries only the potion scripts. urna ask or urna retrieve on a corpus built with a registry model stops with embedder script not found: python/forge/embed_query_model.py. Run them from a repo checkout with the model's dependencies installed. See Known limits and Choose and bring embedding models.
To stamp a file built with your own sentence-transformers model, compute its hash with model_fingerprint (checkout only) and pass the same name as embedding_model and as model_id, because urna search-text fingerprints the model under the manifest name:
import sys
sys.path.insert(0, "python")
from model_fingerprint import compute_model_fingerprint, fingerprint_to_model_hash, resolve_model_dir
name = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
model_hash = fingerprint_to_model_hash(compute_model_fingerprint(resolve_model_dir(name), model_id=name))The functions are documented on Builder pipeline.
forge.embed_default (checkout only)
The lexical floor: stdlib only, no numpy, no model files. Each token maps to a fixed pseudo-random vector derived from its bytes, and a text is the L2-normalized mean of its token vectors. Cosine therefore reflects shared literal tokens only: synonyms do not land close. Use it for tests and for environments with nothing but the standard library.
import sys
sys.path.insert(0, "python")
from forge.embed_default import StaticEmbedder
emb = StaticEmbedder()
emb.embedding_model # "urna-forge-static/v1"| Name | Value |
|---|---|
MODEL_ID | "urna-forge-static" |
STATIC_EMBEDDER_VERSION | "1" |
DEFAULT_DIM | 256 |
DEFAULT_SEED | "urna-forge-static/v1" |
embed_one(text, dim=256, seed=DEFAULT_SEED) | one vector. Text with no alphanumerics maps to a fixed non-zero vector |
StaticEmbedder(dim=256, seed=DEFAULT_SEED) | same members as PotionEmbedder: embedding_model, embedding_dim, fingerprint(), model_hash(), embed_texts(), __call__() |
default_embedder() | returns StaticEmbedder() |
The seed and the dimension are part of the fingerprint, so changing either changes model_hash. The default instance hashes to sha256:75acb1edc57eb9dcb18d989fac562bb2575e20d3de211d849c4dbf649b81d1a1.
The tokenizer splits on alphanumerics, which works poorly for CJK text.
The forge package
import forge (checkout only) exports:
| Name | Is |
|---|---|
PotionEmbedder, potion_embedder | from forge.embed_potion |
default_embedder | the potion factory |
StaticEmbedder | from forge.embed_default |
lexical_embedder | the lexical factory, forge.embed_default.default_embedder |
Embedder protocols
Two shapes of embedder exist, and they are not interchangeable.
| Member | Static protocol: PotionEmbedder, StaticEmbedder | Registry adapters: forge.model_registry.create_embedder(name) (checkout only) |
|---|---|---|
| model name | embedding_model property | embedding_model attribute |
| dimension | embedding_dim property | dim property |
| hash | model_hash() method | model_hash property |
| fingerprint | fingerprint() | fingerprint() |
| embed | embed_texts(texts) returns list[list[float]] | embed_texts(texts, role="document") returns a float32 numpy array of shape (n, dim). Also embed_paths and embed_arrays for media |
| pipeline hook | __call__(specs) | none |
forge.retrieve.retrieve(..., embedder=...) calls model_hash() as a method, so it needs the static protocol. builder.Pipeline(embedder=...) accepts any callable that maps a list of chunk specs to a list of vectors. With a registry adapter, pass role="query" when you embed a query.
The registry itself (presets, dimensions, the remote-code and heavy gates) is on Model registry. The protocol the CLI uses to run an embedder script is on Query embedder protocol.
urna.build
Reference for urna.build, which writes a .urna file from chunks you already embedded, with all 29 parameters, the presets, the chunk schema and its checks.
Builder pipeline (checkout only)
The Python modules that live only in the urna repo checkout, the builder pipeline, model fingerprint, forge.retrieve, graph context and manifest readers.