Errors
The exceptions urna raises in Python, when each one happens, and the exact messages for opening, querying, the model gate and urna.build.
The Rust side of urna raises one exception type, ValueError, for every runtime failure: a missing or corrupt file, a bad query, a model mismatch, a failed build or a failed write. TypeError and OverflowError come only from converting your arguments to the types Rust expects. The checkout-only modules add a few Python exceptions of their own.
Exception types
| Exception | Raised by | When |
|---|---|---|
ValueError | every urna call | any error from Rust: opening, validating, searching, the model gate, spaces, blobs, build checks, the final write |
TypeError | argument conversion | a pathlib.Path to urna.open, chunks that is not a list, provenance that is not a JSON-serializable dict, authors as a str, k as a float or str, a seventh positional argument to build, urna.UrnaFile(...), pickling a file or a hit, an embedding given as a str |
OverflowError | integer conversion | a negative value where Rust expects an unsigned one (byte_start, embedding_dim, blob_bytes(-1), the offsets of chunk_id), or k of 2**31 or more |
ModuleNotFoundError | urna.embed_potion, on first use of the table | numpy or tokenizers is missing: install urna[embed] |
FileNotFoundError | PotionEmbedder, model_fingerprint | the potion table is missing or is a git-lfs pointer; a model directory cannot be resolved |
ImportError | import urna in a checkout | no extension next to python/urna.py |
RuntimeError | builder.Pipeline.emit (checkout) | no chunks; the embedder returned the wrong count or dimension |
IndexError | graph_context.neighbor_context (checkout) | ordinal out of range |
SystemExit | convert_legacy.convert (checkout) | bad legacy input; refusal to write the placeholder hash |
forge.model_registry.RegistryError | the model registry (checkout) | unknown preset, missing dependencies, the remote-code or heavy gate. It subclasses ValueError |
Because the runtime errors share one type, tell them apart by message. The messages below are verbatim from urna 0.5.1.
Open
| Situation | Message |
|---|---|
| file does not exist | No such file or directory (os error 2) |
| path is a directory | Invalid argument (os error 22) |
not a .urna file | magic mismatch: expected [85, 82, 78, 65], got [...] |
Checksum, hash, manifest and contract failures also raise at open, because open validates the whole file. The message is the typed error of the reader. The full list is on Typed errors.
Query
These apply to search, search_ann, search_hybrid, search_graph, search_space and retrieve.
| Situation | Exception and message |
|---|---|
| vector of the wrong length | ValueError: dimension mismatch: expected 256, got 2 |
k of 0 or negative | ValueError: invalid k: 0 |
a str as the query | ValueError: invalid query vector: TypeError: Can't extract `str` to `Vec` |
| a 2-d numpy array | ValueError: invalid query vector: TypeError: only 0-dimensional arrays can be converted to Python scalars |
| empty vector | ValueError: empty query |
| NaN or Inf in the vector | ValueError: NaN or Inf in query |
| all zeros | ValueError: zero-norm query |
k as a float | TypeError: 'float' object cannot be interpreted as an integer |
k of 2**31 | OverflowError: out of range integral type conversion attempted |
search_ann without ef | TypeError: UrnaFile.search_ann() missing 1 required positional argument: 'ef' |
Model gate and spaces
| Call | Message |
|---|---|
retrieve(..., expected_model_hash=X) on a corpus built with Y | model_hash mismatch: the query was embedded with X, but the corpus was built with Y. Results would be cosine-valid but semantically wrong. Pass expected_model_hash=None to bypass this check. |
search_space("vision", ...) on a file without that space | embedding space not found: vision |
search_space(name, ..., expected_model_hash=X) on a space built with Y | model_hash mismatch in space <name>: the query was embedded with X, but the space vectors were embedded with Y |
The retrieve gate runs before the query is parsed, so a mismatch is reported even when the vector is also wrong. Without expected_model_hash there is no check at all: see The model gate.
Blobs
| Call | Message |
|---|---|
blob_bytes(i) when blob i is not inlined, is out of range, or the file has no blobs | blob <i> is not inlined in this file: open the media sidecar named by its blob_refs uri, or rebuild with [output] embed_media |
blob_bytes(-1) | OverflowError: can't convert negative int to unsigned |
Build
urna.build checks its input before it writes anything. The last possible error is the write itself.
| Situation | Message |
|---|---|
| unknown preset | unknown preset: micro (expected exact|compressed|tiny|nano|hybrid) |
| unknown text encoding | unknown text_encoding: gzip (expected raw|zstd) |
| unknown dtype | unknown dtype: bfloat16 (expected float32|float16|int8|int4) |
int4 with a dimension not divisible by 64 | int4 requires (effective) embedding_dim divisible by 64, got 48 |
mrl_dim of 0 or above embedding_dim | mrl_dim must satisfy 0 < mrl_dim <= embedding_dim (64), got 0 |
chunks=[] | manifest invalid: n_chunks must be > 0 |
| an item that is not a dict | chunks[0] is not a dict |
| a missing key | chunks[0] missing embedding |
| embedding of the wrong length | dimension mismatch: expected 64, got 2 |
| NaN or Inf in an embedding | NaN or Inf detected in embeddings |
byte_end smaller than byte_start | invalid input: byte_end (1) < byte_start (5) |
model_hash without the prefix | invalid model_hash: expected 'sha256:<hex>' prefix, got 'abc' |
empty embedding_model | manifest invalid: embedding_model must not be empty |
empty chunker_version | manifest invalid: chunker_version must not be empty |
embedding_dim=0 | manifest invalid: embedding_dim must be > 0 |
blob_data_paths without blob_refs | blob_data_paths requires blob_refs (the 0x17 table parallels the 0x14 order) |
| a blob hash that is not 32 bytes | blob_refs[0] content_hash must be 32 bytes |
chunk_blob_spans of the wrong length | chunk_blob_spans must have one entry per chunk (3), got 1 |
| a space with the wrong number of rows | spaces[0] has 1 rows but the corpus has 3 chunks |
| a space with rows of different lengths | spaces[0] row 2 has dim 4 but expected 8 |
| 16 spaces | at most 15 non-text spaces, got 16 |
| a space given as a tuple | spaces[0] is not a dict |
| two spaces with one name | space_table encode: malformed section payload: section=21 reason=space_table: duplicate name v |
a space model_hash without the prefix | space_table encode: malformed section payload: section=21 reason=space_table: model_hash must be sha256:<hex> |
an int4 space with a dimension not divisible by 64 | spaces[0] int4: invalid input: encode_int4_embeddings: dim=8 must be a nonzero multiple of 64 |
| the parent directory does not exist, or the write fails | write <path>: No such file or directory (os error 2) |
urna.build accepts the zero placeholder model_hash and uppercase hex without an error. The refusal of the placeholder comes at query time, in the CLI. See urna.build.
Handle errors
Catch ValueError around anything that touches a file, and let TypeError surface: it means the calling code passes the wrong type.
import urna
try:
db = urna.open("corpus.urna")
hits = db.retrieve(qvec, 5, expected_model_hash=emb.model_hash())
except ValueError as err:
if str(err).startswith("model_hash mismatch"):
raise SystemExit("this corpus was built with another model") from err
raiseThe wheel's urna command turns any exception into urna: error: <message> on stderr and exit code 1. See The wheel's urna command.
Builder pipeline (checkout only)
The Python modules that live only in the urna repo checkout, the builder pipeline, model fingerprint, forge.retrieve, graph context and manifest readers.
The wheel's urna command
The urna console command installed by the Python wheel, its four read-only verbs, how it differs from the Rust binary, and the PATH collision between the two.