docsv0.5.1

Known limits

Where urna 0.5.1 behaves differently from what you might expect, grouped by building, querying, file format and install, with a workaround for each.

This page lists the places where urna 0.5.1 behaves differently from what its commands, defaults or older documentation suggest, and the known gaps in what it covers. Each entry describes the current behavior and, when there is one, a workaround. The design boundaries (no in-place updates, no metadata filters, no concurrent writers, no encryption) are on the Introduction.

Building

urna build needs a repo checkout

urna build launches the forge, python/tools/urna_forge.py, which lives in the repository and is not in any release artifact: not in the archives, the embedder payload, the npm package, the Homebrew formula or the wheel. Outside a checkout the build stops with urna_forge.py not found (...); run from the repo or install the forge payload, exit 1. There is no forge payload to install. The last screen of urna setup and its --yes output still suggest urna build --spec corpus.toml as a next step.

A checkout also needs two things before the first build: the real potion table (it is stored with Git LFS; run git lfs pull or sh scripts/fetch_potion.sh) and the Python extension at python/_urna.so. The build writes the file through that extension, so a missing or stale one fails at the last stage, after the embeddings are computed.

To build, clone the repository and prepare it as in the Quickstart. Building works on macOS and Linux.

Automatic crf fails on rows without a label

With crf = "auto" in [media], a build fails with AttributeError: 'Row' object has no attribute 'text' (exit 1) as soon as any row has an empty or missing label. This happens even when the gate's hit@1 leg is off. An image_dir source without a labels file hits it on every row.

To work around it, give every row a label (a labels file for image_dir, a label_template in [source.image] for SQLite, CSV and JSONL sources) or set a fixed integer crf.

PDF directory sources always fail

source.kind = "pdf_dir" passes validation and --dry-run, then every build exits 2 with spec error: source.kind=pdf_dir: build via forge_pipeline (pages are temporary).

The working path for PDFs is the older image tool, python python/tools/urna_build_image_corpus.py --pdf, which renders pages as images. It writes no media tables and no named spaces into the file. See Images and PDFs.

The JPEG XL metadata option is ignored

keep_metadata in [media.jxl_transcode] is parsed and never read. Setting it changes nothing in the encode.

Spec truncation skips the model ladder

mrl_dim in [build] truncates the default vectors after checking only that it is above zero and not above the model's dimension (and a multiple of 64 for int4). It does not check the model's validated ladder, so a model that was never trained for truncation, such as potion, accepts any value. dims in [[models]] does check the ladder.

To stay on validated values, pick mrl_dim from the model's ladder in the model registry.

Media profile name is not in the manifest

The name of the [media] profile and its resolved settings are recorded only in <name>.build.lock.json, under resolved_spec.media. The <name>.manifest.json media block holds the encoder record, not the profile name. Read the lock file to see which profile built a corpus.

Hub models are not downloaded by the build

The build computes a registry model's model_hash from the files on disk before it loads the model. When a sentence-transformers model (jina-v5-omni-*, wemm-*) has no local copy, the build fails with a Python TypeError (exit 1), even with URNA_ALLOW_DOWNLOAD=1. The open_clip presets (clip-vit-b32, siglip2) also load with the Hugging Face offline flags set, so their weights must be cached too.

To work around it, download the model beforehand into the Hugging Face cache (under HF_HOME), or point model_path in [[models]], or URNA_MODEL_DIR_<NAME>, at a local copy.

Spec value types are not checked

Validation checks table names, key names, required keys and allowed values. It does not check value types or table shapes. A wrong type fails later as a raw Python exception with a traceback and exit 1, not as spec error: with exit 2. [build] preset and [build] dtype are checked only when the file is written, after the embeddings are computed. --dry-run never opens the source, so missing files, bad SQL, an order_by that is not a total order and join mismatches all pass a dry run.

To catch these early, run a small --sample build into a separate --out-dir before the full build.

Some forge flags are missing from urna build

--strict-env, --seed and --json exist only on the forge script. urna build --strict-env is a usage error, exit 2. To use them, run python python/tools/urna_forge.py --spec ... from the checkout. --seed has no effect today.

Sample builds and row changes move citations

In a corpus built by urna build, each chunk's span is its row position after sorting, and the span is part of the chunk_id. Adding or removing a row shifts the citation of every later row, and a --sample build never shares citations with the full build. Rows sort by the string form of their order_by values, so "10" comes before "2". Changing chunker_version changes every citation.

Registry models have cost cliffs

wemm-2b runs in float16 on Apple GPUs with images resized to 768 pixels on the long side, at about 0.6 images per second. jina-v5-omni-nano has no resize default and embeds images at native resolution, at about 0.3 images per second. Both settings are part of the embedding recipe, so changing either one re-embeds that model's whole cache.

rebuild-only overwrites the previous lock

urna build --rebuild-only compares the new build.lock.json with the one on disk and prints the divergence, then writes the new lock over the old one. The evidence of what diverged is gone after the run.

Copy <name>.build.lock.json and note the file_hash before you run --rebuild-only.

Querying

Hybrid preset files never use BM25

The hybrid preset, which is also the default of urna build, writes an HNSW index and a BM25 index but declares index_type = "hnsw". ask, retrieve, search-text and UrnaFile.retrieve route by index_type, so on these files they always take the HNSW path and never read the BM25 index. urna ask --disclose explain shows it as route: hnsw and bm25=0. No preset and no spec key declares index_type = "hybrid"; only the Rust writer's hybrid() method does.

To use the BM25 index, call UrnaFile.search_hybrid(query_vector, query_text, k, candidates) from Python. On a file that does declare hybrid, UrnaFile.retrieve passes an empty query text, so its BM25 leg adds nothing; the CLI passes your question.

Hybrid search ranks by cosine only

Hybrid search takes the vector candidates and the BM25 candidates, merges them with reciprocal rank fusion, then rescores every candidate by exact cosine and sorts by that score. The fusion order is discarded. BM25 can add a candidate that the vector search missed, but it never lifts a chunk above one with a higher cosine. Hits report score_type = "cosine" even when the file declares hybrid_rrf.

Hybrid search also never falls back to exact search. Without a BM25 index the lexical list is empty and the route still reads hybrid. Without an HNSW index the vector list keeps only candidates rows, so with no BM25 hits and candidates below k it returns fewer than k hits.

Low ef values have no effect

The HNSW search beam is the largest of --ef (or --candidates), k and the ef_construction the file was built with. Files built by urna.build and urna build use ef_construction = 400 by default, so any --ef below 400 runs at 400 and the latency and recall do not change.

To search with a narrower beam, build the file with a lower hnsw_ef_construction in urna.build.

ask and retrieve never use the graph

The chunk graph that urna build writes by default is used only by urna search-graph and UrnaFile.search_graph. ask and retrieve route by index_type, which cannot be graph. Graph search itself returns the same hits as exact search: it seeds from the exact top results, expands along the graph and reranks the union by exact cosine, so the seeds always win.

The Python model gate is opt-in

The CLI checks the query embedder's model name, dimension and model_hash against the file before every text query, and refuses the all-zero placeholder hash. In Python, search, search_ann, search_hybrid and search_graph never check, and retrieve and search_space check only when you pass expected_model_hash. A vector from the wrong model of the same dimension returns results that are valid cosine scores and meaningless.

To keep the check on, pass expected_model_hash=emb.model_hash() to retrieve. See The model gate.

Query embedders ignore spec overrides

For registry models, ask and retrieve build the query embedder from the preset's defaults. Spec settings such as text_query_mode, encode_kwargs, image_prompt, normalize, dtype and device are not carried to query time. For sentence-transformers models the numeric precision enters the model_hash, and it defaults by device (bfloat16 on CUDA, float16 on Apple GPUs, float32 on CPU). A corpus built on one device class and queried on another fails the model gate.

To work around it, set URNA_ST_DTYPE at query time to the precision the corpus was built with.

The default embedder is English

potion-base-8M is distilled from bge-base-en-v1.5. English synonyms land close together (car and automobile at +0.78 cosine), while text in other languages rides English subword rows and the signal is weak (carro and automovel at +0.08). A corpus mostly in another language needs a multilingual model. See Choose and bring embedding models.

BM25 does not segment CJK, Thai or Lao

The BM25 tokenizer lowercases, splits only on characters that are not letters or digits, and drops tokens shorter than two characters. There is no stemming and there are no stop words. Chinese, Japanese, Thai and Lao are written without spaces, so a whole clause becomes one token and matches only an identical clause.

For corpora in those languages, build without BM25 (with_bm25=False).

Text queries start a Python process

ask, retrieve and search-text start a Python process for every call to embed the question. For search-text, which also loads sentence-transformers, that costs about 300 to 500 ms per call. The latency tables on Benchmarks do not include this time.

For many queries, embed in-process from Python and call UrnaFile.search or UrnaFile.retrieve directly.

SigLIP2 text queries can fail offline

The siglip2 text tower loads its tokenizer through a lookup that probes optional files online. In strict offline mode a fresh process can fail that probe even when the model is cached. The image tower and every other preset are not affected. For a sealed offline setup, query siglip2 spaces by image, or use the text towers of wemm-* or jina-v5-omni-*.

search-text fails on corpora built with mrl_dim

ask and retrieve pass --mrl-dim to the query embedder when the manifest records a truncated dimension. urna search-text does not, so on a corpus built with mrl_dim the embedder returns the full dimension and the dimension check fails.

Use urna ask or urna retrieve on those corpora.

cite prints the stored span on media corpora

On a corpus with a media overlay, search hits, ask and retrieve report the blob URI and its byte range. urna cite reads the stored spans section directly, so for the same chunk it prints the original source row, not the blob range. Both identify the same chunk; only the span differs.

Take the span from retrieve when you need the blob range.

File format and hashes

Placeholder model hashes are accepted at build time

urna.build and the Rust writer accept the all-zero placeholder model_hash (sha256: followed by 64 zeros). Only the CLI refuses it, at query time: ask, retrieve and search-text (unless search-text --skip-model-hash-check). Python search methods never check it.

Build with the embedder's real hash, for example model_hash=emb.model_hash().

Adding HNSW changes the citations

Attaching an HNSW index sets index_type = "hnsw" and rerank_policy = "exact" in the search contract, and the search contract is part of content_hash. An exact file and an HNSW file of the same chunks therefore have different content_hash values, and none of their citations resolve against each other. The same goes for exact against tiny, nano or hybrid, and for with_hnsw=True. The BM25 index, the chunk graph, media and named spaces do not move content_hash. The embedding dtype and mrl_dim do, by design.

Choose the preset before you issue citations, and cite against the file you ship.

Hashes prove integrity, not authorship

The header checksum and each section checksum are the first 8 bytes of a SHA-256. file_hash is the SHA-256 of the whole file, and content_hash covers the decoded canonical sections. None of them are keyed or signed: anyone who edits a file can recompute them, so urna validate passing proves the bytes are consistent, not who produced them. Opening an untrusted file still runs the parser, which is bounds-checked and fuzzed. See Security.

To trust a file you received, check its file_hash against a value published through a channel you trust.

urna validate does not decode the indices

urna validate checks the header, every section checksum, the footer hash, the manifest and search contract, the embeddings for NaN or infinity, and every inlined media blob against its hash. It does not decode the HNSW, BM25, graph or media overlay payloads, and it does not scan named-space vectors for NaN. A file whose index payload is malformed but correctly checksummed passes validate and fails when a query opens it.

urna inspect --json and every query verb open the file fully, which decodes and checks those sections.

validate prints OK before it checks inlined media

urna validate prints its OK: line and the header, section, manifest and embedding checks first, then verifies every inlined media blob. A blob that fails its hash leaves the OK: line on stdout and exits 1.

In scripts, test the exit code, not the output.

Installing and releases

Installed binaries answer only potion corpora

The embedder payload that urna setup and the one-line installers lay down carries only the potion query embedder. ask and retrieve on a corpus built with a registry model (clip-vit-b32, siglip2, jina-v5-omni-*, wemm-*) stop with embedder script not found: python/forge/embed_query_model.py (override with --embedder). urna setup cannot fix it, although the explorer's hint suggests running setup.

To query such a corpus, run from the root of a repository checkout with the model's Python dependencies installed in the interpreter urna uses, or pass --embedder with the path to python/forge/embed_query_model.py in a checkout. Run urna stats to see which model a corpus was built with.

search-text needs a checkout

urna search-text embeds with python/embed_query.py, which urna looks for only in a repository checkout (the current directory, its parent, or the checkout of a binary built from source). It also needs sentence-transformers. With an installed binary, run it from a checkout or pass --embedder.

Package channels need a setup step

Homebrew, npm (and bun, pnpm, yarn), cargo install and cargo binstall install the binary alone, and none of them runs urna setup for you. Until setup runs, urna doctor exits 2 when no Python is found, 3 when Python lacks numpy or tokenizers, and 4 when the dependencies are there but the embedder is missing. Setup needs curl on PATH, and either uv or a python3 that can create a venv.

Setup needs a package index

urna setup downloads the embedder payload with curl and then builds its Python env with uv or pip, which reach your configured package index; uv may also download a Python. For an air-gapped machine, the payload can come from a local mirror (URNA_RELEASE_BASE=file:///...), but the env cannot. Point URNA_PYTHON at an interpreter that already has numpy and tokenizers and run urna setup --no-python. Setup accepts only https:// and file:// mirrors, and a stalled download has no overall timeout. See Air-gapped install and queries.

The installer overwrites the binary in place

The shell one-liner copies the new binary over an existing urna without removing it first. macOS can kill a binary overwritten in place (exit 137). Before re-running the one-liner to upgrade on macOS, remove the old binary: rm -f ~/.local/bin/urna, or the same name under URNA_BIN_DIR.

The SBOM and payload have no build provenance

Only the five platform archives carry SLSA build provenance (gh attestation verify <archive> --repo hoffresearch/urna). The embedder payload, which holds Python code urna runs at query time, and the CycloneDX SBOM are covered by the GitHub release attestation only: verify them with gh release verify-asset v0.5.1 <file>. The npm package has no npm provenance, the 0.5.x wheels have no PEP 740 attestations, and the npm wrapper and cargo binstall do not check a checksum beyond HTTPS. See Verify what you installed.

The wheel's urna command shadows the binary

pip install urna installs a Python command also named urna, with four read-only verbs (validate, inspect, stats, search). Whichever urna comes first on PATH wins, and on Linux pip install --user writes to ~/.local/bin, the same directory the one-liner uses. urna --version tells them apart: the Rust binary prints urna 0.5.1, the Python command has no --version.

On Windows, run setup or set the interpreter

When URNA_PYTHON is unset and there is no setup env, urna looks for a project .venv in the Unix layout (.venv/bin/python) and then for python3, which Windows usually lacks. Run urna setup or set URNA_PYTHON. The PowerShell installer supports x86_64 Windows only.

Query commands look for Python in the working directory

To find an interpreter, the query verbs try URNA_PYTHON, then the urna setup venv, then a .venv in the working directory or up to three parents, then python3. The query embedder script is looked up in the repository layout of the working directory before the installed payload. Running ask from a directory you do not control can pick up its Python or its scripts.

Run queries from a directory you trust, or set URNA_PYTHON. See Paths and resolution order.

On this page

Buildingurna build needs a repo checkoutAutomatic crf fails on rows without a labelPDF directory sources always failThe JPEG XL metadata option is ignoredSpec truncation skips the model ladderMedia profile name is not in the manifestHub models are not downloaded by the buildSpec value types are not checkedSome forge flags are missing from urna buildSample builds and row changes move citationsRegistry models have cost cliffsrebuild-only overwrites the previous lockQueryingHybrid preset files never use BM25Hybrid search ranks by cosine onlyLow ef values have no effectask and retrieve never use the graphThe Python model gate is opt-inQuery embedders ignore spec overridesThe default embedder is EnglishBM25 does not segment CJK, Thai or LaoText queries start a Python processSigLIP2 text queries can fail offlinesearch-text fails on corpora built with mrl_dimcite prints the stored span on media corporaFile format and hashesPlaceholder model hashes are accepted at build timeAdding HNSW changes the citationsHashes prove integrity, not authorshipurna validate does not decode the indicesvalidate prints OK before it checks inlined mediaInstalling and releasesInstalled binaries answer only potion corporasearch-text needs a checkoutPackage channels need a setup stepSetup needs a package indexThe installer overwrites the binary in placeThe SBOM and payload have no build provenanceThe wheel's urna command shadows the binaryOn Windows, run setup or set the interpreterQuery commands look for Python in the working directory