Builder pipeline (checkout only)
The Python modules that live only in the urna repo checkout, the builder pipeline, model fingerprint, forge.retrieve, graph context and manifest readers.
The repo's python/ directory has more than the wheel ships. This page covers the modules that import only from a checkout: builder, model_fingerprint, forge.retrieve, graph_context, the forge.forge_manifest readers and convert_legacy.py.
Everything on this page needs a repo checkout
None of these modules is in the urna wheel. pip install urna does not give you builder, model_fingerprint, forge or graph_context. The wheel's own surface is on Module urna.
Set up a checkout
The checkout layout imports urna from python/urna.py and loads the extension from a file next to it, python/_urna.so. Build the extension with the pyo3/extension-module feature:
git clone https://github.com/hoffresearch/urna
cd urna
git lfs pull # the potion table; sh scripts/fetch_potion.sh downloads it without git-lfs
cargo build --release -p urna-python --features pyo3/extension-module
cp target/release/lib_urna.dylib python/_urna.so # macOS
cp target/release/lib_urna.so python/_urna.so # LinuxWithout the feature the library links libpython and can crash under statically linked interpreters such as uv's standalone Python. The ImportError that urna.py raises when it finds no extension suggests a command without the feature: use the one above.
The loader tries from . import _urna first (the wheel layout), then _urna.so, _urna.abi3.so, _urna.dylib and lib_urna.dylib next to urna.py. It has no Windows candidate, so the checkout layout works on macOS and Linux only.
Put python/ on the path before importing, and install numpy and tokenizers for the potion embedder:
import sys
sys.path.insert(0, "python")
import urna
from builder import BuildConfig, Pipeline, chunk_text
from forge.embed_potion import potion_embedderMore on the development setup is on Contributing.
builder
python/builder.py chunks text, embeds it with a cache, and calls urna.build.
ChunkSpec
ChunkSpec(canonical_text: str, source_uri: str, byte_start: int, byte_end: int)A frozen dataclass for one chunk before embedding. spec.chunk_id(chunker_version) returns the id that urna.chunk_id computes for it.
chunk_text
chunk_text(text: str, source_uri: str, *, max_chars: int = 512, overlap: int = 0) -> list[ChunkSpec]Greedy character windows of max_chars characters, with an optional overlap of overlap characters. The spans are UTF-8 byte offsets into text, so urna cite returns spans that point back into the source. Empty text returns []. Raises ValueError when max_chars is 0 or less, or when overlap is negative or not smaller than max_chars.
EmbeddingCache
EmbeddingCache(path: str, model_key: str = "")A SQLite cache of embeddings in the table embeddings_v2, keyed by (chunk_id, model_key). The model is part of the key because a chunk_id does not depend on the model: without it, re-embedding the same corpus with another model would reuse the old vectors.
| Method | Description |
|---|---|
get(chunk_id, dim) | the cached vector or None. Raises ValueError when the cached dimension differs from dim |
put(chunk_id, embedding) | store one vector and commit |
close() | close the database |
An embeddings table from an older version is ignored, never migrated: those chunks are embedded again.
BuildConfig
A mutable dataclass with the arguments Pipeline.emit passes to urna.build. 22 fields, the first five required.
| Field | Type | Default | Passed to urna.build as |
|---|---|---|---|
output_path | str | required | output_path |
embedding_model | str | required | embedding_model |
embedding_dim | int | required | embedding_dim |
chunker_version | str | required | chunker_version |
model_hash | str | required | model_hash |
title | str or None | None | title |
version | str or None | None | version |
description | str or None | None | description |
license | str or None | None | license |
reproducible | bool | True | reproducible. Note that urna.build defaults to False |
preset | str | "exact" | preset |
text_encoding | str or None | None | text_encoding |
dtype | str or None | None | dtype |
mrl_dim | int or None | None | mrl_dim |
with_hnsw | bool or None | None | with_hnsw |
with_bm25 | bool or None | None | with_bm25 |
with_graph | bool | False | with_graph |
graph_top_m | int | 8 | graph_top_m |
drop_overlap | bool | False | forces with_graph=True. It does not change chunking: pass overlap=0 to chunk_text yourself |
hnsw_m | int | 16 | hnsw_m |
hnsw_ef_construction | int | 400 | hnsw_ef_construction |
hnsw_seed | int | 42 | hnsw_seed |
created, authors, blob_refs, blob_data_paths, chunk_blob_spans and spaces have no field. To use them, call urna.build directly. provenance goes through emit.
Pipeline
Pipeline(cfg: BuildConfig, *, embedder, scratch_db: str | None = None)| Member | Description |
|---|---|
embedder | any callable that takes a list of ChunkSpec and returns one vector per spec. PotionEmbedder and StaticEmbedder qualify |
scratch_db | path of an EmbeddingCache. The cache key is embedding_model, a NUL byte, and model_hash |
add(spec), add_many(specs) | queue chunks |
emit(*, provenance=None) | build the file and return output_path |
close() | close the cache |
emit does the following, in order:
- Raises
RuntimeError("pipeline has no chunks")when nothing was added. - Embeds only the chunks missing from the cache. Raises
RuntimeErrorwhen the embedder returns the wrong number of vectors or the wrong dimension. - Deletes an existing file at
output_path. - Calls
urna.buildwith the config.with_graphisTruewhenwith_graphordrop_overlapis set. - Opens the new file and runs
validate()in the same process.
import sys
sys.path.insert(0, "python")
from builder import BuildConfig, Pipeline, chunk_text
from forge.embed_potion import potion_embedder
emb = potion_embedder()
cfg = BuildConfig(
output_path="my_corpus.urna",
embedding_model=emb.embedding_model,
embedding_dim=emb.embedding_dim,
chunker_version="my-chunker/v1",
model_hash=emb.model_hash(),
preset="exact",
)
pipe = Pipeline(cfg, embedder=emb, scratch_db="cache.sqlite")
for source_uri, text in documents: # your (uri, text) pairs
pipe.add_many(chunk_text(text, source_uri))
pipe.emit()
pipe.close()model_fingerprint
python/model_fingerprint.py computes a reproducible model_hash for a sentence-transformers or Hugging Face model directory. It is the only way in the repo to stamp a corpus built with such a model.
Importing it sets HF_HUB_OFFLINE, TRANSFORMERS_OFFLINE and HF_DATASETS_OFFLINE to 1 (when not already set), unless URNA_ALLOW_DOWNLOAD=1.
| Name | Signature | Description |
|---|---|---|
RELEVANT_FILES | tuple | the 10 files hashed, in order: config.json, config_sentence_transformers.json, modules.json, sentence_bert_config.json, tokenizer.json, tokenizer_config.json, special_tokens_map.json, 1_Pooling/config.json, model.safetensors, pytorch_model.bin |
PLACEHOLDER_HASH | str | sha256: followed by 64 zeros |
ModelFingerprint | frozen dataclass | model_id, files_hash, tokenizer_hash, pooling_config_hash, embedding_dim, normalize_embeddings, and .to_dict() |
compute_model_fingerprint | (model_dir, *, model_id=None) -> ModelFingerprint | hashes the files of RELEVANT_FILES that exist. model_id defaults to _name_or_path from config.json, then the directory name. Raises FileNotFoundError when model_dir is not a directory |
fingerprint_to_model_hash | (fp) -> str | sha256: of fp.to_dict() as JSON with sorted keys and no whitespace |
is_placeholder | (model_hash) -> bool | True for the zero placeholder |
hf_cache_snapshot | (model_id) -> Path or None | the snapshot in $HF_HOME/hub (default ~/.cache/huggingface) that refs/main points to, or the only snapshot. None when it is missing or ambiguous |
resolve_model_dir | (model_name_or_path) -> Path | a directory as given, else the Hugging Face cache, else the path sentence-transformers loads from. Raises FileNotFoundError suggesting --model-path |
Pass model_id equal to the embedding_model you write into the file. urna search-text fingerprints the model under the manifest name, so a different model_id gives a different hash and the gate refuses the query.
from model_fingerprint import compute_model_fingerprint, fingerprint_to_model_hash, resolve_model_dir
name = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
cfg.embedding_model = name
cfg.model_hash = fingerprint_to_model_hash(compute_model_fingerprint(resolve_model_dir(name), model_id=name))forge.retrieve
python/forge/retrieve.py embeds a text query and calls UrnaFile.retrieve with the gate on.
retrieve(urnafile, query: str, k: int = 5, embedder=None, verify_model: bool = True) -> list[RetrieveHit]| Parameter | Description |
|---|---|
urnafile | an open UrnaFile |
query | the question, as text |
k | hits to return |
embedder | an embedder with the static protocol. None means potion_embedder() |
verify_model | when True and the embedder has model_hash, passes expected_model_hash=embedder.model_hash() |
build_demo(out_path) builds the Markdown files of python/forge/demo_corpus/ (without its README.md) with Pipeline, potion and the default chunk_text, preset="exact", reproducible=True and chunker_version="forge-demo/1", and returns out_path.
Run the module to build the demo in a temporary directory, ask one question and print cited answers:
python python/forge/retrieve.pygraph_context
neighbor_context(canonical_texts, ordinal: int, *, radius: int = 1, joiner: str = " ") -> strJoins chunk ordinal with its radius neighbors on each side, in file order. It rebuilds context around a hit when the corpus was built without chunk overlap. Raises IndexError when ordinal is out of range.
UrnaFile has no accessor for all canonical texts, so you pass the texts yourself, in file order. The ordinal of a hit is db.chunk_ids().index(hit.chunk_id).
forge.forge_manifest readers
urna build --spec writes a sidecar <name>.manifest.json next to the .urna (see Build artifacts). These functions read that sidecar, not the manifest inside the file, which is UrnaFile.inspect()["manifest"].
| Function | Description |
|---|---|
manifest_items(manifest, need=()) | returns items[]. For a compact manifest (output.provenance = "minimal"), asking in need for a dropped field (image_path, label, media_uri) raises ValueError naming output.provenance = "standard" |
frame_resolver(manifest) | returns a function from an item to its global media stream frame. It uses a media://...#frame=N uri when the item has one, else the item's ordinal through frame_of_row and the media order permutation |
The module also holds the write side the forge uses (write_manifest, redact_path, build_lock, check_lock, MANIFEST_SCHEMA_VERSION = 1, COMPACT_DROPPED).
convert_legacy.py
Converts the pre-v1 SQLite corpus layout (articles, text_blocks, blobs tables) into a v1 .urna:
python python/convert_legacy.py --src legacy.urna --dst out.urna --reproducibleconvert(src, dst, *, reproducible) does the same from Python. It writes one chunk per article with source_uri = legacy://truw_ptbr/<id>, a span from 0 to the UTF-8 length of the text, labels and sources in the provenance, and chunker_version = "legacy/truw_ptbr_v0.1.0", then validates the file. It needs numpy, zstandard, and the sentence-transformers snapshot of the model in the Hugging Face cache (for the fingerprint). Errors exit through SystemExit, including a refusal to write the zero placeholder hash.
Embedders
The offline potion embedder in the urna wheel, the lexical floor and registry adapters in the repo checkout, and the two embedder protocols they follow.
Errors
The exceptions urna raises in Python, when each one happens, and the exact messages for opening, querying, the model gate and urna.build.