docsv0.5.1

Builder pipeline (checkout only)

The Python modules that live only in the urna repo checkout, the builder pipeline, model fingerprint, forge.retrieve, graph context and manifest readers.

The repo's python/ directory has more than the wheel ships. This page covers the modules that import only from a checkout: builder, model_fingerprint, forge.retrieve, graph_context, the forge.forge_manifest readers and convert_legacy.py.

Everything on this page needs a repo checkout

None of these modules is in the urna wheel. pip install urna does not give you builder, model_fingerprint, forge or graph_context. The wheel's own surface is on Module urna.

Set up a checkout

The checkout layout imports urna from python/urna.py and loads the extension from a file next to it, python/_urna.so. Build the extension with the pyo3/extension-module feature:

git clone https://github.com/hoffresearch/urna
cd urna
git lfs pull   # the potion table; sh scripts/fetch_potion.sh downloads it without git-lfs
cargo build --release -p urna-python --features pyo3/extension-module
cp target/release/lib_urna.dylib python/_urna.so   # macOS
cp target/release/lib_urna.so python/_urna.so      # Linux

Without the feature the library links libpython and can crash under statically linked interpreters such as uv's standalone Python. The ImportError that urna.py raises when it finds no extension suggests a command without the feature: use the one above.

The loader tries from . import _urna first (the wheel layout), then _urna.so, _urna.abi3.so, _urna.dylib and lib_urna.dylib next to urna.py. It has no Windows candidate, so the checkout layout works on macOS and Linux only.

Put python/ on the path before importing, and install numpy and tokenizers for the potion embedder:

import sys

sys.path.insert(0, "python")
import urna
from builder import BuildConfig, Pipeline, chunk_text
from forge.embed_potion import potion_embedder

More on the development setup is on Contributing.

builder

python/builder.py chunks text, embeds it with a cache, and calls urna.build.

ChunkSpec

ChunkSpec(canonical_text: str, source_uri: str, byte_start: int, byte_end: int)

A frozen dataclass for one chunk before embedding. spec.chunk_id(chunker_version) returns the id that urna.chunk_id computes for it.

chunk_text

chunk_text(text: str, source_uri: str, *, max_chars: int = 512, overlap: int = 0) -> list[ChunkSpec]

Greedy character windows of max_chars characters, with an optional overlap of overlap characters. The spans are UTF-8 byte offsets into text, so urna cite returns spans that point back into the source. Empty text returns []. Raises ValueError when max_chars is 0 or less, or when overlap is negative or not smaller than max_chars.

EmbeddingCache

EmbeddingCache(path: str, model_key: str = "")

A SQLite cache of embeddings in the table embeddings_v2, keyed by (chunk_id, model_key). The model is part of the key because a chunk_id does not depend on the model: without it, re-embedding the same corpus with another model would reuse the old vectors.

MethodDescription
get(chunk_id, dim)the cached vector or None. Raises ValueError when the cached dimension differs from dim
put(chunk_id, embedding)store one vector and commit
close()close the database

An embeddings table from an older version is ignored, never migrated: those chunks are embedded again.

BuildConfig

A mutable dataclass with the arguments Pipeline.emit passes to urna.build. 22 fields, the first five required.

FieldTypeDefaultPassed to urna.build as
output_pathstrrequiredoutput_path
embedding_modelstrrequiredembedding_model
embedding_dimintrequiredembedding_dim
chunker_versionstrrequiredchunker_version
model_hashstrrequiredmodel_hash
titlestr or NoneNonetitle
versionstr or NoneNoneversion
descriptionstr or NoneNonedescription
licensestr or NoneNonelicense
reproducibleboolTruereproducible. Note that urna.build defaults to False
presetstr"exact"preset
text_encodingstr or NoneNonetext_encoding
dtypestr or NoneNonedtype
mrl_dimint or NoneNonemrl_dim
with_hnswbool or NoneNonewith_hnsw
with_bm25bool or NoneNonewith_bm25
with_graphboolFalsewith_graph
graph_top_mint8graph_top_m
drop_overlapboolFalseforces with_graph=True. It does not change chunking: pass overlap=0 to chunk_text yourself
hnsw_mint16hnsw_m
hnsw_ef_constructionint400hnsw_ef_construction
hnsw_seedint42hnsw_seed

created, authors, blob_refs, blob_data_paths, chunk_blob_spans and spaces have no field. To use them, call urna.build directly. provenance goes through emit.

Pipeline

Pipeline(cfg: BuildConfig, *, embedder, scratch_db: str | None = None)
MemberDescription
embedderany callable that takes a list of ChunkSpec and returns one vector per spec. PotionEmbedder and StaticEmbedder qualify
scratch_dbpath of an EmbeddingCache. The cache key is embedding_model, a NUL byte, and model_hash
add(spec), add_many(specs)queue chunks
emit(*, provenance=None)build the file and return output_path
close()close the cache

emit does the following, in order:

  1. Raises RuntimeError("pipeline has no chunks") when nothing was added.
  2. Embeds only the chunks missing from the cache. Raises RuntimeError when the embedder returns the wrong number of vectors or the wrong dimension.
  3. Deletes an existing file at output_path.
  4. Calls urna.build with the config. with_graph is True when with_graph or drop_overlap is set.
  5. Opens the new file and runs validate() in the same process.
import sys

sys.path.insert(0, "python")
from builder import BuildConfig, Pipeline, chunk_text
from forge.embed_potion import potion_embedder

emb = potion_embedder()
cfg = BuildConfig(
    output_path="my_corpus.urna",
    embedding_model=emb.embedding_model,
    embedding_dim=emb.embedding_dim,
    chunker_version="my-chunker/v1",
    model_hash=emb.model_hash(),
    preset="exact",
)
pipe = Pipeline(cfg, embedder=emb, scratch_db="cache.sqlite")
for source_uri, text in documents:   # your (uri, text) pairs
    pipe.add_many(chunk_text(text, source_uri))
pipe.emit()
pipe.close()

model_fingerprint

python/model_fingerprint.py computes a reproducible model_hash for a sentence-transformers or Hugging Face model directory. It is the only way in the repo to stamp a corpus built with such a model.

Importing it sets HF_HUB_OFFLINE, TRANSFORMERS_OFFLINE and HF_DATASETS_OFFLINE to 1 (when not already set), unless URNA_ALLOW_DOWNLOAD=1.

NameSignatureDescription
RELEVANT_FILEStuplethe 10 files hashed, in order: config.json, config_sentence_transformers.json, modules.json, sentence_bert_config.json, tokenizer.json, tokenizer_config.json, special_tokens_map.json, 1_Pooling/config.json, model.safetensors, pytorch_model.bin
PLACEHOLDER_HASHstrsha256: followed by 64 zeros
ModelFingerprintfrozen dataclassmodel_id, files_hash, tokenizer_hash, pooling_config_hash, embedding_dim, normalize_embeddings, and .to_dict()
compute_model_fingerprint(model_dir, *, model_id=None) -> ModelFingerprinthashes the files of RELEVANT_FILES that exist. model_id defaults to _name_or_path from config.json, then the directory name. Raises FileNotFoundError when model_dir is not a directory
fingerprint_to_model_hash(fp) -> strsha256: of fp.to_dict() as JSON with sorted keys and no whitespace
is_placeholder(model_hash) -> boolTrue for the zero placeholder
hf_cache_snapshot(model_id) -> Path or Nonethe snapshot in $HF_HOME/hub (default ~/.cache/huggingface) that refs/main points to, or the only snapshot. None when it is missing or ambiguous
resolve_model_dir(model_name_or_path) -> Patha directory as given, else the Hugging Face cache, else the path sentence-transformers loads from. Raises FileNotFoundError suggesting --model-path

Pass model_id equal to the embedding_model you write into the file. urna search-text fingerprints the model under the manifest name, so a different model_id gives a different hash and the gate refuses the query.

from model_fingerprint import compute_model_fingerprint, fingerprint_to_model_hash, resolve_model_dir

name = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
cfg.embedding_model = name
cfg.model_hash = fingerprint_to_model_hash(compute_model_fingerprint(resolve_model_dir(name), model_id=name))

forge.retrieve

python/forge/retrieve.py embeds a text query and calls UrnaFile.retrieve with the gate on.

retrieve(urnafile, query: str, k: int = 5, embedder=None, verify_model: bool = True) -> list[RetrieveHit]
ParameterDescription
urnafilean open UrnaFile
querythe question, as text
khits to return
embedderan embedder with the static protocol. None means potion_embedder()
verify_modelwhen True and the embedder has model_hash, passes expected_model_hash=embedder.model_hash()

build_demo(out_path) builds the Markdown files of python/forge/demo_corpus/ (without its README.md) with Pipeline, potion and the default chunk_text, preset="exact", reproducible=True and chunker_version="forge-demo/1", and returns out_path.

Run the module to build the demo in a temporary directory, ask one question and print cited answers:

python python/forge/retrieve.py

graph_context

neighbor_context(canonical_texts, ordinal: int, *, radius: int = 1, joiner: str = " ") -> str

Joins chunk ordinal with its radius neighbors on each side, in file order. It rebuilds context around a hit when the corpus was built without chunk overlap. Raises IndexError when ordinal is out of range.

UrnaFile has no accessor for all canonical texts, so you pass the texts yourself, in file order. The ordinal of a hit is db.chunk_ids().index(hit.chunk_id).

forge.forge_manifest readers

urna build --spec writes a sidecar <name>.manifest.json next to the .urna (see Build artifacts). These functions read that sidecar, not the manifest inside the file, which is UrnaFile.inspect()["manifest"].

FunctionDescription
manifest_items(manifest, need=())returns items[]. For a compact manifest (output.provenance = "minimal"), asking in need for a dropped field (image_path, label, media_uri) raises ValueError naming output.provenance = "standard"
frame_resolver(manifest)returns a function from an item to its global media stream frame. It uses a media://...#frame=N uri when the item has one, else the item's ordinal through frame_of_row and the media order permutation

The module also holds the write side the forge uses (write_manifest, redact_path, build_lock, check_lock, MANIFEST_SCHEMA_VERSION = 1, COMPACT_DROPPED).

convert_legacy.py

Converts the pre-v1 SQLite corpus layout (articles, text_blocks, blobs tables) into a v1 .urna:

python python/convert_legacy.py --src legacy.urna --dst out.urna --reproducible

convert(src, dst, *, reproducible) does the same from Python. It writes one chunk per article with source_uri = legacy://truw_ptbr/<id>, a span from 0 to the UTF-8 length of the text, labels and sources in the provenance, and chunker_version = "legacy/truw_ptbr_v0.1.0", then validates the file. It needs numpy, zstandard, and the sentence-transformers snapshot of the model in the Hugging Face cache (for the fingerprint). Errors exit through SystemExit, including a refusal to write the zero placeholder hash.

On this page