urna.build
Reference for urna.build, which writes a .urna file from chunks you already embedded, with all 29 parameters, the presets, the chunk schema and its checks.
urna.build writes one .urna file from chunks you have already embedded. It does not embed anything: you pass the text, the spans and the vectors, plus the name and model_hash of the model that produced them. It is in the wheel, so it needs no repo checkout.
urna.build(
output_path, embedding_model, embedding_dim, chunker_version, model_hash, chunks, *,
title=None, version=None, created=None, description=None, authors=None, license=None,
provenance=None, reproducible=False, preset="exact", text_encoding=None, dtype=None,
mrl_dim=None, with_hnsw=None, with_bm25=None, with_graph=False, graph_top_m=8,
blob_refs=None, blob_data_paths=None, chunk_blob_spans=None, spaces=None,
hnsw_m=16, hnsw_ef_construction=400, hnsw_seed=42,
) -> str29 parameters: the first six are required and may be passed by position, the other 23 are keyword-only. A seventh positional argument raises TypeError: build() takes 6 positional arguments but 7 were given. The return value is output_path.
Required parameters
| Parameter | Type | Description |
|---|---|---|
output_path | str | where to write. The parent directory must exist. An existing file is overwritten |
embedding_model | str | the model name written to the manifest. Must not be empty. The CLI picks its query embedder from this name |
embedding_dim | int | the length of every embedding. Must be greater than 0 |
chunker_version | str | a label for how you chunked. Must not be empty. It is part of every chunk_id, so changing it changes every citation |
model_hash | str | sha256: followed by 64 hex digits, identifying the exact model that produced the vectors |
chunks | list[dict] | the chunks, at least one. Must be a list: a tuple raises TypeError. See Chunk dicts |
Keyword parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
title | str or None | None | manifest title |
version | str or None | None | manifest version |
created | str or None | None | manifest created. Nothing is written when None. Replaced by the epoch when reproducible=True |
description | str or None | None | manifest description |
authors | list[str] or None | None | manifest authors. A bare str raises TypeError |
license | str or None | None | manifest license |
provenance | dict or None | None | free-form record, serialized with json.dumps into the provenance section ({} when None). Must be a JSON-serializable dict |
reproducible | bool | False | writes created = "1970-01-01T00:00:00Z", replacing any created you pass. Nothing else changes |
preset | str | "exact" | picks the defaults of text_encoding, dtype, with_hnsw and with_bm25. See Presets |
text_encoding | "raw", "zstd" or None | None | encoding of the text sections. None takes the preset's |
dtype | "float32", "float16", "int8", "int4" or None | None | stored embedding type. None takes the preset's. int4 needs the effective dimension divisible by 64 |
mrl_dim | int or None | None | keeps the first mrl_dim components of each vector and L2-normalizes them again, before quantization and HNSW. The file's dimension becomes mrl_dim, and the manifest records mrl_dim and full_dim. Must satisfy 0 < mrl_dim <= embedding_dim |
with_hnsw | bool or None | None | build and attach an HNSW index. Sets index_type = "hnsw" and rerank_policy = "exact". None takes the preset's |
with_bm25 | bool or None | None | build a BM25 index over canonical_text. Sets supports_bm25. Never sets index_type = "hybrid". None takes the preset's |
with_graph | bool | False | build the chunk graph (section 0x0C): edges to the next chunk in both directions, plus up to graph_top_m semantic edges per chunk taken from HNSW. Builds a temporary HNSW when with_hnsw is off. Nothing is written for fewer than 2 chunks |
graph_top_m | int | 8 | most semantic edges per chunk, capped at 2 * hnsw_m |
blob_refs | list[dict] or None | None | the media blob table (section 0x14). See Media and named spaces |
blob_data_paths | list or None | None | files whose bytes are stored inside the .urna (section 0x17), one entry per blob_refs record. Requires blob_refs |
chunk_blob_spans | list[dict] or None | None | the blob span overlay (section 0x16), exactly one entry per chunk |
spaces | list[dict] or None | None | named embedding spaces, at most 15 |
hnsw_m | int | 16 | HNSW M, also the base of the graph edge cap. Not validated |
hnsw_ef_construction | int | 400 | HNSW build beam. It is also the smallest search beam on this file |
hnsw_seed | int | 42 | seed for a deterministic HNSW build |
Presets
A preset sets four defaults. An explicit text_encoding, dtype, with_hnsw or with_bm25 overrides the preset's value. Names are case-sensitive, and an unknown name raises ValueError.
preset | text_encoding | dtype | with_hnsw | with_bm25 | Manifest index_type |
|---|---|---|---|---|---|
exact (default) | raw | float32 | False | False | exact |
compressed | zstd | float16 | False | False | exact |
tiny | zstd | int8 | True | False | hnsw |
nano | zstd | int4 (blocks of 64) | True | False | hnsw |
hybrid | zstd | float32 | True | True | hnsw |
micro is a recipe, not a preset: preset="micro" raises ValueError. Build it with explicit knobs:
urna.build(..., text_encoding="zstd", dtype="int8", mrl_dim=256, with_hnsw=True)Sizes, recall and when to pick each one are on Build presets and Presets and stored precision.
preset="hybrid" writes index_type = "hnsw"
urna.build never declares index_type = "hybrid". A hybrid build has a BM25 section, but the manifest says hnsw, so UrnaFile.retrieve, urna ask, urna retrieve and urna search-text take the HNSW route and never read BM25. Only an explicit UrnaFile.search_hybrid call uses it, and even then the final order is by cosine. See Build presets and Known limits.
Chunk dicts
Each item of chunks is a dict with five keys. Extra keys are ignored.
| Key | Type | Rule |
|---|---|---|
canonical_text | str | the text that is stored, cited and indexed by BM25 |
source_uri | str | free-form. Returned on every hit |
byte_start | int | 0 or more. A negative value raises OverflowError |
byte_end | int | not smaller than byte_start |
embedding | sequence of numbers | length embedding_dim, no NaN or Inf. A list, a tuple or a numpy row works. Should be L2-normalized |
A missing key raises ValueError: chunks[i] missing <key>, and an item that is not a dict raises ValueError: chunks[i] is not a dict.
The file keeps the chunks in the order you pass them. That order is the order of UrnaFile.chunk_ids() and of the chunk graph.
byte_start and byte_end are stored as given, returned as offset_start and offset_end, and hashed into the chunk_id. If you want them to point into a source document, compute them as UTF-8 byte offsets into that document. See Offsets.
Model identity
embedding_model and model_hash tell every reader which model produced the vectors. The CLI uses them to pick a query embedder and to refuse a mismatch. For a file the installed urna ask can answer, embed with the bundled potion table and pass its identity exactly:
from urna.embed_potion import potion_embedder
emb = potion_embedder()
urna.build(path, emb.embedding_model, emb.embedding_dim, "my-chunker/1", emb.model_hash(), chunks)The CLI picks the potion embedder when embedding_model starts with minishlab/potion, then checks the name, the dimension and the hash. Other models are covered on Embedders.
model_hash accepts uppercase hex, but embedders report lowercase, so an uppercase hash never matches at query time.
The zero placeholder hash is accepted at write time
urna.build accepts sha256: followed by 64 zeros. The file builds and validates. The refusal comes later: urna ask, urna retrieve and urna search-text stop on that file with a placeholder error (only search-text --skip-model-hash-check goes past it). Always pass the real hash of your embedder. See Known limits.
Media and named spaces
These parameters attach media blobs and extra embedding spaces. They are excluded from content_hash, so adding them does not change any citation. The concepts are on Media and named spaces.
blob_refs items, all four keys required:
| Key | Type | Rule |
|---|---|---|
content_hash | str | sha256:<64 hex> or the bare 64 hex digits of the blob bytes. Must decode to 32 bytes |
original_uri | str | where the blob lives, or the sidecar that holds it |
byte_len | int | blob size in bytes |
inlined | bool | the bytes are stored in this file |
blob_data_paths items: a str path or None, one per blob_refs record, in the same order. Rust reads the files. An unreadable path raises ValueError: blob_data_paths[i] (<path>): <os error>. Peak memory is about twice the size of the media.
chunk_blob_spans items, all three keys required, one per chunk in chunk order:
| Key | Type | Rule |
|---|---|---|
blob_ref_index | int or None | position in blob_refs, or None for a chunk with no blob |
byte_start | int | start inside the blob |
byte_end | int | end inside the blob |
When the overlay is present, hits and urna cite report these spans instead of the chunk spans.
spaces items:
| Key | Type | Rule |
|---|---|---|
name | str | unique within the file. Query it with UrnaFile.search_space(name, ...) |
model_hash | str | sha256:<hex> of the model that produced these vectors |
dtype | str | optional: "float32" (default), "float16", "int8" or "int4". int4 needs a dimension divisible by 64 |
vectors | list of rows | one row per chunk, all with the same dimension |
Spaces are numbered from 1 in list order. More than 15 raises ValueError: at most 15 non-text spaces, got N.
What build checks
urna.build raises before writing when any of these fail:
preset,text_encodinganddtypeare known names.mrl_dimis between 1 andembedding_dim, and anint4build has an effective dimension divisible by 64.chunksis a non-empty list of dicts with the five keys.- Every embedding has length
embedding_dimand no NaN or Inf. byte_endis not smaller thanbyte_start.embedding_modelandchunker_versionare not empty, andembedding_dimis greater than 0.model_hashhas thesha256:prefix and 64 hex digits.- The blob and space shapes above.
It does not check:
- that embeddings are L2-normalized. The runtime normalizes queries, not stored vectors, so normalize them yourself.
- zero vectors. They are accepted and score 0 against every query.
- duplicate chunks. Use
urna.chunk_idto find them first. - that the spans match the text length.
- the zero placeholder
model_hash, or uppercase hex in it. - the HNSW knob ranges (
hnsw_m=0is accepted).
The messages for each failure are on Errors.
What changes content_hash
content_hash is the first half of every citation, so anything that moves it changes every citation_id in the file. chunk_id depends only on text, source_uri, span and chunker_version, so it stays the same.
| Change | Moves content_hash |
|---|---|
chunk text, source_uri, spans, chunker_version | yes |
dtype (float32, float16, int8, int4), so the compressed, tiny and nano presets | yes |
mrl_dim | yes |
provenance, including its key order | yes |
with_hnsw=True, so the tiny, nano and hybrid presets | yes |
text_encoding (raw or zstd) | no |
with_bm25 alone | no |
with_graph, blobs, spaces | no |
title, created, description and other manifest-only fields | no |
Adding HNSW changes every citation
Turning on HNSW sets index_type and rerank_policy in the search contract, which is hashed into content_hash. The same chunks built with preset="exact" and with with_hnsw=True give two different content_hash values, so citations from one do not resolve against the other. Choose the index before you publish citations. See Known limits.
Writing and reproducibility
- The file is written in one call at the end, straight to
output_path, with no temporary file. A crash during the write can leave a partial file. Do not rebuild over a file another process has open: write to a new path, then move it into place withos.replace. buildholds the GIL for the whole build, HNSW included.- The output is deterministic for the same inputs. With
created=Nonenothing time-dependent is written, so two builds match byte for byte even withoutreproducible.reproducible=Truewrites the epoch ascreated, so itsfile_hashdiffers from a build without it.content_hashis the same in both, becausecreatedis a manifest-only field. with_graphand the HNSW build usehnsw_seed, so they are deterministic too.
Example
import urna
from urna.embed_potion import potion_embedder
docs = [
("notes/offline.md", "The runtime never opens a socket. Queries are answered from the file."),
("notes/citations.md", "Every hit carries a urna:// citation that resolves to the stored text."),
("notes/format.md", "One .urna file holds chunks, embeddings, spans and indices."),
]
emb = potion_embedder()
vectors = emb.embed_texts([text for _, text in docs])
chunks = [
{
"canonical_text": text,
"source_uri": uri,
"byte_start": 0,
"byte_end": len(text.encode("utf-8")),
"embedding": vector,
}
for (uri, text), vector in zip(docs, vectors, strict=True)
]
urna.build(
"notes.urna",
emb.embedding_model,
emb.embedding_dim,
"notes/1",
emb.model_hash(),
chunks,
title="notes",
reproducible=True,
)The same flow step by step, with the query side, is in Use urna from Python.
SearchHit and RetrieveHit
Fields of urna.SearchHit and urna.RetrieveHit, the hit objects urna returns from Python, what each field means, and how to turn a hit into a dict.
Embedders
The offline potion embedder in the urna wheel, the lexical floor and registry adapters in the repo checkout, and the two embedder protocols they follow.