docsv0.5.1

urna.build

Reference for urna.build, which writes a .urna file from chunks you already embedded, with all 29 parameters, the presets, the chunk schema and its checks.

urna.build writes one .urna file from chunks you have already embedded. It does not embed anything: you pass the text, the spans and the vectors, plus the name and model_hash of the model that produced them. It is in the wheel, so it needs no repo checkout.

urna.build(
    output_path, embedding_model, embedding_dim, chunker_version, model_hash, chunks, *,
    title=None, version=None, created=None, description=None, authors=None, license=None,
    provenance=None, reproducible=False, preset="exact", text_encoding=None, dtype=None,
    mrl_dim=None, with_hnsw=None, with_bm25=None, with_graph=False, graph_top_m=8,
    blob_refs=None, blob_data_paths=None, chunk_blob_spans=None, spaces=None,
    hnsw_m=16, hnsw_ef_construction=400, hnsw_seed=42,
) -> str

29 parameters: the first six are required and may be passed by position, the other 23 are keyword-only. A seventh positional argument raises TypeError: build() takes 6 positional arguments but 7 were given. The return value is output_path.

Required parameters

ParameterTypeDescription
output_pathstrwhere to write. The parent directory must exist. An existing file is overwritten
embedding_modelstrthe model name written to the manifest. Must not be empty. The CLI picks its query embedder from this name
embedding_dimintthe length of every embedding. Must be greater than 0
chunker_versionstra label for how you chunked. Must not be empty. It is part of every chunk_id, so changing it changes every citation
model_hashstrsha256: followed by 64 hex digits, identifying the exact model that produced the vectors
chunkslist[dict]the chunks, at least one. Must be a list: a tuple raises TypeError. See Chunk dicts

Keyword parameters

ParameterTypeDefaultDescription
titlestr or NoneNonemanifest title
versionstr or NoneNonemanifest version
createdstr or NoneNonemanifest created. Nothing is written when None. Replaced by the epoch when reproducible=True
descriptionstr or NoneNonemanifest description
authorslist[str] or NoneNonemanifest authors. A bare str raises TypeError
licensestr or NoneNonemanifest license
provenancedict or NoneNonefree-form record, serialized with json.dumps into the provenance section ({} when None). Must be a JSON-serializable dict
reproducibleboolFalsewrites created = "1970-01-01T00:00:00Z", replacing any created you pass. Nothing else changes
presetstr"exact"picks the defaults of text_encoding, dtype, with_hnsw and with_bm25. See Presets
text_encoding"raw", "zstd" or NoneNoneencoding of the text sections. None takes the preset's
dtype"float32", "float16", "int8", "int4" or NoneNonestored embedding type. None takes the preset's. int4 needs the effective dimension divisible by 64
mrl_dimint or NoneNonekeeps the first mrl_dim components of each vector and L2-normalizes them again, before quantization and HNSW. The file's dimension becomes mrl_dim, and the manifest records mrl_dim and full_dim. Must satisfy 0 < mrl_dim <= embedding_dim
with_hnswbool or NoneNonebuild and attach an HNSW index. Sets index_type = "hnsw" and rerank_policy = "exact". None takes the preset's
with_bm25bool or NoneNonebuild a BM25 index over canonical_text. Sets supports_bm25. Never sets index_type = "hybrid". None takes the preset's
with_graphboolFalsebuild the chunk graph (section 0x0C): edges to the next chunk in both directions, plus up to graph_top_m semantic edges per chunk taken from HNSW. Builds a temporary HNSW when with_hnsw is off. Nothing is written for fewer than 2 chunks
graph_top_mint8most semantic edges per chunk, capped at 2 * hnsw_m
blob_refslist[dict] or NoneNonethe media blob table (section 0x14). See Media and named spaces
blob_data_pathslist or NoneNonefiles whose bytes are stored inside the .urna (section 0x17), one entry per blob_refs record. Requires blob_refs
chunk_blob_spanslist[dict] or NoneNonethe blob span overlay (section 0x16), exactly one entry per chunk
spaceslist[dict] or NoneNonenamed embedding spaces, at most 15
hnsw_mint16HNSW M, also the base of the graph edge cap. Not validated
hnsw_ef_constructionint400HNSW build beam. It is also the smallest search beam on this file
hnsw_seedint42seed for a deterministic HNSW build

Presets

A preset sets four defaults. An explicit text_encoding, dtype, with_hnsw or with_bm25 overrides the preset's value. Names are case-sensitive, and an unknown name raises ValueError.

presettext_encodingdtypewith_hnswwith_bm25Manifest index_type
exact (default)rawfloat32FalseFalseexact
compressedzstdfloat16FalseFalseexact
tinyzstdint8TrueFalsehnsw
nanozstdint4 (blocks of 64)TrueFalsehnsw
hybridzstdfloat32TrueTruehnsw

micro is a recipe, not a preset: preset="micro" raises ValueError. Build it with explicit knobs:

urna.build(..., text_encoding="zstd", dtype="int8", mrl_dim=256, with_hnsw=True)

Sizes, recall and when to pick each one are on Build presets and Presets and stored precision.

preset="hybrid" writes index_type = "hnsw"

urna.build never declares index_type = "hybrid". A hybrid build has a BM25 section, but the manifest says hnsw, so UrnaFile.retrieve, urna ask, urna retrieve and urna search-text take the HNSW route and never read BM25. Only an explicit UrnaFile.search_hybrid call uses it, and even then the final order is by cosine. See Build presets and Known limits.

Chunk dicts

Each item of chunks is a dict with five keys. Extra keys are ignored.

KeyTypeRule
canonical_textstrthe text that is stored, cited and indexed by BM25
source_uristrfree-form. Returned on every hit
byte_startint0 or more. A negative value raises OverflowError
byte_endintnot smaller than byte_start
embeddingsequence of numberslength embedding_dim, no NaN or Inf. A list, a tuple or a numpy row works. Should be L2-normalized

A missing key raises ValueError: chunks[i] missing <key>, and an item that is not a dict raises ValueError: chunks[i] is not a dict.

The file keeps the chunks in the order you pass them. That order is the order of UrnaFile.chunk_ids() and of the chunk graph.

byte_start and byte_end are stored as given, returned as offset_start and offset_end, and hashed into the chunk_id. If you want them to point into a source document, compute them as UTF-8 byte offsets into that document. See Offsets.

Model identity

embedding_model and model_hash tell every reader which model produced the vectors. The CLI uses them to pick a query embedder and to refuse a mismatch. For a file the installed urna ask can answer, embed with the bundled potion table and pass its identity exactly:

from urna.embed_potion import potion_embedder

emb = potion_embedder()
urna.build(path, emb.embedding_model, emb.embedding_dim, "my-chunker/1", emb.model_hash(), chunks)

The CLI picks the potion embedder when embedding_model starts with minishlab/potion, then checks the name, the dimension and the hash. Other models are covered on Embedders.

model_hash accepts uppercase hex, but embedders report lowercase, so an uppercase hash never matches at query time.

The zero placeholder hash is accepted at write time

urna.build accepts sha256: followed by 64 zeros. The file builds and validates. The refusal comes later: urna ask, urna retrieve and urna search-text stop on that file with a placeholder error (only search-text --skip-model-hash-check goes past it). Always pass the real hash of your embedder. See Known limits.

Media and named spaces

These parameters attach media blobs and extra embedding spaces. They are excluded from content_hash, so adding them does not change any citation. The concepts are on Media and named spaces.

blob_refs items, all four keys required:

KeyTypeRule
content_hashstrsha256:<64 hex> or the bare 64 hex digits of the blob bytes. Must decode to 32 bytes
original_uristrwhere the blob lives, or the sidecar that holds it
byte_lenintblob size in bytes
inlinedboolthe bytes are stored in this file

blob_data_paths items: a str path or None, one per blob_refs record, in the same order. Rust reads the files. An unreadable path raises ValueError: blob_data_paths[i] (<path>): <os error>. Peak memory is about twice the size of the media.

chunk_blob_spans items, all three keys required, one per chunk in chunk order:

KeyTypeRule
blob_ref_indexint or Noneposition in blob_refs, or None for a chunk with no blob
byte_startintstart inside the blob
byte_endintend inside the blob

When the overlay is present, hits and urna cite report these spans instead of the chunk spans.

spaces items:

KeyTypeRule
namestrunique within the file. Query it with UrnaFile.search_space(name, ...)
model_hashstrsha256:<hex> of the model that produced these vectors
dtypestroptional: "float32" (default), "float16", "int8" or "int4". int4 needs a dimension divisible by 64
vectorslist of rowsone row per chunk, all with the same dimension

Spaces are numbered from 1 in list order. More than 15 raises ValueError: at most 15 non-text spaces, got N.

What build checks

urna.build raises before writing when any of these fail:

  • preset, text_encoding and dtype are known names.
  • mrl_dim is between 1 and embedding_dim, and an int4 build has an effective dimension divisible by 64.
  • chunks is a non-empty list of dicts with the five keys.
  • Every embedding has length embedding_dim and no NaN or Inf.
  • byte_end is not smaller than byte_start.
  • embedding_model and chunker_version are not empty, and embedding_dim is greater than 0.
  • model_hash has the sha256: prefix and 64 hex digits.
  • The blob and space shapes above.

It does not check:

  • that embeddings are L2-normalized. The runtime normalizes queries, not stored vectors, so normalize them yourself.
  • zero vectors. They are accepted and score 0 against every query.
  • duplicate chunks. Use urna.chunk_id to find them first.
  • that the spans match the text length.
  • the zero placeholder model_hash, or uppercase hex in it.
  • the HNSW knob ranges (hnsw_m=0 is accepted).

The messages for each failure are on Errors.

What changes content_hash

content_hash is the first half of every citation, so anything that moves it changes every citation_id in the file. chunk_id depends only on text, source_uri, span and chunker_version, so it stays the same.

ChangeMoves content_hash
chunk text, source_uri, spans, chunker_versionyes
dtype (float32, float16, int8, int4), so the compressed, tiny and nano presetsyes
mrl_dimyes
provenance, including its key orderyes
with_hnsw=True, so the tiny, nano and hybrid presetsyes
text_encoding (raw or zstd)no
with_bm25 aloneno
with_graph, blobs, spacesno
title, created, description and other manifest-only fieldsno

Adding HNSW changes every citation

Turning on HNSW sets index_type and rerank_policy in the search contract, which is hashed into content_hash. The same chunks built with preset="exact" and with with_hnsw=True give two different content_hash values, so citations from one do not resolve against the other. Choose the index before you publish citations. See Known limits.

Writing and reproducibility

  • The file is written in one call at the end, straight to output_path, with no temporary file. A crash during the write can leave a partial file. Do not rebuild over a file another process has open: write to a new path, then move it into place with os.replace.
  • build holds the GIL for the whole build, HNSW included.
  • The output is deterministic for the same inputs. With created=None nothing time-dependent is written, so two builds match byte for byte even without reproducible. reproducible=True writes the epoch as created, so its file_hash differs from a build without it. content_hash is the same in both, because created is a manifest-only field.
  • with_graph and the HNSW build use hnsw_seed, so they are deterministic too.

Example

import urna
from urna.embed_potion import potion_embedder

docs = [
    ("notes/offline.md", "The runtime never opens a socket. Queries are answered from the file."),
    ("notes/citations.md", "Every hit carries a urna:// citation that resolves to the stored text."),
    ("notes/format.md", "One .urna file holds chunks, embeddings, spans and indices."),
]

emb = potion_embedder()
vectors = emb.embed_texts([text for _, text in docs])

chunks = [
    {
        "canonical_text": text,
        "source_uri": uri,
        "byte_start": 0,
        "byte_end": len(text.encode("utf-8")),
        "embedding": vector,
    }
    for (uri, text), vector in zip(docs, vectors, strict=True)
]

urna.build(
    "notes.urna",
    emb.embedding_model,
    emb.embedding_dim,
    "notes/1",
    emb.model_hash(),
    chunks,
    title="notes",
    reproducible=True,
)

The same flow step by step, with the query side, is in Use urna from Python.

On this page