docsv0.5.1

Use urna from Python

Install the urna wheel, build a small .urna file from Python, embed a query offline with potion, search, retrieve cited chunks and resolve a citation.

This guide uses the urna wheel to build a three-chunk file, embed a question with the bundled potion table, search it, get cited text back, and resolve a citation. Everything after pip install runs offline.

Install

pip install "urna[embed]"

You need Python 3.12 or newer. The embed extra adds numpy and tokenizers, which the potion embedder needs. The wheel carries the potion table itself, so there is no urna setup step for Python.

The wheel also installs a command named urna, which is not the Rust CLI. If you use both, read The wheel's urna command first.

Build a small file

urna.build takes chunks you have already embedded. Save this as build_notes.py and run it:

import urna
from urna.embed_potion import potion_embedder

docs = [
    ("notes/offline.md", "The runtime never opens a socket. Queries are answered from the file."),
    ("notes/citations.md", "Every hit carries a urna:// citation that resolves to the stored text."),
    ("notes/format.md", "One .urna file holds chunks, embeddings, spans and indices."),
]

emb = potion_embedder()
vectors = emb.embed_texts([text for _, text in docs])

chunks = [
    {
        "canonical_text": text,
        "source_uri": uri,
        "byte_start": 0,
        "byte_end": len(text.encode("utf-8")),
        "embedding": vector,
    }
    for (uri, text), vector in zip(docs, vectors, strict=True)
]

urna.build(
    "notes.urna",
    emb.embedding_model,
    emb.embedding_dim,
    "notes/1",
    emb.model_hash(),
    chunks,
    title="notes",
    reproducible=True,
)
python build_notes.py

A few things in this call matter later:

  • emb.embedding_model and emb.model_hash() record which model made the vectors. They let the CLI answer this file, and they let you check the query model in Python.
  • "notes/1" is the chunker_version. It is part of every chunk_id, so change it when you change how you chunk.
  • Each chunk is one whole text, so byte_start is 0 and byte_end is its UTF-8 length. If your chunks are pieces of a larger document, compute their byte offsets into that document.
  • reproducible=True writes a fixed created date. Running the script again gives a byte-identical file.

Every parameter, the presets and the checks are on urna.build.

Open a file

import urna

db = urna.open("notes.urna")
print(db.n_embeddings, db.embedding_dim, db.dtype)
print(db.inspect()["manifest"]["embedding_model"])
print(db.model_hash)
print(db.content_hash)
3 256 float32
minishlab/potion-base-8M/v1
sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98
sha256:3b77075fbd3aea4eb06d48771d9ef1586276b904a07524f395dd247c37168b1b

urna.open validates the whole file before it returns and raises ValueError if anything is wrong. It takes a str: wrap a pathlib.Path with str(). There is no close(); the file stays mapped until the object is garbage collected.

The same works for a corpus you downloaded. Check embedding_model first: the bundled potion table can embed queries only for a corpus built with minishlab/potion-base-8M/v1. See Open a corpus you downloaded.

Embed a query

Queries must be embedded with the model that built the file. For a potion corpus, that is the embedder you already have:

from urna.embed_potion import potion_embedder

emb = potion_embedder()
query = emb.embed_texts(["does it need the network?"])[0]   # 256 floats, unit length

The first call loads the table from the wheel, then it is cached for the process. Other models are covered on Embedders.

Search with a vector

search takes any vector of the file's dimension: a list, a tuple or a 1-d numpy array.

for hit in db.search(query, 2):
    print(f"{hit.score:.4f} {hit.source_uri} {hit.index_type}")
0.2535 notes/offline.md exact
0.2190 notes/format.md exact

score is the exact cosine between the query and the stored vector. search compares against every row. On larger files built with HNSW, search_ann(query, k, ef) is the approximate path, and its scores are still exact cosine. The other paths are on UrnaFile.

A search hit carries ids, spans, hashes and a citation, but no text.

Retrieve cited chunks

retrieve searches along the route the file declares and attaches the stored text to each hit. Pass the embedder's hash so a query from the wrong model is refused:

hits = db.retrieve(query, 2, expected_model_hash=emb.model_hash())
for hit in hits:
    print(hit.text)
    print(f"  {hit.citation_id}")
    print(f"  score={hit.score:.4f} ({hit.rerank_source})")
The runtime never opens a socket. Queries are answered from the file.
  urna://sha256:3b77075fbd3aea4eb06d48771d9ef1586276b904a07524f395dd247c37168b1b/sha256:0c46a04662fc15bc1ec1d100dd6f54fbee1fd172d4b34338639376f31ff5e192
  score=0.2535 (full_precision)
One .urna file holds chunks, embeddings, spans and indices.
  urna://sha256:3b77075fbd3aea4eb06d48771d9ef1586276b904a07524f395dd247c37168b1b/sha256:165ebb7b21496454d55635d169568cc86275a8b5f882e7f00b6597c7f3829911
  score=0.2190 (full_precision)

rerank_source is full_precision here because the file stores float32 vectors. On a quantized file it says stored_precision.

Always pass expected_model_hash

In Python the model check is opt-in. Without expected_model_hash, a vector from another model of the same dimension returns hits with valid cosine scores and the wrong meaning. With it, a mismatch raises ValueError: model_hash mismatch: .... The search methods other than retrieve and search_space take no hash at all. See The model gate.

Hits are read-only objects that do not serialize to JSON. Copy the fields you need into a dict, as shown on SearchHit and RetrieveHit.

Resolve a citation

A citation_id is urna://<content_hash>/<chunk_id>. It names both the file content and the chunk, so it stays checkable after you store it.

UrnaFile has no cite method. To check in Python that a stored citation belongs to a file:

def cites_this_file(db, citation_id):
    content_hash, chunk_id = citation_id.removeprefix("urna://").split("/")
    return content_hash == db.content_hash and chunk_id in db.chunk_ids()

To get the text and span back, use the Rust CLI:

$ urna cite notes.urna 'urna://sha256:3b77075fbd3aea4eb06d48771d9ef1586276b904a07524f395dd247c37168b1b/sha256:0c46a04662fc15bc1ec1d100dd6f54fbee1fd172d4b34338639376f31ff5e192'
citation_id:  urna://sha256:3b77075fbd3aea4eb06d48771d9ef1586276b904a07524f395dd247c37168b1b/sha256:0c46a04662fc15bc1ec1d100dd6f54fbee1fd172d4b34338639376f31ff5e192
file:         notes.urna
file_hash:    sha256:aad2eb3df83e94c59efa42e232d07e5bf715f06548b7c0fe2553b8eb91dc6c8e
content_hash: sha256:3b77075fbd3aea4eb06d48771d9ef1586276b904a07524f395dd247c37168b1b
chunk_id:     sha256:0c46a04662fc15bc1ec1d100dd6f54fbee1fd172d4b34338639376f31ff5e192
source_uri:   notes/offline.md
byte_start:   0
byte_end:     69
text:
The runtime never opens a socket. Queries are answered from the file.

Because the file was built with potion's name and hash, the installed CLI can also answer it directly after urna setup: urna ask notes.urna "does it need the network?". See Citations and hashes for what each hash proves.

What needs the checkout

The wheel covers opening, searching, retrieving, validating and building from your own vectors, plus the potion embedder. Everything below lives only in a clone of the repo, with python/ on sys.path:

You wantUseWhere
chunk long documents with real byte spansbuilder.chunk_textcheckout
an embedding cache and a build pipelinebuilder.Pipeline, BuildConfigcheckout
a model_hash for a sentence-transformers modelmodel_fingerprintcheckout
embed with a registry model (wemm, clip, jina, siglip2)forge.model_registry.create_embeddercheckout, plus the model's dependencies
build from a corpus.toml specurna build --speccheckout plus the Rust binary
a zero-dependency lexical embedderforge.embed_defaultcheckout

These are documented on Builder pipeline (checkout only). To build from rows declared in a spec file instead of Python code, see Build from your own rows.

On this page