docsv0.5.1

Glossary

Short definitions of the terms used across the urna docs, from blob and BM25 to top-k and vector database, each linked to the page that explains it.

The terms used across these docs, in alphabetical order. Each entry links to the page that covers it in depth. Code identifiers are spelled the way the file format and the tools spell them.

blob

The bytes of one media asset, such as an image file or a segment of a video stream. The blob_refs section lists each blob by its SHA-256, original URI and length; the bytes live next to the file in <name>.media/, or inside it when the build sets embed_media. See Media and named spaces.

BM25

A lexical ranking method that scores chunks by the query words they contain. urna stores an optional BM25 index in its own section and uses it only in hybrid search, where it can add candidates but never overrides the cosine order. See Search paths and the exact rerank.

canonical text

The stored text of a chunk. It is what the BM25 index reads and what ask, retrieve and cite print, byte for byte. See Citations and hashes.

chunk

The unit urna stores and returns: a canonical text, the source_uri and span it came from, and its embedding. Each chunk is identified by its chunk_id. See The .urna file.

chunk_id

The identity of a chunk: sha256: over the canonical text, the source_uri, the span and the chunker_version. The same chunk built on two machines gets the same chunk_id. See Citations and hashes.

citation

A urna://<content_hash>/<chunk_id> id attached to every hit. urna cite resolves it back to the stored text and span, and refuses a citation whose content_hash does not match the file. See Citations and hashes.

content_hash

A SHA-256 over the six required sections after decoding: chunk ids, canonical texts, spans, embeddings, provenance and search contract. It stays the same when the text is stored raw or compressed, and changes with any chunk, the embedding dtype, truncation or the search contract. See Hashes and citation ids.

corpus

A collection of documents or items prepared for querying. In urna a corpus is one .urna file, plus its media directory when the media is not embedded. See Your first corpus.

cosine

The similarity measure urna scores with, between -1 and 1. Every score urna returns is an exact cosine, recomputed against the stored vectors. It is not a probability that an answer is correct. See Search paths and the exact rerank.

embedding

A vector a model computes to represent a piece of text or an image. urna stores embeddings L2-normalized, at the file's dtype (float32, float16, int8 or int4). See The .urna file.

file_hash

The SHA-256 of the whole .urna file, the same value sha256sum prints. It changes with any byte of the file, including metadata that content_hash ignores. See Hashes and citation ids.

forge

The Python build tooling in the repository (python/forge/ and python/tools/urna_forge.py) that urna build --spec launches. It ships in no release artifact, so building needs a checkout. See Build from your own rows.

hash

A fixed-size value computed over bytes, used to detect changes. urna's checksums and hashes are unkeyed: they prove a file's bytes are intact, not who made it. See Security.

hit@k

Whether an expected item appears in the first k results of a query. The image benchmarks and the media quality gate use hit@1. See Benchmarks.

HNSW

An approximate nearest-neighbor index: a layered graph of vectors that proposes candidates quickly. urna rescores every HNSW candidate by exact cosine before returning it. See Search paths and the exact rerank.

index_type

The field of the manifest and search contract that declares how a file is meant to be searched: exact, hnsw or hybrid. ask, retrieve and search-text pick their route from it. See Search paths and the exact rerank.

L1, L2, L3

The three reproduction levels of a build. L1 is the same top-k results anywhere, L2 is each vector's cosine within 1e-5 on the same device class, and L3 is a byte-identical file_hash, claimable only under a matching build lock. Only L3 has tooling behind it (--rebuild-only and the lock); L1 and L2 are stated targets. See Reproducible builds.

manifest

The JSON block inside a .urna file that records the model, dimension, dtype, metric, index_type, model_hash and capabilities. It is covered by the file hash, not by content_hash. The forge also writes a separate <name>.manifest.json next to the file. See Manifest.

mmap

Memory mapping: the operating system maps a file into the process's address space and reads pages on demand. The runtime maps a .urna file read-only instead of loading it. See The .urna file.

model gate

The check that a query was embedded by the same model the corpus was built with: model name, dimension and model_hash must match the manifest. The CLI always runs it on text queries; in Python it runs only when you pass expected_model_hash. See The model gate.

model_hash

A sha256: fingerprint of the embedding model (its files, tokenizer, pooling, dimension and normalization) recorded in the manifest. The model gate compares it with the query embedder's own hash. See The model gate.

MRL

Matryoshka representation learning: training a model so that the first dimensions of a vector carry most of its meaning. urna's mrl_dim keeps that prefix and re-normalizes it; it pays off only on models trained this way. See Presets and stored precision.

payload

The embedder payload, urna-embedder-payload.tar.gz, that urna setup and the one-line installers unpack into the urna data directory. It holds the potion query embedder, its table and two helper modules; it has no registry-model embedder and no forge. See Installation.

potion

minishlab/potion-base-8M, urna's default embedding model: a static 256-dimension table that runs with numpy and tokenizers, no torch and no network. It is distilled from an English model. See Choose and bring embedding models.

preset

A named bundle of build settings. The file presets exact, compressed, tiny, nano and hybrid set the text encoding, the embedding dtype and the indices; micro is a recipe, not a preset value. Model presets such as potion or siglip2 name entries of the model registry. See Presets and stored precision and Model registry.

provenance

A record of where the data came from. In a .urna file it is a required JSON section, part of content_hash; in a forge build, [output] provenance also sets how much the sidecar manifest records. See Data governance.

quantization

Storing vectors at lower precision to save space: float16, int8 or int4 in urna. Scores on a quantized file are real cosine at the stored precision. See Presets and stored precision.

RAG

Retrieval-augmented generation: an application retrieves context first, then asks a language model to answer with it. urna covers the retrieval step with cited spans; generation belongs to the application. See Give a corpus to an agent.

recall

The share of reference results a search returns, measured against a stated ruler. urna's preset tables measure it against the float32 ranking of near-duplicate queries, which shows rank stability, not answer quality. See Benchmarks.

rerank

Rescoring candidates by exact cosine against the stored vectors. Every non-exact path in urna (HNSW, hybrid, graph) ends with it, so their scores equal the exact scores. See Search paths and the exact rerank.

rerank source

The vectors the rerank reads: full precision when the file stores float32, stored precision when it stores float16, int8 or int4. retrieve reports it as rerank_source and ask --disclose explain prints it. See Search paths and the exact rerank.

search contract

A required section holding metric, score_type, normalize, index_type and rerank_policy. The reader checks it against the manifest, and it is part of content_hash. See The .urna file.

section

A typed block of a .urna file, listed in the section table with its id, encoding, offset, size and checksum. Six sections are required; the others (indices, graph, media, spaces) are optional. See Sections.

sidecar

A file that sits next to a .urna file: the forge's <name>.manifest.json and <name>.build.lock.json, and the <name>.media/ directory. See Build artifacts.

space

A named set of vectors stored in a file alongside the default embeddings, one per extra model or dimension, such as clip-vit-b32 or wemm-2b@256. urna search-space queries one by name. See Media and named spaces.

span

The byte_start and byte_end stored with each chunk's source_uri. Their meaning depends on the builder: byte offsets into the source text, the row position in a forge build, or a range inside a media blob. See Citations and hashes.

spec

The build spec: a TOML, JSON or YAML file (usually corpus.toml) that tells urna build which rows to read, how to render their text, which models to embed with and where to write. See The spec file.

stored precision

A score computed from vectors stored below float32. urna labels such scores "real cosine at stored precision" so a quantized result is never presented as full precision. See Presets and stored precision.

top-k

The k highest-scoring results of a search. Every search verb takes -k, with a default of 10. See Ask and retrieve from the terminal.

vector database

A store that finds items by comparing the vectors computed from their content. urna is one that fits in a single file. See Introduction.

On this page