docsv0.5.1

Data governance

What a .urna file stores, what it leaves out, how to remove or correct content, where processing runs, and which files a build leaves on disk.

A .urna file is meant to be copied: to laptops, edge nodes and machines with no network. When its content includes personal or sensitive data, that design has consequences. This page lists what the file stores, what it does not, how removal works, where processing happens and what a build leaves on disk besides the file. It describes the software; it does not assess your legal obligations.

A .urna file is a copy of the data

A .urna is a datastore, not a cache. It holds the text of every chunk in cleartext, next to embeddings derived from that text. The format has no encryption at rest: zstd is compression, not confidentiality. A file built over personal data is a copy of that data and needs the same controls as its source, including full-disk encryption (FileVault, LUKS or equivalent) on every volume that holds it.

What the file contains

Six sections are required in every file. The others appear only when the build asked for them.

PartWhat it stores
chunks_canonical (0x02)The canonical text of every chunk. This is the text that ask, retrieve and cite return.
chunks_original_spans (0x03)A source_uri and a byte span per chunk. The URI is whatever the builder wrote: a source_uri column or item://<corpus name>/<key> in a forge build, media://... for media rows, any string in a direct urna.build call, local paths included.
chunk_ids (0x01)One sha256: id per chunk, computed over the text, the URI, the span and chunker_version. Anyone holding a candidate text and its source can recompute the id and confirm a match.
embeddings (0x04)One vector per chunk, a derived representation of the text.
provenance (0x05)Any JSON the builder passed. urna build --spec writes {"dataset": <corpus name>, "corpus_input_hash": <sha256>}; urna.build writes {} unless you pass a dict.
search_contract (0x06)The metric, score type, normalization, index type and rerank policy.
Manifestembedding_model, model_hash, chunker_version, dimensions and dtype, plus the optional title, version, created, description, license, authors and any extra keys the builder added.
bm25_index (0x08, optional)The lowercased token vocabulary of the canonical text with term frequencies per chunk: a second copy of the words, without their order.
hnsw_index (0x07), graph_adjacency (0x0C) (optional)Neighbor lists between chunk ordinals. No text.
blob_refs (0x14, optional)Per media record: the SHA-256 of the original bytes, the original URI, the byte length and whether the bytes are inlined.
blob_data (0x17, optional)The media bytes themselves, as encoded by the build, when the spec sets [output] embed_media = true.
space_table (0x15) and bands (optional)Per named space: name, dimension, dtype, model_hash and one vector per chunk (image vectors, for example).

The layout of every section is in sections, and the manifest fields in manifest.

What the file does not contain

  • Encryption, access control or per-chunk permissions. Anyone who can read the file can read all of it.
  • A signature or a verified author. The hashes prove integrity, not origin, and authors is free text. See security.
  • Model weights. The manifest records the model name and its model_hash; the query embedder lives outside the file.
  • A record of queries. The runtime maps the file read-only, so querying never changes it.
  • A delete or tombstone mechanism. Chunks cannot be marked as removed; the edit_journal section id is reserved and nothing writes it.
  • The original source files, with one exception: inlined media in blob_data. For text rows, only the canonical text and the URI are stored.

Remove or correct content

A distributed file cannot be edited in place. To erase or correct a chunk, fix the source rows and rebuild. The rebuild changes the file:

  • content_hash covers the six canonical sections, so any changed chunk gives the new file a new content_hash and a new file_hash.
  • A citation is urna://<content_hash>/<chunk_id>. urna cite rejects a citation whose content_hash does not match the file, so every citation issued from the old build stops resolving against the new one.
  • In a forge build the byte span of a row is its ordinal after sorting, and the span enters the chunk_id. Removing a row shifts the chunk_id of every row after it.

Copies already shipped are not recalled. The runtime never opens a socket, so there is no channel to reach them. Before you distribute a file that holds personal data, plan the removal process:

  • Version the corpus. Each build has its own content_hash and file_hash, which identify it exactly.
  • Publish the hashes of superseded builds, and require operators to pull the current build and delete the old copies.
  • Treat the embeddings and the BM25 vocabulary as derived personal data, in scope for the same requests as the text.
  • Record consent and provenance for third-party content before it goes into a build.

Data-subject rights such as erasure and rectification (GDPR articles 16 and 17, LGPD article 18) cannot be served by editing a copy that has left your hands. For special-category data, such as health records, confirm with counsel that distributing immutable copies is compatible with the rights that apply before you ship.

Citations can change without a content change

Attaching an HNSW index writes index_type = "hnsw" and rerank_policy = "exact" into the canonical search_contract section. An exact build and a tiny, nano or hybrid build of the same rows therefore have different content_hash values and cite differently, even with identical text. Keep the preset fixed across rebuilds when you need old citations to line up with new ones. See known limits.

Some changes leave content_hash and every citation unchanged: the text encoding (raw or zstd), the graph, media sections, named spaces, and manifest-only fields such as title or created. The full table is in citations and hashes.

Where processing happens

Building and querying run on your machine. The Rust runtime never opens a socket and answers from the memory-mapped file. The query embedders run as local Python processes, and the Python embedders and builders set the Hugging Face offline variables unless you opt into downloads with URNA_ALLOW_DOWNLOAD=1. Media encoders (ffmpeg, avifenc, cjxl) run as local subprocesses.

The exceptions are the installers, urna setup and explicit download opt-ins. They are listed in where urna opens a network connection, and the design is explained in offline by construction.

What a build leaves on disk

Personal data can outlive a .urna in these places. Remove them together with the file.

LocationWritten byContents
<output dir>/<name>.manifest.jsonurna build --specPer-item keys and ordinals; with provenance = "standard" or "full" also labels, media URIs and image paths; with "full" also the SQL query and unredacted paths
<output dir>/<name>.build.lock.jsonurna build --specOS and machine, package versions, tool paths and hashes, model hashes, and the resolved spec after ${VAR} expansion (source paths included)
<output dir>/<name>.media/urna build --spec with [media]The encoded media. It is what gets served in sidecar mode, and the build cache when media is inlined
<output dir>/.forge-state/, .tmp/urna build --specMedia stage state; staging for files before the final rename, empty after success
<cache root>/embed/<preset>/*.npzurna build --specThe vectors of every row (text vectors per row, image vectors per unique frame), with .sha256 and .lock files. No text
<cache root>/models/urna build --specmodel_hash probe files
The scratch_db pathbuilder.Pipeline (checkout only)A SQLite cache of embeddings keyed by chunk_id and model
$HF_HOMEsentence-transformers presetsModel weights from the Hugging Face cache
<data root>/urna/forge, <data root>/urna/venvinstallers, urna setupThe embedder payload and the Python env. No corpus data

The output dir is [output] dir (default out) or --out-dir. The cache root is --cache-dir when given, else [output] cache_dir, else URNA_CACHE_DIR, else ${XDG_CACHE_HOME:-~/.cache}/urna. With provenance = "standard" (the default), a path under your home directory is written with ~. With "minimal", items keep only key and ordinal, file names lose their directories, and the SQL query is dropped. The provenance mode changes only the sidecar manifest: the .urna file is byte-identical across the three modes. The data roots are listed in paths; urna setup --uninstall removes the payload and the env, never corpus files.

Corpus licensing

A distributed .urna carries the license of the content it embeds. The manifest license field is free text and nothing enforces it. The repository code is MIT; that license does not cover corpus content. Redistributing a built file means honoring the most restrictive upstream license and keeping its attribution requirements.

The repository's demo datasets combine upstream sources with different licenses. Their bill of materials is in data/demo/Instructions.md. The quickstart corpus and python/forge/demo_corpus are original text released under CC0 1.0. For anything distributed widely, prefer a permissive or CC0 corpus.

On this page