Citations and hashes
How a urna:// citation is built from content_hash and chunk_id, which changes move it, what the offsets mean, and what file_hash covers.
Every hit urna returns carries a citation of the form urna://content_hash/chunk_id. The citation names the corpus content and one chunk inside it, both by SHA-256, so it resolves to the same stored text on any machine that has the same corpus. This page covers the hashes behind it, what changes them, and what the span offsets in a hit mean.
Anatomy of a citation
This is a real citation from the quickstart corpus:
urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be- The first part is the file's
content_hash, a SHA-256 over the six canonical sections. - The second part is the
chunk_id, a SHA-256 over one chunk's text and origin.
Both are written as sha256: followed by 64 lowercase hex digits. Resolve a citation with urna cite:
urna cite examples/quickstart/out/quickstart.urna 'urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be'citation_id: urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
file: examples/quickstart/out/quickstart.urna
file_hash: sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
content_hash: sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df
chunk_id: sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
source_uri: demo/03-citations.md
byte_start: 7
byte_end: 8
text:
because the citation points at content, two people who build the same logical corpus on two machines get the same citation, and a stored corpus and a compressed one cite identically. resolving a citation returns the exact canonical text and the original byte span it came from, which is what lets an agent quote a source it can prove.cite first recomputes the file's content_hash and refuses the citation if it does not match, so a citation never resolves against a different corpus by accident. The text it prints is the stored canonical text of the chunk, not a reopen of the original source file.
chunk_id
The chunk_id is the SHA-256 of this preimage, with every string length-prefixed:
"urna:chunk_id:v1\n"
u32 LE len(canonical_text) canonical_text bytes
u32 LE len(source_uri) source_uri bytes
u64 LE byte_start
u64 LE byte_end
u32 LE len(chunker_version) chunker_version bytesThe embedding is not part of it. The same text from the same place, cut by the same chunker, gets the same chunk_id whatever model embedded it. Changing any of the five inputs (text, uri, either offset, or the manifest's chunker_version) gives a new id.
content_hash
The content_hash is the SHA-256 of the six canonical sections, taken in this fixed order: chunk_ids, chunks_canonical, chunks_original_spans, embeddings, provenance, search_contract. For each section the hash takes the name length (u32 LE), the name, the decoded length (u64 LE) and the decoded bytes.
"Decoded" means after the wire codec: a section stored as zstd or intpack hashes the same as its raw form. Quantized embeddings are the exception, because they are hashed as stored: a float16 corpus and a float32 corpus of the same chunks have different content_hash values. The manifest is not part of the preimage.
file_hash has two meanings
| Value | Covers | Where you see it |
|---|---|---|
| Footer hash | SHA-256 of every byte before the 40-byte footer | Stored as 32 raw bytes in the footer and checked on open; never printed |
Reported file_hash | SHA-256 of the whole file, footer included | validate, stats, inspect, cite, every hit, retrieve JSON |
The reported file_hash is the same number sha256sum or shasum -a 256 prints for the file:
shasum -a 256 examples/quickstart/out/quickstart.urnae4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832 examples/quickstart/out/quickstart.urnaThe header checksum and the per-section checksums are shorter: the first 8 bytes of a SHA-256, stored raw. urna inspect prints a section checksum as 16 hex digits with no prefix. They detect a corrupted header or payload; they are not identifiers.
Use content_hash to ask "is this the same corpus content?" and file_hash to ask "is this the same file, byte for byte?". Two files can share a content_hash and differ in file_hash (for example, a different title in the manifest, or a BM25 index added).
What moves content_hash
Anything that moves content_hash changes every citation into the file, because the hash is the first half of each citation.
| Change | Moves content_hash | Why |
|---|---|---|
Raw vs zstd text, or which text codec the writer picks | No | Codecs decode to identical bytes |
Embedding dtype (float32, float16, int8, int4) | Yes | Quantized bytes are hashed as stored |
mrl_dim truncation | Yes | The truncated vectors are hashed |
Chunk text, source_uri, span, chunker_version | Yes | They are in chunk_ids, chunks_canonical and chunks_original_spans |
| Provenance JSON, including key order | Yes | provenance is canonical |
| Attaching an HNSW index | Yes | It rewrites index_type and rerank_policy, which live in search_contract |
Declaring index_type = "hybrid" | Yes | Same mechanism |
| Adding a BM25 index alone | No | Only a manifest capability changes |
| Graph, blob tables, blob data, space table and bands | No | Not canonical; only manifest flags change |
Manifest-only fields (title, created, capabilities) | No | The manifest is covered by the file hash only |
In practice: rebuilding with a different preset or dtype gives new citations. Changing only the text compression does not.
Adding HNSW changes every citation
Attaching an HNSW index sets index_type = "hnsw" and rerank_policy = "exact", and the writer copies both into the canonical search_contract section. An exact build and an HNSW build of the same chunks therefore have different content_hash values and different citations. This also separates exact from the tiny, nano and hybrid presets, and any build with with_hnsw=True. A comment in the format crate says adding an optional section never invalidates citations; that holds for the graph, blobs and spaces, not for HNSW. See Known limits.
What the offsets mean
byte_start and byte_end come from chunks_original_spans and mean whatever the builder wrote there. urna checks only that byte_end is not below byte_start.
| Built with | Offsets are |
|---|---|
builder.chunk_text (repo checkout) | Real UTF-8 byte offsets into the source text |
| The Python examples in the repo | 0 to the UTF-8 byte length of the chunk text |
urna build --spec (the forge) | The row ordinal: [ordinal, ordinal + 1) |
A corpus with a blob span overlay (0x16) | In hits, a byte range inside the media blob, with the blob's uri as source_uri |
The quickstart corpus is built by the forge, which is why the citation above shows byte_start: 7 and byte_end: 8: the chunk is ordinal 7, the eighth row after sorting by order_by.
Because the forge's ordinal enters the chunk_id, three things follow for spec builds:
- Adding or removing a row shifts the ordinal, and so the citation, of every later row.
- A
--samplebuild renumbers its rows and never shares citations with the full build. - Rows sort by the string value of their
order_bycolumns, so numeric ids sort as text ("10"before"2").
For a blob overlay, the runtime rewrites the spans at open time, so hits and retrieve report the blob range. urna cite reads chunks_original_spans directly and prints the stored span, which for a forge media corpus is the row ordinal. The chunk_id is always computed from the stored span.
What the hashes prove
A matching content_hash and chunk_id prove that the text you are reading is the text the citation was issued for, and a passing urna validate proves the file is internally consistent. All of these are unkeyed SHA-256: they detect corruption, truncation and mismatched corpora. They do not prove who built the file. See Security.
The hashes also make builds comparable. Two builds from the same rows, chunker, model, dtype, provenance and index parameters produce the same content_hash, and the forge pins created by default, so the files can be byte-identical. See Reproducible builds.
The byte-level preimages and golden values are in Hashes and citation ids.
The .urna file
What a .urna file holds, how its header, section table and footer fit together, and what the runtime checks before it answers a query.
Search paths and the exact rerank
The exact, HNSW, hybrid, graph and named-space search paths in urna, how ask and retrieve pick one, and why every returned score is a real cosine.