docsv0.5.1

Citations and hashes

How a urna:// citation is built from content_hash and chunk_id, which changes move it, what the offsets mean, and what file_hash covers.

Every hit urna returns carries a citation of the form urna://content_hash/chunk_id. The citation names the corpus content and one chunk inside it, both by SHA-256, so it resolves to the same stored text on any machine that has the same corpus. This page covers the hashes behind it, what changes them, and what the span offsets in a hit mean.

Anatomy of a citation

This is a real citation from the quickstart corpus:

urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
  • The first part is the file's content_hash, a SHA-256 over the six canonical sections.
  • The second part is the chunk_id, a SHA-256 over one chunk's text and origin.

Both are written as sha256: followed by 64 lowercase hex digits. Resolve a citation with urna cite:

urna cite examples/quickstart/out/quickstart.urna 'urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be'
citation_id:  urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
file:         examples/quickstart/out/quickstart.urna
file_hash:    sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
content_hash: sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df
chunk_id:     sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
source_uri:   demo/03-citations.md
byte_start:   7
byte_end:     8
text:
because the citation points at content, two people who build the same logical corpus on two machines get the same citation, and a stored corpus and a compressed one cite identically. resolving a citation returns the exact canonical text and the original byte span it came from, which is what lets an agent quote a source it can prove.

cite first recomputes the file's content_hash and refuses the citation if it does not match, so a citation never resolves against a different corpus by accident. The text it prints is the stored canonical text of the chunk, not a reopen of the original source file.

chunk_id

The chunk_id is the SHA-256 of this preimage, with every string length-prefixed:

"urna:chunk_id:v1\n"
u32 LE  len(canonical_text)   canonical_text bytes
u32 LE  len(source_uri)       source_uri bytes
u64 LE  byte_start
u64 LE  byte_end
u32 LE  len(chunker_version)  chunker_version bytes

The embedding is not part of it. The same text from the same place, cut by the same chunker, gets the same chunk_id whatever model embedded it. Changing any of the five inputs (text, uri, either offset, or the manifest's chunker_version) gives a new id.

content_hash

The content_hash is the SHA-256 of the six canonical sections, taken in this fixed order: chunk_ids, chunks_canonical, chunks_original_spans, embeddings, provenance, search_contract. For each section the hash takes the name length (u32 LE), the name, the decoded length (u64 LE) and the decoded bytes.

"Decoded" means after the wire codec: a section stored as zstd or intpack hashes the same as its raw form. Quantized embeddings are the exception, because they are hashed as stored: a float16 corpus and a float32 corpus of the same chunks have different content_hash values. The manifest is not part of the preimage.

file_hash has two meanings

ValueCoversWhere you see it
Footer hashSHA-256 of every byte before the 40-byte footerStored as 32 raw bytes in the footer and checked on open; never printed
Reported file_hashSHA-256 of the whole file, footer includedvalidate, stats, inspect, cite, every hit, retrieve JSON

The reported file_hash is the same number sha256sum or shasum -a 256 prints for the file:

shasum -a 256 examples/quickstart/out/quickstart.urna
e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832  examples/quickstart/out/quickstart.urna

The header checksum and the per-section checksums are shorter: the first 8 bytes of a SHA-256, stored raw. urna inspect prints a section checksum as 16 hex digits with no prefix. They detect a corrupted header or payload; they are not identifiers.

Use content_hash to ask "is this the same corpus content?" and file_hash to ask "is this the same file, byte for byte?". Two files can share a content_hash and differ in file_hash (for example, a different title in the manifest, or a BM25 index added).

What moves content_hash

Anything that moves content_hash changes every citation into the file, because the hash is the first half of each citation.

ChangeMoves content_hashWhy
Raw vs zstd text, or which text codec the writer picksNoCodecs decode to identical bytes
Embedding dtype (float32, float16, int8, int4)YesQuantized bytes are hashed as stored
mrl_dim truncationYesThe truncated vectors are hashed
Chunk text, source_uri, span, chunker_versionYesThey are in chunk_ids, chunks_canonical and chunks_original_spans
Provenance JSON, including key orderYesprovenance is canonical
Attaching an HNSW indexYesIt rewrites index_type and rerank_policy, which live in search_contract
Declaring index_type = "hybrid"YesSame mechanism
Adding a BM25 index aloneNoOnly a manifest capability changes
Graph, blob tables, blob data, space table and bandsNoNot canonical; only manifest flags change
Manifest-only fields (title, created, capabilities)NoThe manifest is covered by the file hash only

In practice: rebuilding with a different preset or dtype gives new citations. Changing only the text compression does not.

Adding HNSW changes every citation

Attaching an HNSW index sets index_type = "hnsw" and rerank_policy = "exact", and the writer copies both into the canonical search_contract section. An exact build and an HNSW build of the same chunks therefore have different content_hash values and different citations. This also separates exact from the tiny, nano and hybrid presets, and any build with with_hnsw=True. A comment in the format crate says adding an optional section never invalidates citations; that holds for the graph, blobs and spaces, not for HNSW. See Known limits.

What the offsets mean

byte_start and byte_end come from chunks_original_spans and mean whatever the builder wrote there. urna checks only that byte_end is not below byte_start.

Built withOffsets are
builder.chunk_text (repo checkout)Real UTF-8 byte offsets into the source text
The Python examples in the repo0 to the UTF-8 byte length of the chunk text
urna build --spec (the forge)The row ordinal: [ordinal, ordinal + 1)
A corpus with a blob span overlay (0x16)In hits, a byte range inside the media blob, with the blob's uri as source_uri

The quickstart corpus is built by the forge, which is why the citation above shows byte_start: 7 and byte_end: 8: the chunk is ordinal 7, the eighth row after sorting by order_by.

Because the forge's ordinal enters the chunk_id, three things follow for spec builds:

  • Adding or removing a row shifts the ordinal, and so the citation, of every later row.
  • A --sample build renumbers its rows and never shares citations with the full build.
  • Rows sort by the string value of their order_by columns, so numeric ids sort as text ("10" before "2").

For a blob overlay, the runtime rewrites the spans at open time, so hits and retrieve report the blob range. urna cite reads chunks_original_spans directly and prints the stored span, which for a forge media corpus is the row ordinal. The chunk_id is always computed from the stored span.

What the hashes prove

A matching content_hash and chunk_id prove that the text you are reading is the text the citation was issued for, and a passing urna validate proves the file is internally consistent. All of these are unkeyed SHA-256: they detect corruption, truncation and mismatched corpora. They do not prove who built the file. See Security.

The hashes also make builds comparable. Two builds from the same rows, chunker, model, dtype, provenance and index parameters produce the same content_hash, and the forge pins created by default, so the files can be byte-identical. See Reproducible builds.

The byte-level preimages and golden values are in Hashes and citation ids.

On this page