docsv0.5.1

Hashes and citations

Exact preimages of every integrity value in a .urna file: header and section checksums, the footer hash, file_hash, content_hash, chunk_id and the citation id.

A .urna file carries several SHA-256 values at different widths and over different bytes. This page gives the exact preimage of each one, where the sha256:<hex> form applies, and which changes move content_hash and therefore every citation. For the ideas behind them, see Citations and hashes.

Integrity values

ValueWidthPreimageStored inShown as
header checksum8 bytesSHA-256 over header bytes [0, 72) followed by [80, 128), first 8 bytesheader [72, 80)not shown
section checksum8 bytesSHA-256 over the stored payload [offset, offset + size), padding excluded, first 8 byteseach table entry16 lowercase hex characters, no prefix
footer hash32 bytesSHA-256 over [0, file_size - 40)footernot shown
file_hashsha256:<64 hex>SHA-256 over the whole file, footer includedcomputed on openvalidate, inspect, stats, every search hit
content_hashsha256:<64 hex>see belowcomputed on opensame places, and the first half of every citation

The header and section checksums are 64-bit truncations of SHA-256, stored as raw bytes. They are not sha256:<64 hex> strings. The header preimage skips the checksum field entirely (120 bytes); it is not hashed with the field zeroed.

The footer hash and the reported file_hash are two different numbers. The footer hashes everything before the footer. The reported file_hash hashes the whole file and equals the output of sha256sum:

urna validate quickstart.urna | grep 'File hash'
shasum -a 256 quickstart.urna
  File hash:          sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832  quickstart.urna

Integrity, not authenticity

Every checksum and hash here is unkeyed. A mismatch proves the bytes changed; a match does not prove who wrote the file. Anyone can build a file with valid hashes. See Security.

Physical checksums and content_hash catch different faults. A section checksum is over the stored bytes, so it catches corruption of a compressed payload before decoding. content_hash is over the decoded bytes, so it stays the same when the same content is stored with another codec.

content_hash

content_hash is SHA-256 over the six canonical sections, decoded, in this fixed order: chunk_ids, chunks_canonical, chunks_original_spans, embeddings, provenance, search_contract. For each section:

u32 LE  len(name)
bytes   name (ASCII, for example "chunk_ids")
u64 LE  len(decoded)
bytes   decoded section bytes

The output is sha256:<64 lowercase hex>.

  • "Decoded" means after the wire codec, the dictionary (0x0A) and the dedup expansion (0x0B). A raw file and a zstd file with the same content have the same content_hash.
  • Quantized embeddings are hashed as stored. A float16, int8 or int4 build has a different content_hash from its float32 twin.
  • The manifest is not part of it. Only the footer hash and file_hash cover the manifest.
  • No optional section is part of it.

chunk_id

Each chunk id is SHA-256 over:

"urna:chunk_id:v1\n"
u32 LE  len(canonical_text)   + canonical_text bytes
u32 LE  len(source_uri)       + source_uri bytes
u64 LE  byte_start
u64 LE  byte_end
u32 LE  len(chunker_version)  + chunker_version bytes

The output is sha256:<64 hex>. chunker_version is the manifest field. The Rust function is urna_format::chunk_id(canonical_text, source_uri, byte_start, byte_end, chunker_version). The writer computes the ids; ChunkInput has no id field.

Any change to the text, the URI, the span or chunker_version gives a new chunk id.

Citation id

A citation joins the two hashes under the urna:// scheme:

urna://<content_hash>/<chunk_id>

Both halves keep their sha256: prefix, so a literal citation looks like this:

urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be

Every SearchHit carries it as citation_id. urna cite resolves it back to the stored text and span, and refuses a citation whose content_hash is not the file's.

Where the sha256:hex form applies

ValueForm
chunk_idsha256:<64 hex>
manifest model_hashsha256:<64 hex>, validated; uppercase hex digits pass
reported file_hashsha256:<64 hex>
content_hashsha256:<64 hex>
space model_hash in 0x15only the sha256: prefix is checked
blob content_hash in 0x1432 raw bytes on disk, printed as sha256:<hex>
header checksum, section checksums8 raw bytes
footer hash32 raw bytes

What moves content_hash

A change that moves content_hash changes every citation into the file. A change that only moves file_hash leaves citations valid.

ChangeMoves content_hashWhy
raw vs zstd text, or which text codec winsnocodecs decode byte-identically
embedding dtype (float32, float16, int8, int4)yesquantized bytes are hashed
mrl_dim truncationyestruncated vectors are hashed
chunk text, URI, span, chunker_versionyesthrough 0x01 to 0x03
provenance JSON, key order includedyes0x05 is canonical and keeps insertion order
attaching HNSW (hnsw_index(), presets tiny, nano, hybrid, with_hnsw=True)yesit sets index_type = "hnsw" and rerank_policy = "exact", which the canonical search_contract copies
hybrid() in Rustyesit sets index_type, score_type and rerank_policy
attaching BM25 alonenoonly the manifest flag supports_bm25 changes
graph, blob sections, space table, bandsnooutside the canonical six; they touch only capabilities_ext
manifest-only fields (title, created, capabilities, capabilities_ext, extra keys)nothe manifest is covered by file_hash only

Adding HNSW changes the citations

Attaching an HNSW index rewrites index_type and rerank_policy in the canonical search_contract section, so an exact build and its HNSW twin have different content_hash values and different citations. The same holds between the exact or compressed presets and tiny, nano or hybrid. A comment in urna-format says adding an optional section never invalidates citations; that is true for graph, blobs and spaces, not for HNSW. See Known limits.

Verified with the 0.5.1 crates on one three-chunk corpus: the raw build, its zstd-text twin and its BM25 twin share one content_hash; the HNSW twin has another.

Rust entry points

ItemReturns
UrnaView::file_hash_hex()reported file_hash, whole file
UrnaView::content_hash_hex()content_hash
MmapUrnaFile::file_hash(), MmapUrnaFile::content_hash()the same, computed once at open
urna_format::chunk_id(...)a chunk id
UrnaHeader::compute_checksum(), validate_checksum()header checksum
SectionEntry::compute_checksum(data), validate_checksum(data)section checksum
UrnaFooter::compute_file_hash(bytes)footer hash

MmapUrnaFile::open hashes the whole file twice (the footer check, then file_hash) and decodes the canonical six once for content_hash, so opening costs time linear in the file size.

The manifest fields, including model_hash, are on Manifest.

On this page