Hashes and citations
Exact preimages of every integrity value in a .urna file: header and section checksums, the footer hash, file_hash, content_hash, chunk_id and the citation id.
A .urna file carries several SHA-256 values at different widths and over different bytes. This page gives the exact preimage of each one, where the sha256:<hex> form applies, and which changes move content_hash and therefore every citation. For the ideas behind them, see Citations and hashes.
Integrity values
| Value | Width | Preimage | Stored in | Shown as |
|---|---|---|---|---|
| header checksum | 8 bytes | SHA-256 over header bytes [0, 72) followed by [80, 128), first 8 bytes | header [72, 80) | not shown |
| section checksum | 8 bytes | SHA-256 over the stored payload [offset, offset + size), padding excluded, first 8 bytes | each table entry | 16 lowercase hex characters, no prefix |
| footer hash | 32 bytes | SHA-256 over [0, file_size - 40) | footer | not shown |
file_hash | sha256:<64 hex> | SHA-256 over the whole file, footer included | computed on open | validate, inspect, stats, every search hit |
content_hash | sha256:<64 hex> | see below | computed on open | same places, and the first half of every citation |
The header and section checksums are 64-bit truncations of SHA-256, stored as raw bytes. They are not sha256:<64 hex> strings. The header preimage skips the checksum field entirely (120 bytes); it is not hashed with the field zeroed.
The footer hash and the reported file_hash are two different numbers. The footer hashes everything before the footer. The reported file_hash hashes the whole file and equals the output of sha256sum:
urna validate quickstart.urna | grep 'File hash'
shasum -a 256 quickstart.urna File hash: sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832 quickstart.urnaIntegrity, not authenticity
Every checksum and hash here is unkeyed. A mismatch proves the bytes changed; a match does not prove who wrote the file. Anyone can build a file with valid hashes. See Security.
Physical checksums and content_hash catch different faults. A section checksum is over the stored bytes, so it catches corruption of a compressed payload before decoding. content_hash is over the decoded bytes, so it stays the same when the same content is stored with another codec.
content_hash
content_hash is SHA-256 over the six canonical sections, decoded, in this fixed order: chunk_ids, chunks_canonical, chunks_original_spans, embeddings, provenance, search_contract. For each section:
u32 LE len(name)
bytes name (ASCII, for example "chunk_ids")
u64 LE len(decoded)
bytes decoded section bytesThe output is sha256:<64 lowercase hex>.
- "Decoded" means after the wire codec, the dictionary (
0x0A) and the dedup expansion (0x0B). A raw file and a zstd file with the same content have the samecontent_hash. - Quantized embeddings are hashed as stored. A float16, int8 or int4 build has a different
content_hashfrom its float32 twin. - The manifest is not part of it. Only the footer hash and
file_hashcover the manifest. - No optional section is part of it.
chunk_id
Each chunk id is SHA-256 over:
"urna:chunk_id:v1\n"
u32 LE len(canonical_text) + canonical_text bytes
u32 LE len(source_uri) + source_uri bytes
u64 LE byte_start
u64 LE byte_end
u32 LE len(chunker_version) + chunker_version bytesThe output is sha256:<64 hex>. chunker_version is the manifest field. The Rust function is urna_format::chunk_id(canonical_text, source_uri, byte_start, byte_end, chunker_version). The writer computes the ids; ChunkInput has no id field.
Any change to the text, the URI, the span or chunker_version gives a new chunk id.
Citation id
A citation joins the two hashes under the urna:// scheme:
urna://<content_hash>/<chunk_id>Both halves keep their sha256: prefix, so a literal citation looks like this:
urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4beEvery SearchHit carries it as citation_id. urna cite resolves it back to the stored text and span, and refuses a citation whose content_hash is not the file's.
Where the sha256:hex form applies
| Value | Form |
|---|---|
chunk_id | sha256:<64 hex> |
manifest model_hash | sha256:<64 hex>, validated; uppercase hex digits pass |
reported file_hash | sha256:<64 hex> |
content_hash | sha256:<64 hex> |
space model_hash in 0x15 | only the sha256: prefix is checked |
blob content_hash in 0x14 | 32 raw bytes on disk, printed as sha256:<hex> |
| header checksum, section checksums | 8 raw bytes |
| footer hash | 32 raw bytes |
What moves content_hash
A change that moves content_hash changes every citation into the file. A change that only moves file_hash leaves citations valid.
| Change | Moves content_hash | Why |
|---|---|---|
| raw vs zstd text, or which text codec wins | no | codecs decode byte-identically |
| embedding dtype (float32, float16, int8, int4) | yes | quantized bytes are hashed |
mrl_dim truncation | yes | truncated vectors are hashed |
chunk text, URI, span, chunker_version | yes | through 0x01 to 0x03 |
| provenance JSON, key order included | yes | 0x05 is canonical and keeps insertion order |
attaching HNSW (hnsw_index(), presets tiny, nano, hybrid, with_hnsw=True) | yes | it sets index_type = "hnsw" and rerank_policy = "exact", which the canonical search_contract copies |
hybrid() in Rust | yes | it sets index_type, score_type and rerank_policy |
| attaching BM25 alone | no | only the manifest flag supports_bm25 changes |
| graph, blob sections, space table, bands | no | outside the canonical six; they touch only capabilities_ext |
manifest-only fields (title, created, capabilities, capabilities_ext, extra keys) | no | the manifest is covered by file_hash only |
Adding HNSW changes the citations
Attaching an HNSW index rewrites index_type and rerank_policy in the canonical search_contract section, so an exact build and its HNSW twin have different content_hash values and different citations. The same holds between the exact or compressed presets and tiny, nano or hybrid. A comment in urna-format says adding an optional section never invalidates citations; that is true for graph, blobs and spaces, not for HNSW. See Known limits.
Verified with the 0.5.1 crates on one three-chunk corpus: the raw build, its zstd-text twin and its BM25 twin share one content_hash; the HNSW twin has another.
Rust entry points
| Item | Returns |
|---|---|
UrnaView::file_hash_hex() | reported file_hash, whole file |
UrnaView::content_hash_hex() | content_hash |
MmapUrnaFile::file_hash(), MmapUrnaFile::content_hash() | the same, computed once at open |
urna_format::chunk_id(...) | a chunk id |
UrnaHeader::compute_checksum(), validate_checksum() | header checksum |
SectionEntry::compute_checksum(data), validate_checksum(data) | section checksum |
UrnaFooter::compute_file_hash(bytes) | footer hash |
MmapUrnaFile::open hashes the whole file twice (the footer check, then file_hash) and decodes the canonical six once for content_hash, so opening costs time linear in the file size.
The manifest fields, including model_hash, are on Manifest.
Encodings
The wire encoding registry of .urna v1: ids 0 to 10, zstd limits, the float16, int8 and int4 embedding layouts, the text-codec chooser and intpack.
Manifest
Every field of the .urna v1 manifest, the Capabilities and CapabilitiesExt flags, validation rules, JSON serialization and the additivity rule for new fields.