docsv0.5.1

Compatibility

Which .urna files a 0.5.1 reader accepts, how format v1 grows without breaking old files, legacy NEST files and what makes two builds byte-identical.

Format v1 has been frozen since 0.1.0. Later releases added sections, encodings and manifest fields inside v1 without changing the container. This page states what a 0.5.1 reader accepts, what older readers do with newer files, and the rules that keep both working.

Version numbers

A file carries two pairs of version numbers.

WhereFieldWritten by 0.5.1Bumped when
headerversion_major, version_minor1, 0the header layout changes
manifestformat_version1the binary container changes (header, footer, section table layout)
manifestschema_version1manifest fields or the meaning of the required sections change

A 0.5.1 reader accepts:

CheckAcceptedRejected with
magicURNA, or the legacy NESTMagicMismatch
version_majorexactly 1UnsupportedVersion
version_minor0UnsupportedVersion
format_version1 or lower (0 included)UnsupportedFormatVersion
schema_version1 or lower (0 included)UnsupportedSchemaVersion

A higher version means the file may use something this reader cannot interpret, so the reader refuses it with a typed error instead of guessing. A lower or equal version is accepted on the assumption that the reader still understands older files.

Newer files in this reader

Additions made inside v1 fall into three groups, and a reader treats each one differently.

AdditionWhat a reader that does not know it does
a new optional section idloads the file; the section is checksummed and then ignored
a new wire encoding idrefuses the file with UnsupportedSectionEncoding, naming the section and the encoding
a new manifest fieldkeeps it in the extra map and writes it back unchanged

The reader has no allow-list of section ids. Any id loads when its encoding is legal for its class (vector encodings on 0x04 and the bands, text encodings elsewhere) and its checksum matches. An unknown encoding is never read as raw bytes: it fails loudly.

Unknown keys inside the capabilities object are dropped, which is why new flags go in capabilities_ext or extra (see the additivity rule).

This release's files in older readers

Every reader released before 0.5.0 predates the rename and knows only the NEST magic, so it refuses a file written by 0.5.x at the first check. The format did not change in the rename; only the magic, the chunk-id domain and the names did.

Between releases that share a magic, the rule in the previous section decides. A reader opens a newer file when the file only adds optional sections, and refuses it by encoding id when the file uses an encoding the reader predates. The changelog records the same behavior before the rename: 0.2.x readers skipped the optional sections that 0.3.0 added and refused its new encodings 4, 5, 7, 9 and 10 with a typed error.

Older files in this reader

  • A v0.1-shape file (raw encoding, float32, the six required sections, no optional section) loads unchanged. The 1366-byte golden fixture crates/urna-format/tests/fixtures/golden_v1_minimal.urna is byte-frozen and checked by the test suite, including its file_hash, content_hash and chunk id.
  • Files written by 0.4.0 and earlier carry the magic NEST. The layout and hashes are the same, so the 0.5.1 reader opens, validates and searches them. The writer never emits NEST. The test crates/urna-format/tests/legacy_magic.rs runs on the 0.4.0 golden bytes.
  • The chunk ids inside a NEST file were computed under the old chunk-id domain. They stay valid as stored, and citations into the file resolve, but urna_format::chunk_id and urna.chunk_id do not reproduce them from the same text, because today's preimage starts with urna:chunk_id:v1. Rebuild the corpus if you need ids you can recompute.

Additivity rule

Every addition inside v1 follows these rules, so existing files keep their bytes and their hashes:

  • An optional section stays out of content_hash and out of the required list. Old readers skip it; old citations stay valid. Attaching an HNSW index is the exception, because it also rewrites the canonical search_contract: see what moves content_hash.
  • A new manifest field is an Option that is omitted when unset. A file that does not use it serializes to the same manifest bytes as before, so its file_hash does not move.
  • A new capability flag goes in capabilities_ext or in extra. A new required boolean in capabilities is forbidden: old manifests would fail to parse and every file_hash would change.
  • A new wire encoding gets a new id. Readers that predate it refuse it by id instead of misreading it.

The guards are crates/urna-format/tests/manifest_additivity.rs and crates/urna-format/tests/reserved_ids.rs.

Reserved ids

These ids are claimed in v1 and not used by 0.5.1, so no future feature can reuse them for something else:

  • Sections 0x0D to 0x13 (chunk_scalars, tokenizer_model, edit_journal, repro_manifest, graph_nodes, graph_edge_props, graph_entity_map) and the band 0x30 to 0x3F. 0x09 (embeddings_fp) is read by the runtime but not written.
  • Encodings 6 (frontcode) and 8 (rabitq), refused by the reader today.

The full map is on Sections.

Reproducible builds

The writer is deterministic: the same inputs produce the same bytes. UrnaFileBuilder::reproducible(true) does one thing: it sets the manifest created to "1970-01-01T00:00:00Z" (urna_format::writer::REPRODUCIBLE_CREATED), so a timestamp cannot differ between builds. The builder never fills created on its own, so a build that leaves created unset is byte-identical without the flag too.

Two builds are byte-identical when all of these match:

  • the chunks, in the same order, with the same vectors;
  • the manifest fields;
  • the provenance JSON, key order included;
  • the text encoding and the embedding dtype;
  • the HNSW parameters and seed, when an index is attached.

Identical bytes mean an identical file_hash. For the build tooling above the writer (the build lock and the cache), see Reproducible builds.

On this page