docsv0.5.1

Layout

Byte layout of a .urna v1 file: the 128-byte header, 32-byte section table entries, the manifest, 64-byte aligned payloads and the 40-byte footer.

A .urna file is one binary container: a fixed header, a table of sections, a JSON manifest, the section payloads and a fixed footer. This page gives the byte layout of format v1 as urna-format 0.5.1 writes and reads it. All integers are little-endian and unsigned.

File map

[0, 128)                            header (UrnaHeader, 128 bytes)
[128, 128 + 32 * count)             section table (SectionEntry, 32 bytes each), sorted by id
[manifest_offset, +manifest_size)   manifest JSON, directly after the table, not aligned
[align64(...), ...)                 section payloads, each at a 64-byte aligned offset
[file_size - 40, file_size)         footer (UrnaFooter, 40 bytes), directly after the last payload
  • The writer sorts sections by id, places the manifest at 128 + 32 * count, starts the first payload at the manifest end rounded up to 64, and aligns every next payload to 64.
  • Padding between payloads is zero bytes. It is not part of any section checksum, but the footer hash covers it.
  • The footer follows the last payload byte with no alignment.

The 64-byte alignment lets the runtime read embeddings straight from the memory map with SIMD loads.

Worked example

The quickstart corpus (examples/quickstart/out/quickstart.urna, 17942 bytes, 9 sections) lays out like this:

RangeContent
[0, 128)header
[128, 416)section table, 9 entries
[416, 1027)manifest, 611 bytes
[1027, 1088)zero padding
[1088, 17902)payloads, first at 1088 (0x01 chunk_ids), last ends at 17902 (0x0C graph_adjacency)
[17902, 17942)footer

Its first 128 bytes:

00000000: 5552 4e41 0100 0000 0000 0000 0001 0000  URNA............
00000010: 0c00 0000 0000 0000 0c00 0000 0000 0000  ................
00000020: 1646 0000 0000 0000 8000 0000 0000 0000  .F..............
00000030: 0900 0000 0000 0000 a001 0000 0000 0000  ................
00000040: 6302 0000 0000 0000 b885 9d96 efff 6caf  c.............l.
00000050: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000060: 0000 0000 0000 0000 0000 0000 0000 0000  ................
00000070: 0000 0000 0000 0000 0000 0000 0000 0000  ................

Read it against the table below: magic URNA, version 1.0, flags 0, embedding_dim 256, n_chunks 12, n_embeddings 12, file_size 17942, table at 128 with 9 entries, manifest at 416 with 611 bytes, checksum b8859d96efff6caf, reserved zeros. Run urna inspect --json on any file to see the same values decoded (see urna inspect).

128 bytes, Rust struct UrnaHeader (#[repr(C)], bytemuck::Pod, so the compiler rejects any padding).

OffsetSizeFieldTypeWritten valueReader check
04magic[u8; 4]URNAURNA or the legacy NEST, else MagicMismatch
42version_majoru161must equal 1
62version_minoru160must be 0 or lower, so any minor bump is rejected
84flagsu320not checked (covered by the header checksum)
124embedding_dimu32manifest embedding_dimmust equal the manifest value
168n_chunksu64manifest n_chunksmust equal the manifest value
248n_embeddingsu64number of chunkssets the expected embeddings size
328file_sizeu64total file bytesmust equal the file length, else FileSizeMismatch
408section_table_offsetu64128bounds only
488section_table_countu64number of sectionsoverflow-checked, bounds
568manifest_offsetu64128 + 32 * countbounds
648manifest_sizeu64manifest JSON lengthbounds
728header_checksum[u8; 8]first 8 bytes of SHA-256InvalidHeaderChecksum
8048reserved[u8; 48]zerosnot checked

The header checksum is the first 8 bytes of SHA-256 over header bytes [0, 72) followed by [80, 128): the checksum field is cut out of the preimage, not zeroed. See Hashes and citations.

UrnaHeader::new(embedding_dim, n_chunks, n_embeddings, file_size, section_table_offset, section_table_count, manifest_offset, manifest_size) fills magic and version and computes the checksum.

Section table entry

32 bytes per entry, Rust struct SectionEntry.

OffsetSizeFieldTypeMeaning
04section_idu32id from the section map
44encodingu32wire encoding id, see Encodings
88offsetu64absolute file offset, a multiple of 64
168sizeu64payload length, padding excluded
248checksum[u8; 8]first 8 bytes of SHA-256 over the payload bytes [offset, offset + size)

The checksum covers the physical (encoded) bytes. A compressed section is checked as stored, before decoding.

Manifest

The manifest is UTF-8 JSON between the section table and the first payload. It is not aligned and has no checksum of its own: only the footer hash covers it. Its fields are on Manifest.

40 bytes, Rust struct UrnaFooter.

OffsetSizeFieldTypeMeaning
08footer_sizeu64always 40
832file_hash[u8; 32]SHA-256 over [0, file_size - 40)

The footer field hashes everything before the footer. The file_hash that urna validate, inspect and every search hit report is a different value: SHA-256 of the whole file, footer included, the same as sha256sum file.urna.

Constants

All are public in urna_format (pub use layout::*).

ConstantValueMeaning
URNA_MAGICb"URNA"magic the writer emits
LEGACY_MAGICb"NEST"magic of files from 0.4.0 and earlier; read, never written
URNA_VERSION_MAJOR1header major version
URNA_VERSION_MINOR0header minor version
URNA_FORMAT_VERSION1manifest format_version; bumped when the binary container changes
URNA_SCHEMA_VERSION1manifest schema_version; bumped when manifest fields or required section semantics change
URNA_HEADER_SIZE128header bytes
URNA_SECTION_ENTRY_SIZE32bytes per table entry
URNA_FOOTER_SIZE40footer bytes
SECTION_ALIGNMENT64payload offset alignment
SECTION_PAYLOAD_PREFIX_SIZE12common payload prefix: u32 version + u64 count
SECTION_PAYLOAD_VERSION1version in that prefix

Read order

UrnaView::from_bytes checks a file in this order and stops at the first failure:

StepCheckError
1length is at least 168 bytes (header plus footer)FileTruncated
2magic is URNA or NESTMagicMismatch
3version_major == 1 and version_minor is 0UnsupportedVersion
4header checksumInvalidHeaderChecksum
5file_size equals the file lengthFileSizeMismatch
6section table fits in the bodySectionOffsetOutOfBounds or FileTruncated
7per entry: encoding legal for the section idUnsupportedSectionEncoding
8per entry: offset is a multiple of 64SectionMisaligned
9per entry: payload fits in the bodySectionOffsetOutOfBounds
10per entry: checksumSectionChecksumMismatch
11manifest range fits in the bodyFileTruncated
12manifest JSON parsesJson
13manifest rulesmanifest variants, see Manifest
14footer hash over [0, file_size - 40)FooterHashMismatch
15manifest embedding_dim and n_chunks equal the headerManifestInvalid
16the six required sections are presentMissingRequiredSection
17embeddings dtype matches the encoding, then the exact sizeManifestInvalid, UnsupportedDType, EmbeddingSizeMismatch
18when space_table is present: every listed band present with the exact sizeManifestInvalid, UnsupportedDType, EmbeddingSizeMismatch
19search_contract matches the manifest field by fieldUnsupportedMetric and the other Unsupported* variants

UrnaView::validate_embeddings_values() is a separate call that walks the embeddings for NaN and Inf. MmapUrnaFile::open in urna-runtime runs both, then decodes the index sections. Every error is listed on Typed errors.

What the reader does not enforce

  • Section ids outside the known map load if their encoding is legal for their class and their checksum matches. There is no id allow-list.
  • Duplicate section ids are not rejected. UrnaView::entry returns the first match.
  • Sections that overlap each other, the table or the manifest are not rejected. Only bounds and alignment are checked.
  • flags, reserved and a section_table_offset other than 128 are accepted.
  • The reader does not compare n_embeddings with n_chunks. The runtime catches a mismatch at open as SectionCountMismatch.

The next page maps every section id: Sections.

On this page