docsv0.5.1

Media and named spaces

How a .urna file carries media blobs and extra named vector spaces, each space with its own model_hash and dim, and how to query one.

Besides its text chunks and their default vectors, a .urna file can carry two optional layers: media blobs (the encoded images or video segments the chunks came from) and named spaces (extra vector sets, such as image embeddings from a second model). Both sit in optional sections outside content_hash, so adding them never changes a citation.

The default space and named spaces

Every file has one default space: the vectors in section 0x04, one per chunk, embedded with the model the manifest names in embedding_model. urna ask, urna retrieve, urna search-text and every search* verb except search-space read this space and only this space.

A named space is a second set of vectors over the same chunks, with its own name, model, dim and dtype. A file can hold up to 15. search-space reads one named space and nothing else, so a text query can never be scored against image vectors by accident, and the reverse.

Media blobs

Three sections describe media. The manifest sets capabilities_ext.blobs_present when any of them is written.

SectionNameContent
0x14blob_refsone record per blob: the SHA-256 of the blob's bytes, its original_uri, its byte length, and whether it is inlined
0x16blob_span_overlayone entry per chunk, in chunk order: a blob index and a range inside that blob, or "none"
0x17blob_datathe inlined blob bytes, behind an offset table parallel to 0x14

The records keep their build order, and the overlay points into them by position. An overlay entry of "none" keeps the chunk's own span from section 0x03, so a file can mix media chunks and plain text chunks.

When a file with an overlay is opened, the runtime replaces each pointing chunk's source_uri and span with the blob's original_uri and the range in the overlay. Search hits, urna retrieve and UrnaFile.retrieve report those values. urna cite reads section 0x03 directly and prints the chunk's original span instead: for a forge corpus that is item://<name>/<key> and the row ordinal.

What a range means

The forge writes the overlay in one of two shapes, depending on the media backend (see Tune media compression):

BackendOne blob perRange for a chunk
av1stream segment (an .mp4)the frame index, [frame, frame + 1)
avif, jxl, jxl-transcode, controlimage filethe whole file, [0, byte_len)

The original_uri is media:// followed by the blob's path inside the corpus media directory.

Sidecar or inlined

By default the forge writes the encoded media to a <name>.media/ directory next to the .urna, and the 0x14 records say inlined = false. With [output] embed_media = true the bytes also go into section 0x17 and the file is self-contained. The .media/ directory stays on disk as the build cache, and peak build memory is about twice the media size.

urna media lists the records, and --export writes the inlined ones back to files:

urna media corpus.urna
urna media corpus.urna --export exported/

Each blob is checked against its 0x14 SHA-256 before it is written, and the export stops at the first mismatch. urna validate runs the same check over every inlined blob without writing anything. In Python, UrnaFile.blob_bytes(i) returns one inlined blob.

Named spaces

A named space is one entry in the space table (section 0x15) plus one band section holding its vectors. The manifest sets capabilities_ext.supports_multimodal.

Space table fieldMeaning
space_index1 to 15; the band lives in section 0x20 + space_index (0x21 to 0x2F)
nameunique within the file
dimthe vector length
dtypefloat32, float16, int8 or int4 (int4 needs dim divisible by 64)
model_hashthe identity of the model that produced the vectors, sha256: prefixed
n_vectorsmust equal the number of chunks

Bands are parallel to the chunks: row i of every band belongs to chunk i, so a hit in a named space maps to a chunk, its text and its citation the same way a default-space hit does. At open, each listed band must be present at exactly the size its dim, dtype and count imply, and its values must be finite.

How the forge names spaces

In a build spec, each [[models]] block with image = "space" or text = "space" emits named spaces (see [[models]]):

RoleNo dimsWith dims = [256, 512]
image = "space"<preset><preset>@256, <preset>@512
text = "space"<preset>-text<preset>-text@256, <preset>-text@512

The text = "default" model is the default space, not a named one. For each dim the forge keeps the first dim components of every vector and L2-normalizes them again, and stores the result at the model's space_dtype (default int8). Every space of one preset shares that preset's model_hash, image and text alike.

Listing spaces

urna stats prints a spaces: block with each space's name, dim, dtype, vector count and model_hash. urna inspect --json has a spaces array, and Python has UrnaFile.space_names.

Querying a named space

urna search-space runs an exact scan over one band. It takes a query vector, not text, as a JSON array of floats at the space's dim. With the vector in query.json:

urna search-space corpus.urna "$(cat query.json)" --space clip-vit-b32 -k 10 \
  --expect-model-hash "sha256:<the space's model_hash>"
  • An unknown name fails with embedding space not found: <name>. There is no fallback to the default space.
  • A query of the wrong length fails with a dimension mismatch.
  • --expect-model-hash fails the query when the space's model_hash differs. Without it, the hash is not checked: search-space has no embedder of its own and cannot compute one. In Python the same check is expected_model_hash=.
  • Scores are exact cosine over the whole band, recall = 1.0, and hits report index_type = "space". The rerank source follows the band's dtype, so an int8 space scores at stored precision.
  • A hit's embedding_model field still names the file's default model, not the space's model.

No verb embeds text for a named space: urna ask and urna retrieve always query the default space. Embed the query with the same preset yourself. From a repository checkout, with the preset's packages installed:

import sys
sys.path.insert(0, "python")

import urna
from forge import model_registry

emb = model_registry.create_embedder("clip-vit-b32")
query = emb.embed_texts(["a red dragon over a castle"], role="query")[0].tolist()

db = urna.open("out/cards/cards.urna")
hits = db.search_space("clip-vit-b32", query, 10, expected_model_hash=emb.model_hash)
for hit in hits:
    print(round(hit.score, 4), hit.source_uri, hit.citation_id)

For a space with a dim suffix, such as clip-vit-b32@256 on a model with a ladder, slice the query with model_registry.slice_renorm(vectors, 256) before searching. urna benchmark --space <name> measures the latency of one space.

Outside content_hash

content_hash covers the six required sections only. Blob records, overlay, inlined bytes, space table and bands are excluded, so:

  • Adding media or a named space to a corpus keeps every chunk_id, content_hash and citation.
  • A self-contained file and its sidecar twin cite identically.
  • These sections are still covered by their section checksums and by the file hash, which urna validate checks.

To build a corpus with images, see Images and PDFs.

On this page