urna stats
Reference for urna stats, which prints a .urna file's size, chunk count, dim, dtype, search contract, model, sections, named spaces and hashes.
urna stats prints a one-screen summary of a .urna file: its size, chunk count, embedding dim and dtype, the search contract, the embedding model and its hash, the section sizes, any named spaces, and the two hashes. Use it to find out which model a corpus needs and which search path ask will take.
Usage
urna stats <FILE>Arguments
| Argument | Description |
|---|---|
<FILE> | Path to the .urna file |
Options
| Option | Description |
|---|---|
-h, --help | Print help |
Behavior
stats reads the file into memory and runs the reader's integrity check (checksums, manifest, footer hash, contract) before printing. It does not decode the indices or walk the embeddings for NaN; urna validate does the NaN walk.
Output
To stdout, one field per line:
| Field | Meaning |
|---|---|
file | The path you passed |
size | File size in bytes, and in MB (bytes divided by 1,000,000) |
chunks, embeddings | Chunk and embedding counts from the header |
dim | embedding_dim: the query vector length for the text paths |
mrl_dim, full_dim | Only when the file was built with Matryoshka truncation: the stored prefix dim and the model's full dim. Queries use mrl_dim |
dtype | Stored precision: float32, float16, int8 or int4 |
metric, score_type, normalize | The search contract: ip, cosine (or hybrid_rrf), l2 |
index_type | The declared path ask, retrieve and search-text take: exact, hnsw or hybrid |
rerank | The manifest's rerank_policy: none or exact |
model | The manifest's embedding_model |
model_hash | The model fingerprint the query embedder must match |
chunker | The manifest's chunker_version |
supports_ann, supports_bm25 | Manifest capability flags |
simd_backend | The scoring kernel on this machine (avx2, neon or scalar), not a property of the file |
sections | Section count, then one line per section: id, name, encoding, payload bytes |
spaces | Only when the file has a space table: one line per named space with its dim, dtype, vector count and model_hash |
file_hash | SHA-256 of the whole file |
content_hash | SHA-256 of the canonical sections |
Three details:
- The runtime opens the HNSW and BM25 indices when their sections are present; it does not read
supports_annorsupports_bm25. Read thesectionslist to see what the file carries. index_typeis what routesaskandretrieve(see how ask and retrieve choose a path). A file with a BM25 section andindex_type: hnswnever uses BM25 from those verbs.- Encoding names cover
raw,zstd,float16,int8,int4andintpack; the other text codecs print asunknown. Section names print asunknownfor0x09,0x0A,0x0Band reserved ids.
The Python wheel installs its own urna command, whose stats prints no size and no spaces block. See The wheel's urna command.
Exit codes
| Code | Meaning |
|---|---|
0 | Printed |
1 | File missing or unreadable, or a failed check. Printed to stderr as Error: <message> |
2 | Usage error: a missing argument |
Examples
urna stats examples/quickstart/out/quickstart.urnafile: examples/quickstart/out/quickstart.urna
size: 17942 bytes (0.02 MB)
chunks: 12
embeddings: 12
dim: 256
dtype: float32
metric: ip
score_type: cosine
normalize: l2
index_type: hnsw
rerank: exact
model: minishlab/potion-base-8M/v1
model_hash: sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98
chunker: quickstart/1
supports_ann: true
supports_bm25:true
simd_backend: neon
sections: 9
0x01 chunk_ids encoding=intpack 389 bytes
0x02 chunks_canonical encoding=zstd 1484 bytes
0x03 chunks_original_spans encoding=zstd 154 bytes
0x04 embeddings encoding=raw 12288 bytes
0x05 provenance encoding=zstd 118 bytes
0x06 search_contract encoding=zstd 95 bytes
0x07 hnsw_index encoding=raw 182 bytes
0x08 bm25_index encoding=zstd 1695 bytes
0x0c graph_adjacency encoding=raw 174 bytes
file_hash: sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
content_hash: sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09dfThis corpus was embedded with minishlab/potion-base-8M/v1, so the potion embedder that ships with the installed binary can query it. A corpus built with a registry model (the clip-vit-b32, siglip2, jina-v5-omni or wemm presets) needs a repo checkout and that model's Python dependencies before ask or retrieve can query it with text. See Open a corpus you downloaded.
For every byte of the header and section table, use urna inspect.
urna inspect
Reference for urna inspect, which prints a .urna file's header, section table with checksums, manifest and hashes, as text or as JSON with blobs and spaces.
urna media
Reference for urna media, which lists the media blobs a .urna corpus references and exports the inlined ones to files after checking each SHA-256.