docsv0.5.1

urna stats

Reference for urna stats, which prints a .urna file's size, chunk count, dim, dtype, search contract, model, sections, named spaces and hashes.

urna stats prints a one-screen summary of a .urna file: its size, chunk count, embedding dim and dtype, the search contract, the embedding model and its hash, the section sizes, any named spaces, and the two hashes. Use it to find out which model a corpus needs and which search path ask will take.

Usage

urna stats <FILE>

Arguments

ArgumentDescription
<FILE>Path to the .urna file

Options

OptionDescription
-h, --helpPrint help

Behavior

stats reads the file into memory and runs the reader's integrity check (checksums, manifest, footer hash, contract) before printing. It does not decode the indices or walk the embeddings for NaN; urna validate does the NaN walk.

Output

To stdout, one field per line:

FieldMeaning
fileThe path you passed
sizeFile size in bytes, and in MB (bytes divided by 1,000,000)
chunks, embeddingsChunk and embedding counts from the header
dimembedding_dim: the query vector length for the text paths
mrl_dim, full_dimOnly when the file was built with Matryoshka truncation: the stored prefix dim and the model's full dim. Queries use mrl_dim
dtypeStored precision: float32, float16, int8 or int4
metric, score_type, normalizeThe search contract: ip, cosine (or hybrid_rrf), l2
index_typeThe declared path ask, retrieve and search-text take: exact, hnsw or hybrid
rerankThe manifest's rerank_policy: none or exact
modelThe manifest's embedding_model
model_hashThe model fingerprint the query embedder must match
chunkerThe manifest's chunker_version
supports_ann, supports_bm25Manifest capability flags
simd_backendThe scoring kernel on this machine (avx2, neon or scalar), not a property of the file
sectionsSection count, then one line per section: id, name, encoding, payload bytes
spacesOnly when the file has a space table: one line per named space with its dim, dtype, vector count and model_hash
file_hashSHA-256 of the whole file
content_hashSHA-256 of the canonical sections

Three details:

  • The runtime opens the HNSW and BM25 indices when their sections are present; it does not read supports_ann or supports_bm25. Read the sections list to see what the file carries.
  • index_type is what routes ask and retrieve (see how ask and retrieve choose a path). A file with a BM25 section and index_type: hnsw never uses BM25 from those verbs.
  • Encoding names cover raw, zstd, float16, int8, int4 and intpack; the other text codecs print as unknown. Section names print as unknown for 0x09, 0x0A, 0x0B and reserved ids.

The Python wheel installs its own urna command, whose stats prints no size and no spaces block. See The wheel's urna command.

Exit codes

CodeMeaning
0Printed
1File missing or unreadable, or a failed check. Printed to stderr as Error: <message>
2Usage error: a missing argument

Examples

urna stats examples/quickstart/out/quickstart.urna
file:         examples/quickstart/out/quickstart.urna
size:         17942 bytes (0.02 MB)
chunks:       12
embeddings:   12
dim:          256
dtype:        float32
metric:       ip
score_type:   cosine
normalize:    l2
index_type:   hnsw
rerank:       exact
model:        minishlab/potion-base-8M/v1
model_hash:   sha256:8f2eb91a754b4da59cdd8223d0ba196185fed1bb6f092fd2be0ff02b893b1c98
chunker:      quickstart/1
supports_ann: true
supports_bm25:true
simd_backend: neon
sections:     9
  0x01 chunk_ids                encoding=intpack  389 bytes
  0x02 chunks_canonical         encoding=zstd     1484 bytes
  0x03 chunks_original_spans    encoding=zstd     154 bytes
  0x04 embeddings               encoding=raw      12288 bytes
  0x05 provenance               encoding=zstd     118 bytes
  0x06 search_contract          encoding=zstd     95 bytes
  0x07 hnsw_index               encoding=raw      182 bytes
  0x08 bm25_index               encoding=zstd     1695 bytes
  0x0c graph_adjacency          encoding=raw      174 bytes
file_hash:    sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
content_hash: sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df

This corpus was embedded with minishlab/potion-base-8M/v1, so the potion embedder that ships with the installed binary can query it. A corpus built with a registry model (the clip-vit-b32, siglip2, jina-v5-omni or wemm presets) needs a repo checkout and that model's Python dependencies before ask or retrieve can query it with text. See Open a corpus you downloaded.

For every byte of the header and section table, use urna inspect.

On this page