Rust crates
The urna-format and urna-runtime crates: what to depend on, the builder, the reader, MmapUrnaFile, its search functions, result types and a runnable example.
urna ships two library crates on crates.io. urna-format is the v1 container: layout, manifest, sections, encodings, hashes, reader and writer. urna-runtime memory-maps a file and searches it: exact SIMD search, HNSW, BM25, graph and hybrid candidates, always finished by an exact cosine rerank. The urna CLI and the Python wheel are built on these two crates.
| Crate | Version | Docs |
|---|---|---|
urna-format | 0.5.1 | docs.rs/urna-format |
urna-runtime | 0.5.1 | docs.rs/urna-runtime |
Which crate to depend on
| You want to | Depend on |
|---|---|
open and search a .urna file | urna-runtime (it depends on urna-format) |
| read or verify a file's bytes without searching | urna-format |
| write a file | urna-format |
| write a file with an HNSW or BM25 index | both: the builder takes index payloads that urna-runtime builds |
[dependencies]
urna-format = "0.5.1"
urna-runtime = "0.5.1"Neither crate has Cargo features. Both need Rust 1.85 or newer (edition 2024). The urna CLI crate needs 1.88.
Three things are not in the Rust crates:
- An embedder. Queries are
&[f32]vectors you produce with the model the file was built with. The offline embedders live in the CLI's Python payload and the wheel. - Presets.
exact,compressed,tiny,nanoandhybridare resolved by the Python bridge; in Rust you set the text encoding, the dtype and the indexes yourself (see below). retrieve. PythonUrnaFile.retrieveand the CLIaskandretrievecombine a search call withcanonical_texts()andchunk_ids().
How the crates meet
urna-runtime exposes urna-format in two places: RuntimeError::Format(UrnaError) wraps every format-level failure, and MmapUrnaFile::blob_refs() returns urna_format::BlobRefRecord values. To attach an index to a file you build the payload with urna_runtime::ann::HnswIndex or urna_runtime::bm25::Bm25Index, serialize it with to_bytes(), and pass the bytes to urna_format::UrnaFileBuilder.
Example
Build a three-chunk file, open it and run an exact search. This runs as is against the 0.5.1 crates.
use std::path::Path;
use urna_format::{ChunkInput, Manifest, UrnaFileBuilder};
use urna_runtime::MmapUrnaFile;
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Build a three-chunk file with 4-dim unit vectors.
let manifest = Manifest {
embedding_model: "demo".into(),
embedding_dim: 4,
n_chunks: 3,
chunker_version: "demo-chunker/1".into(),
model_hash: format!("sha256:{}", "ab".repeat(32)),
..Default::default()
};
let mut builder = UrnaFileBuilder::new(manifest);
for (i, text) in ["alpha", "beta", "gamma"].iter().enumerate() {
let mut embedding = vec![0.0f32; 4];
embedding[i] = 1.0;
builder = builder.add_chunk(ChunkInput {
canonical_text: text.to_string(),
source_uri: "notes.txt".into(),
byte_start: (i * 6) as u64,
byte_end: (i * 6 + text.len()) as u64,
embedding,
});
}
builder.write_to_path("demo.urna")?;
// Open it and run an exact search.
let file = MmapUrnaFile::open(Path::new("demo.urna"))?;
let result = file.search(&[0.0, 1.0, 0.0, 0.0], 2)?;
for hit in &result.hits {
println!("{:.4} {} {}", hit.score, hit.source_uri, hit.citation_id);
}
println!("route={} rerank={}", result.explain.route, result.explain.rerank_source.disclosure());
Ok(())
}1.0000 notes.txt urna://sha256:3e9dd79c37c1a1e0e8264a89cfc0df29372250d20bf2ba88bebc45b443950bc6/sha256:2999f54c8525f5e7cedc3c4d16ea7202582734aa25cdc7f33a91c01cadd7926d
0.0000 notes.txt urna://sha256:3e9dd79c37c1a1e0e8264a89cfc0df29372250d20bf2ba88bebc45b443950bc6/sha256:c4a06819ec4c55548dc8e6a8ce45e02377c77785349799f85a9b37303078c271
route=exact rerank=real cosineTo query a corpus someone else built, check the model first and route on the declared index type, the way the CLI does:
use std::path::Path;
use urna_runtime::{MmapUrnaFile, SearchResult};
fn query(path: &Path, model_hash: &str, query: &[f32], k: i32) -> Result<SearchResult, Box<dyn std::error::Error>> {
let file = MmapUrnaFile::open(path)?;
// The query vector must come from the model the corpus was built with.
if file.model_hash() != model_hash {
return Err(format!("corpus model_hash is {}", file.model_hash()).into());
}
let result = match file.declared_index_type() {
"hnsw" => file.search_ann(query, k, 100)?,
_ => file.search(query, k)?,
};
Ok(result)
}The Rust crates never compare model hashes on their own, except search_space when you pass expected_model_hash. See The model gate.
urna-format
Crate root
The root re-exports what most callers need; the rest is under the public modules bytes, chunk, encoding, error, layout, manifest, reader, sections and writer.
| Group | Items |
|---|---|
| build | UrnaFileBuilder, ChunkInput, Manifest, Capabilities, SectionEncoding, EmbeddingDType, chunk_id |
| read | UrnaView |
| layout | everything in layout: UrnaHeader, SectionEntry, UrnaFooter, magic, size and version constants, SECTION_* ids, SECTION_ENCODING_* ids, CANONICAL_SECTIONS, REQUIRED_SECTIONS, OPTIONAL_SECTIONS, section_name, align_up |
| decode sections | decode_chunk_ids, decode_chunks_canonical, decode_chunks_original_spans (OriginalSpan), decode_provenance, decode_search_contract (SearchContract), decode_space_table (SpaceEntry, SPACE_DTYPE_*), decode_blob_refs (BlobRefRecord), decode_blob_span_overlay (BlobSpanEntry, BLOB_REF_NONE), decode_blob_data_table (BlobDataTable), decode_graph_adjacency, parse_csr_parts (Edge, CsrParts, EDGE_TYPE_*, GRAPH_MAX_DEGREE), TxtStreams |
| encode payloads | encode_graph_adjacency, encode_blob_refs, encode_blob_span_overlay, encode_blob_data, encode_space_table, encode_chunks_canonical, f32_to_f16_bytes, f16_bytes_to_f32, encode_int8_embeddings, Int8EmbeddingsView, quantize_f32_to_i8, encode_int4_embeddings, Int4EmbeddingsView, quantize_f32_to_i4, pack_nibbles, nibble_to_i4, int4_blocks_per_row, INT4_BLOCK, expected_embeddings_size |
| errors | UrnaError, Result |
CapabilitiesExt is at urna_format::manifest::CapabilitiesExt and REPRODUCIBLE_CREATED at urna_format::writer::REPRODUCIBLE_CREATED; neither is re-exported at the root.
UrnaFileBuilder
A consuming builder: every method takes self and returns Self. UrnaFileBuilder::new(manifest) starts with provenance {}, reproducible = false, SectionEncoding::Raw and EmbeddingDType::Float32.
| Method | Effect on the manifest | Section written |
|---|---|---|
add_chunk(ChunkInput), add_chunks(iter) | none | feeds 0x01 to 0x04 |
with_provenance(serde_json::Value) | none | 0x05 |
reproducible(bool) | at build, sets created = "1970-01-01T00:00:00Z" | none |
text_encoding(SectionEncoding) | none | Zstd compresses text sections and runs the text-codec chooser |
embedding_dtype(EmbeddingDType) | dtype | 0x04 in the matching encoding |
hnsw_index(Vec<u8>) | index_type = "hnsw", rerank_policy = "exact", supports_ann = true | 0x07, raw |
bm25_index(Vec<u8>) | supports_bm25 = true | 0x08, raw or zstd per text encoding |
hybrid() | index_type = "hybrid", rerank_policy = "exact", score_type = "hybrid_rrf", supports_bm25 = true | none |
graph_adjacency(Vec<u8>) | capabilities_ext.graph_present = true | 0x0C, raw |
blob_refs(Vec<u8>) | capabilities_ext.blobs_present = true | 0x14, raw |
blob_data(Vec<u8>) | blobs_present = true | 0x17, raw |
blob_span_overlay(Vec<u8>) | blobs_present = true | 0x16, raw |
space_table(Vec<u8>) | capabilities_ext.supports_multimodal = true | 0x15, raw |
space_band(space_index: u8, encoding: u32, payload: Vec<u8>) | supports_multimodal = true | 0x20 + space_index, in your encoding |
build_bytes() | validates | returns the file as Vec<u8> |
write_to_path(path) | validates | build_bytes then std::fs::write |
ChunkInput has five public fields: canonical_text: String, source_uri: String, byte_start: u64, byte_end: u64, embedding: Vec<f32>. The writer rejects a wrong vector length (DimensionMismatch), NaN or Inf (InvalidEmbeddingValue) and byte_end below byte_start (InvalidInput), and requires the chunk count to equal manifest.n_chunks. It expects L2-normalized vectors and does not check the norm.
Optional payloads are opaque to the builder. It does not decode them, so a malformed index or blob payload builds and fails at MmapUrnaFile::open.
Attaching HNSW changes content_hash
hnsw_index() rewrites index_type and rerank_policy, which are copied into the canonical search_contract section. A file with HNSW and the same file without it have different content_hash values and different citations. See Hashes and citations and Known limits.
Presets in Rust
The Python presets map to these builder calls:
| Preset | Builder calls |
|---|---|
exact | defaults |
compressed | .text_encoding(SectionEncoding::Zstd).embedding_dtype(EmbeddingDType::Float16) |
tiny | Zstd, Int8, .hnsw_index(...) |
nano | Zstd, Int4 (dimension a multiple of 64), .hnsw_index(...) |
hybrid | Zstd, Float32, .hnsw_index(...), .bm25_index(...) |
The Python hybrid preset never calls .hybrid(), so its files declare index_type = "hnsw". Build the index payloads like this:
use urna_runtime::ann::{HnswIndex, DEFAULT_EF_CONSTRUCTION, DEFAULT_M};
use urna_runtime::bm25::{Bm25Index, DEFAULT_B, DEFAULT_K1};
// vectors: n * dim f32 values, row-major; texts: the canonical texts in chunk order.
let hnsw = HnswIndex::build(vectors, n, dim, DEFAULT_M, DEFAULT_EF_CONSTRUCTION, 42).to_bytes();
let bm25 = Bm25Index::build(&texts, DEFAULT_K1, DEFAULT_B).to_bytes();
let builder = builder.hnsw_index(hnsw).bm25_index(bm25);DEFAULT_M is 16 and DEFAULT_EF_CONSTRUCTION 400; the Rust side has no default seed (Python uses 42). BM25 defaults are k1 = 1.5, b = 0.75. Build the HNSW index from float32 vectors before quantization, as the Python bridge does. Same inputs and seed give the same bytes.
UrnaView
UrnaView::from_bytes(&[u8]) parses and validates a file held in memory, in the order listed on Layout.
| Item | Use |
|---|---|
validate_embeddings_values() | walk 0x04 for NaN and Inf (not run by from_bytes) |
entry(id) | the table entry for a section id, or SectionNotFound |
get_section_data(id) | the stored bytes of a section |
decoded_section(id) | the bytes after codec, dictionary and dedup expansion |
search_contract() | the decoded SearchContract |
file_hash_hex(), content_hash_hex() | the two sha256:<hex> hashes |
len(), is_empty(), raw_bytes() | the underlying slice |
fields header, section_table, manifest, footer | parsed structures |
urna_format::reader::validate_slab_values(encoding, data, n, dim) runs the NaN and Inf walk on any vector slab.
urna-runtime
Crate root
| Group | Items |
|---|---|
| open and search | MmapUrnaFile |
| results | SearchResult, SearchHit, SearchExplain, RerankSourceKind |
| types | DType, SimdBackend, RuntimeError |
| index payloads | ann::HnswIndex, ann::DEFAULT_M, ann::DEFAULT_EF_CONSTRUCTION, bm25::Bm25Index, bm25::DEFAULT_K1, bm25::DEFAULT_B |
| SIMD | simd::detect_backend |
MmapUrnaFile
MmapUrnaFile::open(&Path) memory-maps the file, runs the full UrnaView validation and the NaN walk, decodes chunk ids, spans and every index or table the file carries, and computes file_hash and content_hash. The struct owns the map and has no lifetime parameter; dropping it unmaps the file.
| Method | Returns | Notes |
|---|---|---|
open(&Path) | Result<Self, RuntimeError> | full validation |
embedding_dim(), n_embeddings() | usize | header values |
dtype() | DType | from the manifest |
file_hash(), content_hash() | &str | sha256:<hex> |
model_hash() | &str | manifest value, for your own model check |
declared_index_type(), declared_score_type() | &str | manifest values |
simd_backend() | SimdBackend | detected once per process |
has_ann(), has_bm25(), has_graph(), has_blobs(), has_spaces() | bool | whether that index or table was opened |
has_blob_data() | bool | whether 0x17 was opened |
blob_refs() | Option<&[BlobRefRecord]> | in table order |
blob_bytes(index) | Result<&[u8], RuntimeError> | one inlined blob, sliced from the map |
space_names() | Vec<&str> | named spaces in table order |
chunk_ids() | &[String] | in file order |
canonical_texts() | Result<Vec<String>, RuntimeError> | stored texts in file order; re-parses and re-validates the whole file on each call |
inspect_json() | Result<String, RuntimeError> | the JSON urna inspect --json prints |
revalidate() | Result<(), RuntimeError> | re-runs the parse, the NaN walk and the contract decode |
madvise_cold() | () | advises the OS to drop the mapped pages (Unix); no-op elsewhere |
Search functions
All take k: i32 and have no defaults. Every returned score is a real cosine recomputed by the exact rerank.
| Function | Parameters | Path |
|---|---|---|
search | query: &[f32], k | exact over every vector |
search_ann | query, k, ef_search: usize | HNSW shortlist, exact rerank; exact search when the file has no HNSW |
search_graph | query, k, hops: usize, ef: usize | exact top max(ef, k) seeds, bounded BFS over 0x0C, exact rerank; exact search when the file has no graph |
search_hybrid | query_vec, query_text: &str, k, candidates_per_path: usize | vector shortlist (HNSW, or exact when absent) union BM25 shortlist, exact rerank |
search_space | name: &str, query, k, expected_model_hash: Option<&str> | exact over one named space band |
Queries are checked in this order: k at least 1 (InvalidK), not empty (EmptyQuery), length equals the dimension (DimensionMismatch), no NaN or Inf (InvalidQueryValue), nonzero norm (ZeroNormQuery). The runtime then L2-normalizes the query. search_space first checks the space name (SpaceNotFound) and, when you pass a hash, the space model_hash (SpaceModelMismatch), then runs the same checks against the space dimension.
ef below the file's ef_construction changes nothing
search_ann searches with a beam of the largest of ef_search, k and the ef_construction stored in the file. Python and the forge build with ef_construction = 400, so any ef_search below 400 behaves like 400 on those files. MmapUrnaFile keeps its index private, so you cannot lower it. See Known limits.
search_hybrid ranks by cosine only
search_hybrid fuses the vector and BM25 shortlists with RRF to form the candidate set, then reranks every candidate by exact cosine and discards the RRF order. BM25 can add a candidate but never lifts one above a higher-cosine chunk. Without HNSW, the vector shortlist has exactly candidates_per_path entries, so a value below k returns fewer than k hits. See Known limits.
search_graph returns the same hits as search: its seeds are the exact top max(ef, k) and the rerank uses the same scores. What differs is route, recall and the candidate counts.
Result types
SearchResult:
| Field | Type | Meaning |
|---|---|---|
hits | Vec<SearchHit> | best first |
query_time_ms | f64 | includes validation and building the hits |
index_type | &'static str | exact, hnsw, graph, hybrid or space |
recall | f32 | 1.0 on exact paths, NaN on candidate paths (never estimated) |
truncated | bool | k is below the number of vectors searched |
k_requested | i32 | the k you passed |
k_returned | usize | number of hits |
explain | SearchExplain | route and candidate counts |
SearchHit:
| Field | Type | Meaning |
|---|---|---|
chunk_id | String | sha256:<hex> |
score | f32 | exact cosine |
score_type | &'static str | always "cosine" |
source_uri | String | from 0x03 |
offset_start, offset_end | u64 | span from 0x03, or blob-relative when 0x16 is present |
embedding_model | String | manifest value |
index_type | &'static str | the path that produced the hit |
reranked | bool | true on candidate paths |
file_hash, content_hash | String | sha256:<hex> |
citation_id | String | urna://<content_hash>/<chunk_id> |
SearchExplain (Copy) has route, exact_candidates, ann_candidates, bm25_candidates, graph_candidates, fusion_mode ("none" or "rrf"), rerank_source and recall_estimate.
RerankSourceKind says which vectors the rerank read: FullPrecision (disclosure() gives "real cosine", as_str() gives "full_precision") when the rerank read float32 vectors (the stored dtype, or a 0x09 full-precision slab when one is present), else StoredPrecision ("real cosine at stored precision", "stored_precision").
| Path | route | index_type | recall | reranked |
|---|---|---|---|---|
search | exact | exact | 1.0 | false |
search_ann with HNSW | hnsw | hnsw | NaN | true |
search_graph with a graph | graph | graph | NaN | true |
search_hybrid | hybrid | hybrid | NaN | true |
search_space | exact | space | 1.0 | false |
DType is Float32, Float16, Int8 or Int4, with name(). SimdBackend is Scalar, Avx2 or Neon, with name().
SIMD
The runtime picks a kernel once per process: AVX2 when the CPU has both avx2 and fma, NEON on aarch64, scalar otherwise. Set URNA_FORCE_SCALAR to any value other than 0 to force the scalar path. float32, int8 and int4 have AVX2 and NEON kernels. float16 runs scalar on x86_64; on aarch64 it uses NEON only when the crate was compiled with rustc 1.94 or newer. All paths accumulate in f32, and the int4 kernel gives identical results on all three. See Environment variables.
Every error these crates return is on Typed errors.
The wheel's urna command
The urna console command installed by the Python wheel, its four read-only verbs, how it differs from the Rust binary, and the PATH collision between the two.
Layout
Byte layout of a .urna v1 file: the 128-byte header, 32-byte section table entries, the manifest, 64-byte aligned payloads and the 40-byte footer.