docsv0.5.1

Changelog

Release history of urna from 0.1.0 to 0.5.1, newest first, with the user-visible changes of each release and the work on main since 0.5.1.

The user-visible changes of each urna release, newest first. The complete history, with measurements, test counts and the reasoning behind each change, is docs/CHANGELOG in the repository, and the release artifacts are on the GitHub releases page.

The tagged releases are v0.1.0, v0.2.0, v0.3.0, v0.5.0 and v0.5.1. Version 0.4.0 exists in the changelog but was never tagged. Until 0.5.0 the project was called nest (files .nest, magic NEST); the entries below use the current names.

Unreleased

Changes on main since v0.5.1. None of them changes the behavior of the binary, the Python package or the file format, so the 0.5.1 documentation applies to main as well.

  • The file-size guard in scripts/release_check.sh and CI allows 639 lines per Rust source file.
  • install-test.yml runs inside every release, after the announce step. 0.5.0 and 0.5.1 were tested by hand. It now covers bun, pnpm and yarn, and its wheel job installs the tag's version.
  • Documentation: AGENTS.md rewritten as a contract, a YAML header on every repository doc, llms.txt in the llms.txt format, an afterwork checklist in the pull request template, Homebrew documented as brew tap hoffresearch/urna then brew install urna, and bun, pnpm and yarn documented as install channels for @urna/cli.

0.5.1 (2026-09-26)

The same binaries as 0.5.0, republished so the registry pages show install commands that work: the 0.5.0 pages on crates.io and npm said npm install -g urna, and the PyPI page said cargo install urna-cli. Neither exists.

  • Fixed the Windows one-liner: the release zip holds urna.exe at its root, and install.ps1 looked for it in a subfolder. The script now takes urna.exe wherever the zip puts it. install.ps1 is served from main, so the fix also reaches 0.5.0 installs.
  • The npm package is @urna/cli, and this is the first tag whose npm package CI publishes.
  • Repository docs renamed to docs/BENCH.md, docs/USAGE.md and docs/arc/ARC.toml.
  • pillow joins the forge dependency group in pyproject.toml.

0.5.0 (2026-09-26)

The first tag served by the release pipeline. No format change.

  • The project is urna: repository hoffresearch/urna, binary urna, Python module and wheel urna, environment variables URNA_*, citation scheme urna://, file extension .urna, magic URNA and chunk-id domain urna:chunk_id:v1. The reader still opens files with the legacy NEST magic. Chunk ids in a legacy file were derived under the old domain, so urna.chunk_id does not reproduce them. The layout, section ids, encodings and manifest schema are unchanged (format v1, schema v1).
  • Release channels: archives for five targets on the GitHub release, the Homebrew tap hoffresearch/homebrew-urna, npm @urna/cli, crates.io (urna-format, urna-runtime, and the CLI as urna, so cargo install urna), and the PyPI wheel urna with the potion table bundled.
  • New urna setup: downloads the embedder payload through a system curl child, checks its SHA-256 while streaming, unpacks it into a staging directory, then builds a Python env at <data root>/urna/venv with uv, or python3 -m venv and pip, and runs the doctor checks. New exit codes 10 to 14. Flags --version, --force, --no-payload, --no-python and --uninstall. See urna setup.
  • New urna tui, a terminal explorer with home, corpus, ask and health tabs. A bare urna on a terminal opens it; in a pipe it prints the help and exits 2. See the terminal explorer.
  • cargo install urna --no-default-features builds the engine-only CLI, without setup and tui.
  • The CLI crate needs Rust 1.88. urna-format, urna-runtime and urna-python stay at 1.85.
  • One data-root ladder for the payload: URNA_DATA_DIR, XDG_DATA_HOME, ~/.local/share, %LOCALAPPDATA%, then <exe>/../share. This fixed the Windows one-liner, which put the payload under %LOCALAPPDATA%\urna, a directory the CLI never searched.
  • The interpreter ladder checks the setup venv right after URNA_PYTHON, so an installed urna run from inside another project's .venv still uses its own env.
  • urna doctor failures now name urna setup, and a missing embedder reports where it looked.
  • examples/quickstart/ builds from a clean checkout, and urna --help opens with the five verbs of the loop (build, ask, retrieve, cite, validate).

0.4.0 (2026-09-17, never tagged)

There is no v0.4.0 tag: everything below first shipped in a tagged release with 0.5.0.

  • Declarative builds: urna build --spec corpus.toml describes sources (SQLite, CSV, JSONL, an image directory), media, models and outputs in one file. It launches the forge in python/tools/urna_forge.py, so it runs from a checkout of the repository. See build from your own rows.
  • A model registry with the presets potion, clip-vit-b32, siglip2, jina-v5-omni-nano, jina-v5-omni-small, wemm-2b, wemm-4b and wemm-9b. Presets that run model-repository code need [output] allow_remote_code at build time and URNA_ALLOW_REMOTE_CODE at query time; a manifest can no longer grant it. See the model registry.
  • ask and retrieve pick the query embedder from the manifest's embedding_model: the potion script for potion corpora, the registry script for any other model. The name, dimension and model_hash checks they share with search-text live in one module.
  • Media inside the file: the optional sections blob_refs (0x14), blob_span_overlay (0x16) and blob_data (0x17, from [output] embed_media), urna media [--export DIR], media profiles and the crf = "auto" quality gate. None of these sections enters content_hash. See media and named spaces.
  • Named multimodal spaces: the space_table section (0x15) and one vector band per space (0x20 to 0x2F), urna search-space, spaces in urna stats and urna inspect --json, and urna benchmark --space.
  • A shared embed cache under ${XDG_CACHE_HOME:-~/.cache}/urna, overridable with URNA_CACHE_DIR, and ${VAR} expansion in spec paths.
  • New urna doctor, with typed exit codes 0 and 2 to 6.
  • urna --help groups the verbs as engine and agent.
  • The HNSW build is 1.73x faster on 20k vectors of dimension 384, still single-threaded and byte-deterministic.
  • Hardening: deterministic mutation-fuzz harnesses under cargo test, four cargo-fuzz targets, miri on urna-format, cargo deny and a semver check in CI. The first fuzz runs found and fixed five classes of malformed-file bug.
  • A 0.3.0 reader validates a file with media and named spaces and searches its text embeddings.

0.3.0 (2026-06-10)

Additive within format v1: 0.2 files load unchanged.

  • int4 embeddings (encoding 7): 4-bit codes in blocks of 64 dimensions, one f16 scale per block, so embedding_dim must be divisible by 64. Scores are real cosine at the stored int4 precision. New nano preset: zstd text, int4 embeddings and HNSW.
  • Matryoshka truncation: urna.build(mrl_dim=K) keeps the first K components of each vector and renormalizes before quantization. The manifest records mrl_dim and full_dim. content_hash covers the truncated vectors, so a citation is tied to its mrl_dim.
  • The chunk graph: the optional graph_adjacency section (0x0C), urna search-graph, and urna.build(with_graph=True, graph_top_m=...). The graph does not enter content_hash.
  • New encodings for text sections: intpack (4), zstd_dict (5), fsst (9) and txt_streams (10), with the dictionary (0x0A) and dedup_map (0x0B) sections. They decode to the same bytes, so content_hash does not change. A 0.2 reader rejects a file that uses them, and a build with zstd text (every preset except exact) stores its chunk ids with intpack.

0.2.0 (2026-04-28)

Extends format v1: 0.1 files load unchanged.

  • Encodings: zstd (1) for text sections, float16 (2) and int8 (3) for embeddings. Embeddings are never zstd-compressed, so they can be scored straight from the memory map.
  • Optional hnsw_index (0x07) and bm25_index (0x08) sections. HNSW candidates are always reranked by exact cosine, and the build is deterministic for a given seed.
  • Build presets exact, compressed, tiny and hybrid. See build presets.
  • New verbs and flags: urna search-ann, urna search-text (embeds the query in Python and checks embedding_model, the dimension and model_hash against the manifest), urna benchmark --madvise-cold and urna inspect --json.
  • SIMD dispatch detected at runtime: AVX2 on x86_64, NEON on aarch64, a scalar fallback, and URNA_FORCE_SCALAR to force the fallback.
  • python/model_fingerprint.py computes model_hash from the model files, and --model-path points search-text at a local snapshot.

0.1.0 (2026-04-27)

The first public release. It froze format v1: a change to the container, the hashes, the citation URI or the manifest contract bumps URNA_FORMAT_VERSION or URNA_SCHEMA_VERSION.

  • The container: a 128-byte header, 32-byte section table entries, a 40-byte footer, section payloads aligned to 64 bytes, little-endian integers.
  • Six required sections, all canonical: chunk_ids, chunks_canonical, chunks_original_spans, embeddings, provenance and search_contract.
  • The hashes: 8-byte header and section checksums, a 32-byte footer hash, content_hash over the canonical sections and a domain-separated chunk_id. See hashes.
  • The citation URI urna://<content_hash>/<chunk_id>. urna cite rejects a citation whose content_hash does not match the file.
  • Reproducible builds: the reproducible flag pins created to 1970-01-01T00:00:00Z.
  • Version skew: a reader rejects a higher format or schema version and accepts an equal or lower one. See compatibility.
  • The CLI verbs inspect, validate, stats, search, benchmark and cite, with exact search over raw float32 embeddings.
  • The Python module: urna.open, UrnaFile.search, inspect and validate, urna.build and urna.chunk_id.

On this page