Offline by construction
Where urna touches the network and where it cannot: the runtime has no network stack, queries embed offline, and only installers and setup download.
Answering a query never needs the network. The Rust binary links no network stack, the file is read from a memory map, and the query embedders the CLI runs are offline by default. The network appears only when something is installed: the binary, the embedder payload, the Python environment, or a model you explicitly allow to download. This page lists each of those places. For the step-by-step air-gapped procedure, see Air-gapped install and queries.
What runs at query time
| Piece | Network |
|---|---|
The urna binary (open, validate, every search path, cite, inspect, stats) | None. The Rust crates link no network library; queries are answered from mmap. |
embed_query_potion.py, the embedder for potion corpora | None. It loads the vendored potion table with numpy and tokenizers, no torch, and never opens a socket. |
embed_query_model.py, the embedder for registry-model corpora (checkout only) | It sets the Hugging Face offline variables before loading a model. For the sentence-transformers presets (jina, wemm) a download needs URNA_ALLOW_DOWNLOAD=1 and no local model directory. |
embed_query.py, the search-text default (checkout only) | None by default. Same offline variables, same URNA_ALLOW_DOWNLOAD=1 opt-in. |
urna doctor | None. Its last check is one real potion embed, offline. |
Because the file carries its own vectors, text and indices, nothing else is fetched: no index server, no remote store, no telemetry.
Where sockets exist
Network access in urna happens during installation or on explicit opt-in.
| Where | What it downloads | Through |
|---|---|---|
install.sh | the release archive for the platform, the embedder payload, and a .sha256 for each | curl, from GitHub releases or URNA_RELEASE_BASE |
install.ps1 | the same four files for Windows | Invoke-WebRequest, from GitHub releases or URNA_RELEASE_BASE |
Homebrew, npm (and bun, pnpm, yarn), cargo install, cargo binstall, pip | the package or release archive | the package manager. The npm wrapper fetches the release archive at install time, or on the first run when the install step was skipped. |
urna setup, payload step | urna-embedder-payload.tar.gz and its .sha256 | a curl child process: the binary itself opens no socket. Only https:// and file:// sources, redirects only to https://, two retries, a 20 second connect timeout. |
urna setup, Python step | numpy and tokenizers | uv pip install or pip install, against the package index they are configured for |
| Python embedders and the build forge | model weights from the Hugging Face hub | only with URNA_ALLOW_DOWNLOAD=1 |
The setup screen's "network: curl, only while downloading" covers the payload step. The Python step also reaches a package index, unless you skip it with --no-python and point URNA_PYTHON at an interpreter that already has numpy and tokenizers.
The potion payload
The default embedder is model2vec potion-base-8M, a static token table of about 30 MB (model.safetensors is 30,236,760 bytes). A static table needs no GPU and no model runtime: a query becomes the mean of its token rows, L2-normalized, with numpy. It is deterministic, so the same text gives the same vector on every machine.
The table travels with urna in two ways:
- The embedder payload,
urna-embedder-payload.tar.gzin each release, holds the potion scripts and the table. The one-line installers unpack it;urna setupdownloads and verifies it for every other channel. It lands in<data root>/urna/forge/. - The Python wheel bundles the same table under
urna/models/potion-base-8M/, sourna.embed_potionworks right afterpip install "urna[embed]".
Once the payload and a Python with numpy and tokenizers are on the machine, urna doctor proves the chain with no network: interpreter, dependencies, embedder script, table, and one real embed that must return dim=256 and a model_hash.
The payload carries only the potion embedder. A corpus built with a registry model needs the repository checkout and that model's own dependencies and weights on disk; see Open a corpus you downloaded.
How model downloads are gated
Every Python entry point that can reach the Hugging Face hub (the potion module, the registry embedder, the search-text embedder, the fingerprint helper, the build forge) sets these variables to 1 before it imports a Hugging Face library:
HF_HUB_OFFLINETRANSFORMERS_OFFLINEHF_DATASETS_OFFLINE
They are set with setdefault, so a value already present in your environment wins. If your shell exports HF_HUB_OFFLINE=0, the scripts keep it.
URNA_ALLOW_DOWNLOAD=1 is the opt-in, and only the exact value 1 counts. With it, embed_query.py, the fingerprint helper and the registry's sentence-transformers presets may fetch a model that is not on disk. The registry presets allow the download only when no local model directory resolved (no --model-path, no URNA_MODEL_DIR_<PRESET>, nothing in the cache). The potion embedder never downloads anything: a missing table is an error (exit 3 from the script, a failed check 5 in urna doctor), fixed by urna setup or, in a checkout, git lfs pull.
A separate switch guards model code rather than bytes. Presets that need trust_remote_code (jina, wemm) run only with an explicit opt-in: allow_remote_code in the build spec, URNA_ALLOW_REMOTE_CODE="<preset>" on the query side, and a pinned sha256 for each code file where the registry has pins. A downloaded .urna cannot turn that on by itself.
What offline does not mean
Offline is about sockets, not trust. The query embedder runs Python code from the payload or the checkout, under an interpreter chosen by a fixed ladder that can pick up a nearby .venv; pin URNA_PYTHON in directory trees you do not control. And a .urna file from elsewhere is untrusted input even with no network involved. See Security.
The model gate
Why urna refuses a text query embedded by a different model than the corpus, what model_hash covers, and where the check runs or does not.
Presets and stored precision
How urna stores vectors as float32, float16, int8 or int4, what that costs in recall, and how every score discloses its rerank precision.