docsv0.5.1

Run in Docker

Build the urna Docker image from the repository's Dockerfile, a static binary in a scratch image, and run the engine verbs on a mounted corpus.

The repository ships a Dockerfile that compiles the static Linux binary and copies it into an empty scratch image. The image holds one file, /urna, and no shell, Python or network tools. It serves the engine verbs (validate, inspect, stats and vector search) on a corpus you mount. No image is published: you build it from a checkout.

Build the image

git clone https://github.com/hoffresearch/urna
cd urna
docker build --platform=linux/amd64 -f docker/Dockerfile -t urna .

The build stage uses rust:1-bookworm with musl-tools and runs cargo build --profile dist -p urna --target x86_64-unknown-linux-musl (the dist profile is the release profile with thin LTO, the one the release archives use). The build copies only Cargo.toml, Cargo.lock and crates/, and .dockerignore excludes data/, python/, docs/, target/ and the other large directories, so the git-lfs data and the Python tree never enter the build context.

For an arm64 image, set the TARGET build argument:

docker build -f docker/Dockerfile --build-arg TARGET=aarch64-unknown-linux-musl -t urna .

On Apple silicon, build the aarch64 variant natively as above. Building the amd64 image there goes through QEMU user emulation, which crashes rustc partway through the build.

Run it

The image's entrypoint is the binary, so arguments go straight to urna. Mount the directory that holds your corpus, read-only:

docker run --rm urna --version
docker run --rm -v "$PWD/data:/data:ro" urna validate /data/corpus.urna
docker run --rm -v "$PWD/data:/data:ro" urna stats /data/corpus.urna
docker run --rm -v "$PWD/data:/data:ro" urna cite /data/corpus.urna 'urna://<content_hash>/<chunk_id>'

The binary opens no socket, so the container needs no network: add --network none if your policy asks for it.

What works in the image

VerbIn the image
inspect, validate, stats, media, citeworks
search, search-ann, search-graph, search-space, benchmarkworks, with a query vector you pass as a JSON array
ask, retrieve, search-textfail: no Python to run the query embedder
buildfails: no forge, no Python
doctorexits 2 (no Python interpreter)
setupcompiled in, but has no curl and no Python to work with
tuicompiled in; not a use case for this image

Text queries need an embedder, which needs Python with numpy and tokenizers plus the embedder payload. This Dockerfile does not include them, on purpose: the image stays a single static file. To answer text queries in a container, embed the query outside it and pass the vector to urna search, or use a base image with Python and follow the installation steps inside it.

media --export writes files, so give it a writable mount for the export directory.

For the search verbs and their vector argument, see urna search.

On this page