Images and PDFs
Build an image corpus from a directory or from rows with image paths using urna build, embed pixels into a named space, and index PDF pages.
urna indexes images in two ways. urna build --spec takes a directory of images, or table rows that point at image files, and adds image embeddings as a named space next to the text. PDFs go through a separate tool, python/tools/urna_build_image_corpus.py, which renders each page as an image. Both run from a repository checkout with pillow and an image model's packages installed (torch and open_clip_torch for clip-vit-b32).
A corpus from an image directory
[corpus]
name = "photos"
chunker_version = "photos/1"
[source]
kind = "image_dir"
input_dir = "${PHOTOS}"
labels = "labels.csv"
[media]
profile = "stills"
[[models]]
preset = "potion"
text = "default"
[[models]]
preset = "clip-vit-b32"
image = "space"
[output]
dir = "out/photos"
embed_media = trueexport PHOTOS=/data/photos
urna build --spec photos.toml --dry-run
urna build --spec photos.tomlRelative paths in a spec resolve against the working directory, not the spec's directory.
What the source reads
- Every
.jpg,.jpeg,.png,.bmp,.webp,.tiffand.tiffile underinput_dir, recursively, case-insensitive, sorted by path. One image is one chunk. - The key of a chunk is the image's path relative to
input_dir. Itssource_uriisitem://<corpus name>/<relative path>. labelsis optional: a.csvwhose first two columns areimage_id,label(the header row is skipped), or a JSON object. A label matches the relative path or the bare file stem.- The stored text of a chunk is the relative path, followed by
[label]when there is one.urna ask,urna citeandurna retrievereturn that text.
order_by, derive, joins, [source.text] and [source.image] do not apply to image_dir and are ignored.
Why two models
Every spec needs exactly one model with text = "default". Here potion embeds the stored text (file names and labels) into the default space, which is what urna ask searches. clip-vit-b32 with image = "space" embeds the pixels into a named space called clip-vit-b32. Any preset with an image tower works in that role: clip-vit-b32, siglip2, the jina presets and the wemm presets. potion has no image tower.
Images on table rows
A sqlite, csv or jsonl source can carry one image per row. Declare where the file is and, optionally, its label:
[source.image]
path_template = "${MTG_DATA}/images/{id}.jpg"
label_template = "{name}"Both templates are formatted per row with the row's columns. The file must exist, or the build stops with source.image: file missing for key (...). The row's text comes from [source.text] as usual; the image adds the named space. Declaring path_template is what allows image = "space" models on these sources.
Which pixels are embedded
[embedding.image_input] mode chooses what the image model sees:
| Mode | Embeds | Default |
|---|---|---|
decoded_media | the frames decoded back from the encoded media, so the index describes what the file actually serves | when [media] is present |
source | the original image files | when there is no [media] |
decoded_media without a [media] table is a spec error. Switching modes, or adding or removing [media], changes the embedding recipe, so every model is embedded again once, text-only models included.
Storing the images
Without [media], the file holds the vectors and the stored text but not the images: hits point at item:// URIs, and the pixels stay wherever they were.
With [media], the build encodes the images into <name>.media/ next to the .urna and writes blob records that point each chunk at its encoded image. With [output] embed_media = true, the encoded bytes are also copied into the file, so one .urna carries everything. Every row must have an image once [media] is on: a row without one stops the build with media enabled but N rows have no image (first: <key>).
Rows with byte-identical source images share one encoded frame and one image vector ([media] dedup, on by default). Text vectors stay per row.
Choosing a backend and a quality level is covered in Tune media compression. The layout of blobs and spaces inside the file is in Media and named spaces.
Searching the images
urna ask "..." searches the default space: here, the file names and labels. To search the pixels, query the named space with a vector from the same image model, either with urna search-space or from Python with UrnaFile.search_space. The Python example on Media and named spaces embeds a text query with the clip text tower and searches the clip-vit-b32 space.
PDFs
source.kind = pdf_dir fails at build time
A spec with kind = "pdf_dir" passes --dry-run and validation, then always stops when rows load, with spec error: source.kind=pdf_dir: build via forge_pipeline (pages are temporary) and exit code 2. urna build cannot index PDFs in 0.5.1. See known limits.
Index PDFs with the image tool instead:
python python/tools/urna_build_image_corpus.py \
--input-dir /data/manuals \
--output corpora/manuals.urna \
--dataset manuals \
--pdf \
--model ViT-B-32 --pretrained openaiThe tool renders every page of every .pdf under --input-dir at 150 dpi with PyMuPDF (pip install pymupdf), letterboxes the pages onto one canvas, encodes them, embeds the decoded frames with an open_clip model, and writes one chunk per page. A chunk's stored text is the PDF's relative path and the page number, such as field-guide.pdf page 240.
Pass --model and --pretrained for anything but dermatology: the default model is hf-hub:redlessone/DermLIP_ViT-B-16, a dermatology model. A bare architecture name like ViT-B-32 needs its pretrained tag, or open_clip returns random weights.
| Flag | Default | Effect |
|---|---|---|
--input-dir, --output, --dataset | required | source directory, output .urna, corpus name |
--pdf | off | render PDF pages instead of reading image files |
--model, --pretrained | hf-hub:redlessone/DermLIP_ViT-B-16, none | the open_clip model |
--backend | av1 | av1 stream or avif per page |
--crf | 35 | AV1 rate; --avif-quality (default 35) for avif |
--width | 1024 | canvas width ceiling |
--preset, --dtype | compressed, none | the build preset of the output file and a dtype override |
--labels | none | a JSON map or two-column CSV of labels |
--no-compress | off | embed the rendered pages directly and store no media |
--control | off | a lossless PNG control corpus to measure codec cost against |
--sample, --seed | all, 42 | a random subset of pages |
The tool's --help lists the rest (--gop-policy, --all-intra, --shard-size, --order-similarity, --speed, --pix-fmt, --device, --batch-size, --scratch-db).
How the tool's output differs
The image tool predates the forge and writes a different layout:
- The page vectors are the default space of the file, and the manifest's
embedding_modelis the open_clip model id. There is no text space and no named space. - The encoded media stays in the
<name>.media/directory next to the file. The file has no blob sections and cannot inline the media. Copy the directory with the file. - A
<name>.manifest.jsonrecords the input directory, the model and the pages.
urna ask cannot query these files, because their model is neither potion nor a registry preset. Search them with the matching tool, by text or by image:
python python/tools/urna_search_image.py \
--index corpora/manuals.urna \
--query-text "wiring diagram for the pump" -k 5 \
--model ViT-B-32 --pretrained openaiIt checks the model's model_hash against the file before scoring (--skip-model-check turns that off), and --save-frames DIR writes the matched pages back out as PNG files. Use the same --model and --pretrained as the build.
Next: Tune media compression.
Choose and bring embedding models
Pick an embedding model from the urna registry, keep the build and query sides on the same model, run fully offline with --model-path, and load heavy presets.
Tune media compression
Choose a media profile for an urna image corpus, override single knobs, let the crf gate pick the AV1 rate, or keep JPEGs byte-reversible.