Tune media compression
Choose a media profile for an urna image corpus, override single knobs, let the crf gate pick the AV1 rate, or keep JPEGs byte-reversible.
The [media] table of a build spec decides how urna build --spec encodes the images of a corpus: the codec, the rate, the keyframe policy and the frame order. A profile sets all of them from a measured recipe, and any key you write yourself overrides the profile. Every key is listed in [media].
Pick a profile
| Profile | For | Recipe |
|---|---|---|
| (none) | the defaults | AV1 stream, crf = 35, still tune, GOP probed |
near-dup | visual near-duplicates: reprints, video frames, scans | AV1, cluster ordering, GOP probed per segment |
stills | unique images, one file per image | AVIF, avifenc -q 48, speed 8 |
stills-av1 | unique images in an AV1 stream | AV1 all-intra, still tune |
archive | corpora where no loss is acceptable | JPEG XL byte-reversible JPEG transcode |
retrieval | corpora that serve search and never show pixels | AV1 all-intra, still tune, crf = 50, speed 6 |
retrieval-auto | the same, with the rate chosen by a search-quality check | AV1 all-intra, still tune, crf = "auto" on a utility floor |
What the recipes were measured on, as recorded in python/forge/media_profiles.py:
stills: on 38,627 trading-card images (2026-09-12), about 1.20 GB against 1.37 GB for the all-intra AV1 stream, 13% smaller at a matched SSIMULACRA2 mean of 61.96 on a 2,048-image sample. Building is slower: the image embedding pass ran 4 to 10 times slower, because frames decode one AVIF file at a time.retrieval: on the same cards (2026-09-03), 533 MB self-contained, 7.46 times smaller than the JPEG sources, with no measurable loss in text-to-image hit@1 over 100 queries.retrieval-auto: on the same cards (2026-09-12), embedding drift p10 fell from 0.932 at crf 40 to 0.829 at crf 60 and would have refused every rung, while hit@1 over 100 queries did not move up to crf 50. That is why this profile gates on search quality and turns the drift and visual floors off.near-dup: 29% fewer bytes on a corpus of same-artwork reprints, from cluster ordering plus inter coding.
Resolved values
A profile fills in defaults before your explicit keys are applied. These are the values each profile resolves to (bold marks what the profile sets):
| Key | (none) | near-dup | stills | stills-av1 | archive | retrieval | retrieval-auto |
|---|---|---|---|---|---|---|---|
backend | av1 | av1 | avif | av1 | jxl-transcode | av1 | av1 |
crf | 35 | 35 | 48 | 35 | 35 (unused) | 50 | "auto" |
tune | still | still | still (unused) | still | still (unused) | still | still |
speed | 8 | 8 | 8 | 8 | 8 (unused) | 6 | 6 |
gop | auto | auto | auto (unused) | intra | auto (unused) | intra | intra |
order | none | cluster | none | none | none | none | none |
quality.visual_floor_p10 | 60.0 | 60.0 | 60.0 | 60.0 | 60.0 | 60.0 | -1e9 |
quality.visual_floor_min | 45.0 | 45.0 | 45.0 | 45.0 | 45.0 | 45.0 | -1e9 |
quality.drift_floor_p10 | 0.95 | 0.95 | 0.95 | 0.95 | 0.95 | 0.95 | -1.0 |
quality.crf_ladder | 25 to 50, step 5 | same | same | same | same | same | [40, 45, 50, 55, 60] |
quality.utility_floor_hit1 | -1.0 | -1.0 | -1.0 | -1.0 | -1.0 | -1.0 | 0.0 |
quality.utility_tol | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.02 |
width (1024), fps (1), pix_fmt (yuv420p), shard_size (2048) and dedup (true) keep their defaults under every profile.
Override single keys
An explicit key always wins over the profile. The quality table merges key by key, so you can keep a profile's floors and change one:
[media]
profile = "retrieval-auto"
[media.quality]
utility_tol = 0.0Check what a spec resolves to before encoding anything. The dry run prints the media line:
urna build --spec corpus.toml --dry-runmedia: avif crf=48 tune=still speed=8 order=none dedup=TrueBackends
| Backend | Output in <name>.media/ | Tools | Keys that apply |
|---|---|---|---|
av1 | <name>-av1.mp4, or one <name>-av1-NNN.mp4 per shard | ffmpeg and ffprobe with libsvtav1 | width, crf (-crf), speed (-preset, 0 to 13), tune, fps, gop, shard_size, order, pix_fmt |
avif | <name>-avif/NNNNNN.avif | avifenc, avifdec | width, crf (as avifenc -q), speed, pix_fmt (--yuv) |
jxl | <name>-jxl/NNNNNN.jxl | cjxl, plus djxl to embed decoded frames | none; lossless of the source pixels |
jxl-transcode | <name>-jxl/NNNNNN.jxl, or the original file for sources it copies | cjxl, djxl | jxl_transcode.* |
control | <name>-png/NNNNNN.png | none | width |
width is a ceiling: the canvas is the smaller of width and the median source width, so images are never upscaled. yuv444p works on avif only; on av1 the encoder falls back to 4:2:0 and the build stops. The encoders run with fixed parallelism (lp=2 for SVT-AV1, -j 8 for avifenc), so the output bytes do not depend on the machine's core count.
On av1, gop = "auto" encodes up to 32 sample frames both ways and keeps the smaller result, preferring all-intra when an inter win costs more than 2.0 SSIMULACRA2 points. With more unique frames than shard_size, the stream is split and each shard runs its own probe.
order = "similarity" or "cluster" puts similar images next to each other so inter coding has something to predict. It runs an extra, uncached embedding pass over the source images with the [media.cluster] space model. Only av1 applies the order: on the other backends the pass runs and its result is dropped, so leave order = "none" there.
Let the gate choose the crf
crf = "auto" encodes a sample of the corpus at every rung of quality.crf_ladder and keeps the largest crf that passes every enabled check. It needs:
backend = "av1"(orprofile = "stills-av1",retrieval-auto). On any other backend it is a spec error.- An image model in the spec (
image = "space"); the first one is used unlessquality.gate_modelnames another. ffmpegandssimulacra2onPATH(brew install jpeg-xlprovidesssimulacra2). Withoutssimulacra2the build stops.
The crf gate needs a label on every image
The gate reads one label per unique image, and an image without a label crashes the build with AttributeError: 'Row' object has no attribute 'text' (exit 1), even when the utility check is off. Give every row a label ([source] labels for image_dir, [source.image] label_template for table sources), or set a fixed crf. See known limits.
The sample
Images are grouped by the quality.buckets heuristics: resolution (longest side under 512, under 1024, or larger), entropy, has_text, alpha and source_format. Up to sample_per_bucket images (default 12) are taken evenly from each group. Each rung encodes the sample as an all-intra AV1 stream at that crf, whatever gop says, decodes it, and scores the frames.
The three checks
| Check | Passes when | Turned off by |
|---|---|---|
| visual | every group's SSIMULACRA2 p10 is at least visual_floor_p10 (60.0), and the lowest score overall is at least visual_floor_min (45.0) | very low floors, as retrieval-auto sets |
| drift | the p10 of cosine(source embedding, decoded embedding) under the gate model is at least drift_floor_p10 (0.95) | a negative drift_floor_p10 |
| utility | text-to-image hit@1 on the decoded sample is at least max(utility_floor_hit1, hit1_source - utility_tol) | a negative utility_floor_hit1 (the default) |
The utility check turns each sampled image's label into a query with utility_query_template (default "{label}"), embeds it with the gate model's text tower, and asks whether the image's own frame comes first. hit1_source is the same measure on the undecoded source images. utility_queries limits how many sampled images become queries (0 means all). The gate model needs a text tower whose dim matches its image dim: clip-vit-b32, siglip2, the jina and the wemm presets.
The default visual and drift floors come from the card corpus at 488x680 (2,048-image sample, still tune, speed 6): crf 30 measured SSIMULACRA2 p10 65.3, minimum 58.6 and drift p10 0.967 and passes; crf 35 measured p10 55.7 and fails. With the defaults, a similar corpus gets crf 30.
The result
The largest passing crf wins. When no rung passes, the build uses the smallest crf in the ladder and prints [forge] warning: no ladder crf met the floors (...); using smallest crf N. The gate never stops a build on quality.
The manifest records the whole run under media.crf_auto: the groups, the sampled images, every rung's scores and which checks passed, the chosen crf and the gate model's model_hash. --resume reuses an earlier choice while the rows and the [media] table are unchanged.
retrieval-auto turns the visual and drift checks off, but the gate still runs ssimulacra2 on every sampled frame at every rung, so the tool is still required.
Lossless: jxl and jxl-transcode
These are the only lossless backends. jxl keeps the source pixels. jxl-transcode (the archive profile) repacks each JPEG so the original file can be rebuilt byte for byte.
Key under [media.jxl_transcode] | Default | Effect |
|---|---|---|
on_unsupported_jpeg | "copy-source" | for a source that is not a reversibly transcodable JPEG: error stops the build, copy-source stores the original bytes, lossless-jxl stores a lossless JPEG XL of the pixels |
verify_roundtrip | true | rebuilds each JPEG with djxl and requires its SHA-256 to equal the source |
keep_metadata | false | read and never used |
Each file's outcome lands in the manifest as media.decisions[] with transcode, copied or lossless and whether it was verified.
keep_metadata has no effect
keep_metadata = true is accepted and ignored: cjxl always runs without metadata options. See known limits.
Where the choices are recorded
The <name>.manifest.json media block holds what the encoder did: backend, crf or quality, speed, canvas, keyframe interval, tool versions, byte counts, compression_ratio against the original source files (comparable across backends), the gate report and the deduplication counts. The full resolved [media] table, profile name included, is in <name>.build.lock.json under resolved_spec.media.
The profile name is not in the manifest
The manifest records the resolved values but not which profile produced them. Read the profile from the build lock. See known limits.
To see how the chosen media affects search quality on your own data, measure it: python/tools/urna_model_bench.py reports stability, codec drift and task utility for a forge corpus, as separate numbers. For how media is stored inside the file, see Media and named spaces.
Images and PDFs
Build an image corpus from a directory or from rows with image paths using urna build, embed pixels into a named space, and index PDF pages.
Serve a corpus over HTTP
Run the FastAPI and Flask examples from the urna repo to answer questions over HTTP with cited chunks, and what to change before serving a real corpus.