docsv0.5.1

Tune media compression

Choose a media profile for an urna image corpus, override single knobs, let the crf gate pick the AV1 rate, or keep JPEGs byte-reversible.

The [media] table of a build spec decides how urna build --spec encodes the images of a corpus: the codec, the rate, the keyframe policy and the frame order. A profile sets all of them from a measured recipe, and any key you write yourself overrides the profile. Every key is listed in [media].

Pick a profile

ProfileForRecipe
(none)the defaultsAV1 stream, crf = 35, still tune, GOP probed
near-dupvisual near-duplicates: reprints, video frames, scansAV1, cluster ordering, GOP probed per segment
stillsunique images, one file per imageAVIF, avifenc -q 48, speed 8
stills-av1unique images in an AV1 streamAV1 all-intra, still tune
archivecorpora where no loss is acceptableJPEG XL byte-reversible JPEG transcode
retrievalcorpora that serve search and never show pixelsAV1 all-intra, still tune, crf = 50, speed 6
retrieval-autothe same, with the rate chosen by a search-quality checkAV1 all-intra, still tune, crf = "auto" on a utility floor

What the recipes were measured on, as recorded in python/forge/media_profiles.py:

  • stills: on 38,627 trading-card images (2026-09-12), about 1.20 GB against 1.37 GB for the all-intra AV1 stream, 13% smaller at a matched SSIMULACRA2 mean of 61.96 on a 2,048-image sample. Building is slower: the image embedding pass ran 4 to 10 times slower, because frames decode one AVIF file at a time.
  • retrieval: on the same cards (2026-09-03), 533 MB self-contained, 7.46 times smaller than the JPEG sources, with no measurable loss in text-to-image hit@1 over 100 queries.
  • retrieval-auto: on the same cards (2026-09-12), embedding drift p10 fell from 0.932 at crf 40 to 0.829 at crf 60 and would have refused every rung, while hit@1 over 100 queries did not move up to crf 50. That is why this profile gates on search quality and turns the drift and visual floors off.
  • near-dup: 29% fewer bytes on a corpus of same-artwork reprints, from cluster ordering plus inter coding.

Resolved values

A profile fills in defaults before your explicit keys are applied. These are the values each profile resolves to (bold marks what the profile sets):

Key(none)near-dupstillsstills-av1archiveretrievalretrieval-auto
backendav1av1avifav1jxl-transcodeav1av1
crf3535483535 (unused)50"auto"
tunestillstillstill (unused)stillstill (unused)stillstill
speed88888 (unused)66
gopautoautoauto (unused)intraauto (unused)intraintra
ordernoneclusternonenonenonenonenone
quality.visual_floor_p1060.060.060.060.060.060.0-1e9
quality.visual_floor_min45.045.045.045.045.045.0-1e9
quality.drift_floor_p100.950.950.950.950.950.95-1.0
quality.crf_ladder25 to 50, step 5samesamesamesamesame[40, 45, 50, 55, 60]
quality.utility_floor_hit1-1.0-1.0-1.0-1.0-1.0-1.00.0
quality.utility_tol0.00.00.00.00.00.00.02

width (1024), fps (1), pix_fmt (yuv420p), shard_size (2048) and dedup (true) keep their defaults under every profile.

Override single keys

An explicit key always wins over the profile. The quality table merges key by key, so you can keep a profile's floors and change one:

[media]
profile = "retrieval-auto"

[media.quality]
utility_tol = 0.0

Check what a spec resolves to before encoding anything. The dry run prints the media line:

urna build --spec corpus.toml --dry-run
media:  avif crf=48 tune=still speed=8 order=none dedup=True

Backends

BackendOutput in <name>.media/ToolsKeys that apply
av1<name>-av1.mp4, or one <name>-av1-NNN.mp4 per shardffmpeg and ffprobe with libsvtav1width, crf (-crf), speed (-preset, 0 to 13), tune, fps, gop, shard_size, order, pix_fmt
avif<name>-avif/NNNNNN.avifavifenc, avifdecwidth, crf (as avifenc -q), speed, pix_fmt (--yuv)
jxl<name>-jxl/NNNNNN.jxlcjxl, plus djxl to embed decoded framesnone; lossless of the source pixels
jxl-transcode<name>-jxl/NNNNNN.jxl, or the original file for sources it copiescjxl, djxljxl_transcode.*
control<name>-png/NNNNNN.pngnonewidth

width is a ceiling: the canvas is the smaller of width and the median source width, so images are never upscaled. yuv444p works on avif only; on av1 the encoder falls back to 4:2:0 and the build stops. The encoders run with fixed parallelism (lp=2 for SVT-AV1, -j 8 for avifenc), so the output bytes do not depend on the machine's core count.

On av1, gop = "auto" encodes up to 32 sample frames both ways and keeps the smaller result, preferring all-intra when an inter win costs more than 2.0 SSIMULACRA2 points. With more unique frames than shard_size, the stream is split and each shard runs its own probe.

order = "similarity" or "cluster" puts similar images next to each other so inter coding has something to predict. It runs an extra, uncached embedding pass over the source images with the [media.cluster] space model. Only av1 applies the order: on the other backends the pass runs and its result is dropped, so leave order = "none" there.

Let the gate choose the crf

crf = "auto" encodes a sample of the corpus at every rung of quality.crf_ladder and keeps the largest crf that passes every enabled check. It needs:

  • backend = "av1" (or profile = "stills-av1", retrieval-auto). On any other backend it is a spec error.
  • An image model in the spec (image = "space"); the first one is used unless quality.gate_model names another.
  • ffmpeg and ssimulacra2 on PATH (brew install jpeg-xl provides ssimulacra2). Without ssimulacra2 the build stops.

The crf gate needs a label on every image

The gate reads one label per unique image, and an image without a label crashes the build with AttributeError: 'Row' object has no attribute 'text' (exit 1), even when the utility check is off. Give every row a label ([source] labels for image_dir, [source.image] label_template for table sources), or set a fixed crf. See known limits.

The sample

Images are grouped by the quality.buckets heuristics: resolution (longest side under 512, under 1024, or larger), entropy, has_text, alpha and source_format. Up to sample_per_bucket images (default 12) are taken evenly from each group. Each rung encodes the sample as an all-intra AV1 stream at that crf, whatever gop says, decodes it, and scores the frames.

The three checks

CheckPasses whenTurned off by
visualevery group's SSIMULACRA2 p10 is at least visual_floor_p10 (60.0), and the lowest score overall is at least visual_floor_min (45.0)very low floors, as retrieval-auto sets
driftthe p10 of cosine(source embedding, decoded embedding) under the gate model is at least drift_floor_p10 (0.95)a negative drift_floor_p10
utilitytext-to-image hit@1 on the decoded sample is at least max(utility_floor_hit1, hit1_source - utility_tol)a negative utility_floor_hit1 (the default)

The utility check turns each sampled image's label into a query with utility_query_template (default "{label}"), embeds it with the gate model's text tower, and asks whether the image's own frame comes first. hit1_source is the same measure on the undecoded source images. utility_queries limits how many sampled images become queries (0 means all). The gate model needs a text tower whose dim matches its image dim: clip-vit-b32, siglip2, the jina and the wemm presets.

The default visual and drift floors come from the card corpus at 488x680 (2,048-image sample, still tune, speed 6): crf 30 measured SSIMULACRA2 p10 65.3, minimum 58.6 and drift p10 0.967 and passes; crf 35 measured p10 55.7 and fails. With the defaults, a similar corpus gets crf 30.

The result

The largest passing crf wins. When no rung passes, the build uses the smallest crf in the ladder and prints [forge] warning: no ladder crf met the floors (...); using smallest crf N. The gate never stops a build on quality.

The manifest records the whole run under media.crf_auto: the groups, the sampled images, every rung's scores and which checks passed, the chosen crf and the gate model's model_hash. --resume reuses an earlier choice while the rows and the [media] table are unchanged.

retrieval-auto turns the visual and drift checks off, but the gate still runs ssimulacra2 on every sampled frame at every rung, so the tool is still required.

Lossless: jxl and jxl-transcode

These are the only lossless backends. jxl keeps the source pixels. jxl-transcode (the archive profile) repacks each JPEG so the original file can be rebuilt byte for byte.

Key under [media.jxl_transcode]DefaultEffect
on_unsupported_jpeg"copy-source"for a source that is not a reversibly transcodable JPEG: error stops the build, copy-source stores the original bytes, lossless-jxl stores a lossless JPEG XL of the pixels
verify_roundtriptruerebuilds each JPEG with djxl and requires its SHA-256 to equal the source
keep_metadatafalseread and never used

Each file's outcome lands in the manifest as media.decisions[] with transcode, copied or lossless and whether it was verified.

keep_metadata has no effect

keep_metadata = true is accepted and ignored: cjxl always runs without metadata options. See known limits.

Where the choices are recorded

The <name>.manifest.json media block holds what the encoder did: backend, crf or quality, speed, canvas, keyframe interval, tool versions, byte counts, compression_ratio against the original source files (comparable across backends), the gate report and the deduplication counts. The full resolved [media] table, profile name included, is in <name>.build.lock.json under resolved_spec.media.

The profile name is not in the manifest

The manifest records the resolved values but not which profile produced them. Read the profile from the build lock. See known limits.

To see how the chosen media affects search quality on your own data, measure it: python/tools/urna_model_bench.py reports stability, codec drift and task utility for a forge corpus, as separate numbers. For how media is stored inside the file, see Media and named spaces.

On this page