[media]
Reference for [media] in a urna build spec: backends, every knob, profiles and their resolved values, the crf gate, ordering, and jxl transcoding.
[media] tells the build to encode the source images and store them with the corpus, so a hit can return the image it came from. Its presence turns the media stage on. This page lists every key of [media] and its three subtables, the profiles, and the crf gate behind crf = "auto". For choosing settings, see Tune media compression.
[media]
profile = "stills" # a measured starting point; explicit keys below still win
dedup = true
[output]
embed_media = true # store the encoded bytes inside the .urnaWhen the media stage runs
The stage runs when the spec has [media] and the rows carry images (kind = "image_dir", or a [source.image] path_template). With [media] and no images at all, the stage is skipped. When some rows have an image and others do not, the build stops with a traceback (exit 1) ending in:
ForgeError: media enabled but N rows have no image (first: <key>)The stage encodes each unique image once (identical source files share a frame when dedup = true), writes the result to <name>.media/ beside the .urna, and records the backend, the encoder settings and the sizes in the manifest's media block. The written .urna references the media as blobs, with a span per chunk pointing at its frame or file. With [output] embed_media = true the bytes are also stored inside the .urna.
[media]
| Key | Type | Default | Allowed | Meaning |
|---|---|---|---|---|
profile | string | "" | "" or a profile name | Fills in knob values before the explicit keys apply |
backend | string | "av1" | av1, avif, jxl, jxl-transcode, control | The encoder. See backends |
width | integer | 1024 | positive | Maximum canvas width. See canvas |
crf | integer or "auto" | 35 | any integer, or "auto" | Quality for av1 (-crf) and avif (avifenc -q). "auto" runs the crf gate, av1 only |
tune | string | "still" | still, default | SVT-AV1 tune for av1 |
speed | integer | 8 | encoder range | av1 preset (0 to 13) or avifenc --speed |
fps | integer | 1 | positive | Frame timestamps of the av1 stream. One frame per unique image |
pix_fmt | string | "yuv420p" | yuv420p, yuv444p | Chroma format for av1 and avif |
shard_size | integer | 2048 | 0 or more | av1 only: split the stream into one file per this many unique frames. 0 never splits |
gop | string | "auto" | auto, intra, inter | av1 only: keyframe policy. See GOP |
order | string | "none" | none, similarity, cluster | Reorder frames so similar images sit together. Applied by av1 only. See ordering |
dedup | boolean | true | Rows with byte-identical source images share one frame and one image vector | |
quality | table | [media.quality], read only with crf = "auto" | ||
cluster | table | [media.cluster], read only when order is not none | ||
jxl_transcode | table | [media.jxl_transcode], read only by backend = "jxl-transcode" |
Validation checks profile, backend, gop, order, tune, that crf is an integer or "auto", and the rules below. It does not check width, speed, fps, pix_fmt or shard_size: a wrong value is used as given and can fail during encoding with a Python traceback (exit 1).
Validation errors (exit 2):
media.profile: unknown '<p>' (valid: [...]), at parse time.media.crf: "auto" gates the av1 ladder only (backend = "av1", or profile = "stills-av1"); the avif and jxl backends take a fixed crf.media.quality.gate_model: crf=auto needs an image model, andmedia.quality.gate_model '<g>' must be a spec model with image=space.media.cluster.space: order=<o> needs an image model, andmedia.cluster.space '<s>' must be a spec model with image=space.output.embed_media requires a [media] section (there is nothing to inline).
Backends
| Backend | Produces | Files in <name>.media/ | One blob per |
|---|---|---|---|
av1 | An AV1 video stream (SVT-AV1 in ffmpeg), one frame per unique image | <name>-av1.mp4, or <name>-av1-000.mp4, <name>-av1-001.mp4 and so on when sharded | Stream file; the chunk's span is its frame index |
avif | One AVIF per image (avifenc) | <name>-avif/000000.avif and on | File |
jxl | One lossless JPEG XL per source image (cjxl -d 0), no canvas | <name>-jxl/000000.jxl and on | File |
jxl-transcode | One JPEG XL per source JPEG that rebuilds the original bytes exactly (cjxl --lossless_jpeg=1) | <name>-jxl/000000.jxl, or 000000.<ext> for a copied source | File |
control | One lossless PNG per image at the canvas size | <name>-png/000000.png and on | File |
Required tools: ffmpeg and ffprobe with libsvtav1 for av1 (and for the crf gate and the GOP probe); avifenc and avifdec for avif; cjxl for the jxl backends, plus djxl to verify transcodes and to embed decoded JPEG XL frames; ssimulacra2 for crf = "auto". A missing tool fails when the stage runs, for example jxl backend needs 'cjxl' on PATH: brew install jpeg-xl. The manifest's source_bytes is always the size of the original files, so compression_ratio compares across backends.
Which knobs apply
| Knob | av1 | avif | jxl | jxl-transcode | control |
|---|---|---|---|---|---|
width (canvas) | yes | yes | no | no | yes |
crf | -crf | -q | no | no | no |
crf = "auto" | yes | spec error | spec error | spec error | spec error |
speed | -preset | --speed | no | no | no |
tune | yes | no | no | no | no |
fps, gop, shard_size | yes | no | no | no | no |
order | applied | computed, then dropped | computed, then dropped | computed, then dropped | computed, then dropped |
pix_fmt | yes | --yuv 420 or 444 | no | no | no |
[media.jxl_transcode] | no | no | no | yes | no |
Canvas
av1, avif and control fit every image onto one canvas, letterboxed. The build opens up to 128 evenly spaced images and takes the median width and the median aspect ratio. The canvas width is the smaller of width and that median width, the height follows from the median aspect, and both are rounded down to even numbers. width is a ceiling: images are never upscaled past the median source width. The JPEG XL backends keep each source image as it is.
Encoder details
tune = "still"applies only to all-intraav1encodes. The tune number is probed on the local SVT-AV1 build; when it has none, the build prints[forge] warning: local SVT-AV1 has no Still Picture tune; using default tuneon stderr and recordstune_resolved: null.pix_fmt = "yuv444p"onav1: if the encoder writes 4:2:0 anyway, the build stops (encoder wrote yuv420p for a requested yuv444p ...). Onavif, a value other than the two allowed ones fails with a PythonKeyError.- SVT-AV1 runs with
lp=2andavifencwith 8 jobs, so the output bytes do not depend on the machine's core count.
GOP
gop sets the keyframe interval of the av1 stream:
intra: every frame is a keyframe (keyint 1), so any frame decodes on its own. The manifest recordsgop.policy = "flag".inter: a keyframe every 16 frames with scene-change detection off. Smaller when neighbouring images are similar; decoding a frame means decoding from its keyframe.auto: a probe encodes up to 32 sampled frames both ways at the target crf and keeps the smaller result, ties going to intra. An inter win is overruled when its meanssimulacra2score is more than 2.0 below the intra arm. Withoutssimulacra2the probe compares bytes only and warns. The probe runs once per stream, or once per shard when the stream is sharded.
Ordering
order puts similar images next to each other before encoding, which lets inter frames pay off:
similarity: a greedy nearest-neighbour walk starting from the first image.cluster: greedy clustering in row order; an image joins a cluster when its cosine to the cluster's running centroid reaches[media.cluster] threshold. Clusters are written largest first.
Both use the source-image embeddings of the [media.cluster] space model, from an extra embedding pass that is not cached. Only av1 applies the order. The other backends compute it, paying for that pass, and then drop it. The manifest records the order and the permutation.
Profiles
A profile fills in knob values before the explicit keys apply, so an explicit key always wins. The quality table merges key by key: an explicit [media.quality] key overrides the profile's value for that key only. Resolved values, with the profile's own settings marked by an asterisk:
| Knob | no profile | near-dup | stills | stills-av1 | archive | retrieval | retrieval-auto |
|---|---|---|---|---|---|---|---|
backend | av1 | av1 | avif* | av1 | jxl-transcode* | av1 | av1 |
crf | 35 | 35 | 48* | 35 | 35 (unused) | 50* | "auto"* |
tune | still | still* | still (unused) | still* | still (unused) | still* | still* |
speed | 8 | 8 | 8* | 8 | 8 (unused) | 6* | 6* |
gop | auto | auto* | auto (unused) | intra* | auto (unused) | intra* | intra* |
order | none | cluster* | none | none | none | none | none |
quality.visual_floor_p10 | 60.0 | 60.0 | 60.0 | 60.0 | 60.0 | 60.0 | -1e9* |
quality.visual_floor_min | 45.0 | 45.0 | 45.0 | 45.0 | 45.0 | 45.0 | -1e9* |
quality.drift_floor_p10 | 0.95 | 0.95 | 0.95 | 0.95 | 0.95 | 0.95 | -1.0* |
quality.crf_ladder | 25 to 50, step 5 | same | same | same | same | same | 40, 45, 50, 55, 60* |
quality.utility_floor_hit1 | -1.0 | -1.0 | -1.0 | -1.0 | -1.0 | -1.0 | 0.0* |
quality.utility_tol | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.02* |
Every other knob keeps its default under every profile.
near-dup: for near-duplicate images (reprints, video frames, scans). Cluster ordering plus per-segment GOP lets inter coding pay where images repeat. Per-segment probing needs the stream to be sharded (more unique frames thanshard_size); otherwise there is one probe for the whole stream.stills: unique images, one AVIF per image at quality 48, speed 8. Each image decodes on its own with no video decode. Image-space models embed decoded frames one AVIF at a time, which makes that embedding step slower than on theav1stream.stills-av1: the all-intraav1stream with the still tune. Use it when you wantcrf = "auto".archive: byte-reversible JPEG repacking with a verified round trip, for corpora where no loss is acceptable.retrieval: for corpora that serve search and whose images nobody looks at: all-intra, still tune, fixedcrf = 50, speed 6.retrieval-auto:retrievalwithcrf = "auto"gated on search quality alone. The visual and drift checks are off; every rung must keep text-to-image hit@1 within 0.02 of the source images. It needs a gate model with a text tower and a label on every row.
The profile name is recorded in the build lock (resolved_spec.media.profile), not in the manifest. The manifest records the resolved backend settings.
[media.quality]
Read only when crf = "auto".
| Key | Type | Default | Meaning |
|---|---|---|---|
strategy | string | "stratified" | Not read |
buckets | list of strings | ["resolution", "entropy", "has_text", "alpha", "source_format"] | Image traits that form the sampling strata. An unknown name fails when the gate runs: ValueError: media.quality.buckets: unknown bucket '<n>' (exit 1) |
sample_per_bucket | integer | 12 | Images sampled per stratum, evenly spaced |
visual_floor_p10 | float | 60.0 | Minimum 10th percentile of ssimulacra2 in every stratum |
visual_floor_min | float | 45.0 | Minimum ssimulacra2 over the whole sample |
drift_floor_p10 | float | 0.95 | Minimum 10th percentile of cosine(source embedding, decoded embedding). Negative turns the drift check off |
gate_model | string | "" | Model that embeds the sample. "" means the first image = "space" model. Must be a spec model with image = "space" |
crf_ladder | list of integers | [25, 30, 35, 40, 45, 50] | The crf values tried, in ascending order |
utility_floor_hit1 | float | -1.0 | Minimum text-to-image hit@1, from 0 to 1. Negative turns the utility check off |
utility_queries | integer | 0 | Number of queries, evenly spaced over the sample. 0 means one per sampled image |
utility_query_template | string | "{label}" | Query text per image. Must contain {label}, and {label} is the only placeholder allowed |
utility_tol | float | 0.0 | Allowed hit@1 loss against the source images, from 0 to 1 |
The utility_* keys are validated only when utility_floor_hit1 is 0 or more: utility_floor_hit1 and utility_tol must lie within 0 to 1, utility_queries must be 0 or more, the template must contain {label}, and the gate model must have a text tower (media.quality.utility_floor_hit1: gate model '<g>' has no text tower).
The default floors were set on trading-card images at 488 by 680 pixels (4:2:0, 2048-image sample, av1 still tune, speed 6). There, crf 30 passed (ssimulacra2 p10 65.3, minimum 58.6, drift p10 0.967) and crf 35 failed on p10 (55.7), so the defaults choose crf 30 for that corpus.
The crf gate
With crf = "auto" the build picks the crf itself, per corpus, before encoding. It needs backend = "av1", an image = "space" model, ffmpeg, ssimulacra2 (media.crf="auto" needs ssimulacra2 on PATH: brew install jpeg-xl otherwise) and pillow, plus a label on every row.
crf = "auto" fails on rows without a label
The gate reads one label per unique image, and a row whose label is empty or missing makes the build crash with AttributeError: 'Row' object has no attribute 'text' (exit 1). This happens even with the utility check off. An image_dir source without a labels file, a source without [source.image] label_template, or a template that renders empty for any row all hit it. Give every row a label, or set a fixed crf. See Known limits.
How it runs:
- Sample. Every unique image gets a stratum from
buckets:resolution(longest side under 512 is small, under 1024 medium, else large),entropy(variance of the gradient magnitude on a 128-pixel grayscale thumbnail: under 100 low, under 1000 mid, else high),has_text(more than 10% of gradient magnitudes above 40),alpha(the image has transparency) andsource_format(the file extension). Up tosample_per_bucketevenly spaced images per stratum form the sample. - For each crf in
crf_ladder, from low to high: encode the sample as an all-intraav1stream with the spec'sspeed,pix_fmtandtune, decode it, score each frame withssimulacra2against the letterboxed source, and embed the decoded frames with the gate model. - Check each rung:
- Visual: every stratum's
ssimulacra2p10 is at leastvisual_floor_p10, and the sample's minimum is at leastvisual_floor_min. - Drift: off when
drift_floor_p10is negative; else the p10 of cosine(source, decoded) is at leastdrift_floor_p10. - Utility, when
utility_floor_hit1is 0 or more: one query per sampled image (orutility_queriesof them), rendered fromutility_query_templatewith the row's label and embedded by the gate model's text tower. hit@1 counts queries whose own image ranks first among the sample's decoded frames. The rung passes when that hit@1 is at least the larger ofutility_floor_hit1and (hit@1 on the source images minusutility_tol). - A rung passes when every enabled check passes.
- Visual: every stratum's
- Choose the highest passing crf. When none passes, the gate uses the lowest crf on the ladder and prints
[forge] warning: no ladder crf met the floors (...); using smallest crf Non stdout. The gate never fails the build.
The manifest records the whole run under media.crf_auto: the buckets, the sampled indices, each rung's scores and pass flags, the chosen crf, the warning and the gate model's model_hash, plus the utility settings when that check ran. The chosen value is also media.crf. --resume reuses an earlier choice when the rows and [media] are unchanged.
retrieval-auto turns the visual checks off but still requires ssimulacra2 and still scores every sampled frame at every rung.
[media.cluster]
Read only when order is similarity or cluster.
| Key | Type | Default | Meaning |
|---|---|---|---|
space | string | "" | Model whose source-image embeddings drive the order. "" means the first image = "space" model. Must be a spec model with image = "space" |
threshold | float | 0.92 | cluster only: cosine to a cluster's running centroid needed to join it |
[media.jxl_transcode]
Read only by backend = "jxl-transcode".
| Key | Type | Default | Meaning |
|---|---|---|---|
on_unsupported_jpeg | string | "copy-source" | What to do with a source that cannot be transcoded reversibly (not a JPEG, refused by cjxl, or a failed round trip): error stops the build, copy-source stores the original bytes as NNNNNN.<ext>, lossless-jxl stores a lossless JPEG XL of the pixels (cjxl -d 0). Any other value fails validation |
verify_roundtrip | boolean | true | Rebuild each JPEG with djxl and require its SHA-256 to equal the source's. Needs djxl |
keep_metadata | boolean | false | Parsed and never used |
Each file's outcome is recorded in the manifest as media.decisions[], with file, action (transcode, copied or lossless) and verified.
keep_metadata has no effect
keep_metadata is accepted and changes the media state key, but no encoder reads it: cjxl runs without metadata options whatever its value. Do not rely on it to keep or drop EXIF, ICC or XMP data. See Known limits.
The next table is [build] and [output].
[[models]]
Reference for [[models]] in a urna build spec: model roles, dims, space_dtype, per-model knobs, and which vector spaces each model writes into the file.
[build] and [output]
Reference for [build] and [output] in a urna build spec: the engine preset, dtype, graph and mrl_dim of space 0, output modes, provenance and the cache.