docsv0.5.1

[media]

Reference for [media] in a urna build spec: backends, every knob, profiles and their resolved values, the crf gate, ordering, and jxl transcoding.

[media] tells the build to encode the source images and store them with the corpus, so a hit can return the image it came from. Its presence turns the media stage on. This page lists every key of [media] and its three subtables, the profiles, and the crf gate behind crf = "auto". For choosing settings, see Tune media compression.

[media]
profile = "stills"     # a measured starting point; explicit keys below still win
dedup = true

[output]
embed_media = true     # store the encoded bytes inside the .urna

When the media stage runs

The stage runs when the spec has [media] and the rows carry images (kind = "image_dir", or a [source.image] path_template). With [media] and no images at all, the stage is skipped. When some rows have an image and others do not, the build stops with a traceback (exit 1) ending in:

ForgeError: media enabled but N rows have no image (first: <key>)

The stage encodes each unique image once (identical source files share a frame when dedup = true), writes the result to <name>.media/ beside the .urna, and records the backend, the encoder settings and the sizes in the manifest's media block. The written .urna references the media as blobs, with a span per chunk pointing at its frame or file. With [output] embed_media = true the bytes are also stored inside the .urna.

[media]

KeyTypeDefaultAllowedMeaning
profilestring"""" or a profile nameFills in knob values before the explicit keys apply
backendstring"av1"av1, avif, jxl, jxl-transcode, controlThe encoder. See backends
widthinteger1024positiveMaximum canvas width. See canvas
crfinteger or "auto"35any integer, or "auto"Quality for av1 (-crf) and avif (avifenc -q). "auto" runs the crf gate, av1 only
tunestring"still"still, defaultSVT-AV1 tune for av1
speedinteger8encoder rangeav1 preset (0 to 13) or avifenc --speed
fpsinteger1positiveFrame timestamps of the av1 stream. One frame per unique image
pix_fmtstring"yuv420p"yuv420p, yuv444pChroma format for av1 and avif
shard_sizeinteger20480 or moreav1 only: split the stream into one file per this many unique frames. 0 never splits
gopstring"auto"auto, intra, interav1 only: keyframe policy. See GOP
orderstring"none"none, similarity, clusterReorder frames so similar images sit together. Applied by av1 only. See ordering
dedupbooleantrueRows with byte-identical source images share one frame and one image vector
qualitytable[media.quality], read only with crf = "auto"
clustertable[media.cluster], read only when order is not none
jxl_transcodetable[media.jxl_transcode], read only by backend = "jxl-transcode"

Validation checks profile, backend, gop, order, tune, that crf is an integer or "auto", and the rules below. It does not check width, speed, fps, pix_fmt or shard_size: a wrong value is used as given and can fail during encoding with a Python traceback (exit 1).

Validation errors (exit 2):

  • media.profile: unknown '<p>' (valid: [...]), at parse time.
  • media.crf: "auto" gates the av1 ladder only (backend = "av1", or profile = "stills-av1"); the avif and jxl backends take a fixed crf.
  • media.quality.gate_model: crf=auto needs an image model, and media.quality.gate_model '<g>' must be a spec model with image=space.
  • media.cluster.space: order=<o> needs an image model, and media.cluster.space '<s>' must be a spec model with image=space.
  • output.embed_media requires a [media] section (there is nothing to inline).

Backends

BackendProducesFiles in <name>.media/One blob per
av1An AV1 video stream (SVT-AV1 in ffmpeg), one frame per unique image<name>-av1.mp4, or <name>-av1-000.mp4, <name>-av1-001.mp4 and so on when shardedStream file; the chunk's span is its frame index
avifOne AVIF per image (avifenc)<name>-avif/000000.avif and onFile
jxlOne lossless JPEG XL per source image (cjxl -d 0), no canvas<name>-jxl/000000.jxl and onFile
jxl-transcodeOne JPEG XL per source JPEG that rebuilds the original bytes exactly (cjxl --lossless_jpeg=1)<name>-jxl/000000.jxl, or 000000.<ext> for a copied sourceFile
controlOne lossless PNG per image at the canvas size<name>-png/000000.png and onFile

Required tools: ffmpeg and ffprobe with libsvtav1 for av1 (and for the crf gate and the GOP probe); avifenc and avifdec for avif; cjxl for the jxl backends, plus djxl to verify transcodes and to embed decoded JPEG XL frames; ssimulacra2 for crf = "auto". A missing tool fails when the stage runs, for example jxl backend needs 'cjxl' on PATH: brew install jpeg-xl. The manifest's source_bytes is always the size of the original files, so compression_ratio compares across backends.

Which knobs apply

Knobav1avifjxljxl-transcodecontrol
width (canvas)yesyesnonoyes
crf-crf-qnonono
crf = "auto"yesspec errorspec errorspec errorspec error
speed-preset--speednonono
tuneyesnononono
fps, gop, shard_sizeyesnononono
orderappliedcomputed, then droppedcomputed, then droppedcomputed, then droppedcomputed, then dropped
pix_fmtyes--yuv 420 or 444nonono
[media.jxl_transcode]nononoyesno

Canvas

av1, avif and control fit every image onto one canvas, letterboxed. The build opens up to 128 evenly spaced images and takes the median width and the median aspect ratio. The canvas width is the smaller of width and that median width, the height follows from the median aspect, and both are rounded down to even numbers. width is a ceiling: images are never upscaled past the median source width. The JPEG XL backends keep each source image as it is.

Encoder details

  • tune = "still" applies only to all-intra av1 encodes. The tune number is probed on the local SVT-AV1 build; when it has none, the build prints [forge] warning: local SVT-AV1 has no Still Picture tune; using default tune on stderr and records tune_resolved: null.
  • pix_fmt = "yuv444p" on av1: if the encoder writes 4:2:0 anyway, the build stops (encoder wrote yuv420p for a requested yuv444p ...). On avif, a value other than the two allowed ones fails with a Python KeyError.
  • SVT-AV1 runs with lp=2 and avifenc with 8 jobs, so the output bytes do not depend on the machine's core count.

GOP

gop sets the keyframe interval of the av1 stream:

  • intra: every frame is a keyframe (keyint 1), so any frame decodes on its own. The manifest records gop.policy = "flag".
  • inter: a keyframe every 16 frames with scene-change detection off. Smaller when neighbouring images are similar; decoding a frame means decoding from its keyframe.
  • auto: a probe encodes up to 32 sampled frames both ways at the target crf and keeps the smaller result, ties going to intra. An inter win is overruled when its mean ssimulacra2 score is more than 2.0 below the intra arm. Without ssimulacra2 the probe compares bytes only and warns. The probe runs once per stream, or once per shard when the stream is sharded.

Ordering

order puts similar images next to each other before encoding, which lets inter frames pay off:

  • similarity: a greedy nearest-neighbour walk starting from the first image.
  • cluster: greedy clustering in row order; an image joins a cluster when its cosine to the cluster's running centroid reaches [media.cluster] threshold. Clusters are written largest first.

Both use the source-image embeddings of the [media.cluster] space model, from an extra embedding pass that is not cached. Only av1 applies the order. The other backends compute it, paying for that pass, and then drop it. The manifest records the order and the permutation.

Profiles

A profile fills in knob values before the explicit keys apply, so an explicit key always wins. The quality table merges key by key: an explicit [media.quality] key overrides the profile's value for that key only. Resolved values, with the profile's own settings marked by an asterisk:

Knobno profilenear-dupstillsstills-av1archiveretrievalretrieval-auto
backendav1av1avif*av1jxl-transcode*av1av1
crf353548*3535 (unused)50*"auto"*
tunestillstill*still (unused)still*still (unused)still*still*
speed888*88 (unused)6*6*
gopautoauto*auto (unused)intra*auto (unused)intra*intra*
ordernonecluster*nonenonenonenonenone
quality.visual_floor_p1060.060.060.060.060.060.0-1e9*
quality.visual_floor_min45.045.045.045.045.045.0-1e9*
quality.drift_floor_p100.950.950.950.950.950.95-1.0*
quality.crf_ladder25 to 50, step 5samesamesamesamesame40, 45, 50, 55, 60*
quality.utility_floor_hit1-1.0-1.0-1.0-1.0-1.0-1.00.0*
quality.utility_tol0.00.00.00.00.00.00.02*

Every other knob keeps its default under every profile.

  • near-dup: for near-duplicate images (reprints, video frames, scans). Cluster ordering plus per-segment GOP lets inter coding pay where images repeat. Per-segment probing needs the stream to be sharded (more unique frames than shard_size); otherwise there is one probe for the whole stream.
  • stills: unique images, one AVIF per image at quality 48, speed 8. Each image decodes on its own with no video decode. Image-space models embed decoded frames one AVIF at a time, which makes that embedding step slower than on the av1 stream.
  • stills-av1: the all-intra av1 stream with the still tune. Use it when you want crf = "auto".
  • archive: byte-reversible JPEG repacking with a verified round trip, for corpora where no loss is acceptable.
  • retrieval: for corpora that serve search and whose images nobody looks at: all-intra, still tune, fixed crf = 50, speed 6.
  • retrieval-auto: retrieval with crf = "auto" gated on search quality alone. The visual and drift checks are off; every rung must keep text-to-image hit@1 within 0.02 of the source images. It needs a gate model with a text tower and a label on every row.

The profile name is recorded in the build lock (resolved_spec.media.profile), not in the manifest. The manifest records the resolved backend settings.

[media.quality]

Read only when crf = "auto".

KeyTypeDefaultMeaning
strategystring"stratified"Not read
bucketslist of strings["resolution", "entropy", "has_text", "alpha", "source_format"]Image traits that form the sampling strata. An unknown name fails when the gate runs: ValueError: media.quality.buckets: unknown bucket '<n>' (exit 1)
sample_per_bucketinteger12Images sampled per stratum, evenly spaced
visual_floor_p10float60.0Minimum 10th percentile of ssimulacra2 in every stratum
visual_floor_minfloat45.0Minimum ssimulacra2 over the whole sample
drift_floor_p10float0.95Minimum 10th percentile of cosine(source embedding, decoded embedding). Negative turns the drift check off
gate_modelstring""Model that embeds the sample. "" means the first image = "space" model. Must be a spec model with image = "space"
crf_ladderlist of integers[25, 30, 35, 40, 45, 50]The crf values tried, in ascending order
utility_floor_hit1float-1.0Minimum text-to-image hit@1, from 0 to 1. Negative turns the utility check off
utility_queriesinteger0Number of queries, evenly spaced over the sample. 0 means one per sampled image
utility_query_templatestring"{label}"Query text per image. Must contain {label}, and {label} is the only placeholder allowed
utility_tolfloat0.0Allowed hit@1 loss against the source images, from 0 to 1

The utility_* keys are validated only when utility_floor_hit1 is 0 or more: utility_floor_hit1 and utility_tol must lie within 0 to 1, utility_queries must be 0 or more, the template must contain {label}, and the gate model must have a text tower (media.quality.utility_floor_hit1: gate model '<g>' has no text tower).

The default floors were set on trading-card images at 488 by 680 pixels (4:2:0, 2048-image sample, av1 still tune, speed 6). There, crf 30 passed (ssimulacra2 p10 65.3, minimum 58.6, drift p10 0.967) and crf 35 failed on p10 (55.7), so the defaults choose crf 30 for that corpus.

The crf gate

With crf = "auto" the build picks the crf itself, per corpus, before encoding. It needs backend = "av1", an image = "space" model, ffmpeg, ssimulacra2 (media.crf="auto" needs ssimulacra2 on PATH: brew install jpeg-xl otherwise) and pillow, plus a label on every row.

crf = "auto" fails on rows without a label

The gate reads one label per unique image, and a row whose label is empty or missing makes the build crash with AttributeError: 'Row' object has no attribute 'text' (exit 1). This happens even with the utility check off. An image_dir source without a labels file, a source without [source.image] label_template, or a template that renders empty for any row all hit it. Give every row a label, or set a fixed crf. See Known limits.

How it runs:

  1. Sample. Every unique image gets a stratum from buckets: resolution (longest side under 512 is small, under 1024 medium, else large), entropy (variance of the gradient magnitude on a 128-pixel grayscale thumbnail: under 100 low, under 1000 mid, else high), has_text (more than 10% of gradient magnitudes above 40), alpha (the image has transparency) and source_format (the file extension). Up to sample_per_bucket evenly spaced images per stratum form the sample.
  2. For each crf in crf_ladder, from low to high: encode the sample as an all-intra av1 stream with the spec's speed, pix_fmt and tune, decode it, score each frame with ssimulacra2 against the letterboxed source, and embed the decoded frames with the gate model.
  3. Check each rung:
    • Visual: every stratum's ssimulacra2 p10 is at least visual_floor_p10, and the sample's minimum is at least visual_floor_min.
    • Drift: off when drift_floor_p10 is negative; else the p10 of cosine(source, decoded) is at least drift_floor_p10.
    • Utility, when utility_floor_hit1 is 0 or more: one query per sampled image (or utility_queries of them), rendered from utility_query_template with the row's label and embedded by the gate model's text tower. hit@1 counts queries whose own image ranks first among the sample's decoded frames. The rung passes when that hit@1 is at least the larger of utility_floor_hit1 and (hit@1 on the source images minus utility_tol).
    • A rung passes when every enabled check passes.
  4. Choose the highest passing crf. When none passes, the gate uses the lowest crf on the ladder and prints [forge] warning: no ladder crf met the floors (...); using smallest crf N on stdout. The gate never fails the build.

The manifest records the whole run under media.crf_auto: the buckets, the sampled indices, each rung's scores and pass flags, the chosen crf, the warning and the gate model's model_hash, plus the utility settings when that check ran. The chosen value is also media.crf. --resume reuses an earlier choice when the rows and [media] are unchanged.

retrieval-auto turns the visual checks off but still requires ssimulacra2 and still scores every sampled frame at every rung.

[media.cluster]

Read only when order is similarity or cluster.

KeyTypeDefaultMeaning
spacestring""Model whose source-image embeddings drive the order. "" means the first image = "space" model. Must be a spec model with image = "space"
thresholdfloat0.92cluster only: cosine to a cluster's running centroid needed to join it

[media.jxl_transcode]

Read only by backend = "jxl-transcode".

KeyTypeDefaultMeaning
on_unsupported_jpegstring"copy-source"What to do with a source that cannot be transcoded reversibly (not a JPEG, refused by cjxl, or a failed round trip): error stops the build, copy-source stores the original bytes as NNNNNN.<ext>, lossless-jxl stores a lossless JPEG XL of the pixels (cjxl -d 0). Any other value fails validation
verify_roundtripbooleantrueRebuild each JPEG with djxl and require its SHA-256 to equal the source's. Needs djxl
keep_metadatabooleanfalseParsed and never used

Each file's outcome is recorded in the manifest as media.decisions[], with file, action (transcode, copied or lossless) and verified.

keep_metadata has no effect

keep_metadata is accepted and changes the media state key, but no encoder reads it: cjxl runs without metadata options whatever its value. Do not rely on it to keep or drop EXIF, ICC or XMP data. See Known limits.

The next table is [build] and [output].

On this page