docsv0.5.1

[source]

Reference for [source] in a urna build spec: sqlite, csv, jsonl and image_dir sources, text and image templates, joins, derive, and row identity.

[source] says where the rows come from and how each row becomes one chunk: its stored text, its optional image, its key and its position. This page also covers [embedding.image_input], which chooses what image-space models embed.

[source]

KeyTypeDefaultKindsRequiredMeaning
kindstring""allyessqlite, csv, jsonl or image_dir (pdf_dir is accepted but cannot build, see below)
dbpath""sqliteyesSQLite database, opened read-only
querystring""sqliteyesSQL that returns the rows. No ORDER BY needed
pathpath""csv, jsonlyesThe file to read
order_bystring or list of strings[]sqlite, csv, jsonlyesColumns that order the rows and form the row key. A single string is read as a one-item list. The combined key must be unique
input_dirpath""image_diryesFolder searched recursively for images
labelspath""image_dirnoLabel file for the images
texttable{ template = "" }sqlite, csv, jsonlno[source.text]
imagetable{}sqlite, csv, jsonlno[source.image]
joinsarray of tables[]sqliteno[[source.joins]]
derivetable{}sqlite, csv, jsonlno[source.derive]

Paths resolve against the build's current directory and expand ${VAR} and a leading ~/ (see the spec file). A key that does not apply to the declared kind is ignored without a message.

Validation errors:

MessageWhen
source.kind: must be one of ['csv', 'image_dir', 'jsonl', 'pdf_dir', 'sqlite']kind missing or unknown
source.db/source.query: requiredsqlite without db or query
source.path: requiredcsv or jsonl without path
source.order_by: required for total ordering (RFC-0 N1)sqlite, csv or jsonl without order_by
source.input_dir: requiredimage_dir without input_dir

Kinds

sqlite

Runs query on the database opened as file:<db>?mode=ro. Each result row becomes a map of its columns. Joins run next, then derive, then sorting and templates.

csv

Reads path with a header row. Every value is a string.

jsonl

Reads path as one JSON object per non-blank line.

image_dir

Finds every .jpg, .jpeg, .png, .bmp, .webp, .tiff and .tif file under input_dir (extensions match in any case), sorted by path. Each image is one row: the key is its path relative to input_dir, the stored text is that path, or <path> [<label>] when it has a label, and the source_uri is item://<name>/<path>. order_by, text, image, joins and derive are ignored.

pdf_dir

Accepted by validation. Loading the rows always fails.

pdf_dir always fails

A spec with kind = "pdf_dir" passes --dry-run, then every build stops when rows load, with exit 2:

spec error: source.kind=pdf_dir: build via forge_pipeline (pages are temporary)

The message points back at the build itself; there is no other forge entry point to use. PDF corpora are built today with the separate image tool in the checkout, python/tools/urna_build_image_corpus.py --pdf. See Images and PDFs and Known limits.

Labels for image_dir

labels is either a .csv file or a JSON object:

  • A .csv file: the first row is a header and is skipped; each following row maps its first column (image id) to its second (label).
  • Any other extension is read as a JSON object mapping image id to label.

An image takes the label keyed by its relative path, else the one keyed by its file name without extension. A label file that does not exist fails with a Python FileNotFoundError (exit 1).

[source.text]

KeyTypeDefaultRequiredMeaning
templatestring""noPython format string over the row's columns. The rendered text is the chunk's stored, citable text and the input to text embedding

Rendering rules:

  • A column the row does not have renders as an empty string.
  • The template is trimmed and split into lines. A line whose placeholders all render empty is dropped, so optional fields leave no blank lines.
  • On each kept line: empty () and empty [ ] or [,] are removed, a trailing , ] becomes ], an em dash (U+2014) at the start or end of the line is removed, and runs of whitespace collapse to one space.
  • Lines are joined with newlines.

An empty template, or a template that renders to nothing for a row, falls back to identity-only text: the row's label when [source.image] label_template gives one, else its key. The manifest records text_quality = "template" when a template is set and "identity-only" when it is not.

[source.text]
template = """
{name} ({category})
{description}
"""

[source.image]

Attaches one image file to each row of a sqlite, csv or jsonl source.

KeyTypeDefaultRequiredMeaning
path_templatestring""noFormat string over the row that renders the image path. Missing columns render empty
label_templatestring""noFormat string over the row that renders the image label, trimmed

When path_template is set, every row's rendered path must be an existing file, or loading stops with spec error: source.image: file missing for key (...): <path> (exit 2). The file's SHA-256 is taken at load and used for deduplication.

The label enters the row's input hash, the manifest items, the identity-only text fallback, and the utility queries of the crf gate.

Setting path_template is what allows a sqlite, csv or jsonl source to use image = "space" models. Without it, validation fails with models.<preset>.image=space: source declares no images.

[[source.joins]]

sqlite only. Each entry adds columns from a second query to the main rows.

KeyTypeDefaultRequiredMeaning
querystringnoneyesSQL run on the same read-only connection
onstringnoneyesColumn that both the main query and this query return

Both keys have no default: an entry without one fails with a Python TypeError (exit 1), not a spec error.

[[source.joins]]
query = "SELECT sku, category, brand FROM product_meta"
on = "sku"

How a join applies:

  1. The join query's rows are indexed by their on value. When two join rows share a value, the later one wins.
  2. Each main row whose on value is in the index gains the join row's columns. A column the main row already has keeps the main row's value.
  3. Joins run in order, before [source.derive].

Join errors (exit 2):

  • spec error: source.joins: join query has no column '<on>'
  • spec error: source.joins on '<on>': 0 of N rows matched - check the key name and its sqlite type on both sides, when the join returned rows and none matched. A TEXT key never matches an INTEGER key; cast one side.

A partial match is not an error. It prints [forge] join on '<on>': H/N rows matched on stdout.

[source.derive]

sqlite, csv and jsonl. Each entry adds a computed column:

[source.derive]
image_stem = "basename_stem(image_url)"
name_lower = "lower(name)"
HelperResult
basename_stem(column)Drops a ?query suffix, keeps the file name, drops its last extension. https://cdn.example.com/a/b/card-12.jpg?v=3 becomes card-12
lower(column)The value in lower case

Derived columns are computed per row after joins and before sorting, so they work in order_by, templates and path_template. The expression must be exactly helper(column). Errors fire when rows load (exit 2): spec error: source.derive: unsupported expression '<expr>' (use helper(column)) and spec error: source.derive: unknown helper '<h>' (valid: basename_stem, lower).

Row identity

These rules fix which chunk each row becomes and what its citation depends on.

  • Key: the order_by values as strings, joined with a vertical bar. For image_dir, the relative path.
  • Order: rows sort by the tuple of their order_by values as strings. Numeric ids sort as text, so 10 comes before 2. Zero-pad them to sort numerically.
  • Uniqueness: two rows with the same key stop the load with spec error: source.order_by [...] is not a total order (duplicate key (...)); append a unique column (RFC-0 N1).
  • Ordinal: the row's position after sorting (and after --sample, which renumbers). The chunk's span is [ordinal, ordinal + 1).
  • source_uri: a non-empty column named source_uri in the row, else item://<corpus.name>/<key>.
  • Item input hash: SHA-256 over the stored text's digest, the image's SHA-256 (or none), the label and chunker_version.
  • corpus_input_hash: SHA-256 over the item input hashes in row order.

The span enters the chunk_id, so inserting or deleting a row changes the chunk_id of every row sorted after it. A --sample build renumbers ordinals and never shares citations with the full build. See citations and hashes.

Deduplication: rows whose source images are byte-identical share one frame and one image vector. This happens when [media] dedup = true (the default) or when the spec has no [media]. Text vectors are always computed per row.

[embedding.image_input]

Chooses what image = "space" models embed.

KeyTypeDefaultRequiredMeaning
modestring""nosource, decoded_media, or "". "" means decoded_media when the spec has [media], else source
  • source embeds the original image files.
  • decoded_media embeds the frames decoded from the encoded media, so the vectors match what the corpus stores. It adds the decoder settings (backend, canvas, crf, pixel format) to those models' recipes.

Errors: embedding.image_input.mode: source | decoded_media for any other value, and embedding.image_input.mode=decoded_media requires [media].

The resolved mode enters every model's embedding recipe, text-only models included. Changing it, or adding or removing [media] while mode is "", recomputes every model's vectors once. Other keys under [embedding] are ignored without a message.

On this page