docsv0.5.1

urna cite

Reference for urna cite, which resolves a urna://content_hash/chunk_id citation into the chunk's stored text and source span after checking the content_hash.

urna cite resolves a urna://content_hash/chunk_id citation against a .urna file. It checks that the citation's content_hash matches the file, finds the chunk, and prints its stored canonical text, its source span and the file's hashes. It needs no Python and no embedder.

Usage

urna cite <FILE> <CITATION>

Arguments

ArgumentDescription
<FILE>Path to the .urna file
<CITATION>urna://<content_hash>/<chunk_id> URI

Quote the citation in the shell. The citation_id from any hit, from ask and from retrieve JSON is accepted as is.

Options

OptionDescription
-h, --helpPrint help

Behavior

  1. Parses the citation. It must start with urna:// and contain a / after the content hash. The sha256: prefix on the content hash part is optional; the chunk id must be written in full, sha256: included. Anything after a further / is ignored.
  2. Reads the file and runs the reader's integrity check (header, section checksums, manifest, footer hash, contract).
  3. Recomputes the file's content_hash and stops if it differs from the citation's. A citation never resolves against a different corpus.
  4. Looks up the chunk_id in chunk_ids and stops if it is not there.
  5. Prints the chunk's canonical text and its span from chunks_original_spans.

The printed text is the stored canonical text, the same bytes the search verbs, ask and retrieve return. cite never reopens the original source file. What byte_start and byte_end mean depends on the builder: see what the offsets mean.

cite reads the stored span directly. On a media corpus with a blob span overlay (0x16), search hits and retrieve report the blob's uri and a byte range inside the blob, while cite prints the stored source_uri and span, which for a forge build is the row ordinal.

The whole file is read into memory for this check; it is not memory-mapped.

Output

To stdout, one field per line, then the text:

FieldMeaning
citation_idThe citation as you passed it
fileThe path you passed
file_hashSHA-256 of the whole file
content_hashThe file's content_hash, which matched the citation
chunk_idThe resolved chunk
source_uriWhere the chunk came from
byte_start, byte_endThe stored span
text:Followed by the canonical text on the next lines

Exit codes

CodeMeaning
0The citation resolved
1The citation does not start with urna:// or has no chunk id; content_hash mismatch; chunk id not found; file missing or unreadable; a failed integrity check. Printed to stderr as Error: <message>
2Usage error: a missing argument

Examples

Resolve a citation from the quickstart corpus:

urna cite examples/quickstart/out/quickstart.urna 'urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be'
citation_id:  urna://sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
file:         examples/quickstart/out/quickstart.urna
file_hash:    sha256:e4d5f8907faad38c192dc6929e36dbf16db558410b2dae5abb4d65f10f508832
content_hash: sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df
chunk_id:     sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be
source_uri:   demo/03-citations.md
byte_start:   7
byte_end:     8
text:
because the citation points at content, two people who build the same logical corpus on two machines get the same citation, and a stored corpus and a compressed one cite identically. resolving a citation returns the exact canonical text and the original byte span it came from, which is what lets an agent quote a source it can prove.

A citation issued for other content fails before any lookup:

urna cite examples/quickstart/out/quickstart.urna 'urna://sha256:0000000000000000000000000000000000000000000000000000000000000000/sha256:b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be'
Error: content_hash mismatch: citation says sha256:0000000000000000000000000000000000000000000000000000000000000000 but file is sha256:1147b2560863331b21bd9d60fe6bdd99507dc34e17108444dc38194f8e6f09df

A chunk id without its sha256: prefix is not found:

Error: chunk_id b5dfeb09a643f6f0afde3f361316dde3ea503ff86c525d297e6ba79b5781a4be not found in file

What goes into each hash is in Citations and hashes.

On this page