docsv0.5.1

SearchHit and RetrieveHit

Fields of urna.SearchHit and urna.RetrieveHit, the hit objects urna returns from Python, what each field means, and how to turn a hit into a dict.

The search methods of UrnaFile return SearchHit objects, and UrnaFile.retrieve returns RetrieveHit objects. Both are read-only records with one attribute per field.

SearchHit

Returned by search, search_ann, search_hybrid, search_graph and search_space. urna.SearchHit is the class (it prints as builtins.SearchHitPy), so isinstance(hit, urna.SearchHit) works. 12 fields:

FieldTypeMeaning
chunk_idstrsha256:<64 hex> identity of the chunk
scorefloatexact cosine between the query and the stored vector. It reads the full-precision vectors when the file has them, else the stored dtype
score_typestralways "cosine"
source_uristrthe source_uri given at build time
offset_startintstart of the stored span, see Offsets
offset_endintend of the stored span
embedding_modelstrthe manifest model of the corpus. On search_space hits this is still the text model, not the space's model
index_typestrthe path that produced the hit: "exact", "hnsw", "hybrid", "graph" or "space". A method that falls back to exact search reports "exact"
rerankedboolFalse for exact and space hits, True for hnsw, hybrid and graph hits
file_hashstrthe corpus file_hash
content_hashstrthe corpus content_hash
citation_idstrurna://<content_hash>/<chunk_id>, which urna cite resolves

A SearchHit carries no text. To get the stored text with the hit, call retrieve, or resolve the citation_id with urna cite.

RetrieveHit

Returned by UrnaFile.retrieve. urna.RetrieveHit is the class (it prints as builtins.RetrieveHitPy). 11 fields, the same set the urna retrieve verb emits as JSON:

FieldTypeMeaning
chunk_idstras in SearchHit
scorefloatas in SearchHit: the exact cosine rerank value
score_typestralways "cosine"
source_uristras in SearchHit
offset_startintas in SearchHit
offset_endintas in SearchHit
citation_idstras in SearchHit
textstrthe stored canonical text of the chunk, the same bytes urna cite returns. It is never a reread of the original source
file_hashstras in SearchHit
content_hashstras in SearchHit
rerank_sourcestr"full_precision" when the rerank read float32 vectors, "stored_precision" when it read float16, int8 or int4 vectors with no full-precision copy. The same on every hit of one call

A RetrieveHit has no embedding_model, index_type or reranked. rerank_source tells you whether a score on a quantized file is the cosine of the original vectors or of their stored approximation. Presets and stored precision explains the difference.

Offsets

offset_start and offset_end hold whatever the builder wrote into byte_start and byte_end. They are byte offsets into the source only when the builder made them so:

Built byWhat the offsets mean
builder.chunk_text (checkout)UTF-8 byte offsets of the chunk inside the text passed to it
urna.build with your own chunk dictswhatever you passed. The examples in the repo pass 0 and the byte length of the chunk text
urna build --spec (the forge)the row ordinal: ordinal to ordinal + 1
a file with a blob span overlay (0x16)byte ranges inside the referenced blob, rewritten by the overlay

Do not slice the original document with these offsets unless you know which builder made the file. Citations and hashes covers the same point for citations.

Object behavior

Hits are plain read-only records:

  • repr() is the default <builtins.SearchHitPy object at 0x...>, with no field values.
  • == compares identity, so two identical searches return hits that are not equal. Compare chunk_id or citation_id instead.
  • They cannot be pickled (TypeError: cannot pickle 'builtins.SearchHitPy' object), have no __dict__ and no to_dict().
  • They are not JSON serializable. Web frameworks cannot return them directly.

Copy the fields you need into a dict:

SEARCH_FIELDS = (
    "chunk_id", "score", "score_type", "source_uri", "offset_start", "offset_end",
    "embedding_model", "index_type", "reranked", "file_hash", "content_hash", "citation_id",
)
RETRIEVE_FIELDS = (
    "chunk_id", "score", "score_type", "source_uri", "offset_start", "offset_end",
    "citation_id", "text", "file_hash", "content_hash", "rerank_source",
)


def hit_to_dict(hit, fields):
    return {name: getattr(hit, name) for name in fields}


rows = [hit_to_dict(h, RETRIEVE_FIELDS) for h in db.retrieve(qvec, 5)]

Serve a corpus over HTTP uses the same pattern.

On this page