SearchHit and RetrieveHit
Fields of urna.SearchHit and urna.RetrieveHit, the hit objects urna returns from Python, what each field means, and how to turn a hit into a dict.
The search methods of UrnaFile return SearchHit objects, and UrnaFile.retrieve returns RetrieveHit objects. Both are read-only records with one attribute per field.
SearchHit
Returned by search, search_ann, search_hybrid, search_graph and search_space. urna.SearchHit is the class (it prints as builtins.SearchHitPy), so isinstance(hit, urna.SearchHit) works. 12 fields:
| Field | Type | Meaning |
|---|---|---|
chunk_id | str | sha256:<64 hex> identity of the chunk |
score | float | exact cosine between the query and the stored vector. It reads the full-precision vectors when the file has them, else the stored dtype |
score_type | str | always "cosine" |
source_uri | str | the source_uri given at build time |
offset_start | int | start of the stored span, see Offsets |
offset_end | int | end of the stored span |
embedding_model | str | the manifest model of the corpus. On search_space hits this is still the text model, not the space's model |
index_type | str | the path that produced the hit: "exact", "hnsw", "hybrid", "graph" or "space". A method that falls back to exact search reports "exact" |
reranked | bool | False for exact and space hits, True for hnsw, hybrid and graph hits |
file_hash | str | the corpus file_hash |
content_hash | str | the corpus content_hash |
citation_id | str | urna://<content_hash>/<chunk_id>, which urna cite resolves |
A SearchHit carries no text. To get the stored text with the hit, call retrieve, or resolve the citation_id with urna cite.
RetrieveHit
Returned by UrnaFile.retrieve. urna.RetrieveHit is the class (it prints as builtins.RetrieveHitPy). 11 fields, the same set the urna retrieve verb emits as JSON:
| Field | Type | Meaning |
|---|---|---|
chunk_id | str | as in SearchHit |
score | float | as in SearchHit: the exact cosine rerank value |
score_type | str | always "cosine" |
source_uri | str | as in SearchHit |
offset_start | int | as in SearchHit |
offset_end | int | as in SearchHit |
citation_id | str | as in SearchHit |
text | str | the stored canonical text of the chunk, the same bytes urna cite returns. It is never a reread of the original source |
file_hash | str | as in SearchHit |
content_hash | str | as in SearchHit |
rerank_source | str | "full_precision" when the rerank read float32 vectors, "stored_precision" when it read float16, int8 or int4 vectors with no full-precision copy. The same on every hit of one call |
A RetrieveHit has no embedding_model, index_type or reranked. rerank_source tells you whether a score on a quantized file is the cosine of the original vectors or of their stored approximation. Presets and stored precision explains the difference.
Offsets
offset_start and offset_end hold whatever the builder wrote into byte_start and byte_end. They are byte offsets into the source only when the builder made them so:
| Built by | What the offsets mean |
|---|---|
builder.chunk_text (checkout) | UTF-8 byte offsets of the chunk inside the text passed to it |
urna.build with your own chunk dicts | whatever you passed. The examples in the repo pass 0 and the byte length of the chunk text |
urna build --spec (the forge) | the row ordinal: ordinal to ordinal + 1 |
| a file with a blob span overlay (0x16) | byte ranges inside the referenced blob, rewritten by the overlay |
Do not slice the original document with these offsets unless you know which builder made the file. Citations and hashes covers the same point for citations.
Object behavior
Hits are plain read-only records:
repr()is the default<builtins.SearchHitPy object at 0x...>, with no field values.==compares identity, so two identical searches return hits that are not equal. Comparechunk_idorcitation_idinstead.- They cannot be pickled (
TypeError: cannot pickle 'builtins.SearchHitPy' object), have no__dict__and noto_dict(). - They are not JSON serializable. Web frameworks cannot return them directly.
Copy the fields you need into a dict:
SEARCH_FIELDS = (
"chunk_id", "score", "score_type", "source_uri", "offset_start", "offset_end",
"embedding_model", "index_type", "reranked", "file_hash", "content_hash", "citation_id",
)
RETRIEVE_FIELDS = (
"chunk_id", "score", "score_type", "source_uri", "offset_start", "offset_end",
"citation_id", "text", "file_hash", "content_hash", "rerank_source",
)
def hit_to_dict(hit, fields):
return {name: getattr(hit, name) for name in fields}
rows = [hit_to_dict(h, RETRIEVE_FIELDS) for h in db.retrieve(qvec, 5)]Serve a corpus over HTTP uses the same pattern.
UrnaFile
Reference for urna.UrnaFile, the read-only handle on a .urna file, with its 13 properties and 11 methods for search, retrieve, blobs and validation.
urna.build
Reference for urna.build, which writes a .urna file from chunks you already embedded, with all 29 parameters, the presets, the chunk schema and its checks.