Documentation
¶
Overview ¶
Package search is ContentKit's keyword retrieval over content_search_documents: exact names and aliases, native-script prefixes and bounded typos for every language, with the host's eligibility join applied inside every route.
Index ¶
- Constants
- func DeleteKeywordDocuments(ctx context.Context, db Executor, schema string, keys []DocumentKey) error
- func EligibilityJoin(e *Eligibility, args pgx.NamedArgs) (join, priority string, err error)
- func MarkDirty(ctx context.Context, db Executor, schema string, marks []DirtyMark) error
- func QuoteSchema(schema string) (string, error)
- func UpsertKeywordDocuments(ctx context.Context, db Executor, schema string, docs []KeywordDocument) error
- type Candidate
- type DirtyMark
- type DocumentKey
- type DocumentSink
- type Eligibility
- type Executor
- type Hit
- type KeywordDocument
- type Options
- type PublishedDocument
- type RRFContribution
- type RRFHit
- type RRFKey
- type RRFOptions
- type RRFTraceHit
- type RRFTraceKey
- type Result
Constants ¶
const MaxCandidateLimit = 10_000
MaxCandidateLimit bounds one route's document window.
Variables ¶
This section is empty.
Functions ¶
func DeleteKeywordDocuments ¶
func DeleteKeywordDocuments(ctx context.Context, db Executor, schema string, keys []DocumentKey) error
DeleteKeywordDocuments removes the documents; unknown keys are ignored.
func EligibilityJoin ¶
func EligibilityJoin(e *Eligibility, args pgx.NamedArgs) (join, priority string, err error)
EligibilityJoin renders the host eligibility query as the lateral join every route applies to a candidate row aliased sd, binding its named args. It returns the join clause (empty without eligibility) and the priority expression to select.
func MarkDirty ¶
MarkDirty queues documents for the worker. Hosts call it in the transaction that changes the content; the queue trigger assigns each row a new revision.
func QuoteSchema ¶
QuoteSchema validates and quotes a schema identifier for embedding in SQL.
func UpsertKeywordDocuments ¶
func UpsertKeywordDocuments(ctx context.Context, db Executor, schema string, docs []KeywordDocument) error
UpsertKeywordDocuments derives every lexical projection from structured input for a batch of documents of any tenant, kind or language.
Types ¶
type Candidate ¶
type Candidate struct {
contentref.ContentRef
Language string
Score float32
}
Candidate is one document proposed by another retrieval source.
type DirtyMark ¶
type DirtyMark struct {
DocumentKey
Deleted bool
Reason string
}
DirtyMark asks the worker to rebuild (or delete) one document.
type DocumentKey ¶
type DocumentKey struct {
contentref.ContentRef
Language string
}
DocumentKey identifies one keyword document: a content reference in one language. The language is metadata of the document, never of the content.
type DocumentSink ¶
type DocumentSink interface {
Upsert(ctx context.Context, doc PublishedDocument) error
Delete(ctx context.Context, key DocumentKey, version int64) error
}
DocumentSink receives every keyword document the worker publishes or deletes, after the document is written and before its dirty row is acknowledged. Delivery is at-least-once: a failing call keeps the row queued (with a new Version) while the keyword index still commits, so an unavailable sink never blocks keyword search. Implementations must be bounded and atomically apply only versions newer than the last version of each DocumentKey. Keep that version after deletion as a tombstone, so late upserts cannot resurrect a deleted document and late deletes cannot erase a newer upsert. A remote call may finish after returning an error or timeout; the worker's transaction lock cannot serialize those remote effects. These external writes do not share the Postgres transaction. Implementations may durably enqueue work; ContentKit ships no sink implementation.
type Eligibility ¶
Eligibility is trusted host SQL joined laterally to every candidate document (alias sd: tenant_id, content_kind, content_id, content_version_id, language). It returns no row when the document is not eligible for this request, or one row with the column priority (integer: preferred document among a content item's equal matches, lower first). Ownership, access, publication and every requested version trait must hold on that one row; sibling documents never make each other eligible.
type Hit ¶
type Hit struct {
contentref.ContentRef
Language string
// Priority orders same-item documents at equal match; lower is preferred.
Priority int32
Score float32
}
Hit is one matched document.
func Eligible ¶
func Eligible(ctx context.Context, pool *pgxpool.Pool, opts Options, candidates []Candidate) ([]Hit, error)
Eligible returns the candidates that exist as documents of the request tenant and language and pass the host filter and eligibility join, in their input order with Priority filled. Candidates another source proposes are never trusted with eligibility: this is the same join every keyword route runs.
type KeywordDocument ¶
type KeywordDocument struct {
DocumentKey
Title string
Aliases []string
Keywords []string
}
KeywordDocument is the host's canonical search input for one document. Derived search text belongs to ContentKit; callers never write it. An empty Title deletes the document. Aliases are alternate names; Keywords are contextual terms (for example trusted tags), not descriptions.
func (KeywordDocument) Validate ¶ added in v0.58.4
func (d KeywordDocument) Validate() error
Validate checks a document before it crosses into the keyword index. An empty title is a deletion, so only its key needs validation.
type Options ¶
type Options struct {
Schema string
Tenant string
Language string
// ContentKinds limits the kinds searched; empty means every kind.
ContentKinds []string
Limit int
// FilterSQL is trusted host SQL appended as `AND (<FilterSQL>)` on sd;
// FilterArgs binds its pgx '@name' placeholders. Reserved names: tenant,
// language, q, prefix, limit, kinds, candidates.
FilterSQL string
FilterArgs map[string]any
Eligibility *Eligibility
}
Options selects one tenant, language and window of a keyword request.
type PublishedDocument ¶
type PublishedDocument struct {
KeywordDocument
// Version orders deliveries of one DocumentKey: the dirty-queue revision
// that published the document. Retries may repeat it; sinks atomically
// ignore equal or lower versions across both upserts and deletes.
Version int64
}
PublishedDocument is one keyword document as delivered to a DocumentSink.
type RRFContribution ¶
type RRFHit ¶
func FuseRRF ¶
func FuseRRF(lists [][]RRFKey, opts RRFOptions) []RRFHit
FuseRRF fuses multiple ranked lists (best-first) into one ranked list.
type RRFKey ¶
type RRFKey struct {
contentref.ContentKey
Language string
}
RRFKey identifies one document across fused lists.
type RRFOptions ¶
type RRFOptions struct {
// K is the stabilizer constant; higher K flattens rank differences.
// Defaults to 60 when <= 0.
K int
// Weights applied to each list. Empty => all 1.0.
Weights []float32
}
RRF (Reciprocal Rank Fusion) combines ranked lists without relying on raw score calibration: score(doc) = Σ weight_i / (k + rank_i), rank 1-based.
type RRFTraceHit ¶
type RRFTraceHit struct {
Hit RRFHit `json:"-"`
Key RRFTraceKey `json:"key"`
Score float32 `json:"score"`
Contributions []RRFContribution `json:"contributions"`
}
func FuseRRFWithTrace ¶
func FuseRRFWithTrace(lists [][]RRFKey, opts RRFOptions) ([]RRFTraceHit, error)
FuseRRFWithTrace returns the same ordered hits as FuseRRF plus the exact source-list contributions used to compute each score.
type RRFTraceKey ¶
type Result ¶
type Result struct {
// Hits are documents ordered by score, content kind, content id and version.
Hits []Hit
// Truncated reports that a route filled its SQL window or that more scored
// documents existed than Limit: documents beyond the window were never ranked.
Truncated bool
}
Result is one language's scored document window.
func KeywordSearch ¶
func KeywordSearch(ctx context.Context, pool *pgxpool.Pool, query string, opts Options) (Result, error)
KeywordSearch preserves names and aliases, including native scripts. Exact names precede aliases, token/prefix matches, then conservative one-edit typos. PostgreSQL retrieves bounded candidates; Go never scans the document catalog. Host filters and the eligibility join run inside every route before its limit.