Documentation
¶
Overview ¶
Package scriptindex is the managed-script consumer of the shared indexjobs framework (#1370). It registers a Source/Sink pair under source_kind = "scripts" so scripts are embedded off the request path, and a person asking for what they want done reaches the automation that does it rather than needing the words its author used.
A script is embedded as a set of chunks (#2027): its description card, then its source in pieces each within the provider's input, so the reasoning in its comments and the tables its SQL names are found as well as its description, and no part of it is trimmed. The vectors live in script_embedding_chunks, one row per chunk, the way knowledge pages keep theirs (000097); search scores each script by its best chunk. SourceID is the script id, and an item id is "<script id>:<chunk>".
Whether a script owes an embedding is read from two hashes on its row: the hash of what it is indexed on now (index_text_hash, written by every save) and the hash of what its chunks were built from (index_embedded_hash, written here), with the model that built them (index_model).
Every enabled script is indexed regardless of lifecycle status: the store's search applies the discoverable-status filter at query time, so the index covers what any caller can rank. A disabled script is never embedded and never counted as missing coverage.
Index ¶
- Constants
- func Chunks(s *script.Script, maxBytes int) []string
- func Corpus(s *script.Script) string
- func RegisterConsumer(reg interface{ ... }, db *sql.DB, currentModel string, maxInputBytes int) error
- type Sink
- func (s *Sink) Coverage(ctx context.Context) (indexjobs.Coverage, error)
- func (s *Sink) FindGaps(ctx context.Context) ([]string, error)
- func (*Sink) Kind() string
- func (s *Sink) ListExisting(ctx context.Context, key indexjobs.Key) (map[string]indexjobs.Vector, error)
- func (s *Sink) StampExpected(ctx context.Context, key indexjobs.Key, _ int) error
- func (s *Sink) Upsert(ctx context.Context, key indexjobs.Key, rows []indexjobs.Vector) error
- func (s *Sink) UpsertBatch(ctx context.Context, key indexjobs.Key, rows []indexjobs.Vector) error
- type Source
- type Store
- func (s *Store) Coverage(ctx context.Context, currentModel string) (indexed, expected int, err error)
- func (s *Store) FindGaps(ctx context.Context, currentModel string) ([]string, error)
- func (s *Store) GetIndexed(ctx context.Context, id string) (*script.Script, error)
- func (s *Store) ListVectors(ctx context.Context, scriptID string) (map[string]indexjobs.Vector, error)
- func (s *Store) ReplaceVectors(ctx context.Context, scriptID string, rows []indexjobs.Vector) error
- func (s *Store) Stamp(ctx context.Context, scriptID, model string) error
- func (s *Store) UpsertVectors(ctx context.Context, scriptID string, rows []indexjobs.Vector) error
Constants ¶
const SourceKind = "scripts"
SourceKind is the indexjobs source_kind this package serves.
Variables ¶
This section is empty.
Functions ¶
func Chunks ¶ added in v1.140.0
Chunks splits a script into the units its vectors are built from, each within maxBytes so no part of it is trimmed before it is embedded (#2027): the card first, then the source, each source chunk headed by the script's title so it is placed in the same space a query about the script is. The source is cut where a definition, a load or a comment block starts at the left margin, so a comment stays with the code it explains. A maxBytes below textchunk.MinViableChunkBytes disables splitting.
Types ¶
type Sink ¶
type Sink struct {
// contains filtered or unexported fields
}
Sink implements indexjobs.Sink for the scripts kind over the script's chunk table. currentModel is the provider model the gap query diffs stored chunk sets against, so a model swap re-embeds scripts built by the previous one.
func NewSink ¶
NewSink returns a Sink backed by the given store. currentModel is the embedding provider's model identifier (embedding.ModelName); "" on a provider that does not name its model.
func (*Sink) Coverage ¶
Coverage reports enabled scripts whose chunk set is current against all enabled scripts. ExpectedKnown is true: every enabled script converges.
func (*Sink) FindGaps ¶
FindGaps returns enabled script ids whose chunk set is missing, was built from text the script no longer has, or was built by another model.
func (*Sink) ListExisting ¶
func (s *Sink) ListExisting(ctx context.Context, key indexjobs.Key) (map[string]indexjobs.Vector, error)
ListExisting returns the script's chunk vectors keyed by item id for the worker's dedup pass, so an edit re-embeds only the chunks whose text moved.
func (*Sink) StampExpected ¶
StampExpected records that the script's chunk set was built by the current model from the text LoadItems read. The worker calls it only after a successful pass; a failure here is non-fatal, since the next sweep finds the script again and the dedup pass reuses its unchanged chunks.
type Source ¶
type Source struct {
// contains filtered or unexported fields
}
Source implements indexjobs.Source for the scripts kind. A unit is one enabled script (SourceID = script id) and yields one item per chunk of it: the card, then the source (script.IndexChunks).
func NewSource ¶
NewSource returns a Source backed by the given store, chunking each script to maxInputBytes per item.
func (*Source) LoadItems ¶
LoadItems returns one item per chunk of the script, in chunk order, and records the hash of the text they were cut from for StampExpected. A script disabled or deleted between enqueue and claim yields an empty slice (a clean completion that clears its vectors), per the Source contract.
func (*Source) OnSucceeded ¶
OnSucceeded is a no-op: the ranked search reads the chunk table on every query, so there is no in-memory cache to refresh after a backfill.
type Store ¶
type Store struct {
// contains filtered or unexported fields
}
Store reads scripts and writes their chunk vectors for the indexjobs scripts consumer. It is separate from the request-path script store: it touches only script_embedding_chunks and the index_embedded_hash and index_model markers, and is scoped to the backfill path. The request-path store records index_text_hash on every save; when it moves, this Store rebuilds the script's chunks, reusing the ones whose text did not.
func (*Store) Coverage ¶
func (s *Store) Coverage(ctx context.Context, currentModel string) (indexed, expected int, err error)
Coverage returns the enabled scripts whose chunk set is current (indexed) and all enabled scripts (expected).
func (*Store) FindGaps ¶
FindGaps returns the ids of enabled scripts owing an embedding. A save that lands while a script is being embedded leaves the hashes apart once that pass stamps, so the script is found here again.
func (*Store) GetIndexed ¶ added in v1.140.0
GetIndexed returns the fields an enabled script is indexed on: everything script.IndexCorpus reads, and no other. A field composed into the text but missing here would leave the worker hashing a different document from the one the save hashed, so the script would be found owing an embedding on every sweep. A script disabled or deleted since it was enqueued yields errNotIndexable.
func (*Store) ListVectors ¶
func (s *Store) ListVectors(ctx context.Context, scriptID string) (map[string]indexjobs.Vector, error)
ListVectors returns the script's chunk vectors keyed by item id, for the worker's text-hash and model dedup pass. A script with no chunks yields an empty map, so every chunk is embedded.
func (*Store) ReplaceVectors ¶ added in v1.140.0
ReplaceVectors writes the script's chunk set atomically: every supplied row is upserted and every other chunk deleted, so a script whose source shrank keeps no vector for code it no longer has. An empty set deletes every chunk, which is how a disabled script is cleared. The script's updated_at is left alone: a background embed is not an edit, and the portal listing orders on it.