scriptindex

package
v1.140.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 6, 2026 License: Apache-2.0 Imports: 13 Imported by: 0

Documentation

Overview

Package scriptindex is the managed-script consumer of the shared indexjobs framework (#1370). It registers a Source/Sink pair under source_kind = "scripts" so scripts are embedded off the request path, and a person asking for what they want done reaches the automation that does it rather than needing the words its author used.

A script is embedded as a set of chunks (#2027): its description card, then its source in pieces each within the provider's input, so the reasoning in its comments and the tables its SQL names are found as well as its description, and no part of it is trimmed. The vectors live in script_embedding_chunks, one row per chunk, the way knowledge pages keep theirs (000097); search scores each script by its best chunk. SourceID is the script id, and an item id is "<script id>:<chunk>".

Whether a script owes an embedding is read from two hashes on its row: the hash of what it is indexed on now (index_text_hash, written by every save) and the hash of what its chunks were built from (index_embedded_hash, written here), with the model that built them (index_model).

Every enabled script is indexed regardless of lifecycle status: the store's search applies the discoverable-status filter at query time, so the index covers what any caller can rank. A disabled script is never embedded and never counted as missing coverage.

Index

Constants

View Source
const SourceKind = "scripts"

SourceKind is the indexjobs source_kind this package serves.

Variables

This section is empty.

Functions

func Chunks added in v1.140.0

func Chunks(s *script.Script, maxBytes int) []string

Chunks splits a script into the units its vectors are built from, each within maxBytes so no part of it is trimmed before it is embedded (#2027): the card first, then the source, each source chunk headed by the script's title so it is placed in the same space a query about the script is. The source is cut where a definition, a load or a comment block starts at the left margin, so a comment stays with the code it explains. A maxBytes below textchunk.MinViableChunkBytes disables splitting.

func Corpus added in v1.140.0

func Corpus(s *script.Script) string

Corpus is everything a script is indexed on: its card and its source. Its hash decides whether a saved version must be embedded again.

func RegisterConsumer added in v1.140.0

func RegisterConsumer(reg interface {
	Register(indexjobs.Source, indexjobs.Sink) error
}, db *sql.DB, currentModel string, maxInputBytes int,
) error

RegisterConsumer registers the scripts Source/Sink pair, chunking each script to maxInputBytes per item.

Types

type Sink

type Sink struct {
	// contains filtered or unexported fields
}

Sink implements indexjobs.Sink for the scripts kind over the script's chunk table. currentModel is the provider model the gap query diffs stored chunk sets against, so a model swap re-embeds scripts built by the previous one.

func NewSink

func NewSink(store *Store, currentModel string) *Sink

NewSink returns a Sink backed by the given store. currentModel is the embedding provider's model identifier (embedding.ModelName); "" on a provider that does not name its model.

func (*Sink) Coverage

func (s *Sink) Coverage(ctx context.Context) (indexjobs.Coverage, error)

Coverage reports enabled scripts whose chunk set is current against all enabled scripts. ExpectedKnown is true: every enabled script converges.

func (*Sink) FindGaps

func (s *Sink) FindGaps(ctx context.Context) ([]string, error)

FindGaps returns enabled script ids whose chunk set is missing, was built from text the script no longer has, or was built by another model.

func (*Sink) Kind

func (*Sink) Kind() string

Kind reports the scripts source kind.

func (*Sink) ListExisting

func (s *Sink) ListExisting(ctx context.Context, key indexjobs.Key) (map[string]indexjobs.Vector, error)

ListExisting returns the script's chunk vectors keyed by item id for the worker's dedup pass, so an edit re-embeds only the chunks whose text moved.

func (*Sink) StampExpected

func (s *Sink) StampExpected(ctx context.Context, key indexjobs.Key, _ int) error

StampExpected records that the script's chunk set was built by the current model from the text LoadItems read. The worker calls it only after a successful pass; a failure here is non-fatal, since the next sweep finds the script again and the dedup pass reuses its unchanged chunks.

func (*Sink) Upsert

func (s *Sink) Upsert(ctx context.Context, key indexjobs.Key, rows []indexjobs.Vector) error

Upsert replaces the script's chunk set with the supplied rows, removing any chunk its current text no longer produces.

func (*Sink) UpsertBatch

func (s *Sink) UpsertBatch(ctx context.Context, key indexjobs.Key, rows []indexjobs.Vector) error

UpsertBatch writes one batch of the embed pass in place, leaving the script's other chunks alone so partial progress survives a failure.

type Source

type Source struct {
	// contains filtered or unexported fields
}

Source implements indexjobs.Source for the scripts kind. A unit is one enabled script (SourceID = script id) and yields one item per chunk of it: the card, then the source (script.IndexChunks).

func NewSource

func NewSource(store *Store, maxInputBytes int) *Source

NewSource returns a Source backed by the given store, chunking each script to maxInputBytes per item.

func (*Source) Kind

func (*Source) Kind() string

Kind reports the scripts source kind.

func (*Source) LoadItems

func (s *Source) LoadItems(ctx context.Context, sourceID string) ([]indexjobs.Item, error)

LoadItems returns one item per chunk of the script, in chunk order, and records the hash of the text they were cut from for StampExpected. A script disabled or deleted between enqueue and claim yields an empty slice (a clean completion that clears its vectors), per the Source contract.

func (*Source) OnSucceeded

func (*Source) OnSucceeded(string)

OnSucceeded is a no-op: the ranked search reads the chunk table on every query, so there is no in-memory cache to refresh after a backfill.

type Store

type Store struct {
	// contains filtered or unexported fields
}

Store reads scripts and writes their chunk vectors for the indexjobs scripts consumer. It is separate from the request-path script store: it touches only script_embedding_chunks and the index_embedded_hash and index_model markers, and is scoped to the backfill path. The request-path store records index_text_hash on every save; when it moves, this Store rebuilds the script's chunks, reusing the ones whose text did not.

func NewStore

func NewStore(db *sql.DB) *Store

NewStore returns a Store over the given database.

func (*Store) Coverage

func (s *Store) Coverage(ctx context.Context, currentModel string) (indexed, expected int, err error)

Coverage returns the enabled scripts whose chunk set is current (indexed) and all enabled scripts (expected).

func (*Store) FindGaps

func (s *Store) FindGaps(ctx context.Context, currentModel string) ([]string, error)

FindGaps returns the ids of enabled scripts owing an embedding. A save that lands while a script is being embedded leaves the hashes apart once that pass stamps, so the script is found here again.

func (*Store) GetIndexed added in v1.140.0

func (s *Store) GetIndexed(ctx context.Context, id string) (*script.Script, error)

GetIndexed returns the fields an enabled script is indexed on: everything script.IndexCorpus reads, and no other. A field composed into the text but missing here would leave the worker hashing a different document from the one the save hashed, so the script would be found owing an embedding on every sweep. A script disabled or deleted since it was enqueued yields errNotIndexable.

func (*Store) ListVectors

func (s *Store) ListVectors(ctx context.Context, scriptID string) (map[string]indexjobs.Vector, error)

ListVectors returns the script's chunk vectors keyed by item id, for the worker's text-hash and model dedup pass. A script with no chunks yields an empty map, so every chunk is embedded.

func (*Store) ReplaceVectors added in v1.140.0

func (s *Store) ReplaceVectors(ctx context.Context, scriptID string, rows []indexjobs.Vector) error

ReplaceVectors writes the script's chunk set atomically: every supplied row is upserted and every other chunk deleted, so a script whose source shrank keeps no vector for code it no longer has. An empty set deletes every chunk, which is how a disabled script is cleared. The script's updated_at is left alone: a background embed is not an edit, and the portal listing orders on it.

func (*Store) Stamp added in v1.140.0

func (s *Store) Stamp(ctx context.Context, scriptID, model string) error

Stamp records that the script's chunk set was built by model from the text LoadItems read for it. When LoadItems read nothing (the script was disabled or deleted), there is nothing built to record.

func (*Store) UpsertVectors

func (s *Store) UpsertVectors(ctx context.Context, scriptID string, rows []indexjobs.Vector) error

UpsertVectors writes one batch of chunk vectors in place, leaving every chunk outside the batch alone.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL