backfill

package
v0.49.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 1, 2026 License: Apache-2.0, MIT Imports: 21 Imported by: 0

Documentation

Index

Constants

View Source
const (
	// DefaultPreviewBatch is the number of spans one batch reads, previews
	// and writes back in a single transaction.
	DefaultPreviewBatch = 500
	// DefaultPreviewPause is the idle time between batches, so a backfill
	// over a large corpus shares the database with live derive and reads.
	DefaultPreviewPause = 200 * time.Millisecond
)

Variables

This section is empty.

Functions

This section is empty.

Types

type PreviewOptions added in v0.49.0

type PreviewOptions struct {
	// Store serves the rows without previews and takes the previews back.
	Store storage.PreviewBackfiller

	// SessionID restricts the scan to one session; empty means every
	// session.
	SessionID string

	// BatchSize bounds every read and every write transaction. Zero
	// means DefaultPreviewBatch.
	BatchSize int

	// Pause is the sleep between batches; zero or negative means none. The
	// CLI defaults it to DefaultPreviewPause.
	Pause time.Duration

	// DryRun scans and counts but writes nothing.
	DryRun bool

	Logf func(format string, args ...any)

	// AfterBatch, when set, runs after each batch has been applied (or
	// counted, in a dry run) and before the pause. It exists so a caller
	// can observe or interrupt progress at a batch boundary.
	AfterBatch func(progress PreviewProgress)
}

PreviewOptions configures one run of the span preview backfill.

type PreviewProgress added in v0.49.0

type PreviewProgress struct {
	// Batch is the rows the batch read.
	Batch []storage.SpanBackfillRow
	// Backfilled is the running total of rows written (or, in a dry
	// run, rows that would have been).
	Backfilled int
	// Last is the keyset cursor the next batch starts after.
	Last storage.SpanBackfillCursor
}

PreviewProgress is the state of a run after one batch.

type PreviewResult added in v0.49.0

type PreviewResult struct {
	// Batches is the number of non-empty pages read.
	Batches int `json:"batches"`
	// Backfilled is the number of spans whose previews were written
	// (counted but not written in a dry run).
	Backfilled int `json:"backfilled"`
	// Last is the key of the last span the run reached; zero when the
	// scan found nothing.
	Last storage.SpanBackfillCursor `json:"last"`
	// DryRun echoes the option so a report is self-describing.
	DryRun bool `json:"dry_run"`
}

PreviewResult summarizes a run.

func Previews added in v0.49.0

func Previews(ctx context.Context, opts PreviewOptions) (*PreviewResult, error)

Previews fills the stored input_preview / output_preview of spans derived before the preview columns existed, from the payload already on the row, as derive.PreviewBlocks would have at write time. It never re-derives and never touches the payload, content_hash or derive_seq.

The scan is keyset-paged on (session_id, trace_id, span_id) with one bounded transaction per batch and a pause between batches, so it can run against a live deployment. It is idempotent and resumable: a row that has previews no longer matches the selection, so an interrupted run picks up where it stopped when re-run, and a completed run is a no-op. The run ends when a page comes back empty.

type TranscriptUploadOptions added in v0.16.0

type TranscriptUploadOptions struct {
	// ProjectDir is one Claude Code project directory
	// (~/.claude/projects/<flattened-cwd>) holding <session>.jsonl
	// files and per-session subagents/ directories.
	ProjectDir string

	// SessionIDs filters which sessions upload; empty means every
	// .jsonl in the directory.
	SessionIDs []string

	// IngestURL is the tapes-ingest base URL.
	IngestURL string

	// HarnessID tags the session envelope; Claude Code transcripts are
	// "claude".
	HarnessID string

	Verbose bool
	Logf    func(format string, args ...any)
}

TranscriptUploadOptions configures a harness-transcript upload run.

type TranscriptUploadResult added in v0.16.0

type TranscriptUploadResult struct {
	Sessions int      `json:"sessions"`
	Files    int      `json:"files"`
	Uploaded int      `json:"uploaded"`
	Deduped  int      `json:"deduped"`
	Failed   int      `json:"failed"`
	Failures []string `json:"failures,omitempty"`
}

TranscriptUploadResult summarizes an upload run.

func UploadTranscripts added in v0.16.0

func UploadTranscripts(ctx context.Context, opts TranscriptUploadOptions) (*TranscriptUploadResult, error)

UploadTranscripts pushes harness transcripts (main + subagents) into the tapes raw layer via POST /v1/ingest/transcript. Idempotent: the server dedups by content version, so re-running uploads nothing for unchanged files and a new version for grown ones.

type WireTraceOptions added in v0.16.0

type WireTraceOptions struct {
	// CapturesDir is the paperd wire-trace root holding turn-* bundles
	// (request.json + response.sse + meta.json per turn).
	CapturesDir string

	// IngestURL is the base URL of a tapes-ingest server, e.g.
	// "http://127.0.0.1:8090". The backfill POSTs one envelope per
	// captured turn to {IngestURL}/v1/ingest.
	IngestURL string

	// SessionIDs filters replay to bundles whose captured
	// X-Tapes-Harness-Session-Id matches; empty replays everything.
	SessionIDs []string

	// DryRun parses and reduces every bundle but skips the POST.
	DryRun bool

	// Verbose logs each turn's outcome to Logf.
	Verbose bool

	// Logf receives progress output; defaults to a no-op.
	Logf func(format string, args ...any)
}

WireTraceOptions configures a wire-trace → ingest backfill run.

type WireTraceResult added in v0.16.0

type WireTraceResult struct {
	Scanned  int      `json:"scanned"`
	Posted   int      `json:"posted"`
	RawOnly  int      `json:"raw_only"`
	Skipped  int      `json:"skipped"`
	Failed   int      `json:"failed"`
	Failures []string `json:"failures,omitempty"`
}

WireTraceResult summarizes a wire-trace backfill run.

func WireTrace added in v0.16.0

func WireTrace(ctx context.Context, opts WireTraceOptions) (*WireTraceResult, error)

WireTrace replays paperd wire-trace capture bundles through a tapes-ingest server, reconstructing the envelope tapes-extproc would have dispatched live: verbatim request bytes, response reduced with the same shared capture reducer, meta rebuilt from the bundle, and the session block recovered from the captured X-Tapes-* headers.

The whole flow is idempotent end to end: the raw layer dedupes on (org, request_id) — the turn directory name, stable across re-runs — and node inserts are content-addressed ON CONFLICT DO NOTHING, so turns that already landed via live capture reproduce their existing nodes byte for byte and only gain the previously-missing raw row.

Only provider chat-completion calls are replayed (…/v1/messages); auxiliary traffic in the trace (count_tokens, tapes API reads) is skipped.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL