diff

package
v0.1.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 7, 2026 License: MIT Imports: 25 Imported by: 0

Documentation

Overview

Package diff implements the format-agnostic keyed row diff: an in-memory hash join over two Sources.

Algorithm (in-memory mode):

pass 1  scan left, build keyHash → rowHash table (~17 B/row)
pass 2  scan right, probe: miss → added; hash mismatch → changed
        (store the right row); match → unchanged
pass 3  only if anything differed: rescan left to attribute changed
        columns and collect removed-key examples

Identical inputs (the CI hot path) do exactly one scan of each side and stores nothing but the table.

All passes run on every core: sources deliver rows concurrently (parquet decodes row groups in parallel, CSV parses records in a worker pool; see source.ParallelScanner) and the join table is striped 64 ways.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func InferKey

func InferKey(left, right source.Source, renames map[string]string) (string, error)

InferKey picks a key column for the pair, or errors with guidance. renames maps right-side column names to the left-side names they are compared under (--rename), so a renamed column is a key candidate like any other.

func ParseBudget

func ParseBudget(spec string, rows int64) (int64, error)

ParseBudget turns a --max-diff spec ("1000" or "0.5%") into an absolute row budget against a row count.

func ResolveColumns

func ResolveColumns(left, right source.Schema, opts Options) (keyNames []string, keyTypes []source.Type, valNames []string, valTypes []source.Type, err error)

ResolveColumns reports the key and compared-column layout a diff of these schemas will use: the export writers build their file schemas from it. Coerced columns take the comparison-domain type (int-vs-float → float64).

func WriteSnapshot

func WriteSnapshot(src source.Source, path string, opts Options) (int64, error)

WriteSnapshot scans src and writes its hash manifest to path.

Types

type ColumnChange

type ColumnChange struct {
	Column string `json:"column"`
	Left   string `json:"left"`
	Right  string `json:"right"`
}

ColumnChange is one changed cell in an example row.

type ColumnStat

type ColumnStat struct {
	// Changed counts the row pairs that differ in this column (differences
	// inside --tolerance are not counted).
	Changed int64 `json:"changed"`
	// MatchRate is the share of compared row pairs that agree in this
	// column, in [0,1].
	MatchRate float64 `json:"match_rate"`
	// MaxAbsDiff/MeanAbsDiff size the differences of a numeric column
	// (int64 and float only, over pairs where neither side is NULL).
	MaxAbsDiff  float64 `json:"max_abs_diff,omitempty"`
	MeanAbsDiff float64 `json:"mean_abs_diff,omitempty"`
	Numeric     bool    `json:"numeric,omitempty"`
	// contains filtered or unexported fields
}

ColumnStat is one compared column's change profile.

type Options

type Options struct {
	Keys          []string // key column names (required)
	IgnoreColumns []string // excluded from comparison
	// Rename maps a right-side column name to the left-side name it should
	// be compared under, so a renamed column is an ordinary compared column
	// instead of an added/removed pair.
	Rename map[string]string
	// Keyless diffs without a key: whole rows are matched as a multiset, so
	// leftovers on either side are added/removed and Changed is always 0.
	// Mutually exclusive with Keys, OnDup and Tolerance.
	Keyless bool
	// Mask lists columns whose values are replaced by a stable short token
	// everywhere they would be displayed or exported. Display-only: the
	// comparison still uses the real values.
	Mask []string
	// Where records the row filter the caller applied to both inputs, in
	// rendered form. The engine never evaluates it, since filtering happens
	// in the source layer, but it is recorded in snapshots and reported, so a
	// baseline can never be compared against a differently-filtered file.
	Where string
	// MaxDiff is the CI budget ("1000" or "0.5%"). Setting it lets the run
	// stop as soon as the budget is provably exceeded, marking the Result
	// Aborted with partial counts. The caller still decides the verdict.
	MaxDiff string
	Limit   int // max examples kept per category (default 10)
	Threads int // concurrency (default GOMAXPROCS)
	// Summary skips per-column change attribution and example rows: counts
	// and exit code only. This drops the whole third pass: two scans total,
	// same as the identical-inputs fast path.
	Summary bool
	// Mode selects the join strategy: "auto" (default) picks in-memory for
	// inputs whose key table fits comfortably in RAM and streaming above
	// that; "memory" and "stream" force one. Streaming (a grace hash join
	// spilling hash pairs to TempDir) keeps peak memory bounded by
	// partition size instead of input size.
	Mode string
	// TempDir is where streaming mode spills (default os.TempDir()).
	TempDir string
	// OnDup selects duplicate-key handling: "error" (default) fails the
	// diff; "warn" keeps each key's first occurrence per side and reports
	// how many rows were set aside; "match" pairs duplicate keys' rows as
	// multisets (identical rows cancel, leftovers are added/removed, no
	// change attribution inside a group; see dupmatch.go).
	OnDup string
	// FloatPrecision, when > 0, rounds float comparisons to that many
	// decimal digits before hashing and comparing (hash-consistent
	// alternative to an epsilon tolerance).
	FloatPrecision int
	// IgnoreCase folds string comparisons to lower case.
	IgnoreCase bool
	// Trim strips leading and trailing whitespace from string comparisons.
	Trim bool
	// TimestampPrecision truncates timestamp comparisons to "s", "ms" or
	// "us" (empty or "us" = exact microseconds).
	TimestampPrecision string
	// Tolerance, when set, is the epsilon applied to every numeric column;
	// ColumnTolerance overrides it per column. Rows whose every difference
	// is inside tolerance are counted as WithinTolerance instead of Changed.
	// Requires full diff mode (not Summary, not a snapshot).
	Tolerance       *Tolerance
	ColumnTolerance map[string]Tolerance
	// Sink, when set, receives every differing row (added/removed/changed)
	// with typed values as the diff runs, with no extra scans. Implementations
	// must be safe for concurrent calls. Incompatible with Summary.
	Sink RowSink
	// Progress, when set, is updated as the diff runs (phase + rows seen).
	Progress *Progress
}

Options configures a row diff.

type Progress

type Progress struct {
	// contains filtered or unexported fields
}

Progress carries live counters a caller can render (atomically updated, once per batch, so the overhead is negligible).

func (*Progress) Snapshot

func (p *Progress) Snapshot() (string, int64)

Snapshot returns the current phase name and row count.

type Result

type Result struct {
	Schema schema.Diff `json:"schema"`

	LeftRows  int64 `json:"left_rows"`
	RightRows int64 `json:"right_rows"`
	Added     int64 `json:"added"`
	Removed   int64 `json:"removed"`
	Changed   int64 `json:"changed"`
	Unchanged int64 `json:"unchanged"`

	// Aborted is set when the run stopped early because --max-diff was
	// provably exceeded: the verdict was already fixed, so the remaining
	// work was skipped. Column attribution and examples are then absent.
	Aborted bool `json:"aborted,omitempty"`
	// PartialCounts distinguishes the two ways a run aborts: true when the
	// probe scan itself was cancelled mid-flight, so every count is a lower
	// bound rather than a total; false when the counts are complete and only
	// the attribution pass was skipped.
	PartialCounts bool `json:"partial_counts,omitempty"`
	// AbortReason explains an Aborted run, for the report.
	AbortReason string `json:"abort_reason,omitempty"`

	// Comparison names the settings that loosened this comparison
	// (normalizations, tolerances); empty for an exact diff.
	Comparison string `json:"comparison,omitempty"`

	// Masked lists the columns whose values were replaced by a token in this
	// result (--mask). Their values are absent from examples and exports.
	Masked []string `json:"masked,omitempty"`

	// Keyless marks a result produced without a key (--keyless): examples
	// name whole rows rather than keys, and Changed is always 0.
	Keyless bool `json:"keyless,omitempty"`

	// Filter is the row filter both inputs were reduced by (--where), if
	// any. Every count is then of the filtered universe.
	Filter string `json:"filter,omitempty"`
	// FilesPruned counts the data files partition pruning skipped entirely.
	FilesPruned int `json:"files_pruned,omitempty"`

	// WithinTolerance counts row pairs whose every difference fell inside
	// --tolerance. They are excluded from Changed and reported separately so
	// the counts stay honest about what was compared.
	WithinTolerance int64 `json:"within_tolerance,omitempty"`

	// ColumnChanges counts, per column, how many changed rows changed in
	// that column.
	ColumnChanges map[string]int64 `json:"column_changes,omitempty"`

	// ColumnStats carries the same counts plus a match rate and, for numeric
	// columns, the size of the differences.
	ColumnStats map[string]*ColumnStat `json:"column_stats,omitempty"`

	// DupsLeft/DupsRight count rows set aside under --on-dup warn (the
	// first occurrence of each key stays in the diff).
	DupsLeft  int64 `json:"dups_left,omitempty"`
	DupsRight int64 `json:"dups_right,omitempty"`
	// DupKeys counts the keys that appear more than once on the left under
	// --on-dup match, and DupRows the rows in them. (Keys duplicated only
	// on the right are matched as multisets too, but are not counted here.)
	DupKeys int64 `json:"dup_keys,omitempty"`
	DupRows int64 `json:"dup_rows,omitempty"`

	AddedExamples   []string     `json:"added_examples,omitempty"`
	RemovedExamples []string     `json:"removed_examples,omitempty"`
	ChangedExamples []RowExample `json:"changed_examples,omitempty"`
	DupExamples     []string     `json:"dup_examples,omitempty"`
}

Result is the complete diff outcome.

func DiffAgainstSnapshot

func DiffAgainstSnapshot(snapPath string, right source.Source, opts Options) (*Result, error)

DiffAgainstSnapshot compares a live source (right side) against a snapshot baseline (left side).

func Run

func Run(left, right source.Source, opts Options) (*Result, error)

Run executes the diff.

func (*Result) ComparedRows

func (r *Result) ComparedRows() int64

ComparedRows is the number of row pairs present on both sides: denominator of every per-column match rate.

func (*Result) FinishStats

func (r *Result) FinishStats()

FinishStats derives the per-column match rate and mean difference from the accumulated totals. It is idempotent, so a caller that adjusts the row counts after the diff (folding back rows skipped as shared between two snapshots of one table) can simply call it again.

func (*Result) RowsSame

func (r *Result) RowsSame() bool

RowsSame reports whether the compared rows are identical.

func (*Result) Same

func (r *Result) Same() bool

Same reports whether the inputs are identical (schema and rows). An aborted run is never same: it stopped because too many rows differed.

type RowExample

type RowExample struct {
	Key     string         `json:"key"`
	Columns []ColumnChange `json:"columns"`
}

RowExample is one example changed row.

type RowSink

type RowSink interface {
	WriteDiffRow(status byte, key, left, right []source.Value) error
}

RowSink receives full diff rows during the run. status is 'a', 'r' or 'c'; key always holds the key values; left/right hold the compared columns' values for the side(s) that have the row (nil otherwise). Values are only valid during the call.

type Tolerance

type Tolerance struct {
	Abs float64 `json:"abs,omitempty"`
	Rel float64 `json:"rel,omitempty"`
}

Tolerance is an epsilon for numeric columns: two values are tolerably equal when they differ by at most Abs, or relatively by at most Rel (of the larger magnitude). Either bound alone is enough; a zero bound is off.

func ParseTolerance

func ParseTolerance(spec string) (string, Tolerance, error)

ParseTolerance parses one --tolerance spec:

0.01               absolute, every numeric column
0.01,rel=0.001     absolute and relative
rel=1e-6           relative only
price=0.01         one column
price=0.01,rel=0.001

An empty column name means "every numeric column".

func (Tolerance) String

func (t Tolerance) String() string

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL