Documentation
¶
Overview ¶
Package diff implements the format-agnostic keyed row diff: an in-memory hash join over two Sources.
Algorithm (in-memory mode):
pass 1 scan left, build keyHash → rowHash table (~17 B/row)
pass 2 scan right, probe: miss → added; hash mismatch → changed
(store the right row); match → unchanged
pass 3 only if anything differed: rescan left to attribute changed
columns and collect removed-key examples
Identical inputs (the CI hot path) do exactly one scan of each side and stores nothing but the table.
All passes run on every core: sources deliver rows concurrently (parquet decodes row groups in parallel, CSV parses records in a worker pool; see source.ParallelScanner) and the join table is striped 64 ways.
Index ¶
- func InferKey(left, right source.Source, renames map[string]string) (string, error)
- func ParseBudget(spec string, rows int64) (int64, error)
- func ResolveColumns(left, right source.Schema, opts Options) (keyNames []string, keyTypes []source.Type, valNames []string, ...)
- func WriteSnapshot(src source.Source, path string, opts Options) (int64, error)
- type ColumnChange
- type ColumnStat
- type Options
- type Progress
- type Result
- type RowExample
- type RowSink
- type Tolerance
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func InferKey ¶
InferKey picks a key column for the pair, or errors with guidance. renames maps right-side column names to the left-side names they are compared under (--rename), so a renamed column is a key candidate like any other.
func ParseBudget ¶
ParseBudget turns a --max-diff spec ("1000" or "0.5%") into an absolute row budget against a row count.
func ResolveColumns ¶
func ResolveColumns(left, right source.Schema, opts Options) (keyNames []string, keyTypes []source.Type, valNames []string, valTypes []source.Type, err error)
ResolveColumns reports the key and compared-column layout a diff of these schemas will use: the export writers build their file schemas from it. Coerced columns take the comparison-domain type (int-vs-float → float64).
Types ¶
type ColumnChange ¶
type ColumnChange struct {
Column string `json:"column"`
Left string `json:"left"`
Right string `json:"right"`
}
ColumnChange is one changed cell in an example row.
type ColumnStat ¶
type ColumnStat struct {
// Changed counts the row pairs that differ in this column (differences
// inside --tolerance are not counted).
Changed int64 `json:"changed"`
// MatchRate is the share of compared row pairs that agree in this
// column, in [0,1].
MatchRate float64 `json:"match_rate"`
// MaxAbsDiff/MeanAbsDiff size the differences of a numeric column
// (int64 and float only, over pairs where neither side is NULL).
MaxAbsDiff float64 `json:"max_abs_diff,omitempty"`
MeanAbsDiff float64 `json:"mean_abs_diff,omitempty"`
Numeric bool `json:"numeric,omitempty"`
// contains filtered or unexported fields
}
ColumnStat is one compared column's change profile.
type Options ¶
type Options struct {
Keys []string // key column names (required)
IgnoreColumns []string // excluded from comparison
// Rename maps a right-side column name to the left-side name it should
// be compared under, so a renamed column is an ordinary compared column
// instead of an added/removed pair.
Rename map[string]string
// Keyless diffs without a key: whole rows are matched as a multiset, so
// leftovers on either side are added/removed and Changed is always 0.
// Mutually exclusive with Keys, OnDup and Tolerance.
Keyless bool
// Mask lists columns whose values are replaced by a stable short token
// everywhere they would be displayed or exported. Display-only: the
// comparison still uses the real values.
Mask []string
// Where records the row filter the caller applied to both inputs, in
// rendered form. The engine never evaluates it, since filtering happens
// in the source layer, but it is recorded in snapshots and reported, so a
// baseline can never be compared against a differently-filtered file.
Where string
// MaxDiff is the CI budget ("1000" or "0.5%"). Setting it lets the run
// stop as soon as the budget is provably exceeded, marking the Result
// Aborted with partial counts. The caller still decides the verdict.
MaxDiff string
Limit int // max examples kept per category (default 10)
Threads int // concurrency (default GOMAXPROCS)
// Summary skips per-column change attribution and example rows: counts
// and exit code only. This drops the whole third pass: two scans total,
// same as the identical-inputs fast path.
Summary bool
// Mode selects the join strategy: "auto" (default) picks in-memory for
// inputs whose key table fits comfortably in RAM and streaming above
// that; "memory" and "stream" force one. Streaming (a grace hash join
// spilling hash pairs to TempDir) keeps peak memory bounded by
// partition size instead of input size.
Mode string
// TempDir is where streaming mode spills (default os.TempDir()).
TempDir string
// OnDup selects duplicate-key handling: "error" (default) fails the
// diff; "warn" keeps each key's first occurrence per side and reports
// how many rows were set aside; "match" pairs duplicate keys' rows as
// multisets (identical rows cancel, leftovers are added/removed, no
// change attribution inside a group; see dupmatch.go).
OnDup string
// FloatPrecision, when > 0, rounds float comparisons to that many
// decimal digits before hashing and comparing (hash-consistent
// alternative to an epsilon tolerance).
FloatPrecision int
// IgnoreCase folds string comparisons to lower case.
IgnoreCase bool
// Trim strips leading and trailing whitespace from string comparisons.
Trim bool
// TimestampPrecision truncates timestamp comparisons to "s", "ms" or
// "us" (empty or "us" = exact microseconds).
TimestampPrecision string
// Tolerance, when set, is the epsilon applied to every numeric column;
// ColumnTolerance overrides it per column. Rows whose every difference
// is inside tolerance are counted as WithinTolerance instead of Changed.
// Requires full diff mode (not Summary, not a snapshot).
Tolerance *Tolerance
ColumnTolerance map[string]Tolerance
// Sink, when set, receives every differing row (added/removed/changed)
// with typed values as the diff runs, with no extra scans. Implementations
// must be safe for concurrent calls. Incompatible with Summary.
Sink RowSink
// Progress, when set, is updated as the diff runs (phase + rows seen).
Progress *Progress
}
Options configures a row diff.
type Progress ¶
type Progress struct {
// contains filtered or unexported fields
}
Progress carries live counters a caller can render (atomically updated, once per batch, so the overhead is negligible).
type Result ¶
type Result struct {
Schema schema.Diff `json:"schema"`
LeftRows int64 `json:"left_rows"`
RightRows int64 `json:"right_rows"`
Added int64 `json:"added"`
Removed int64 `json:"removed"`
Changed int64 `json:"changed"`
Unchanged int64 `json:"unchanged"`
// Aborted is set when the run stopped early because --max-diff was
// provably exceeded: the verdict was already fixed, so the remaining
// work was skipped. Column attribution and examples are then absent.
Aborted bool `json:"aborted,omitempty"`
// PartialCounts distinguishes the two ways a run aborts: true when the
// probe scan itself was cancelled mid-flight, so every count is a lower
// bound rather than a total; false when the counts are complete and only
// the attribution pass was skipped.
PartialCounts bool `json:"partial_counts,omitempty"`
// AbortReason explains an Aborted run, for the report.
AbortReason string `json:"abort_reason,omitempty"`
// Comparison names the settings that loosened this comparison
// (normalizations, tolerances); empty for an exact diff.
Comparison string `json:"comparison,omitempty"`
// Masked lists the columns whose values were replaced by a token in this
// result (--mask). Their values are absent from examples and exports.
Masked []string `json:"masked,omitempty"`
// Keyless marks a result produced without a key (--keyless): examples
// name whole rows rather than keys, and Changed is always 0.
Keyless bool `json:"keyless,omitempty"`
// Filter is the row filter both inputs were reduced by (--where), if
// any. Every count is then of the filtered universe.
Filter string `json:"filter,omitempty"`
// FilesPruned counts the data files partition pruning skipped entirely.
FilesPruned int `json:"files_pruned,omitempty"`
// WithinTolerance counts row pairs whose every difference fell inside
// --tolerance. They are excluded from Changed and reported separately so
// the counts stay honest about what was compared.
WithinTolerance int64 `json:"within_tolerance,omitempty"`
// ColumnChanges counts, per column, how many changed rows changed in
// that column.
ColumnChanges map[string]int64 `json:"column_changes,omitempty"`
// ColumnStats carries the same counts plus a match rate and, for numeric
// columns, the size of the differences.
ColumnStats map[string]*ColumnStat `json:"column_stats,omitempty"`
// DupsLeft/DupsRight count rows set aside under --on-dup warn (the
// first occurrence of each key stays in the diff).
DupsLeft int64 `json:"dups_left,omitempty"`
DupsRight int64 `json:"dups_right,omitempty"`
// DupKeys counts the keys that appear more than once on the left under
// --on-dup match, and DupRows the rows in them. (Keys duplicated only
// on the right are matched as multisets too, but are not counted here.)
DupKeys int64 `json:"dup_keys,omitempty"`
DupRows int64 `json:"dup_rows,omitempty"`
AddedExamples []string `json:"added_examples,omitempty"`
RemovedExamples []string `json:"removed_examples,omitempty"`
ChangedExamples []RowExample `json:"changed_examples,omitempty"`
DupExamples []string `json:"dup_examples,omitempty"`
}
Result is the complete diff outcome.
func DiffAgainstSnapshot ¶
DiffAgainstSnapshot compares a live source (right side) against a snapshot baseline (left side).
func (*Result) ComparedRows ¶
ComparedRows is the number of row pairs present on both sides: denominator of every per-column match rate.
func (*Result) FinishStats ¶
func (r *Result) FinishStats()
FinishStats derives the per-column match rate and mean difference from the accumulated totals. It is idempotent, so a caller that adjusts the row counts after the diff (folding back rows skipped as shared between two snapshots of one table) can simply call it again.
type RowExample ¶
type RowExample struct {
Key string `json:"key"`
Columns []ColumnChange `json:"columns"`
}
RowExample is one example changed row.
type RowSink ¶
RowSink receives full diff rows during the run. status is 'a', 'r' or 'c'; key always holds the key values; left/right hold the compared columns' values for the side(s) that have the row (nil otherwise). Values are only valid during the call.
type Tolerance ¶
Tolerance is an epsilon for numeric columns: two values are tolerably equal when they differ by at most Abs, or relatively by at most Rel (of the larger magnitude). Either bound alone is enough; a zero bound is off.
func ParseTolerance ¶
ParseTolerance parses one --tolerance spec:
0.01 absolute, every numeric column 0.01,rel=0.001 absolute and relative rel=1e-6 relative only price=0.01 one column price=0.01,rel=0.001
An empty column name means "every numeric column".