gitparse

package
v3.97.9 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 24, 2026 License: AGPL-3.0 Imports: 17 Imported by: 1

README

gitparse

gitparse turns a git repository's history into a stream of *Diff values, one for each changed file in each commit. It is the front door for the git based sources: everything TruffleHog scans out of a repository's history comes through this package.

It works by running the git binary and reading its text output, rather than reading git's object files directly. Parsing text is not pretty, but git log --patch already handles renames, binary files, path filters and merges, and rewriting all of that ourselves would be a much bigger thing to get wrong.

What comes out

A Diff is one file's added content inside one commit, along with the details of that commit:

diffChan, err := parser.RepoPath(ctx, repoPath, "", "", nil, false)
for diff := range diffChan {
    diff.Commit.Hash    // the commit this change belongs to
    diff.PathB          // the path of the changed file
    diff.LineStart      // where in the file the hunk starts
    diff.ReadCloser()   // the added lines, and nothing else
}

Only added lines are kept. Removed and unchanged lines are thrown away, because a secret that was deleted was already there in the commit that added it, and that commit is in the stream too.

Commit details are attached to every diff. A commit with no file changes at all, such as a merge or an empty commit, is still sent once on its own, because its message, author and notes can hold a secret even when no file changed.

Diff content goes to one of two writers, chosen by the caller. buffer_writer keeps it in memory and is the default. buffered_file_writer spills to disk past a size and is turned on with UseCustomContentWriter. Repositories with big files want the second one.

How the log is produced

There are two ways, and the caller picks.

The normal way is one git log --patch for the whole repository. Simple, fast, and fine for most repositories.

The lower memory way is turned on with UseLowMemoryScan, which the CLI exposes as --git-low-memory-scan. It exists because one git log --patch keeps a little state for every commit it walks and never gives any of it back, so on a repository with millions of commits it grows until the machine kills it.

That mode splits the work in two:

  1. List the commits. git rev-list prints the hashes we want and nothing else. The hashes come back in groups, so the next step can start before the whole history has been walked.

  2. Make the patches. Each group of hashes is handed to its own short lived git log process, which reads them on standard input. Each process only holds state for its own group and then exits, so this step costs the same no matter how long the history is.

The results of all those processes are joined back into the single channel the caller sees, so nothing above this package needs to know which way was used.

Things that look like details but are not

Each of these came from a test or a measurement that said otherwise.

Listing uses git rev-list, not git log. They can both print hashes, but git log sets up git's diff machinery as soon as a diff option is present, and then compares files in every commit just to decide which commits to print. rev-list never does that. This is the step that uses the most memory on long histories, so the difference matters, and it is large.

Diff options stay out of the listing step. Options like --diff-filter=AM live with the patch options, not the commit picking options. rev-list rejects them anyway. Leaving them out means the listing picks up a few extra commits whose changes are all filtered away, and the git log that makes the patches drops those commits itself, exactly as the single command form does. The end result is the same commits, worked out in the cheaper place.

Patches come from git log --no-walk=unsorted, not git show. They print the same thing for ordinary commits, but only git log applies --diff-filter to whole commits, so using git show here would scan more commits than the normal way does. unsorted matters too: plain --no-walk re-sorts each group by commit date, which scrambles the order across groups.

Hashes go in on standard input, not as arguments. A command line has a length limit, and on Windows it is short enough to cap a group at a few dozen commits. Standard input has no such limit, so groups can be thousands of commits and full hashes can be used instead of shortened ones. Fewer, bigger groups mean fewer git processes to start.

Path filters go to both steps. The listing needs them to pick the same commits the single command would have shown, and the patch step needs them to decide which files show up.

The last commit of a stream needs finishing off. The parser normally completes a commit when it sees the next one begin. The last commit never gets that, so it is completed at the end of the stream instead. Without this, a commit with no file changes was dropped whenever it landed at the end, which used to be rare with one long stream and became common once the log was split into groups.

Stopping early

A scan can stop before the diffs run out, when it hits its maximum depth or reaches its base commit. A Go channel cannot tell you that its reader has gone away, so the scan cancels its context on the way out, and that is what shuts down the git processes still producing diffs.

Cancelling also keeps the listing step short. It only ever runs a few groups ahead of what is being consumed, because the channels between the steps are unbuffered, so a scan that stops early never walks the whole history.

Tests

go test ./pkg/gitparse/

The tests in lowmemory_test.go hold the lower memory mode to one rule: it must produce exactly what the single command form produces, with the same commits, paths and content, wherever the group boundaries land. They run with tiny group sizes on purpose, since a group size of one puts every commit at the end of a stream at once, which is where the awkward cases live.

Documentation

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type Commit

type Commit struct {
	Hash      string
	Author    string
	Committer string
	Date      time.Time
	Message   strings.Builder
	Size      int // in bytes
	// contains filtered or unexported fields
}

Commit contains commit header info and diffs.

type Diff

type Diff struct {
	PathB     string
	LineStart int
	IsBinary  bool

	Commit *Commit
	// contains filtered or unexported fields
}

Diff contains the information about a file diff in a commit. It abstracts the underlying content representation, allowing for flexible handling of diff content. The use of contentWriter enables the management of diff data either in memory or on disk, based on its size, optimizing resource usage and performance.

func (*Diff) Len added in v3.66.2

func (d *Diff) Len() int

Len returns the length of the storage.

func (*Diff) ReadCloser added in v3.66.2

func (d *Diff) ReadCloser() (io.ReadCloser, error)

ReadCloser returns a ReadCloser for the contentWriter.

type Option added in v3.27.0

type Option func(*Parser)

Option is used for adding options to Config.

func UseCustomContentWriter added in v3.67.2

func UseCustomContentWriter() Option

UseCustomContentWriter sets useCustomContentWriter option.

func UseLowMemoryScan added in v3.97.5

func UseLowMemoryScan() Option

UseLowMemoryScan sets the scan to optimize for limited memory at the cost of speed (currently up to 9%)

func WithMaxCommitSize added in v3.27.0

func WithMaxCommitSize(maxCommitSize int64) Option

WithMaxCommitSize sets maxCommitSize option. Commits larger than maxCommitSize will be put in the commit channel and additional diffs will be added to a new commit.

func WithMaxDiffSize added in v3.27.0

func WithMaxDiffSize(maxDiffSize int64) Option

WithMaxDiffSize sets maxDiffSize option. Diffs larger than maxDiffSize will be truncated.

func WithWaitDelay added in v3.93.2

func WithWaitDelay(waitDelay time.Duration) Option

WithWaitDelay sets the waitDelay option. This specifies how long to wait after context cancellation before forcefully killing git processes.

type ParseState added in v3.45.1

type ParseState int
const (
	Initial ParseState = iota
	CommitLine
	MergeLine
	AuthorLine
	AuthorDateLine
	CommitterLine
	CommitterDateLine
	MessageStartLine
	MessageLine
	MessageEndLine
	NotesStartLine
	NotesLine
	NotesEndLine
	DiffLine
	ModeLine
	IndexLine
	FromFileLine
	ToFileLine
	BinaryFileLine
	HunkLineNumberLine
	HunkContentLine
	ParseFailure
)

func (ParseState) String added in v3.45.1

func (state ParseState) String() string

type Parser added in v3.27.0

type Parser struct {
	// contains filtered or unexported fields
}

Parser sets values used in GitParse.

func NewParser added in v3.27.0

func NewParser(options ...Option) *Parser

NewParser creates a GitParse config from options and sets defaults.

func (*Parser) FromReader added in v3.31.3

func (c *Parser) FromReader(ctx context.Context, stdOut io.Reader, diffChan chan *Diff, isStaged bool)

func (*Parser) RepoPath added in v3.27.0

func (c *Parser) RepoPath(
	ctx context.Context,
	source string,
	head string,
	base string,
	excludedGlobs []string,
	isBare bool,
) (chan *Diff, error)

RepoPath parses the output of the `git log` command for the `source` path. The Diff chan will return diffs in the order they are parsed from the log, though the diffs are generated using `git show` in groups.

head and base are commit hashes or refs. When base is non-empty the log is restricted to the range base..head (commits reachable from head but not from base), which is the diff-scan contract behind `--since-commit`. The range is computed by git itself so that it is independent of commit dates and merge topology. An empty base means a full-history scan of head (or of --all when head is also empty). An empty head with a non-empty base means every commit reachable from any ref but not from base (`--all ^base`), which is what `--since-commit X` without `--branch` has always covered; callers that want a single-branch diff pass the head explicitly.

func (*Parser) Staged added in v3.43.0

func (c *Parser) Staged(ctx context.Context, source string) (chan *Diff, error)

Staged parses the output of the `git diff` command for the `source` path.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL