proctree

package
v1.55.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 11, 2026 License: MIT Imports: 11 Imported by: 0

Documentation

Overview

Package proctree enumerates and kills process trees.

It exists because a process group is not a process tree. Setpgid isolates a spawned leader in its own group, but a descendant that calls setsid() - which is what Node's `detached: true` does, and therefore what Claude Code's CLI Bash tool does - leaves that group and becomes unreachable by kill(-pgid), along with everything beneath it. Reaping such a descendant requires walking ppid links, not signalling a group.

The package knows nothing about runs, steps, or configuration. It is deliberately small so its blast radius is auditable: every per-pid kill and group kill is guarded by a freshly re-read start time, and both also refuse a protected pid - pid 0 and 1, the current process, its ancestors, and the leader of its own process group.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func Kill

func Kill(procs []Proc)

Kill SIGKILLs each process whose recorded start time still matches the one the kernel reports, skipping anything that cannot be verified.

Verification is the whole point. Between the sample that produced procs and this call, the kernel can recycle a pid onto an unrelated process; signalling it would be precisely the collateral damage this package exists to prevent. If the check cannot be made, nothing is killed - refusing to reap is always recoverable, killing the wrong process is not.

Verification is targeted rather than a full process listing on purpose. This runs on every command teardown in the harness, including short-lived git subprocesses, and a full `ps -A` costs tens of milliseconds (hundreds under the race detector) with roughly a thousand processes on the box. Paying that per command would be a self-inflicted version of the slowdown this package exists to fix. With nothing to kill - the overwhelmingly common case - Kill costs nothing at all.

func KillGroups

func KillGroups(groups []int, recorded []Proc)

KillGroups SIGKILLs each group whose leader pid is still the same process that sampling recorded.

A group kill is the highest-blast-radius operation here: a stale pgid does not signal one wrong process, it signals every member of whatever group that pid now leads. Pids are recycled, and the gap between sampling and reaping can be long - a step in the motivating incident ran 86 minutes, and a persisted record can be days old - so an unverified pgid is a licence to SIGKILL an unrelated group.

The leader's sampled start time comes from recorded, which works because a setsid() escapee leads its own group and is therefore always sampled into the descendant union alongside it. A group with no recorded leader fails closed and is skipped; the start-time-guarded per-pid Kill still covers its members.

Protected pids are refused before verification for the same reason the per-pid kill refuses them: no walk that starts from a Setpgid-isolated leader can reach our own ancestry today, but the cost of a future one that does is the daemon signalling its own group, so the guard belongs on the wider-blast-radius operation rather than only on the narrower one.

func ReapRecord

func ReapRecord(rec Record)

ReapRecord kills whatever still survives from a recorded tree.

If the leader pid is alive with a matching start time the tree still belongs to a live command, so nothing is killed: another daemon owns it. Otherwise the leader is gone and any recorded descendant that is still running with a matching start time is an orphan, which is precisely what this recovers.

Recorded groups go through KillGroups, which re-verifies each group leader's start time before signalling. That guard matters most here: a record survives daemon restarts and can be days old, so a bare pgid list would be a licence to SIGKILL a whole group that a recycled pid has since led.

func RemoveRecord

func RemoveRecord(dir string, leaderPID int)

RemoveRecord deletes a leader's record. It is a no-op if the file is gone.

func WriteRecord

func WriteRecord(dir string, rec Record) error

WriteRecord persists rec as <dir>/<leaderPID>.json.

The write goes to a temp file and is then renamed, so a daemon killed mid-write leaves either the old record or the new one, never a half-parsed one that recovery would have to guess about.

Types

type Proc

type Proc struct {
	PID     int
	PPID    int
	PGID    int
	Started time.Time
}

Proc is one row of a process snapshot.

func Descendants

func Descendants(snap []Proc, leaderPID int) []Proc

Descendants returns every process below leaderPID in snap, excluding the leader itself.

Two independent criteria are used, because either one alone has a blind spot:

  • Transitive ppid links catch a setsid() escapee and everything beneath it, which a group-based reap cannot reach.
  • Shared pgid catches a process whose intermediate parent already exited, at which point the kernel rewrites its ppid to 1 and the trail to the leader is gone.

The pgid criterion is applied only when the leader is itself a group leader (PGID == PID), which ConfigureShellCommand guarantees via Setpgid. Without that check, a leader that inherited the daemon's group would drag every sibling process - including the daemon - into the result.

func Snapshot

func Snapshot() ([]Proc, error)

Snapshot lists every visible process.

type Record

type Record struct {
	LeaderPID   int       `json:"leader_pid"`
	LeaderStart time.Time `json:"leader_started_at"`
	Descendants []Proc    `json:"descendants,omitempty"`
	Groups      []int     `json:"groups,omitempty"`
}

Record is the on-disk trail a live process tree leaves behind so a restarted daemon can finish reaping it.

It exists because the in-memory descendant union dies with the daemon. If the daemon is SIGKILLed - which is exactly what the OOM killer does when leaked worker pools exhaust the host - every leaked descendant becomes unattributable the moment its parent exits and the kernel rewrites its ppid to 1. The record is the only thing that can still name them afterwards.

The recorded start times are what make the record safe to act on later. LeaderStart protects the whole-tree ownership decision, while each descendant's start time protects both its per-pid kill and any group kill it leads. Pids are recycled, so bare pid or pgid lists read minutes or days after the fact would be a licence to kill strangers.

func ReadRecords

func ReadRecords(dir string) ([]Record, error)

ReadRecords loads every record in dir. A missing directory is not an error: it just means no tree was ever tracked under this root.

Unparseable files are skipped rather than failing the sweep. A truncated record is the expected artifact of an abrupt daemon death, and it must not stop the intact records from being recovered.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL