Documentation
¶
Index ¶
- Constants
- Variables
- func ActiveCustomRules() int
- func BatchBytesWithPrivacyFilter(ctx context.Context, inputs []NamedBlob) ([][]byte, error)
- func Bytes(b []byte) []byte
- func BytesWithPrivacyFilter(ctx context.Context, b []byte) []byte
- func ConfigFingerprint() string
- func ConfigureCustomRules(cfg CustomRulesConfig)
- func ConfigurePII(cfg PIIConfig)
- func ConfigurePrivacyFilter(cfg OPFConfig)
- func ConfigurePrivacyFilterWithRuntime(cfg OPFConfig, rt opfRuntime)
- func ConfigureScanners(cfg ScannersConfig) error
- func IsKnownOPFCategory(name string) bool
- func IsLineDelimited(content []byte) bool
- func JSONLContent(content string) (string, error)
- func JSONLContentWithPrivacyFilter(ctx context.Context, content string) (string, error)
- func OPFBreakerTripped() bool
- func OPFCommand() string
- func OPFEnabled() bool
- func OPFMisconfiguredNoCategories() bool
- func ResetOPFConfigForTest()
- func SetScannerDegradedForTest(v bool)
- func String(s string) string
- func StringWithPrivacyFilter(ctx context.Context, s string) string
- func SumProseLeafBytes(inputs []NamedBlob) int
- func WithScannerDegradedSole(t testingTB)
- type CustomRulesConfig
- type NamedBlob
- type OPFConfig
- type PIICategory
- type PIIConfig
- type Pack
- type RedactedBytes
- type Rule
- type Sample
- type ScannersConfig
- type Span
Constants ¶
const RedactedPlaceholder = "REDACTED"
RedactedPlaceholder is the replacement text used for redacted secrets.
const RedactorsDirName = "redactors"
RedactorsDirName is the .entire subdirectory used for user-defined rule packs.
Variables ¶
var ErrOPFNoEnabledCategories = errors.New(
"redaction.openai_privacy_filter is enabled but no detection category is enabled; " +
"enable at least one category (e.g. \"private_person\": true) under " +
"redaction.openai_privacy_filter.categories in .entire/settings.json " +
"(or .entire/settings.local.json), or set enabled: false")
ErrOPFNoEnabledCategories is returned by BatchBytesWithPrivacyFilter — the trailer-stamping path — when OPF is enabled but the effective category set is empty (categories omitted, {}, all-false, or only unknown keys). The model scan cannot run in this state; succeeding silently would let the caller attest OPF ran (Entire-OPF-Applied trailer) when it never did. The non-batched entry points (detectOPF, JSONLContentWithPrivacyFilter) still fall back silently in this state because their callers do not attest that OPF ran.
var ErrRedactionIncomplete = errors.New("redaction incomplete")
ErrRedactionIncomplete marks content that redaction flagged but could not fully rewrite: the splice missed and the structural re-encode failed too.
It exists to keep the "fail rather than certify" contract holding at the call sites that fall back to plain-text redaction when JSONLBytes errors. That fallback is for content that is not JSONL at all. Content carrying this sentinel parsed fine, and the plain-text fallback would scan the very variant-encoded spelling the verification caught the splice missing, so those callers must fail the write instead. Match with errors.Is.
var ErrScannerDegraded = errors.New("redaction scanner degraded: goredact scan failed with betterleaks disabled")
ErrScannerDegraded is returned by JSONLBytes/JSONLBytesWithPrivacyFilter when the goredact engine failed at runtime and betterleaks is disabled: the sole scanner produced no coverage, so transcript writes must fail rather than persist under-scanned content. Callers distinguish it from JSON-parse errors with errors.Is and must NOT fall back to redact.Bytes.
Functions ¶
func ActiveCustomRules ¶ added in v0.10.3
func ActiveCustomRules() int
ActiveCustomRules reports how many user-defined rules (inline + pack) actually compiled and are live in the detection pipeline — as opposed to how many were configured, which silently diverges when a regex fails to compile.
func BatchBytesWithPrivacyFilter ¶ added in v0.7.8
BatchBytesWithPrivacyFilter redacts N blobs with a single OPF inference call instead of N. Returns redacted bytes in input order (output[i] is the redaction of inputs[i]).
Failure semantics — fail-closed: any error from the OPF runtime returns a non-nil error. Callers running this for privacy-critical operations (e.g. the pre-push rewrite) must abort rather than proceed with partially-redacted content. The per-blob JSONLContentWithPrivacyFilter falls back to the regex-only pipeline (the eight always-on/opt-in layers, no OPF) on batch failure; this batched variant intentionally does not, because the only caller (cross-blob walker) needs an explicit signal that OPF did not finish.
When OPF is unconfigured, disabled, or the per-process circuit breaker has tripped, returns regex-only output for every blob with no error. This matches the existing non-batched paths and keeps the caller's hot-path code clean. Enabled with zero effective categories is different: that returns ErrOPFNoEnabledCategories, because the caller is about to stamp the Entire-OPF-Applied trailer and a silent regex-only success here would make that attestation false. Enabled with a nil runtime errors for the same reason. Both fail-closed checks run BEFORE the breaker check so the guarantee is unconditional — a tripped breaker must not downgrade a misconfiguration back into silent regex-only success.
func BytesWithPrivacyFilter ¶ added in v0.7.8
BytesWithPrivacyFilter augments Bytes with the OpenAI Privacy Filter for raw (non-JSONL) byte content. Used by checkpoint write paths that handle metadata files which may or may not be JSONL.
func ConfigFingerprint ¶ added in v0.10.3
func ConfigFingerprint() string
ConfigFingerprint returns a stable hash over everything that affects the output of String and JSONLContent (the eight regex layers). Callers that cache redacted output use it to tell when a cached result was produced under different rules and must be discarded.
Two things it deliberately does NOT cover, because it cannot:
- The OpenAI Privacy Filter. OPF is a network-backed layer whose output is not a pure function of local config, so cached output must never be shared between OPF and non-OPF paths. Callers that cache are responsible for keying on which pipeline produced the bytes.
- The betterleaks ruleset version, which the library does not expose. A caller that caches across CLI upgrades must mix in the binary's own version, which moves whenever the vendored ruleset does.
Safe for concurrent use.
func ConfigureCustomRules ¶ added in v0.6.2
func ConfigureCustomRules(cfg CustomRulesConfig)
ConfigureCustomRules compiles user-defined redaction rules and stores the result for use by redact.String(). Sample-validation runs here too, so failures surface the next time any process initializes redaction.
Call once at process startup after loading settings. Thread-safe.
func ConfigurePII ¶ added in v0.5.1
func ConfigurePII(cfg PIIConfig)
ConfigurePII sets the global PII redaction configuration. Pre-compiles patterns so the hot path (String → detectPII) does no compilation. Call once at startup after loading settings. Thread-safe.
func ConfigurePrivacyFilter ¶ added in v0.7.8
func ConfigurePrivacyFilter(cfg OPFConfig)
ConfigurePrivacyFilter sets the global OPF configuration and constructs the default shell-out runtime. Call once at process startup after loading settings. Thread-safe. Subsequent calls replace the previous configuration and reset the circuit breaker (a new config might fix what broke the prior one).
func ConfigurePrivacyFilterWithRuntime ¶ added in v0.7.8
func ConfigurePrivacyFilterWithRuntime(cfg OPFConfig, rt opfRuntime)
ConfigurePrivacyFilterWithRuntime is the test-only variant that takes an explicit runtime instead of constructing one.
func ConfigureScanners ¶ added in v0.10.3
func ConfigureScanners(cfg ScannersConfig) error
ConfigureScanners installs the scanner selection and eagerly constructs the goredact engine when enabled. Call once at process startup after loading settings; an error means the caller must fail the operation — proceeding would run a scanner set the configuration did not choose. Thread-safe; replaces prior config and resets the degradation state.
func IsKnownOPFCategory ¶ added in v0.7.8
IsKnownOPFCategory reports whether name is one of the OPF native labels the CLI knows how to tag and render. Exported so the settings layer can reject typos at parse time — silent zero-detection of a privacy category would leave users thinking they're protected when they're not.
func IsLineDelimited ¶ added in v0.10.3
IsLineDelimited reports whether JSONLContent will redact content line by line rather than as a single JSON value.
Callers that split content and redact the pieces separately MUST check this first. The line path composes -- redact(A+B) == redact(A)+redact(B) for newline-terminated A -- because each line is redacted in isolation. The single-JSON-value path does NOT: it is field-aware across the whole document, so redacting a fragment of it instead falls back to raw regex and entropy detection over partial JSON, which is the identifier corruption redactSingleJSONValue exists to avoid.
A filename is not a safe proxy for this. OpenCode writes a single JSON object to the same full.jsonl path that other agents write JSONL to.
func JSONLContent ¶
JSONLContent parses each line as JSON to determine which string values need redaction, then performs targeted replacements on the raw JSON bytes. Lines with no secrets are returned unchanged, preserving original formatting.
For multi-line JSON content (e.g., pretty-printed single JSON objects like OpenCode export), the function first attempts to parse the entire content as a single JSON value. This ensures field-aware redaction (which skips ID fields) is used instead of falling back to entropy-based detection on raw text lines, which would corrupt high-entropy identifiers.
Large content is sharded across goroutines; output is byte-identical either way. See jsonlContent.
func JSONLContentWithPrivacyFilter ¶ added in v0.7.8
JSONLContentWithPrivacyFilter augments JSONLContent with the OpenAI Privacy Filter via batched inference. Walks the content twice: pass 1 collects unique prose-shaped leaves into a single RedactBatch call; pass 2 applies the eight regex layers per leaf plus the cached OPF spans for that leaf. One OPF shell-out covers the whole transcript instead of one per leaf — without batching, a typical 500-leaf transcript would take many minutes per commit.
Falls back to the plain JSONLContent flow when OPF is unconfigured, the breaker is tripped, no categories are enabled, or the batch call errors.
func OPFBreakerTripped ¶ added in v0.7.8
func OPFBreakerTripped() bool
OPFBreakerTripped reports whether the per-process OPF circuit breaker has been tripped — i.e. an OPF invocation failed at some point during this process's lifetime. The pre-push rewrite uses this to detect when OPF silently fell back to regex-only mid-rewrite and abort before CAS-ing the new ref; otherwise the rewritten commits would carry the Entire-OPF-Applied: true trailer despite containing only regex-only content, and the next push would skip them.
func OPFCommand ¶ added in v0.7.8
func OPFCommand() string
OPFCommand returns the configured OPF binary command, or the default when OPF is unconfigured. Used by error messages so the user sees the exact command they need to fix.
func OPFEnabled ¶ added in v0.7.8
func OPFEnabled() bool
OPFEnabled reports whether the OpenAI Privacy Filter is configured and turned on for this process. Callers gate pre-push rewrite work on this: when false, the pre-push hook pushes the local regex-only checkpoint branch verbatim with no extra processing. Independent of the circuit breaker — a tripped breaker still reports Enabled=true because the runtime config didn't change; the rewrite logic itself handles the breaker by short-circuiting per-commit OPF calls.
func OPFMisconfiguredNoCategories ¶ added in v0.10.1
func OPFMisconfiguredNoCategories() bool
OPFMisconfiguredNoCategories reports whether OPF is enabled for this process with an empty effective category set — the state described by ErrOPFNoEnabledCategories. The pre-push flow checks this twice: when resolving the run/skip decision (so the user is never prompted to run a scan that cannot run) and again at the start of the rewrite (so direct callers fail closed too). Explicit opt-outs (ENTIRE_OPF=no, prompt_default: never) still resolve to skip and push regex-only content without the trailer.
func ResetOPFConfigForTest ¶ added in v0.7.8
func ResetOPFConfigForTest()
ResetOPFConfigForTest clears OPF configuration and the circuit breaker. Test-only.
func SetScannerDegradedForTest ¶ added in v0.10.3
func SetScannerDegradedForTest(v bool)
SetScannerDegradedForTest flips the scanner degradation flag. Runtime scan errors are engineered out (see detectGoredact), so tests cannot reach this state organically. Exported for redact's own sentinel tests and for dependent packages' tests of the checkpoint write paths that consume ErrScannerDegraded. Test-only.
func String ¶
String replaces secrets and PII in s using layered detection:
- Entropy-based: high-entropy alphanumeric sequences (threshold 4.5)
- Pattern-based: scanner engines selected via ConfigureScanners — betterleaks regex rules (several hundred known secret formats) and/or the goredact engine; betterleaks-only when unconfigured
- Provider token prefixes: deterministic prefix rules for credential formats betterleaks misses in isolation (e.g. Supabase sb_secret_)
- Credentialed URIs: URLs containing userinfo passwords
- Database connection strings: JDBC, keyword DSNs, and semicolon strings
- User-defined custom rules: configured via ConfigureCustomRules
- Bounded credential key/value pairs: DB_PASSWORD=...
- PII detection: email, phone, address patterns (only when configured via ConfigurePII)
A string is redacted if ANY method flags it.
func StringWithPrivacyFilter ¶ added in v0.7.8
StringWithPrivacyFilter augments String with the OpenAI Privacy Filter. Use only at condensation/export boundaries; per-turn writes must use String to avoid the OPF shell-out cost inside the agent loop.
func SumProseLeafBytes ¶ added in v0.7.8
SumProseLeafBytes returns the cumulative byte size of prose-shaped (has-space) leaves across inputs — the upper bound on what BatchBytesWithPrivacyFilter would send to OPF inference.
Callers use this to enforce a cap before paying the OPF cost: a push with 100MB of mostly-structural JSON has tens of KB of actual leaves; a push with 100MB of dense prose has hundreds of MB. The blob-byte size doesn't tell you which without looking inside.
Returns a CONSERVATIVE UPPER BOUND on what would go to OPF — same has-space gate and JSONL/JSON parse with whole-content fallback as the collector inside BatchBytesWithPrivacyFilter, BUT this function does NOT deduplicate identical leaves across blobs. The actual batch sent to OPF dedups by leaf-text, so a push with many repeated leaves will report higher byte counts here than OPF actually sees. Callers using this for cap enforcement get an over-strict bound, which is safe (false positives possible, false negatives impossible).
func WithScannerDegradedSole ¶ added in v0.10.3
func WithScannerDegradedSole(t testingTB)
WithScannerDegradedSole configures a goredact-only scanner set and flips the degradation flag so JSONLBytes/JSONLBytesWithPrivacyFilter return ErrScannerDegraded, registering a cleanup that restores the betterleaks default. Named With (not Set) per Go convention because it registers cleanup. Callers must NOT use t.Parallel — the scanner state is process-global. Test-only.
Types ¶
type CustomRulesConfig ¶ added in v0.6.2
type CustomRulesConfig struct {
// Inline maps a label (used only in logs/diagnostics) to a Go RE2 regex
// string. Failed compilations are logged to Logger and dropped.
Inline map[string]string
// Packs are pre-parsed rule packs (see LoadPacks). Per-rule regex
// compilation failures are logged and dropped; sample mismatches are
// logged but do not drop the rule.
Packs []*Pack
// Logger receives this package's diagnostics (rule compile errors,
// sample mismatches). The CLI injects the logger its entry point
// initialized so warnings reach .entire/logs/ instead of the
// process-default stderr logger, which hook contexts swallow.
// nil means slog.Default(), preserving standalone behavior for
// library consumers and tests.
Logger *slog.Logger
}
CustomRulesConfig configures inline custom_redactions and parsed rule packs.
type NamedBlob ¶ added in v0.7.8
NamedBlob is one input to BatchBytesWithPrivacyFilter. Name drives redaction shape: a ".jsonl" or ".json" suffix triggers JSON-aware leaf extraction (string values inside the parsed structure); any other suffix treats the whole content as a single leaf.
Content is the raw blob bytes. The blob's redacted output appears at the same index in the function's return slice.
type OPFConfig ¶ added in v0.7.8
type OPFConfig struct {
Enabled bool
Categories map[string]bool
Command string // path or name of the opf binary; "" defaults to "opf"
Timeout int // seconds; 0 defaults to 30
// Logger receives OPF runtime-failure diagnostics. nil means
// slog.Default(); the CLI injects its entry-point-initialized logger
// so warnings land in .entire/logs/.
Logger *slog.Logger
// contains filtered or unexported fields
}
OPFConfig configures the optional OpenAI Privacy Filter detection layer. Defaults are applied by ConfigurePrivacyFilter; callers should pass values straight from settings without local normalization.
type PIICategory ¶ added in v0.5.1
type PIICategory string
PIICategory identifies a category of personally identifiable information.
const ( PIIEmail PIICategory = "email" PIIPhone PIICategory = "phone" PIIAddress PIICategory = "address" )
type PIIConfig ¶ added in v0.5.1
type PIIConfig struct {
// Enabled globally enables/disables PII redaction.
// When false, no PII patterns are checked (secrets still redacted).
Enabled bool
// Categories maps each PII category to whether it is enabled.
// Missing keys default to false (disabled).
Categories map[PIICategory]bool
// CustomPatterns allows teams to define additional regex patterns.
// Each key is a label used in the replacement token (uppercased),
// and each value is a regex pattern string.
// Example: {"employee_id": `EMP-\d{6}`} produces [REDACTED_EMPLOYEE_ID].
CustomPatterns map[string]string
// Logger receives diagnostics (invalid CustomPatterns regexes). nil
// means slog.Default(); the CLI injects its entry-point-initialized
// logger so warnings land in .entire/logs/.
Logger *slog.Logger
// contains filtered or unexported fields
}
PIIConfig controls which PII categories are detected and redacted.
type Pack ¶ added in v0.6.2
type Pack struct {
Name string `json:"name" yaml:"name"`
Version string `json:"version" yaml:"version"`
Description string `json:"description,omitempty" yaml:"description"`
Rules []Rule `json:"rules" yaml:"rules"`
// contains filtered or unexported fields
}
Pack is a versioned bundle of redaction rules loaded from a single file under .entire/redactors/. Both YAML and JSON encodings are accepted; the schema is identical.
func LoadPacks ¶ added in v0.6.2
LoadPacks discovers and parses all rule packs in dir within fsys, including any subdirectories (so the conventional .entire/redactors/local/ path for personal/uncommitted rules is picked up automatically). Files with the extensions .yaml, .yml, and .json are considered packs; other files are ignored. A missing directory is treated as "no packs configured" and returns no error. Per-file parse errors are warned to logger (nil means slog.Default()) and the file is skipped — never fatal — so one bad file does not silence the rest.
Taking an fs.FS rather than a directory path is what lets the CLI hand this the .entire root's FS, so pack discovery is confined to .entire like every other read there, while this package keeps no dependency on the CLI. Names in warnings and in a pack's SourcePath are therefore relative to fsys (redactors/local/foo.yaml), not absolute.
Soft caps: files larger than maxPackFileBytes are skipped with a warning, and discovery stops after maxPackFiles parsed packs. The trust boundary is "user owns repo," so these are runaway-input guards, not security limits.
func ParsePack ¶ added in v0.6.2
ParsePack decodes a single pack file. sourcePath is used both to pick the encoding (YAML by default; JSON only when the extension is .json) and to enforce that the pack's `name` matches the filename stem.
Precondition: sourcePath must be a vetted local file path (the production caller is LoadPacks, which only invokes ParsePack with paths produced by WalkDir under the configured .entire/redactors/ directory). Callers passing arbitrary or remote paths must enforce their own trust model — ParsePack does not sanitize sourcePath beyond reading its extension.
type RedactedBytes ¶ added in v0.5.5
type RedactedBytes struct {
// contains filtered or unexported fields
}
RedactedBytes represents transcript data that has been through secret redaction. Consumers that require pre-redacted input (e.g., compact.Compact, checkpoint stores) accept this type to enforce the contract at compile time.
Produced by JSONLBytes (primary constructor) or trusted wrappers for data previously persisted by checkpoint writers.
func AlreadyRedacted ¶ added in v0.5.5
func AlreadyRedacted(data []byte) RedactedBytes
AlreadyRedacted wraps transcript bytes known to already be redacted by a prior write path. Use this ONLY for trusted sources such as persisted checkpoint transcripts or controlled test fixtures. For fresh transcript input, use JSONLBytes.
func JSONLBytes ¶
func JSONLBytes(b []byte) (RedactedBytes, error)
JSONLBytes redacts secrets in JSONL-formatted byte content and returns the result as RedactedBytes, certifying the output has been through redaction. Returns ErrScannerDegraded when the goredact scanner degraded while betterleaks is disabled.
func JSONLBytesWithPrivacyFilter ¶ added in v0.7.8
func JSONLBytesWithPrivacyFilter(ctx context.Context, b []byte) (RedactedBytes, error)
JSONLBytesWithPrivacyFilter augments JSONLBytes with the OpenAI Privacy Filter. Use only at condensation/export boundaries; per-turn writes must use JSONLBytes. Returns ErrScannerDegraded when the goredact scanner degraded while betterleaks is disabled.
func (RedactedBytes) Bytes ¶ added in v0.5.5
func (r RedactedBytes) Bytes() []byte
Bytes returns the underlying byte slice.
func (RedactedBytes) Len ¶ added in v0.5.5
func (r RedactedBytes) Len() int
Len returns the number of bytes in the redacted payload.
type Rule ¶ added in v0.6.2
type Rule struct {
ID string `json:"id" yaml:"id"`
Description string `json:"description,omitempty" yaml:"description"`
Regex string `json:"regex" yaml:"regex"`
Samples []Sample `json:"samples,omitempty" yaml:"samples"`
}
Rule is a single redaction rule within a Pack.
type Sample ¶ added in v0.6.2
type Sample struct {
Input string `json:"input" yaml:"input"`
Redacted bool `json:"redacted" yaml:"redacted"`
}
Sample is a self-test entry for a Rule. The runner asserts whether the rule's regex matching `Input` matches the `Redacted` expectation.
type ScannersConfig ¶ added in v0.10.3
ScannersConfig selects which scanner engines run in detectAllLayers.
type Span ¶ added in v0.7.8
Span is a redaction region returned by an opfRuntime, with BYTE-offset boundaries against the input text. OPF itself reports character (rune) offsets via its JSON output; the shell-out adapter translates those to byte offsets before returning Spans so callers can slice []byte input directly without re-walking runes.