Documentation
¶
Overview ¶
Package sarif provides Draugr's result currency: a pragmatic model of SARIF 2.1.0 findings, plus merge and deduplication. Every scanner normalizes its output to a Report; the engine merges reports and the result can be serialized to standard SARIF JSON for GitHub / Azure DevOps / GitLab.
Index ¶
- Constants
- Variables
- func AgreementNote(others []Observation, counted Severity) string
- func DerivedHelpURI(ruleID string) string
- func LineHash(lines []string, startLine int) string
- func ProvesAbsence(method string) bool
- func SameRepository(a, b string) bool
- type CallFrame
- type CallPath
- type Consulted
- type Correlation
- type Counts
- type Escalation
- type Field
- type Layer
- type Level
- type Location
- type MarshalOptions
- type Match
- type Observation
- type Package
- type Provenance
- type Reachability
- type ReachabilityState
- type Remediation
- type Report
- func (r Report) Counts() Counts
- func (r Report) Dedup() Report
- func (r Report) HelpURI(ruleID string) string
- func (r Report) Highest() Level
- func (r Report) HighestSeverity() Severity
- func (r Report) MarshalSARIF() ([]byte, error)
- func (r Report) MarshalSARIFWith(opts MarshalOptions) ([]byte, error)
- type RepositoryRef
- type Result
- func (r Result) Correlated() bool
- func (r Result) Fingerprint() string
- func (r Result) Imported() bool
- func (r Result) Remediation() Remediation
- func (r Result) Severity(floor Severity) Severity
- func (r Result) SilencedInSource() bool
- func (r *Result) StampLineHash(lines []string)
- func (r Result) Suppressed() bool
- func (r Result) VulnerabilityID() string
- type Rule
- type Severity
- type Suppression
- type Taxon
Constants ¶
const ( // OriginSaga is a rule in this project's own descriptor. The default reading of an empty // Origin, so a report written before imported claims existed still means what it said. OriginSaga = "saga" // OriginVEX is a statement imported from a document somebody else wrote. OriginVEX = "vex" // OriginTool is a suppression the author wrote into the source and the scanner honored, a Semgrep // `nosem`, a `# noqa`, a linter's inline pragma. // // The weakest of the three, and kept apart for that reason. A descriptor rule was reviewed by // whoever owns the descriptor and a supplier's claim is answerable by the supplier; this one // was written by whoever was editing the file, possibly to get a build green, and nothing // about it went past a second person. Counting it with the others would let the weakest form // of acceptance hide inside the strongest. OriginTool = "tool" )
Where a suppression's decision came from, for Suppression.Origin.
const ( // MethodCallGraph followed calls from an entry point to the vulnerable symbol. Bounded by what // the language lets a static analysis see, which every such tool publishes and none of which // is nothing. MethodCallGraph = "call-graph" // MethodDataFlow followed the data as well as the calls, so it can find a path unreachable // where a call graph finds one. MethodDataFlow = "data-flow" // MethodFrameworkHeuristic read a framework's own conventions about which handlers are wired. // A statement about configuration rather than about code, and a guess where the two disagree. MethodFrameworkHeuristic = "framework-heuristic" // MethodImportCheck saw whether the vulnerable package is referenced anywhere in the tree, // which is presence with extra steps. MethodImportCheck = "import-check" )
Methods an analyzer may report, strongest evidence first.
Four things are sold as reachability and they are not the same claim. The difference is entirely about the negative: saying a flaw is reachable is a positive claim somebody can go and check, and saying it is unreachable is an absence claim worth exactly what the analysis cannot see.
const LineHashKey = "primaryLocationLineHash/v1"
LineHashKey is the partial-fingerprint name for a finding's content hash.
GitHub code scanning's own name, deliberately. It is what GitHub reads to decide that an alert in this run is the same alert as one in the last, so emitting it under any other name would mean computing the right thing and having nobody consume it. Versioned by GitHub's convention: the suffix changes if the algorithm does, and consumers compare like with like.
const Version = "2.1.0"
Version is the SARIF specification version Draugr emits.
Variables ¶
var Severities = []string{ string(SeverityCritical), string(SeverityHigh), string(SeverityMedium), string(SeverityLow), }
Severities lists the bands a gate may be set to, most to least severe.
Functions ¶
func AgreementNote ¶ added in v0.121.1
func AgreementNote(others []Observation, counted Severity) string
AgreementNote is the line under a finding saying which other scanners found it, and where they disagree about how bad it is.
Said rather than hidden, because two tools agreeing is itself a signal, and because a reader who enabled a second scanner should be able to see it working, without this the row looks exactly like a run with one scanner and the second appears to have found nothing.
Here rather than beside one renderer because the scan report and the diff both draw it, and a reader who has learned the line in a terminal meets the same words in a pull-request comment.
A rating is shown only when it differs from the one being counted. Where the scanners agree, repeating the same numbers on every row is noise a reader has to look past; where they disagree, it is the one thing this line is carrying that they could not get anywhere else. The full record is in the JSON and the SARIF either way.
func DerivedHelpURI ¶ added in v0.35.0
DerivedHelpURI maps identifiers with a stable, publicly resolvable home to their advisory. It's the fallback for scanners that publish no rule metadata; anything unrecognized gets "".
func LineHash ¶ added in v0.104.0
LineHash is a content fingerprint for a finding at a line.
The point is identity across runs, which the ordinary Fingerprint deliberately does not provide: that one hashes the line *number*, so adding an import at the top of a file makes every finding below it look new. This hashes what the code says rather than where it sits, so a finding survives edits elsewhere in the file.
Normalized per line, leading and trailing whitespace removed. So reformatting and reindenting do not churn it either. Returns "" when there is nothing to hash, which is honest: absent means "no content-based identity", and a fabricated one would be worse than none.
**What it does not survive: an edit inside the context window.** Nearby lines are part of the identity, so inserting a line immediately above a finding changes it. That is inherent rather than a shortcoming to fix, without context, every bare `}` in a repository would share one fingerprint. And it is the same trade CodeQL makes. The case that matters is an edit elsewhere in the file, which is the common one and the one that used to invalidate everything below it.
A finding on the first lines of a file is more exposed to this, because there is nothing above it to include and anything inserted there lands inside the window.
func ProvesAbsence ¶ added in v0.123.0
ProvesAbsence reports whether a method's unreachable verdict is strong enough to lower a band.
The one place this is decided. A call graph that finds no path has looked for one; a framework heuristic that finds no route has read a configuration file. Only the first is evidence that the vulnerable code does not run, and a band lowered by the second is a finding somebody stops looking at on the strength of a guess.
An unrecognized or absent method is not strong. That is the conservative direction and the right default for a field an analyzer wrote: a tool that will not say how it decided has not earned a de-escalation, and one added later gets the strict reading until somebody classifies it here.
func SameRepository ¶ added in v0.85.0
SameRepository reports whether two repository references name the same repository.
Tolerant on purpose: a descriptor may say `https://host/org/repo.git`, `git@host:org/repo.git` or `.`, and a CI environment says `org/repo`. Comparing them literally would answer "different" for the same repository and quietly drop its findings, which is worse than the problem this is used to solve.
Each reference is reduced to its path, case-insensitively, without scheme, credentials, port or `.git`. Two match when they are equal, or when the shorter is a trailing run of whole segments of the longer, so a CI variable saying `org/repo` still matches a descriptor's clone URL.
The whole path, not its last two segments. A forge may nest groups arbitrarily, and keeping only the tail makes `payments/backend/api` and `platform/backend/api` the same repository. Nothing errors when that happens: the two sets of findings merge, and whichever is processed last decides what the report says about a file both of them contain.
Two segments is the floor for a suffix match. A bare `api` does not say which `api`, and accepting it would re-open the same collapse from the other end.
Types ¶
type CallFrame ¶ added in v0.101.0
type CallFrame struct {
// Function is the function called, qualified as its language qualifies it.
Function string `json:"function"`
// Package is the package the function belongs to, when the analyzer reports one.
Package string `json:"package,omitempty"`
// Module is the module the package belongs to, when the analyzer reports one.
Module string `json:"module,omitempty"`
// File and Line locate the call. Relative to the module root, as the analyzer reported it.
File string `json:"file,omitempty"`
Line int `json:"line,omitempty"`
}
CallFrame is one call in a CallPath.
type CallPath ¶ added in v0.101.0
type CallPath struct {
// Frames are the calls making up the route, starting in this project's own code.
Frames []CallFrame `json:"frames"`
}
CallPath is one route from this project's code to a vulnerable symbol, ordered caller first so it reads the way a stack trace is read.
type Consulted ¶ added in v0.107.0
type Consulted struct {
// Signal names the dataset: "kev" or "epss".
Signal string `json:"signal"`
// AsOf is the day the copy was obtained, as YYYY-MM-DD. Empty when the caller supplied a
// file with no fetch to record, which is itself worth seeing.
AsOf string `json:"asOf,omitempty"`
// Entries is how many records the dataset held. A feed that loaded and turned out to be
// empty answers every lookup with "not listed", which is indistinguishable from a working
// one unless somebody can see the count.
Entries int `json:"entries,omitempty"`
// Threshold is the EPSS probability at or above which a finding was raised. Zero for KEV, which
// has no threshold. Being on it is the whole signal.
Threshold float64 `json:"threshold,omitempty"`
// ThresholdFrom names what set that threshold: the default, a descriptor key, or a flag.
//
// A number without its source is the one part of an escalation a reader cannot check. Every other
// input says where it came from, the dataset names itself and the day its copy was obtained,
// while the line a score was measured against arrives anonymous, and the answer to "who decided
// 0.5" decides whether a band is a policy or an accident. It is also the value most worth
// governing: a threshold reachable by the team being gated is a gate that can be loosened without
// a record.
ThresholdFrom string `json:"thresholdFrom,omitempty"`
}
Consulted is one exploitability dataset a run had available.
Deliberately not the full provenance: where the copy came from and what its checksum was belong to the human-facing report, which has room for them. What a consumer of the evidence needs to explain a band is narrower. Which signals could have fired, and as of when.
type Correlation ¶ added in v0.103.0
type Correlation struct {
// AlsoFoundBy is what the other scanners said about this same flaw, ordered by tool name.
// Set on the finding that is counted, and empty on the ones that are not.
AlsoFoundBy []Observation `json:"alsoFoundBy,omitempty"`
// CountedUnder names the scanner whose finding this one is counted under, and is what makes
// this finding evidence rather than a count. Empty on the finding that is counted.
CountedUnder string `json:"countedUnder,omitempty"`
}
Correlation records that more than one scanner reported the same flaw.
A relationship rather than a merge, and for the reason the rest of this model is: nothing is deleted. Both scanners' findings stay in the report with their own rule ids, their own severity and their own account of what they saw, because the disagreement between two scanners is the reason to run two, and a merge that keeps one opinion throws away the thing you were paying for.
What changes is the counting. One finding of the group is counted; the others are evidence. Reporting four vulnerabilities as eight is the failure this exists to fix, and it is the test buyers are told to run: point several scanners at one target and count the tickets.
Deliberately not part of Fingerprint. Which finding of a group happens to be counted is a fact about this run's scanner selection, not about the finding, and folding it in would make a diff churn on the day somebody adds a second scanner.
type Escalation ¶ added in v0.56.0
type Escalation struct {
// From is the severity before enrichment, the scanner's own rating, and the one the report still
// displays, because it is what the scanner actually said.
From Severity `json:"from"`
// To is the severity the finding was *ranked* as. Enrichment feeds the priority matrix
// rather than rewriting what the scanner reported, so this is the value the band was
// computed from and the one that explains a P1 sitting on a "high" row.
To Severity `json:"to"`
// Signal names the dataset that fired: "kev" or "epss".
Signal string `json:"signal"`
// Detail is the specific fact, e.g. "on KEV" or "EPSS 0.87".
Detail string `json:"detail"`
// AsOf is the day the data was fetched, as YYYY-MM-DD. Without it the claim is "KEV said
// so", which is not something a reader can check or reproduce.
AsOf string `json:"asOf,omitempty"`
// AlsoMatched are the datasets that applied to this finding without being the one that set
// its rating.
//
// Only one signal can raise a severity, so without this the others leave no trace at all, and the
// one that loses is always the same one. KEV outranks EPSS wherever both fire, so anything
// counting how often a dataset reached a finding reads EPSS as having done less than it did, by
// an amount nothing in the record reveals.
AlsoMatched []Match `json:"alsoMatched,omitempty"`
}
Escalation records that exploitability data raised a finding's severity, and on what grounds.
Deliberately not part of Fingerprint: a finding is the same finding whether or not a feed moved it, and folding this in would make every diff churn on the day EPSS reprices a CVE.
type Field ¶ added in v0.51.0
Field is one statement in a Provenance entry.
A slice of pairs rather than a map: rendering needs a stable order, and alphabetical is the wrong one. It puts "coverage" before "benchmark". The scanner knows which matters most to a reader, so it decides.
type Layer ¶ added in v0.96.0
type Layer struct {
// DiffID is the layer's content digest, as the image records it.
DiffID string `json:"diffId,omitempty"`
// Index is the layer's position from the bottom of the image, starting at 0.
Index int `json:"index"`
// Of is the number of layers in the image.
Of int `json:"of,omitempty"`
// CreatedBy is the build instruction that produced the layer, verbatim from the image's own
// history. "RUN /bin/sh -c apt-get install …". It names the line to change, which is more use
// than a layer digest and more honest than a guessed base image.
CreatedBy string `json:"createdBy,omitempty"`
}
Layer identifies the image layer a finding's package came from, and the build step that made it.
Index is the layer's position in the image, counting from the bottom, and Of is how many layers there are. "3 of 9" is interpretable where a bare digest is not, and the low indices are the inherited ones.
type Level ¶
type Level string
Level is the severity of a result, mirroring SARIF's result.level.
const ( LevelError Level = "error" LevelWarning Level = "warning" LevelNote Level = "note" LevelNone Level = "none" )
The SARIF result levels.
func ParseLevel ¶ added in v0.64.0
ParseLevel converts a user-supplied gate level, rejecting anything it does not recognize.
Rejecting matters more than it looks. An unknown level ranks 0, and every finding is at least 0, so a typo, or a plausible-sounding value like "high", silently turns a gate into "fail on anything at all" rather than failing to parse. A flag either does something or says why not.
type Location ¶
type Location struct {
URI string `json:"uri,omitempty"`
StartLine int `json:"startLine,omitempty"`
}
Location points at where a finding was observed.
type MarshalOptions ¶ added in v0.35.0
type MarshalOptions struct {
// Compact drops what only a human reads, indentation, and the rule prose relayed from the
// scanner, while keeping the report valid SARIF.
//
// It exists for a consumer that is going to *act* on the report rather than read it,
// typically an agent paying for every byte of context. Rule descriptions and remediation
// text are the bulk of a Draugr report (61% of it, measured on this repo), and a reader
// that can follow a link doesn't need them inlined. So helpUri survives compaction and the
// prose doesn't: keep the pointer, drop the paragraphs.
Compact bool
// AutomationID is written to runs[].automationDetails.id, which is how a consumer tells two
// analyses of one commit apart. Empty writes no automationDetails at all, which is what a
// report carried before this existed and what a scan with nothing to name still produces.
//
// See report.AutomationID for what Draugr puts in it, and for what GitHub code scanning does
// to an upload that carries none.
AutomationID string
// ToolVersion is the Draugr that produced the run, written to the run's tool driver. Empty
// writes no version, which is what a report built outside a Draugr run carries.
ToolVersion string
}
MarshalOptions tunes how a report is serialized. The zero value is the default: indented, with everything a person or an editor might want.
type Match ¶ added in v0.115.0
type Match struct {
// Signal names the dataset: "kev" or "epss".
Signal string `json:"signal"`
// Detail is the specific fact, e.g. "on KEV" or "EPSS 0.87".
Detail string `json:"detail"`
// AsOf is the day the data was fetched, as YYYY-MM-DD.
AsOf string `json:"asOf,omitempty"`
}
Match is a dataset that applied to a finding.
The same three facts an Escalation carries about the dataset that won, minus the severities: a match that did not set the rating moved nothing, and naming a from and a to for it would describe a change that never happened.
type Observation ¶ added in v0.103.0
type Observation struct {
// Tool is the scanner that made this observation.
Tool string `json:"tool"`
// RuleID is what that scanner called it, which is not always what the counted finding is called.
// And is the id an exclusion may already be written against.
RuleID string `json:"ruleId,omitempty"`
// Severity and Score are that scanner's own rating, kept whether or not it agrees. The
// report decides what is worth showing; the record does not get to be selective.
Severity Severity `json:"severity,omitempty"`
Score float64 `json:"score,omitempty"`
}
Observation is another scanner's account of a flaw already counted.
Complete rather than a name, because the interesting case is disagreement. Two scanners rating the same CVE medium and low have said something neither says alone. They draw on different advisory sources, and the gap is a statement about coverage. A reader who only learns that a second tool "also found it" cannot see that, and the console has nothing to show them.
type Package ¶ added in v0.90.0
type Package struct {
// Name is the package as its ecosystem names it, e.g. "flask".
Name string `json:"name"`
// Version is what is installed.
Version string `json:"version,omitempty"`
// FixedVersion is the first release that resolves the finding. Empty when there is none. Which is
// a different and more alarming answer than "unknown", and the reason this is reported rather
// than inferred from the absence of a fix.
FixedVersion string `json:"fixedVersion,omitempty"`
// PURL is the package URL, e.g. "pkg:pypi/flask@0.12.2". The one identifier that is portable
// across ecosystems and the one every consumer of this asked for first.
PURL string `json:"purl,omitempty"`
// Ecosystem is the package manager the name belongs to, as the ecosystem calls itself: "pip",
// "npm", "gem". A name alone is ambiguous. There is a `request` on npm and a `requests` on PyPI,
// and neither is the other.
Ecosystem string `json:"ecosystem,omitempty"`
}
Package is the dependency a finding is about.
Scanners have always known this and only ever said it in prose. Trivy's message reads "Package: flask\nFixed Version: 0.12.3", which is a fact formatted for a human and unavailable to anything else. Three things wanted it and could not have it: GitLab's dependency and container reports, which require a structured name and version; correlating a run's findings with the SBOM it produced alongside them; and a VEX statement that can say which package within a product carries a vulnerability rather than only that the product does.
Deliberately not part of Fingerprint. The same flaw in the same package at the same location is the same finding whether or not the scanner told us which package. So adding this must not split a finding in two the day a scanner starts reporting it.
type Provenance ¶ added in v0.51.0
type Provenance struct {
// Tool is the scanner that produced this account.
Tool string `json:"tool"`
// Version is the scanner's version as the engine resolved it. The same value that goes into its
// cache key, so the evidence and the cache cannot disagree about what ran. Empty when the scanner
// does not report one.
Version string `json:"version,omitempty"`
// Fields are the scanner's own statements about the run, in the order it considers useful.
//
// Untyped, because the interesting ones are domain knowledge, "benchmark", "coverage", "scope".
// And this package is the finding currency for every scanner Draugr will ever have. It should not
// learn what a CIS benchmark is to carry the fact that one was applied.
Fields []Field `json:"fields,omitempty"`
}
Provenance is one scanner's account of a run it performed.
A slice on Report rather than a map of fields, because a control can be served by more than one scanner and each has its own answer, two scanners auditing a cluster apply two different benchmarks. Flattening them into one map keeps whichever was written last, silently, which is the failure this type exists to prevent.
func (Provenance) Describe ¶ added in v0.51.0
func (p Provenance) Describe() string
Describe returns the fields as "key: value" pairs, for a reporter with one line to spend.
Punctuated, because a value is not always a phrase. "benchmark CIS 1.9" reads whichever way it is written, and "coverage this repository has no go.mod, so its findings carry no verdict" reads as a sentence that starts with a word left over from somewhere else. The colon is what says the first word names the rest rather than beginning it.
func (Provenance) Repository ¶ added in v0.64.0
func (p Provenance) Repository() (RepositoryRef, bool)
Repository extracts the repository this provenance entry describes, if it describes one.
type Reachability ¶ added in v0.101.0
type Reachability struct {
// State is the verdict: reachable, unreachable, or unknown.
State ReachabilityState `json:"state"`
// Analyzer names the tool that decided, e.g. "govulncheck".
Analyzer string `json:"analyzer"`
// Method is how it decided, one of the Method constants below. Buyers are told to reject
// reachability claims that do not say how they were reached, and they are right to: a call
// graph and a framework heuristic are both called reachability and are not the same evidence.
// ProvesAbsence is where that difference is acted on.
Method string `json:"method,omitempty"`
// Symbols are the vulnerable functions the advisory names, whether or not they are called.
// Reported for both reachable and unreachable, because "which functions would have to be
// called" is what makes an unreachable verdict checkable by hand.
Symbols []string `json:"symbols,omitempty"`
// Paths are the routes found from this project's code to the vulnerable code, each ordered
// caller first. Empty unless State is reachable.
Paths []CallPath `json:"paths,omitempty"`
// RankedAs is the severity the priority band was computed from, set only when the state
// moved it. The report still displays the scanner's own severity; this explains a P3 sitting
// on a "high" row, the way Escalation.To explains a P1 on a "medium" one.
RankedAs Severity `json:"rankedAs,omitempty"`
// AsOf is the day the analysis ran, as YYYY-MM-DD. A reachability verdict describes one
// revision of the code and stops being true when the code changes, so a claim without a date
// is not one a reader can check.
AsOf string `json:"asOf,omitempty"`
}
Reachability records whether this project's code can reach a dependency's vulnerable code, and on what evidence.
The third member of a family, and it inherits both of the family's rules. Suppression records that a person decided to count a finding for less; Escalation records evidence that it deserves more; this records evidence about whether the code can be reached at all. Like Escalation it feeds the priority matrix and never rewrites the severity the scanner reported, and like Escalation it is deliberately not part of Fingerprint. A finding is the same finding whether or not analysis moved it, and folding this in would churn every diff on the day a call is added or removed.
Unlike Suppression it never excuses a finding. An inference is not a decision, and a finding that disappears because a call graph did not find a path has no author to ask about it.
func (*Reachability) RankAt ¶ added in v0.101.0
func (r *Reachability) RankAt(base Severity) Severity
RankAt returns the severity a finding's priority band should be computed from, given what reachability analysis concluded.
Only an unreachable verdict moves anything, and only from a method that proves absence. Reachable does not raise: severity already assumes the vulnerable code runs, so treating a confirmed call as an escalation would count the same assumption twice. Unknown moves nothing by definition. It is the absence of a finding, not one.
A weak method's unreachable verdict still travels and is still shown. What it does not do is lower the band, because the reader deciding whether to act on it needs to make that call themselves rather than have it made by a guess.
type ReachabilityState ¶ added in v0.101.0
type ReachabilityState string
ReachabilityState says whether this project's own code can reach a dependency's vulnerable code. Three values rather than two, because the third is the one that keeps the other two honest.
const ( // ReachabilityReachable means analysis found a route from this project's code to the // vulnerable code, and the route is recorded in Paths. ReachabilityReachable ReachabilityState = "reachable" // ReachabilityUnreachable means analysis covered this dependency and found no such route. // // It is a statement about how the code is called today, not a claim that the flaw is // harmless: reflection, dynamic dispatch and code generation all defeat a call graph, and // the route appears the day somebody writes the call. That is why this ranks a finding // lower and never removes it. ReachabilityUnreachable ReachabilityState = "unreachable" // ReachabilityUnknown means no analysis covered this dependency. // // The whole reason the state is explicit. An analyzer that never looked at a dependency produces // exactly the same silence as one that looked and found nothing, and treating the two alike turns // "we did not check" into "you are fine". Which is the failure this codebase refuses everywhere // else. A tool that cannot say which of the two it means may only report this. ReachabilityUnknown ReachabilityState = "unknown" )
type Remediation ¶ added in v0.96.0
type Remediation string
Remediation says who can resolve a finding and by what kind of action. It is the answer to "can I do something about this", which is the question a reader has after the severity.
const ( // RemediationUpgrade: a version that fixes it exists, in something the reader controls. RemediationUpgrade Remediation = "upgrade" // RemediationUpstream: nothing fixes it where it is, but the thing underneath can move, an // operating system release past end of service life, whose successor is the fix. One action, and // it resolves everything in that layer at once. RemediationUpstream Remediation = "upstream" // RemediationExternal: the surface is operated by somebody else. Still found, still reported, // still counted, never presented as something to go and fix, because telling a reader to change a // file on a control plane they cannot reach is worse than saying nothing. RemediationExternal Remediation = "external" // RemediationNone: no fix is published anywhere, and the thing it is in is the reader's. // Mitigation or acceptance, rather than an upgrade. RemediationNone Remediation = "none" )
The kinds of remediation, most actionable first.
type Report ¶
type Report struct {
Tool string `json:"tool,omitempty"`
Results []Result `json:"results"`
// Rules is the metadata the scanner published about the rules it applied, keyed by rule id. A
// result names its rule; the rule is what explains it. Carrying this through is what lets a
// reader, in a terminal, an editor, or a pull request. Find out what "DS-0002" means. Not every
// scanner publishes it, so entries may be missing.
Rules map[string]Rule `json:"rules,omitempty"`
// Provenance is what the scanners said about the run itself, as opposed to what they found:
// which standard was applied, how much of it could be decided, what the scan was scoped to.
//
// A finding answers "what is wrong". Evidence also has to answer "what was measured, and against
// what". And that was not recorded anywhere, so a compliance report could not say which benchmark
// produced it. One entry per tool run; see Provenance.
Provenance []Provenance `json:"provenance,omitempty"`
// Decided are the classifications this run reached a verdict on, whether or not a finding
// resulted. A taxon here with no finding means the scanner looked and found nothing wrong.
//
// The distinction this exists for: a scanner that reports nothing about a control has either
// examined it and been satisfied, or never examined it at all, and those mean opposite things.
// Without this, a report that says "1 of 2 scanners found it" is guessing that the other
// dissented, when far more often the other simply does not check that control.
//
// "Decided" rather than "examined" on purpose: a check a scanner looked at and could not settle,
// a CIS control that requires human judgement. Is not a dissent either.
Decided []Taxon `json:"decided,omitempty"`
// Consulted names the exploitability datasets this run had loaded, whether or not any of
// them moved a finding.
//
// The same argument as Decided, one level up. An Escalation is written only when a signal raised
// a severity, so its absence on a finding covers three unrelated cases: the CVE is not in the
// dataset, it is but the finding was already at the top band, and the dataset was never loaded. A
// consumer explaining a priority to somebody cannot tell "not on KEV" from "KEV was not
// consulted" without this. And silence reads as the second, which makes the whole ranking look
// like it came from nowhere.
Consulted []Consulted `json:"consulted,omitempty"`
}
Report is a set of findings, normalized to SARIF semantics. Tool names the primary scanner; when a report carries results from several tools (after Merge), each Result keeps its own Tool.
func FromSARIF ¶
FromSARIF parses standard SARIF 2.1.0 JSON into a Report, flattening all runs and setting each result's Tool from its run's driver name.
func Merge ¶
Merge combines reports into one, deduplicating results by fingerprint and preserving first-seen order. Each result's Tool is backfilled from its source report when unset.
func (Report) Dedup ¶
Dedup returns a copy with exact-duplicate results removed, preserving first-seen order.
func (Report) HelpURI ¶ added in v0.35.0
HelpURI returns where a reader can look up ruleID: what the scanner published, or a URL derived from a well-known identifier scheme when it published nothing. Empty when we can't say. A wrong link is worse than none.
func (Report) Highest ¶
Highest returns the most severe level present, or LevelNone when there are no results.
func (Report) HighestSeverity ¶ added in v0.97.0
HighestSeverity returns the most severe band present, or an empty severity when there are no results to judge.
This is what the gate compares, and it is deliberately the same number the report prints. The SARIF level and the severity band are two different ladders: a finding carrying a CVSS score takes its band from the score, so a scanner that reports a 7.8 as `warning` still shows as `high`. Judging the level instead lets a finding the report calls high pass a gate the reader believes is set to catch it, the verdict and the page disagreeing about the same finding.
func (Report) MarshalSARIF ¶
MarshalSARIF serializes the report to standard SARIF 2.1.0 JSON as a single "Draugr" run, with each result's originating scanner recorded in its property bag ("tool").
func (Report) MarshalSARIFWith ¶ added in v0.35.0
func (r Report) MarshalSARIFWith(opts MarshalOptions) ([]byte, error)
MarshalSARIFWith is MarshalSARIF with explicit options.
type RepositoryRef ¶ added in v0.64.0
type RepositoryRef struct {
// URL is the repository as the descriptor named it.
URL string `json:"url"`
// Revision is the commit that was scanned. Empty when git could not be asked.
Revision string `json:"revision,omitempty"`
// Uncommitted counts files in the working copy. Not part of what was scanned unless
// WorkingTree is set, in which case they are precisely what was.
Uncommitted int `json:"uncommitted,omitempty"`
// WorkingTree reports that the scan read the checkout on disk rather than a commit, so the
// result is not reproducible from the revision alone.
WorkingTree bool `json:"workingTree,omitempty"`
}
RepositoryRef is the repository a scan read, and the commit it read.
Recorded by repository scanners as Provenance fields rather than as a typed member of Report, because Provenance is already the channel for "what this scanner says about its own run" and this package should not grow a field per domain fact. Parsed back out here so every consumer, console, markdown, HTML, JSON. Reads it the same way instead of each learning the key names.
func RepositoriesIn ¶ added in v0.64.0
func RepositoriesIn(reports []Report) []RepositoryRef
RepositoriesIn collects the distinct repository/revision pairs the reports recorded.
Keyed on the pair rather than the scanner: five controls reading one commit is one fact. When two controls disagree, both are kept, each repository scanner checks out independently, so on a branch that moves mid-scan they can genuinely read different commits, and collapsing that would be an assumption presented as evidence.
func (RepositoryRef) Short ¶ added in v0.64.0
func (r RepositoryRef) Short() string
Short renders the revision the way a human refers to a commit, with git's own "+" for a tree that has moved past it.
type Result ¶
type Result struct {
// Tool is the scanner that produced the finding.
Tool string `json:"tool,omitempty"`
RuleID string `json:"ruleId"`
Level Level `json:"level"`
Message string `json:"message"`
Location Location `json:"location,omitempty"`
// Score is the finding's numeric CVSS-style severity (0–10), sourced from the SARIF
// "security-severity" property. HasScore reports whether a score was present; without
// one, normalized Severity falls back to Level.
Score float64 `json:"score,omitempty"`
HasScore bool `json:"-"`
// Priority is the computed action band (P1–P4) for this finding, stamped by the engine
// from the component's risk classification. Empty when prioritization is not configured.
Priority string `json:"priority,omitempty"`
// Historical marks a finding that describes a commit rather than the current tree.
//
// A history scan reports the path a secret had in the commit that introduced it, and after a
// rename that path no longer exists. Unmarked, the finding reads as something already cleaned up.
// Which is exactly backwards, because a credential in history is still fetchable by anyone who
// can clone and still needs rotating, whatever the tree looks like now.
Historical bool `json:"historical,omitempty"`
// PriorityFloor explains a band the component's classification alone does not account for.
//
// Some findings are not bounded by where the component sits. A leaked credential is valid
// wherever it is valid, a cloud account, a registry, an artifact store. And git history is often
// readable by more people than the service is reachable by, so an `internal` component can
// understate who can obtain the thing. Where a control says so, the band it produces is not
// damped below a floor.
//
// Set only when the floor actually raised the band, so a reader asking "why is this P2 on a
// supporting internal component" has the answer in the report rather than in the source.
PriorityFloor string `json:"priorityFloor,omitempty"`
// Control names the check this finding came from, sca, sast, secrets. Stamped when a run's
// per-control reports are merged into one document.
//
// The merged file is the only thing a downstream consumer sees, and until a finding carries
// its control that consumer cannot tell two controls apart: one rule id reported by two
// checks is two separate things to do, and grouping by rule id alone silently makes it one.
// The control is also what a reader is being asked to act on ("your dependency scan found
// this"), which no other field states.
Control string `json:"control,omitempty"`
// Component names the part of the application this finding belongs to, stamped by the engine
// from the component whose scan produced it. Empty for a project-scoped control, which has
// no one component to attribute to.
//
// A location alone is ambiguous the moment a descriptor has two components: three components have
// three go.mod files, and two can carry the same path. It is also what makes the priority
// checkable. The band is computed from the component's declared exposure and criticality, so a
// report showing the band without naming the component states a conclusion and withholds its
// premise.
Component string `json:"component,omitempty"`
// Exposure and Criticality are that component's declared classification. How reachable it is to
// an attacker, and how much depends on it. The half of the priority calculation that comes from
// the descriptor rather than from the scanner.
//
// Carried on the finding because naming the component is not the same as stating its premise.
// Every other input to a band already travels with it: the scanner's score, the dataset that
// raised it and the day that copy was obtained, the analyzer that ranked it down and how. A
// report that records all of those and then requires a descriptor to be fetched before the two
// remaining inputs can be read is one an auditor cannot settle from the evidence in front of
// them. And the descriptor they would fetch is the one in the repository today, not necessarily
// the one that produced this band.
//
// Empty for a project-scoped finding, which belongs to no one component, and for a component that
// declares neither. Which Draugr reads as public and critical, so that an unclassified component
// surfaces rather than hides.
//
// Strings rather than the descriptor's own types: `saga` imports this package, so the
// dependency cannot run the other way. These carry `saga.Exposure` and `saga.Criticality`
// values.
Exposure string `json:"exposure,omitempty"`
Criticality string `json:"criticality,omitempty"`
// Labels are the component's own metadata, carried so a consumer holding many components can
// narrow to the ones somebody is answerable for.
//
// The organization's vocabulary rather than Draugr's: nothing here reads a key or attaches a
// meaning to one, and no key is privileged. A team that files by squad, by tier, by regime or
// by all three is describing its own shape, and a tool that decided what `team` meant would be
// describing a different one.
//
// They never reach a verdict, and never reach the console, the markdown report or a
// pull-request comment. Those answer what to fix for somebody who already knows the work is
// theirs. Filtering a fleet is a question asked where there is a fleet.
Labels map[string]string `json:"labels,omitempty"`
// Repository is the repository this finding was found in, for a component that has more than one,
// or a fragment that contributed one from somewhere else.
//
// Paths are rewritten repository-relative so a finding can be anchored to a file, which means
// two repositories that share a path share everything else about a finding. Without this they
// are one finding, and the second repository's copy is discarded on the way in.
Repository string `json:"repository,omitempty"`
// PartialFingerprints are SARIF's own mechanism for identity across runs, keyed by name.
//
// Distinct from Fingerprint, and the difference is what each is for. Fingerprint deduplicates
// *within* a run and hashes the line number to do it, which is exactly right there and useless
// across runs: adding an import at the top of a file changes the line of every finding below
// it. These hash what the code says rather than where it sits.
//
// Empty for a finding with no file and line. A vulnerable dependency is identified by its
// package, which is already on the finding. Absent means "no content-based identity", and a
// fabricated one would be worse than none.
PartialFingerprints map[string]string `json:"partialFingerprints,omitempty"`
// Package identifies the dependency a finding is about, when it is about one.
//
// Nil for a finding that is not: a SAST rule is about a line of code, an IaC check about a
// resource. Set by the scanners that know, which is those reading a package manifest or an
// image's installed set.
Package *Package `json:"package,omitempty"`
// Image is the container image a finding is about, for the scanners that scan one.
//
// It travels on the finding rather than being recovered from the component afterwards,
// because a component may hold several images: recovering it later can only produce one
// answer for all of them, which would be right whenever there is one image and silently
// wrong the moment there are two.
Image string `json:"image,omitempty"`
// OperatingSystem is the OS whose package set contains the finding, e.g. "debian 12".
//
// Set only for findings in an image's OS layer, which is what distinguishes a vulnerable
// system package from a vulnerable application dependency sitting on top of it. Empty for a
// finding in a language ecosystem, where there is no OS answer to give.
OperatingSystem string `json:"operatingSystem,omitempty"`
// OSEndOfLife marks a finding whose operating system release no longer receives security updates
// from its vendor, Trivy's EOSL, End Of Service Life.
//
// It changes what the finding means. On a supported release, "no fix available" is a state
// that will end when the vendor publishes one; on a release past end of life, no fix is ever
// coming, and upgrading the release is the only action that resolves it. That is usually the
// highest-leverage move available, because it resolves every finding in the OS layer at once.
OSEndOfLife bool `json:"osEndOfLife,omitempty"`
// ProviderOperated marks a finding about a surface somebody else runs, the control plane of a
// managed Kubernetes cluster being the case it exists for.
//
// Set from what the descriptor declares, never guessed: whether a cluster is managed is a
// fact about a contract, not something visible in what a scanner reads. It is the same
// argument that puts exposure and criticality in the descriptor.
ProviderOperated bool `json:"providerOperated,omitempty"`
// BuiltUpstream marks a finding inside something somebody else publishes, an image this team only
// runs, or a repository it uses and does not maintain.
//
// It changes the action rather than the severity. A vulnerable library in an image this team
// builds is theirs to upgrade; the same library in an image they only run is fixed by taking
// a newer image, and telling them to upgrade the library is advice they cannot act on. A
// denied license in the dependency tree of a repository they do not publish is the same
// shape: not a license they chose, and not one they can swap out.
//
// Stamped by the engine from what the job's target declares, so a scanner cannot forget it.
BuiltUpstream bool `json:"builtUpstream,omitempty"`
// Layer is the image layer the finding's package arrived in. Nil for anything that is not an
// image finding, and for an image whose scanner did not report one.
//
// It is the only reliable answer to "is this mine or inherited". An image records nothing about
// what it was built FROM. The name is not in there, so a base image cannot be named, and where a
// multi-layer base ends is not knowable either. The layer, and the build step that created it,
// are facts; anything further is inference and has to be labeled as such.
Layer *Layer `json:"layer,omitempty"`
// Suppression is set when a Saga exclusion matched this finding. A suppressed result is
// reported but not counted: it does not reach Counts, the verdict, or the fix-first list.
// Nil for an active finding.
Suppression *Suppression `json:"suppression,omitempty"`
// Escalation is set when exploitability enrichment raised this finding's severity, and says
// which signal did it. Nil when nothing moved it.
//
// Suppression's twin: one records a decision to count a finding for less, this one records
// evidence that it deserves more. Both answer the same question, a report that states a
// conclusion and withholds its premise is a hint rather than evidence, and "critical because CISA
// observed this being exploited, as of a date you can check" is the premise.
Escalation *Escalation `json:"escalation,omitempty"`
// Reachability is set when a reachability analyzer covered this finding's dependency, and
// says whether this project's code can actually reach the vulnerable code. Nil when no
// analyzer ran; ReachabilityUnknown when one ran but could not tell.
//
// Only ever set on a finding about a dependency, so it is nil wherever Package is.
Reachability *Reachability `json:"reachability,omitempty"`
// Correlation is set when more than one scanner reported this same flaw: on the finding that
// is counted it names who else found it, and on the others it names the one they are counted
// under. Nil when only one scanner reported it.
Correlation *Correlation `json:"correlation,omitempty"`
}
Result is a single finding.
func (Result) Correlated ¶ added in v0.103.0
Correlated reports whether another scanner's finding is the one being counted for this flaw.
True only on the copies that are not counted. The finding that is counted has a Correlation too, naming who else found it. And is very much still a finding.
func (Result) Fingerprint ¶
Fingerprint is a stable identifier for deduplication: two results with the same fingerprint are considered the same finding.
func (Result) Imported ¶ added in v0.100.0
Imported reports whether the decision to suppress came from outside this project.
Worth asking separately from Suppressed: the two are counted apart in the report, because "we accepted this" and "our supplier says it does not apply" are different answers to the auditor's question, and a total that merges them can only answer the weaker one.
func (Result) Remediation ¶ added in v0.96.0
func (r Result) Remediation() Remediation
Remediation classifies what can be done about the finding.
Computed rather than stored, from facts each recorded by whoever knows them: the scanner reports a fixed version, the scanner reports end of service life, the descriptor declares who operates the surface. Storing the conclusion as well would let it disagree with its premises.
func (Result) Severity ¶ added in v0.5.0
Severity resolves a finding's normalized severity, in the order the SARIF/prioritization design prescribes:
- the finding's numeric score (CVSS / SARIF security-severity), if present;
- otherwise its SARIF level;
- then raised to floor if floor is more severe (a control-default floor, e.g. a leaked secret is never "low"). Pass an empty floor for no floor.
func (Result) SilencedInSource ¶ added in v0.104.0
SilencedInSource reports whether a suppression came from a comment in the code rather than from a decision anybody recorded.
func (*Result) StampLineHash ¶ added in v0.104.0
StampLineHash records a content fingerprint on a result, if there is one to record.
func (Result) Suppressed ¶ added in v0.42.0
Suppressed reports whether this finding was excluded, by a Saga rule or by an imported claim.
func (Result) VulnerabilityID ¶ added in v0.103.0
VulnerabilityID is the vulnerability this finding is about, or "" when it is not about one.
A rule id belongs to whoever emitted it, and two scanners reporting the same advisory do not spell it the same way. Anything asking "which vulnerability is this", matching a supplier's VEX statement, writing one, recognizing the same flaw found twice. Has to ask that question rather than compare rule ids, or it silently answers for one scanner and not the other.
Empty for a finding that is not a vulnerability at all. A leaked credential and a misconfigured security group are real findings with no advisory to be un-affected by.
type Rule ¶ added in v0.35.0
type Rule struct {
// Name is a human-readable identifier (SARIF reportingDescriptor.name), where the id
// itself is opaque.
Name string `json:"name,omitempty"`
// ShortDescription is a single sentence; FullDescription is a paragraph.
ShortDescription string `json:"shortDescription,omitempty"`
FullDescription string `json:"fullDescription,omitempty"`
// Help is remediation guidance, often Markdown.
Help string `json:"help,omitempty"`
// HelpURI points at the rule's documentation or advisory.
HelpURI string `json:"helpUri,omitempty"`
// Taxa are the shared classifications this rule implements, a CIS benchmark control, a CWE. Empty
// when the scanner claims none.
//
// This is what makes two tools' findings recognizable as being about the same thing, and it is
// deliberately not the rule id. An id belongs to whoever emitted it: `draugr/cis/5.1.1` and
// `kube-bench/cis/5.1.1` are two tools' accounts, and collapsing them into one id. As they were.
// Makes provenance unrecoverable. A taxon is the vocabulary both are speaking, so the
// correspondence is stated rather than inferred from a string collision.
//
// SARIF models exactly this with `taxonomies` and `taxa`, so a consumer that has never heard
// of Draugr can group by CIS control, and a third-party scanner can participate by emitting
// the same references.
Taxa []Taxon `json:"taxa,omitempty"`
}
Rule is what a scanner says about one of its rules, beyond the bare id. Every field is optional: scanners vary widely in how much they publish, and an absent field is normal.
type Severity ¶ added in v0.5.0
type Severity string
Severity is Draugr's normalized, cross-control severity ladder. Unlike Level (the SARIF wire values error/warning/note), Severity is what prioritization ranks on: it splits a numeric CVSS-style score into four bands so a dependency CVE, a leaked secret, and an IaC misconfiguration can share one ordered list. See docs/concepts.md (prioritization).
const ( SeverityCritical Severity = "critical" SeverityHigh Severity = "high" SeverityMedium Severity = "medium" SeverityLow Severity = "low" )
Severity bands, from most to least severe.
func ParseSeverity ¶ added in v0.97.0
ParseSeverity reads a severity band from a flag or a descriptor.
It also accepts the SARIF levels the gate used to take, mapped to the band each one means, so a pipeline written against the older vocabulary keeps working rather than failing at the point it is least convenient. What it will not do is guess: an unrecognized word is an error, because a threshold nobody can parse silently becoming the default is a gate that passes for a reason its author never chose.
func (Severity) Deescalate ¶ added in v0.101.0
Deescalate returns the next-lower severity band; low is already the minimum. Used by reachability analysis to rank a finding one band down when nothing can reach it.
One band, not several, and never to nothing. A call graph is evidence about how the code is called today, not a proof that the flaw cannot be triggered. Reflection and dynamic dispatch are invisible to it, and the call that makes it reachable can be written tomorrow. Ranking it below the things that are reachable is the useful part; making it disappear would be a different and unsupported claim.
type Suppression ¶ added in v0.42.0
type Suppression struct {
// Kind is the SARIF suppression kind. Draugr writes "external": the decision came from the
// Saga, not from an annotation in the source.
Kind string `json:"kind"`
// Justification is the reason the Saga gave. Required by Draugr even though SARIF allows
// it to be absent.
Justification string `json:"justification"`
// AcceptedBy is who decided this was acceptable, when the exclusion said. Empty means the
// suppression is unattributed, reported as such, because "who decided" is half the question an
// auditor is asking and a blank is an answer worth seeing.
AcceptedBy string `json:"acceptedBy,omitempty"`
// Expires is when the acceptance lapses, as YYYY-MM-DD. Empty means it does not.
Expires string `json:"expires,omitempty"`
// VEXStatus is what the exclusion declared this suppression means as a claim about the product:
// not_affected, affected, or fixed. Empty when the exclusion said nothing, which is the common
// case and reports as affected. The reading that is never an overstatement.
VEXStatus string `json:"vexStatus,omitempty"`
// VEXJustification is why the product is not affected, from VEX's fixed vocabulary. Set
// only alongside a not_affected status.
VEXJustification string `json:"vexJustification,omitempty"`
// Origin says who decided. Empty or "saga" is a rule in the descriptor; "vex" is a statement
// somebody else made and this project chose to accept.
//
// The distinction is the whole point of reading a supplier's document rather than retyping it.
// A `not_affected` you copied into `config.exclude` becomes indistinguishable from a decision
// you made and are answerable for; kept apart, the report can say the analysis was theirs, and
// an auditor can ask them rather than you.
Origin string `json:"origin,omitempty"`
// Author is who asserted an imported claim, the VEX document's author. Empty for a descriptor
// rule, where AcceptedBy already answers "who decided".
Author string `json:"author,omitempty"`
// Asserted is when an imported claim was made, as the document wrote it. A supplier's
// statement is made on a date they chose, and a year-old `not_affected` about a package that
// has moved on is worth seeing as old rather than reading as current.
Asserted string `json:"asserted,omitempty"`
// Source is the descriptor or fragment the exclusion was written in, or the document an
// imported claim came from. Empty when the whole descriptor is one file, where naming it
// would be noise.
//
// Splitting exclusions across files is only safe if the report can still say which file
// authorized each one, otherwise composition trades a long descriptor for an unanswerable one,
// which is the worse of the two.
Source string `json:"source,omitempty"`
}
Suppression records that a finding was excluded, and why.
type Taxon ¶ added in v0.55.0
type Taxon struct {
// Taxonomy names the scheme, e.g. "CIS-Kubernetes" or "CWE".
Taxonomy string `json:"taxonomy"`
// ID is the identifier within it, e.g. "5.1.1" or "79".
ID string `json:"id"`
// Name is a short human label, when the scanner supplies one.
Name string `json:"name,omitempty"`
// Version is the taxonomy revision the id belongs to, e.g. "cis-1.12". A check number means
// a different thing between benchmark revisions, so an id without one is ambiguous.
Version string `json:"version,omitempty"`
}
Taxon is a classification a rule implements, in some published taxonomy.