Documentation
¶
Overview ¶
Package redeliver recovers GitHub App webhook deliveries the hub's own downtime lost. GitHub keeps the record of a failed delivery — an operator can always redeliver one by hand from the App's Advanced tab — but it never retries one automatically, so a hub that stayed down or unreachable for a while would otherwise never see the pushes that happened in that window. On start and then hourly, a webhook-mode hub asks the App API which of the last 72 hours' deliveries still failed and asks GitHub to redeliver each one.
Redelivering costs nothing extra when the operator's own endpoint is the thing that is down: GitHub still accepts the redeliver request and simply fails it again. What must be bounded is how many of MaxAttempts a GUID spends while that is true, because "the hub keeps answering with an error" must not look identical to "GitHub keeps failing this one delivery for some other reason". reachable answers that from evidence already in the same listing: any answered delivery attempt — a 2xx, or an application-level rejection like 401 or 503, but not a gateway/tunnel non-answer (502, 504, 530, or no status at all) — including this GUID's own latest attempt. A GUID's own evidence is identified by delivery attempt id, not by comparing timestamps: GitHub's own delivered_at is whole seconds on GitHub's clock, while the hub's own recorded attempt time is local and sub-second, so the two tie whenever a redeliver lands in the same GitHub-clock second as the attempt that requested it — which a fast round trip makes the common case, not the rare one. Comparing the id of the GUID's latest listed attempt against the id that was latest the last time this hub acted on it sidesteps clock resolution entirely: any answered attempt with a different id is new evidence, whatever second it landed in. Without that evidence the sweep still redelivers, since the delivery may succeed even though nothing has recently, but it does not spend one of MaxAttempts on it.
That still leaves one case unbounded by MaxAttempts: the tunnel itself is down, so nothing the App API lists was ever answered at all, and every hourly redelivery creates a new attempt with a fresh delivered_at that never ages out of the 72-hour listing window on its own. FirstDeliveredAt closes this: it is set once, from the GUID's earliest attempt this hub has seen, and never overwritten, so a GUID is abandoned once 72 hours have passed since then regardless of how fresh its latest attempt looks.
A redelivered event arrives at the hub's normal webhook route and deduplicates there by delivery ID exactly like any other delivery: this package needs no dedup of its own, only the memory of how many counted attempts it has already spent.
Index ¶
Constants ¶
const ( // DefaultAPIBaseURL is GitHub's public API. A test points this at an // httptest server instead. DefaultAPIBaseURL = "https://api.github.com" // DefaultInterval is how often the sweep repeats once started, and also // the minimum gap this package enforces between two counted or // uncounted redeliver calls for the same GUID. The sweep always runs // once immediately as well, on daemon start. DefaultInterval = time.Hour // MaxDeliveryAge bounds how far back the sweep lists deliveries, and also // how long a GUID it cannot count an attempt against survives: once // MaxDeliveryAge has passed since a GUID's FirstDeliveredAt, it is // abandoned outright, whether or not it ever spent an attempt. This is a // hard bound: nothing in this package ever widens it. MaxDeliveryAge = 72 * time.Hour // MaxAttempts bounds how many counted attempts the sweep spends asking // GitHub to redeliver the same GUID before narrating it abandoned and // never touching it again. An attempt only counts while reachable // reports evidence the endpoint is answering at all. MaxAttempts = 3 // Retention is how long a redelivery record survives in the store after // its last attempt, whether it is still retrying, succeeded, or was // abandoned. Retention = 7 * 24 * time.Hour // MaxBackoff bounds the exponential retreat after a failed sweep, and // caps any GitHub-recommended Retry-After or X-RateLimit-Reset wait too. // It must exceed Interval — a shorter cap would make a persistent // failure retry more often than a healthy sweep does, hammering an // outage. MaxBackoff = 6 * time.Hour )
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Options ¶
type Options struct {
// Client performs the GitHub App API requests.
Client *http.Client
// APIBaseURL is GitHub's API origin. Empty means DefaultAPIBaseURL.
APIBaseURL string
// AppID and PrivateKeyPEM mint the App JWT the sweep authenticates with,
// the same key material the hub's installation verifier signs with.
AppID int64
PrivateKeyPEM []byte
// Store is the sweep's memory of prior attempts per delivery GUID.
Store hub.WebhookRedeliveryStore
// Interval is the gap between sweeps, and the minimum gap this package
// enforces between two redeliver calls for the same GUID. Zero means
// DefaultInterval.
Interval time.Duration
// Now and Sleep exist so a test drives time instead of waiting for it.
Now func() time.Time
Sleep func(context.Context, time.Duration) error
// Narrate receives one line per redelivery, one per abandonment, one for
// a sweep that failed before it could finish, and one per sweep in which
// any redeliver call went uncounted for lack of reachability evidence.
// Nil discards every line.
Narrate func(narrate.Line)
}
Options configures a Sweeper. Store, AppID and PrivateKeyPEM are required for the sweep to do anything; everything else has a working default.
type Status ¶
type Status struct {
LastSweepAt *time.Time
Redelivered int
Abandoned int
// Uncounted is how many redeliver calls this pass made without evidence
// the endpoint is reachable, so they were not spent against MaxAttempts.
// A sustained non-zero value is what makes an ongoing outage visible in
// `wb daemon status` even though nothing is being abandoned for it.
Uncounted int
// LastFailureAt and LastFailureClass are sticky: a later clean sweep
// does not clear them. An operator asking "has this ever failed, and
// how" gets an answer that survives the next successful pass. Neither is
// set for a sweep cut short by ctx ending — a shutdown is not a failure.
LastFailureAt *time.Time
LastFailureClass string
}
Status is the sweep's most recent completed-pass summary, read by `wb daemon status`. LastSweepAt and LastFailureAt are pointers so JSON omits them before anything has happened yet, rather than rendering the zero time.
type Sweeper ¶
type Sweeper struct {
// contains filtered or unexported fields
}
Sweeper is the missed-webhook recovery loop.
func (*Sweeper) Run ¶
Run sweeps once immediately — the "on start" half of the requirement — then again after every delay Sweep returns, until ctx ends. It returns nil on a cancelled context: a stopped sweeper is a normal shutdown.
func (*Sweeper) Status ¶
Status returns the most recent completed sweep's summary, the zero value before the first one finishes.
func (*Sweeper) Sweep ¶
Sweep performs one pass and returns how long to wait before the next one: the configured interval normally, a GitHub-recommended wait (capped at MaxBackoff) when a 403 or 429 carried one, or the current exponential backoff otherwise. It never returns an error and never panics on one: a sweep that cannot reach GitHub is narrated in one line and retried later, which is the whole point of a background recovery loop outliving a transient GitHub outage. A sweep cut short by ctx ending is not narrated as a failure at all — sweepOnce already resets that case to a nil error.