Documentation
¶
Overview ¶
Package crawl visits a web application in headless Chrome, like one visitor clicking through it, and records what it saw: the pages, their forms, and every document and API request the pages made, as a HAR recording. stampede generate --crawl feeds the recording to the AI pipeline in place of a HAR file captured by hand.
The crawler only follows links (GET requests) on the start URL's host. It never submits a form, skips links that look destructive (logout, delete, remove, unsubscribe), and every request the pages make passes a host policy, so it cannot wander onto other sites.
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Form ¶
type Form struct {
Page string `json:"page"`
Method string `json:"method"`
Action string `json:"action"`
Fields []string `json:"fields"`
}
Form is a form found on a page; it is never submitted.
type Options ¶
type Options struct {
Start string
// MaxPages bounds the pages visited (default 30); MaxDepth the clicks
// away from Start (default 3).
MaxPages int
MaxDepth int
// Allow vets every request URL; nil allows only Start's host.
Allow func(*url.URL) bool
// PageTimeout bounds each page load (default 20s); Settle is how long
// to wait after the load for API calls (default 800ms).
PageTimeout time.Duration
Settle time.Duration
Logger *slog.Logger
}
Options configures a crawl.
type Page ¶
type Page struct {
URL string `json:"url"`
Title string `json:"title"`
Status int `json:"status"`
Depth int `json:"depth"`
}
Page is a visited page.
type Result ¶
type Result struct {
Pages []Page `json:"pages"`
Forms []Form `json:"forms"`
// HAR is a HAR 1.2 recording of the documents and API calls.
HAR []byte `json:"-"`
// Requests is the number of entries in HAR.
Requests int `json:"requests"`
}
Result is what a crawl saw.