Documentation
¶
Overview ¶
Package readability extracts the main article and its metadata from HTML.
Use Parse for HTML from an io.Reader. Use ParseNode for a tree parsed with golang.org/x/net/html. Use IsProbablyReaderable or IsProbablyReaderableNode when you only need a fast readerability check.
Parse and the readerability functions use defaults when called without options. Pass functional options to change selected settings.
Article.Content and Article.Node can contain unsafe HTML. The package does not sanitize them. Sanitize extracted HTML before you add it to a web page.
Index ¶
- Variables
- func IsProbablyReaderable(input string, opts ...ReaderableOption) bool
- func IsProbablyReaderableNode(root *html.Node, opts ...ReaderableOption) bool
- type Article
- type Option
- func WithAllowedVideoRegex(pattern *regexp.Regexp) Option
- func WithCharThreshold(threshold int) Option
- func WithClassesToPreserve(classes ...string) Option
- func WithDebug(debug bool) Option
- func WithDisableJSONLD(disable bool) Option
- func WithKeepClasses(keep bool) Option
- func WithLinkDensityModifier(modifier float64) Option
- func WithLogger(logger *slog.Logger) Option
- func WithMaxElemsToParse(max int) Option
- func WithNbTopCandidates(count int) Option
- type ReaderableOption
- type TooManyElementsError
Examples ¶
Constants ¶
This section is empty.
Variables ¶
var ( // ErrNoContent means that extraction did not produce article content. ErrNoContent = engine.ErrNoContent // ErrNoBody means that the supplied HTML tree does not have a body element. ErrNoBody = engine.ErrNoBody // ErrInvalidURL means that the nonempty page URL is not an absolute HTTP or // HTTPS URL with a host. ErrInvalidURL = engine.ErrInvalidURL )
Functions ¶
func IsProbablyReaderable ¶
func IsProbablyReaderable(input string, opts ...ReaderableOption) bool
IsProbablyReaderable reports whether input is likely to contain an article. It applies a fast heuristic and does not extract the article.
Example ¶
package main
import (
"fmt"
"github.com/ryanfowler/readability"
)
func main() {
source := `<article><p>short</p></article>`
fmt.Println(readability.IsProbablyReaderable(source))
}
Output: false
func IsProbablyReaderableNode ¶
func IsProbablyReaderableNode(root *html.Node, opts ...ReaderableOption) bool
IsProbablyReaderableNode reports whether a parsed HTML tree is likely to contain an article. It does not change root.
Example ¶
package main
import (
"fmt"
"strings"
"github.com/ryanfowler/readability"
"golang.org/x/net/html"
)
func main() {
source := `<article><p>An article body with useful prose.</p></article>`
document, err := html.Parse(strings.NewReader(source))
if err != nil {
fmt.Println("error:", err)
return
}
fmt.Println(readability.IsProbablyReaderableNode(document))
}
Output: false
Types ¶
type Article ¶
type Article struct {
// Title is the article title.
Title string `json:"title"`
// Byline identifies the article author.
Byline string `json:"byline"`
// Dir is the text direction, such as "ltr" or "rtl".
Dir string `json:"dir"`
// Lang is the article language from the document metadata.
Lang string `json:"lang"`
// Content is the processed inner HTML of Node. It is not sanitized.
Content string `json:"content"`
// Node is the processed article element. Its inner HTML is Content when the
// package returns the Article. Node is not included in JSON because an
// html.Node contains cyclic links.
Node *html.Node `json:"-"`
// TextContent is the article text. It uses one space for each normal
// whitespace sequence. It retains whitespace in preformatted elements.
TextContent string `json:"textContent"`
// Length is the length of TextContent in UTF-16 code units.
Length int `json:"length"`
// Excerpt is the article description or a short extract from the content.
Excerpt string `json:"excerpt"`
// SiteName is the name of the source site.
SiteName string `json:"siteName"`
// PublishedTime is the publication time from the document metadata. The
// package does not change its source format.
PublishedTime string `json:"publishedTime"`
}
Article contains the extracted article and its metadata.
Content and Node can contain unsafe HTML. Sanitize them before you add them to a web page. Metadata fields are empty when the source does not supply a value.
func Parse ¶
Parse reads HTML from input and extracts an article.
pageURL can be empty. If it is not empty, it must be an absolute HTTP or HTTPS URL with a host. Parse uses pageURL and the document base URL to resolve relative links and media URLs.
With no options, Parse uses the Mozilla defaults. Parse returns input read errors directly. Other errors support errors.Is with ErrNoBody, ErrInvalidURL, or ErrNoContent. Parse can also return a *TooManyElementsError.
Example ¶
package main
import (
"fmt"
"strings"
"github.com/ryanfowler/readability"
)
func main() {
const source = `<html>
<head><title>Hello</title></head>
<body><article><h1>Hello</h1><p>This is a useful article paragraph.</p></article></body>
</html>`
article, err := readability.Parse(
strings.NewReader(source),
"https://example.com/",
readability.WithCharThreshold(0),
)
if err != nil {
fmt.Println("error:", err)
return
}
fmt.Println(article.Title)
}
Output: Hello
func ParseNode ¶
ParseNode extracts an article from a parsed HTML tree.
root can be a complete document or a tree with a body root. ParseNode does not change root. The caller must not change root while ParseNode uses it. pageURL can be empty; a nonempty value must be an absolute HTTP or HTTPS URL with a host. With no options, ParseNode uses the Mozilla defaults.
Example ¶
package main
import (
"fmt"
"strings"
"github.com/ryanfowler/readability"
"golang.org/x/net/html"
)
func main() {
const source = `<html>
<head><title>News</title></head>
<body><article><p>An article body with useful prose.</p></article></body>
</html>`
document, err := html.Parse(strings.NewReader(source))
if err != nil {
fmt.Println("error:", err)
return
}
article, err := readability.ParseNode(
document,
"https://example.com/news",
readability.WithCharThreshold(0),
)
if err != nil {
fmt.Println("error:", err)
return
}
fmt.Println(article.Title)
}
Output: News
type Option ¶ added in v0.1.1
type Option func(*options)
Option configures article extraction. Options are applied in order, so a later option overrides an earlier option for the same setting.
func WithAllowedVideoRegex ¶ added in v0.1.1
WithAllowedVideoRegex sets the pattern used to identify video URLs that cleanup can retain. A nil pattern selects the built-in allowlist.
func WithCharThreshold ¶ added in v0.1.1
WithCharThreshold sets the minimum result length in UTF-16 code units. The package retries extraction with less strict cleanup when a result is shorter. Zero prevents retries.
func WithClassesToPreserve ¶ added in v0.1.1
WithClassesToPreserve sets the CSS classes retained during class cleanup. It has no effect when WithKeepClasses(true) is also used.
func WithDebug ¶ added in v0.1.1
WithDebug controls additional verbose log records. A non-nil logger is required to receive these records.
func WithDisableJSONLD ¶ added in v0.1.1
WithDisableJSONLD controls whether metadata extraction from JSON-LD is disabled.
func WithKeepClasses ¶ added in v0.1.1
WithKeepClasses controls whether all CSS classes are retained.
func WithLinkDensityModifier ¶ added in v0.1.1
WithLinkDensityModifier changes the link-density limits used by cleanup rules to remove a candidate.
func WithLogger ¶ added in v0.1.1
WithLogger sets the logger that receives extraction records. A nil logger turns logging off. The package does not use the global slog logger.
func WithMaxElemsToParse ¶ added in v0.1.1
WithMaxElemsToParse sets the maximum number of HTML elements accepted during extraction. Zero removes the limit.
func WithNbTopCandidates ¶ added in v0.1.1
WithNbTopCandidates sets the number of top article candidates to compare.
type ReaderableOption ¶ added in v0.1.1
type ReaderableOption func(*readerableOptions)
ReaderableOption configures the fast readerability heuristic. Options are applied in order.
func WithMinContentLength ¶ added in v0.1.1
func WithMinContentLength(length int) ReaderableOption
WithMinContentLength sets the minimum candidate length in UTF-16 code units.
func WithMinScore ¶ added in v0.1.1
func WithMinScore(score float64) ReaderableOption
WithMinScore sets the score that a document must exceed to be considered readerable.
type TooManyElementsError ¶
type TooManyElementsError struct {
// Count is the number of elements in the document.
Count int
// Max is the configured maximum number of elements.
Max int
}
TooManyElementsError reports that a document exceeds the limit set by WithMaxElemsToParse.
func (*TooManyElementsError) Error ¶
func (e *TooManyElementsError) Error() string