readability

package module
v0.1.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jul 19, 2026 License: MIT Imports: 14 Imported by: 0

README

readability

Go Reference

readability extracts the main article and its metadata from HTML. It is an idiomatic Go port of Mozilla Readability.

The package can return these values:

  • Processed article HTML
  • Plain article text
  • The article title, author, excerpt, site name, language, text direction, and publication time
  • A parsed html.Node for the article

[!WARNING] Article.Content and Article.Node can contain unsafe HTML. The package does not sanitize these values. Sanitize the HTML before you add it to a web page.

Requirements

  • Go 1.23 or a later version

Installation

go get github.com/ryanfowler/readability

Quick start

Pass the HTML source and the page URL to Parse:

package main

import (
    "fmt"
    "log"

    "github.com/ryanfowler/readability"
)

func main() {
    source := `<html>
<head><title>Example article</title></head>
<body><article><p>This is the article text.</p></article></body>
</html>`

    article, err := readability.Parse(source, "https://example.com/news/1", nil)
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(article.Title)
    fmt.Println(article.TextContent)
}

Use the source page URL when it is available. The package uses this URL and the document <base> element to resolve relative links and media URLs. The page URL can be empty. A nonempty page URL must be an absolute HTTP or HTTPS URL with a host.

A nil options pointer selects the default options.

Check a page before extraction

IsProbablyReaderable applies a fast heuristic. It does not extract the article. Use it when you only need to know if a page is likely to contain an article:

if readability.IsProbablyReaderable(source, nil) {
    article, err := readability.Parse(source, pageURL, nil)
    if err != nil {
        log.Printf("extraction failed: %v", err)
        return
    }
    fmt.Println(article.Title)
}

A true result is not a guarantee that extraction will succeed. A false result does not mean that the HTML is invalid.

Use a parsed HTML tree

Use ParseNode if you already parsed the HTML with golang.org/x/net/html. This prevents a second parse of the source:

doc, err := html.Parse(strings.NewReader(source))
if err != nil {
    // Handle the error.
}

if readability.IsProbablyReaderableNode(doc, nil) {
    article, err := readability.ParseNode(doc, pageURL, nil)
    if err != nil {
        log.Printf("extraction failed: %v", err)
        return
    }
    fmt.Println(article.Title)
}

The node functions accept a complete document or a tree that has a body root. They do not change the supplied tree. You can use the same tree in later calls. Do not change the tree while either node function uses it.

Configure extraction

Start with DefaultOptions when you want to change one or more extraction options:

options := readability.DefaultOptions()
options.CharThreshold = 100
options.MaxElemsToParse = 50_000

article, err := readability.Parse(source, pageURL, &options)

Do not use a partial struct literal unless you need its zero values. A non-nil Options value supplies the full configuration. The exception is AllowedVideoRegex: a nil value keeps the built-in video allowlist.

Option Default Function
MaxElemsToParse 0 Sets the maximum number of HTML elements. Zero removes the limit.
NbTopCandidates 5 Sets the number of top article candidates to compare.
CharThreshold 500 Retries extraction when the result is shorter than this value. Zero prevents retries.
ClassesToPreserve []string{"page"} Lists the CSS classes to retain when class cleanup is active.
KeepClasses false Retains all CSS classes when true.
DisableJSONLD false Prevents metadata extraction from JSON-LD when true.
AllowedVideoRegex Built-in allowlist Identifies video URLs that cleanup can retain.
LinkDensityModifier 0 Changes the link-density limits that remove a candidate.
Logger nil Receives extraction log records. Nil turns logs off.
Debug false Adds verbose debug data to log records when true.

CharThreshold controls extraction retries. If a result is too short, the package retries with less strict removal and cleanup rules. If all results are too short, the package returns the longest nonempty result.

Set Options.Logger to a *slog.Logger to receive logs. The logger handler controls the log level and output. The package does not use the global slog logger.

Use DefaultReaderableOptions to change the readerability heuristic:

options := readability.DefaultReaderableOptions()
options.MinContentLength = 200
likely := readability.IsProbablyReaderable(source, &options)

Character counts

These values use UTF-16 code units:

  • Article.Length
  • Options.CharThreshold
  • ReaderableOptions.MinContentLength

This rule matches JavaScript String.length and Mozilla Readability. Most characters count as one unit. A character outside the Basic Multilingual Plane, such as many emoji, counts as two units.

Handle errors

Use errors.Is with the package error values:

  • ErrNoBody: the input does not contain a required body element.
  • ErrInvalidURL: the page URL is not valid.
  • ErrNoContent: extraction did not produce article content.

Use errors.As to inspect an element-limit error:

_, err := readability.Parse(source, pageURL, &options)
if err != nil {
    var limitErr *readability.TooManyElementsError
    switch {
    case errors.As(err, &limitErr):
        log.Printf("document has %d elements; limit is %d", limitErr.Count, limitErr.Max)
    case errors.Is(err, readability.ErrNoContent):
        log.Print("no article content found")
    default:
        log.Printf("extraction failed: %v", err)
    }
}

Output safety

Treat all extracted data as untrusted input.

  • Sanitize Article.Content before you add it to a web page.
  • Sanitize or validate Article.Node before you render it.
  • Escape plain-text metadata for its output context.
  • Set MaxElemsToParse when you process HTML from an untrusted source and need an element limit.

Compatibility

This package tracks Mozilla Readability.js at commit 08be6b4bdb204dd333c9b7a0cfbc0e730b257252. The repository pins Mozilla's official 130-case test corpus in tests/readability-js.

Go parses and serializes the HTML instead of a browser DOM. As a result, attribute order and other HTML serialization details can differ.

License

The project is licensed under the MIT License. It contains code derived from Mozilla Readability and Arc90 Readability. The pinned Mozilla test fixtures retain their original license.

Documentation

Overview

Package readability extracts the main article and its metadata from HTML.

Use Parse for an HTML string. Use ParseNode for a tree that was parsed with golang.org/x/net/html. Use IsProbablyReaderable or IsProbablyReaderableNode when you only need a fast readerability check.

Pass nil for an options pointer to use the defaults. To change selected options, call DefaultOptions or DefaultReaderableOptions first. Then change the returned value.

Article.Content and Article.Node can contain unsafe HTML. The package does not sanitize them. Sanitize extracted HTML before you add it to a web page.

Index

Examples

Constants

This section is empty.

Variables

View Source
var (
	// ErrNoContent means that extraction did not produce article content.
	ErrNoContent = errors.New("readability: no content")
	// ErrNoBody means that the supplied HTML tree does not have a body element.
	ErrNoBody = errors.New("readability: document has no body")
	// ErrInvalidURL means that the nonempty page URL is not an absolute HTTP or
	// HTTPS URL with a host.
	ErrInvalidURL = errors.New("readability: invalid URL")
)

Functions

func IsProbablyReaderable

func IsProbablyReaderable(input string, options *ReaderableOptions) bool

IsProbablyReaderable reports whether input is likely to contain an article.

This function applies a fast heuristic. It does not extract the article. It returns false if it cannot parse a document body. Pass nil for options to use the defaults.

Example
package main

import (
	"fmt"

	"github.com/ryanfowler/readability"
)

func main() {
	source := `<article><p>short</p></article>`
	fmt.Println(readability.IsProbablyReaderable(source, nil))
}
Output:
false

func IsProbablyReaderableNode

func IsProbablyReaderableNode(root *html.Node, options *ReaderableOptions) bool

IsProbablyReaderableNode reports whether a parsed HTML tree is likely to contain an article.

This function applies a fast heuristic. It does not extract the article. root can be a complete document or a tree with a body root. The function returns false if root is nil or has no body element. It does not change root. The caller must not change root while the function uses it. Pass nil for options to use the defaults.

Example
package main

import (
	"fmt"
	"strings"

	"github.com/ryanfowler/readability"
	"golang.org/x/net/html"
)

func main() {
	source := `<article><p>An article body with useful prose.</p></article>`
	document, err := html.Parse(strings.NewReader(source))
	if err != nil {
		fmt.Println("error:", err)
		return
	}

	fmt.Println(readability.IsProbablyReaderableNode(document, nil))
}
Output:
false

Types

type Article

type Article struct {
	// Title is the article title.
	Title string `json:"title"`
	// Byline identifies the article author.
	Byline string `json:"byline"`
	// Dir is the text direction, such as "ltr" or "rtl".
	Dir string `json:"dir"`
	// Lang is the article language from the document metadata.
	Lang string `json:"lang"`
	// Content is the processed inner HTML of Node. It is not sanitized.
	Content string `json:"content"`
	// Node is the processed article element. Its inner HTML is Content when the
	// package returns the Article. Node is not included in JSON because an
	// html.Node contains cyclic links.
	Node *html.Node `json:"-"`
	// TextContent is the article text. It uses one space for each normal
	// whitespace sequence. It retains whitespace in preformatted elements.
	TextContent string `json:"textContent"`
	// Length is the length of TextContent in UTF-16 code units.
	Length int `json:"length"`
	// Excerpt is the article description or a short extract from the content.
	Excerpt string `json:"excerpt"`
	// SiteName is the name of the source site.
	SiteName string `json:"siteName"`
	// PublishedTime is the publication time from the document metadata. The
	// package does not change its source format.
	PublishedTime string `json:"publishedTime"`
}

Article contains the extracted article and its metadata.

Content and Node can contain unsafe HTML. Sanitize them before you add them to a web page. Metadata fields are empty when the source does not supply a value.

func Parse

func Parse(input, pageURL string, options *Options) (*Article, error)

Parse extracts an article from input.

pageURL can be empty. If it is not empty, it must be an absolute HTTP or HTTPS URL with a host. Parse uses pageURL and the document base URL to resolve relative links and media URLs.

Pass nil for options to use the defaults. Parse returns an error that supports errors.Is with ErrNoBody, ErrInvalidURL, or ErrNoContent. It can also return a *TooManyElementsError.

Example
package main

import (
	"fmt"

	"github.com/ryanfowler/readability"
)

func main() {
	const source = `<html>
<head><title>Hello</title></head>
<body><article><h1>Hello</h1><p>This is a useful article paragraph.</p></article></body>
</html>`

	options := readability.DefaultOptions()
	options.CharThreshold = 0
	article, err := readability.Parse(source, "https://example.com/", &options)
	if err != nil {
		fmt.Println("error:", err)
		return
	}

	fmt.Println(article.Title)
}
Output:
Hello

func ParseNode

func ParseNode(root *html.Node, pageURL string, options *Options) (*Article, error)

ParseNode extracts an article from a parsed HTML tree.

root can be a complete document or a tree with a body root. ParseNode does not change root. The caller must not change root while ParseNode uses it.

pageURL can be empty. If it is not empty, it must be an absolute HTTP or HTTPS URL with a host. ParseNode uses pageURL and the document base URL to resolve relative links and media URLs.

Pass nil for options to use the defaults. ParseNode returns an error that supports errors.Is with ErrNoBody, ErrInvalidURL, or ErrNoContent. It can also return a *TooManyElementsError.

Example
package main

import (
	"fmt"
	"strings"

	"github.com/ryanfowler/readability"
	"golang.org/x/net/html"
)

func main() {
	const source = `<html>
<head><title>News</title></head>
<body><article><p>An article body with useful prose.</p></article></body>
</html>`

	document, err := html.Parse(strings.NewReader(source))
	if err != nil {
		fmt.Println("error:", err)
		return
	}

	options := readability.DefaultOptions()
	options.CharThreshold = 0
	article, err := readability.ParseNode(
		document,
		"https://example.com/news",
		&options,
	)
	if err != nil {
		fmt.Println("error:", err)
		return
	}

	fmt.Println(article.Title)
}
Output:
News

type Options

type Options struct {
	// MaxElemsToParse is the maximum number of HTML elements that extraction
	// accepts. Zero removes the limit.
	MaxElemsToParse int
	// NbTopCandidates is the number of top article candidates to compare.
	NbTopCandidates int
	// CharThreshold is the minimum result length in UTF-16 code units. The
	// package retries extraction if the result is shorter. Each retry uses less
	// strict cleanup rules. If all results are too short, the package returns the
	// longest nonempty result. Zero prevents retries.
	CharThreshold int
	// ClassesToPreserve lists the CSS classes to retain during class cleanup.
	// This field has no effect when KeepClasses is true.
	ClassesToPreserve []string
	// KeepClasses retains all CSS classes when it is true.
	KeepClasses bool
	// DisableJSONLD prevents metadata extraction from JSON-LD when it is true.
	DisableJSONLD bool
	// AllowedVideoRegex identifies video URLs that cleanup can retain. A nil
	// value selects the built-in allowlist.
	AllowedVideoRegex *regexp.Regexp
	// LinkDensityModifier changes the link-density limits that the cleanup rules
	// use to remove a candidate.
	LinkDensityModifier float64
	// Logger receives extraction log records. A nil value turns logs off. The
	// package does not use the global slog logger.
	Logger *slog.Logger
	// Debug enables additional verbose log records. Logger must be non-nil to
	// receive these records.
	Debug bool
}

Options controls article extraction.

Pass nil to Parse or ParseNode to use the defaults. To change selected fields, first call DefaultOptions and then change the returned value. A non-nil Options value supplies all options. Zero values have a function. A nil AllowedVideoRegex is the exception; it selects the built-in allowlist.

func DefaultOptions

func DefaultOptions() Options

DefaultOptions returns an Options value with the Mozilla defaults.

type ReaderableOptions

type ReaderableOptions struct {
	// MinScore is the score that the document must exceed.
	MinScore float64
	// MinContentLength is the minimum candidate length in UTF-16 code units.
	MinContentLength int
}

ReaderableOptions controls the fast readerability heuristic.

Pass nil to a readerability function to use the defaults. To change selected fields, first call DefaultReaderableOptions and then change the returned value. A non-nil ReaderableOptions value supplies all options.

func DefaultReaderableOptions

func DefaultReaderableOptions() ReaderableOptions

DefaultReaderableOptions returns a ReaderableOptions value with the Mozilla defaults.

type TooManyElementsError

type TooManyElementsError struct {
	// Count is the number of elements in the document.
	Count int
	// Max is the configured maximum number of elements.
	Max int
}

TooManyElementsError reports that a document exceeds MaxElemsToParse.

func (*TooManyElementsError) Error

func (e *TooManyElementsError) Error() string

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL