readability

package module
v0.6.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 17, 2026 License: MIT Imports: 21 Imported by: 0

README

Go-Readability V2

This is a fork of codeberg.org/readeck/go-readability, branch v2, taken at commit b18540d99ebf105cd67122585a0a41ec299b70bc. Readeck's implementation is itself a fork of github.com/go-shiori/go-readability, originally written by Radhi Fadlillah and maintained by Felipe Martin and GitHub contributors.

The inherited implementation matches the Mozilla Readability.js 0.6.0 baseline, plus the Go forks' additional fixes and performance improvements. Going forward, we strive to mirror Mozilla's original JavaScript Readability implementation as it evolves.

Our philosophy is bring your own HTML. The library extracts readable article content and metadata from an io.Reader or an existing html.Node. Everything outside the core extraction library has been removed: URL fetching, request modifiers, the CLI and HTTP server, network helper scripts, and diagnostic parser logging. The caller controls how HTML is acquired and how work is scheduled.

Reader decoding, DOM extraction, parser options, metadata and date getters, and HTML/plain-text rendering are retained.

The library is single-threaded. Each extraction runs on the calling goroutine, with no internal extraction workers or worker pool. It is suitable for servers running many engines in parallel: give each engine its own parser and input DOM, and let the server control concurrency.

Usage

Version 0.6.0 uses the module path github.com/markusmobius/go-readabilityV2 and requires Go 1.26 or newer. Supply HTML from your own reader or DOM:

package main

import (
	"fmt"
	"log"
	"net/url"
	"os"

	readability "github.com/markusmobius/go-readabilityV2"
)

func main() {
	srcFile, err := os.Open("index.html")
	if err != nil {
		log.Fatal(err)
	}
	defer srcFile.Close()

	baseURL, err := url.Parse("https://example.com/path/to/article")
	if err != nil {
		log.Fatal(err)
	}
	article, err := readability.FromReader(srcFile, baseURL)
	if err != nil {
		log.Fatal(err)
	}

	fmt.Printf("Found article with title %q\n\n", article.Title())
	// Print the parsed, cleaned-up HTML markup of the article.
	if err := article.RenderHTML(os.Stdout); err != nil {
		log.Fatal(err)
	}
}

The base URL resolves relative links and related metadata; it is never fetched. Use FromDocument for an existing DOM, or NewParser to customize extraction. CheckDocument provides the fast readability check without full extraction.

Quality and Performance

Measured on 2026-09-16 using all 983 labeled pages from content-extractor-benchmark, the same pinned corpus used for Go-DomDistiller and Rust-DomDistiller. The Codeberg row uses the unmodified v2.1.2 source at the commit above and its original dependencies. The Go fork and Rust rows use version 0.6.0.

Extraction Quality
Extractor Precision Recall F1 Accuracy
Codeberg Go-Readability v2.1.2 0.8705 0.8862 0.8783 0.8774
Go-ReadabilityV2 0.8705 0.8862 0.8783 0.8774
Rust-Readability 0.8705 0.8862 0.8783 0.8774

All three engines have exactly the same quality counts: TP 2,601, FN 334, FP 387 and TN 2,561. The same seven empty extractions are scored as empty text, not skipped. Scores use the benchmark's case-sensitive snippet matching and globally aggregated counts, not token-level scoring or metadata-quality scores.

The Go fork and Rust match exactly on all 983 pages for text, HTML, title, byline, excerpt, site name, image URL, favicon, language and errors. Codeberg's original dependencies use x/net 0.41.0 and x/text 0.26.0; the fork uses 0.59.0 and 0.42.0. With the original dependencies, 387 pages differ only in the HTML field; all other compared fields and quality counts agree. A separate Codeberg control using the fork's dependency versions matches all ten fields exactly on 983/983 pages. Thus equal quality is not a claim of byte-identical HTML across different parser versions.

Extraction Time

Median time per complete 983-page pass; ranges show the measured minimum and maximum across all 102 measured passes per engine. Speedup is Go-ReadabilityV2's median time divided by each engine's median time.

Extractor Median Range Speedup vs Go-ReadabilityV2
Codeberg Go-Readability v2.1.2 2,044 ms 1,854-2,792 ms 0.96x
Go-ReadabilityV2 1,955 ms 1,811-2,712 ms 1.00x
Rust-Readability 960 ms 876-1,521 ms 2.04x

Rust's median speedup was 2.04x over Go-ReadabilityV2 and 2.13x over Codeberg.

Measured on an AMD Ryzen AI 7 PRO 350 under Linux/WSL2, pinned to one logical CPU, with Go 1.27.1 and Rust 1.98.1 release builds. Each engine received two warmup passes followed by 102 measured passes, cycling all six engine orders 17 times. The dependency-aligned Codeberg control is for correctness only; the timed Codeberg row retains its original dependency versions.

Timing includes extraction from pre-parsed DOMs, plain-text rendering and snippet scoring. It excludes file I/O, decoding, initial HTML/URL parsing, startup, IPC and the extra HTML/metadata collection used for exact-output checks. Timing varies between runs: these are single-machine measurements, not a guaranteed speedup or an end-to-end reader/network benchmark.

Raw samples, dependency graphs and source/binary fingerprints are retained in benchmark-results-0.6.0-crate.json. The runner and reproduction instructions are maintained with Rust-Readability.

License

The upstream MIT LICENSE is retained unchanged.

Documentation

Overview

Package readability is a Go package that find the main readable content from a HTML page. It works by removing clutter like buttons, ads, background images, script, etc.

This package is based from Readability.js by Mozilla, and written line by line to make sure it looks and works as similar as possible. This way, hopefully all web page that can be parsed by Readability.js are parse-able by go-readability as well.

Index

Examples

Constants

This section is empty.

Variables

View Source
var ErrTimestampMissing = errors.New("timestamp not found in document")

Functions

func CheckDocument

func CheckDocument(doc *html.Node) bool

CheckDocument checks whether the document is readable without parsing the whole thing. It's the wrapper for `Parser.CheckDocument()` and useful if you only use the default parser.

Types

type Article

type Article struct {
	// Node is the top-level container of cleaned-up article content. It may be nil if there were
	// errors or if article content was blank.
	Node *html.Node
	// contains filtered or unexported fields
}

Article is the final readable content of a parsed article.

func FromDocument

func FromDocument(doc *html.Node, pageURL *nurl.URL) (Article, error)

FromDocument parses an document and returns the readable content. It's the wrapper or `Parser.ParseDocument()` and useful if you only want to use the default parser.

func FromReader

func FromReader(input io.Reader, pageURL *nurl.URL) (Article, error)

FromReader parses an `io.Reader` and returns the readable content. It's the wrapper or `Parser.Parse()` and useful if you only want to use the default parser.

Example
source := "<title>Example article</title><article><p>" +
	strings.Repeat("A detailed article, with context and supporting explanations. ", 20) +
	"</p></article>"
baseURL, _ := url.Parse("https://example.com/path/to/article")
article, err := FromReader(strings.NewReader(source), baseURL)
if err != nil {
	panic(err)
}

fmt.Println("Title:", article.Title())
fmt.Println("Content:", article.Node != nil)
Output:
Title: Example article
Content: true

func (Article) Byline

func (a Article) Byline() string

func (Article) Excerpt

func (a Article) Excerpt() string

func (Article) Favicon

func (a Article) Favicon() string

func (Article) ImageURL

func (a Article) ImageURL() string

func (Article) Language

func (a Article) Language() string

func (Article) ModifiedTime

func (a Article) ModifiedTime() (time.Time, error)

ModifiedTime is the time when the article was modified. If no timestamp was found in the article metadata, the error will be ErrTimestampMissing.

func (Article) PublishedTime

func (a Article) PublishedTime() (time.Time, error)

PublishedTime is the time when the article was published. If no timestamp was found in the article metadata, the error will be ErrTimestampMissing.

func (Article) RenderHTML

func (a Article) RenderHTML(w io.Writer) error

func (Article) RenderText

func (a Article) RenderText(w io.Writer) error

func (Article) SiteName

func (a Article) SiteName() string

func (Article) Title

func (a Article) Title() string

type Parser

type Parser struct {
	// MaxElemsToParse is the max number of nodes supported by this
	// parser. Default: 0 (no limit)
	MaxElemsToParse int
	// NTopCandidates is the number of top candidates to consider when
	// analysing how tight the competition is among candidates.
	NTopCandidates int
	// CharThresholds is the default number of chars an article must
	// have in order to return a result
	CharThresholds int
	// ClassesToPreserve are the classes that readability sets itself.
	ClassesToPreserve []string
	// KeepClasses specify whether the classes should be stripped or not.
	KeepClasses bool
	// TagsToScore is element tags to score by default.
	TagsToScore []string
	// DisableJSONLD determines if metadata in JSON+LD will be extracted
	// or not. Default: false.
	DisableJSONLD bool
	// AllowedVideoRegex is a regular expression that matches video URLs that should be
	// allowed to be included in the article content. If undefined, it will use default filter.
	AllowedVideoRegex *regexp.Regexp
	// contains filtered or unexported fields
}

Parser is the parser that parses the page to get the readable content.

func NewParser

func NewParser() Parser

NewParser returns new Parser which set up with default value.

func (*Parser) CheckDocument

func (ps *Parser) CheckDocument(doc *html.Node) bool

CheckDocument checks whether the document is readable without parsing the whole thing.

func (*Parser) Parse

func (ps *Parser) Parse(input io.Reader, pageURL *nurl.URL) (Article, error)

Parse parses a reader and find the main readable content.

func (*Parser) ParseAndMutate

func (ps *Parser) ParseAndMutate(doc *html.Node, pageURL *nurl.URL) (Article, error)

ParseAndMutate is like ParseDocument, but mutates doc during parsing.

func (*Parser) ParseDocument

func (ps *Parser) ParseDocument(doc *html.Node, pageURL *nurl.URL) (Article, error)

ParseDocument parses the specified document and find the main readable content.

Directories

Path Synopsis
internal
re2go
Code generated by re2go 4.2, DO NOT EDIT.
Code generated by re2go 4.2, DO NOT EDIT.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL