pdfx

package
v0.5.1-rc.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 1, 2026 License: Apache-2.0 Imports: 8 Imported by: 0

Documentation

Overview

Package pdfx is the LOCAL RUNG of the document ladder: PDF text, extracted in this process, out of this binary, with no service to call and nothing to install.

The ladder's shape is the reason this package is one small file. A document arrives and something has to read it; the rungs go from cheapest to dearest — an embedded text layer here, then OCR, then a vision model looking at the page as a picture. The overwhelming majority of PDFs a person hands a coding session (a spec, a paper, an invoice, an API reference) carry a text layer, and reading it is microseconds of CPU. Paying a vision model to look at a picture of text the file already contains is the expensive answer to a question nobody had to ask.

So the engine here is github.com/AOShei/go-fast-pdf: pure Go, zero dependencies, MIT, compiled straight into the static binary. The rejected alternatives were not worse at reading PDFs — they were worse at SHIPPING: PDFOxide wants a runtime .so beside the binary and extractous wants Tesseract installed, and a rung of the ladder that fails on a machine where nobody ran apt-get is not a local rung at all.

The seam is deliberately one function wide. Extract(path) is the entire contract this package owes the rest of the program, which is what makes the engine swappable: a better pure-Go extractor lands as an edit to this file and nothing above it moves.

THREE OUTCOMES, NOT TWO. A caller that only knows "text or error" cannot tell a broken file from a scanned one, and those two need opposite responses from the model — the first is a dead end, the second is a handoff to the ladder's OCR rungs. So a PDF that parses cleanly but carries no text layer gets its own sentinel (ErrNoTextLayer, with the page count attached) rather than an empty string and a shrug.

Index

Constants

This section is empty.

Variables

View Source
var ErrNoTextLayer = errors.New("pdf has no text layer")

ErrNoTextLayer is the scanned-PDF outcome: the file parsed, it has pages, and every one of them is whitespace. Callers word it for their own audience with errors.Is; NoTextLayerError carries the page count for the ones that want to say how big the document was.

Functions

func Extract

func Extract(path string) (string, error)

Extract returns the document's text, with a light page marker between pages.

The marker is there because a page boundary is real information — "page 3" is how a person and a model both refer to a place in a document — and it is omitted on a single-page file because a marker that never varies is noise.

Three returns, one each for the three outcomes: text and nil, "" and an error wrapping ErrNoTextLayer for a scanned file, "" and a plain error for a file that would not parse.

func ExtractPages added in v0.4.0

func ExtractPages(path string, first, last int) (string, error)

ExtractPages returns the text of the pages named, first through last inclusive and one-based, with Extract's own page markers between them. The document is parsed whole — a PDF has no random access into its pages without one — and only the pages named are rendered, so a caller asking for two pages of a scanned manual still pays the parse but reads three pages and not four hundred.

A range that names no page this document has is an error that says how many pages there are, and pages that parse to nothing at all are the same scanned-file outcome Extract reports.

Types

type NoTextLayerError

type NoTextLayerError struct {
	Pages int
}

NoTextLayerError is ErrNoTextLayer with the page count attached.

func (*NoTextLayerError) Error

func (e *NoTextLayerError) Error() string

func (*NoTextLayerError) Unwrap

func (e *NoTextLayerError) Unwrap() error

Unwrap makes errors.Is(err, ErrNoTextLayer) the test callers write, so the page count is available to whoever wants it and invisible to whoever does not.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL