xml

package
v1.3.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 25, 2026 License: Apache-2.0 Imports: 7 Imported by: 0

README

encoding/xml

A pull reader for Office Open XML parts: worksheets, shared strings, document bodies, styles. It reads one token at a time from an io.Reader or a byte slice, hands back names, attributes and text as slices into its own buffer, and allocates nothing in steady state. A caller copies only what it keeps.

It is a filter, not a mapper. encoding/xml in the standard library decodes any document into any struct; this package reads the handful of elements an Office consumer wants out of a part that is mostly something else, and skips the rest at scanning speed.

Why not the standard library

encoding/xml allocates for every token: a Name with two strings, an []Attr, a copied CharData. Reading a worksheet with forty thousand cells through Decoder.Token allocates half a million times before the caller keeps anything, and Unmarshal on top of it adds reflection. Measured on the same worksheet, 1.5 MB, 2000 rows of 20 cells, linux/amd64:

ns/op MB/s allocs/op
Decoder.RawToken loop 59,363,299 25.6 364,082
Decoder.Token loop 79,809,134 19.0 548,110
Reader.Next loop 6,953,800 218.1 0
Reader.Skip over the sheet body 5,943,703 255.2 0
Unmarshal into row/cell structs 131,751,407 11.5 839,876
Decodable into the same structs 12,616,078 120.2 75,782 (the strings kept)
Decodable into typed cells 11,471,666 132.2 0
Decodable into typed cells, Children loops, one per method 12,397,310 122.4 0
Decodable into typed cells, four Children loops in one method 20,500,697 74.0 166,002

And on a shared-strings part, 20,000 strings with entities and rich-text runs:

MB/s allocs/op
Unmarshal 15.0 436,058
Reader, one string per <si> 200.0 20,000

go test -bench . ./encoding/xml reproduces the table.

The same benchmarks under TinyGo 0.42.0 on the same machine. TinyGo does not count objects, so only bytes per operation are meaningful there, and the standard library's figures are its own encoding/xml compiled by TinyGo:

MB/s B/op
Decoder.RawToken loop 22.3 24,485,984
Reader.Next loop 139.7 1,017
Reader.Skip over the sheet body 294.0 632
Unmarshal into row/cell structs 9.1 53,796,336
Decodable into typed cells 84.2 63,939
Decodable into typed cells, Children loops, one per method 50.2 10,688,355
Unmarshal of shared strings 15.3 25,600,384
Reader, one string per <si> 133.6 803,004

The bytes on the Reader rows are the buffer and the kept strings, sized once; the iterator row shows what the per-loop closure contexts cost when the loop is per cell, which is why a TinyGo decoder writes the explicit loop.

Reading

r := xml.NewReader(part, xml.Options{})   // or xml.NewBytesReader(data, ...)
for {
	k, err := r.Next()
	if err != nil {
		return err
	}
	if k == xml.EOF {
		break
	}
	if k == xml.StartElement && r.NameIs("sheetData") {
		err = r.Decode(&sheet)          // sheet implements Decodable
	}
}

On a StartElement the caller can:

  • read Name, LocalName, Prefix, and any attribute with Attr(name) or all of them with NextAttr. The attributes were indexed as the tag was scanned, so Attr compares against the element's few names and copies nothing;
  • walk the direct children with NextChild, which skips whatever the caller does not consume;
  • take the element's own text with ElementText, entities decoded, children skipped;
  • capture the whole subtree, tags included, with RawElement;
  • discard the subtree with Skip, which scans it raw: tags, quotes and depth only, no attribute indexing and no end-tag matching inside.

NextChild is the shape of every decoder: a loop over the children with a switch on the name, and nothing to do for the elements it does not name. Children is the same loop as a range statement, and Tokens is Next as one; both end silently on an error, which Err reports after the loop.

for name := range r.Children(r.Element()) {
	switch string(name) {
	case "v":
		v, err := r.ElementText()
		...
	}
}
if err := r.Err(); err != nil {
	return err
}

Keep one Children loop per function. A range-over-func loop costs nothing only while the compiler inlines the iterator at the call site, and it stops doing that two or three nested loops in. Measured on the worksheet above, the typed decode written as four nested Children loops in one method allocates four times per cell and runs at 74 MB/s; the same decode with one loop per method allocates nothing and runs at 122 MB/s, against 132 MB/s for the explicit NextChild loops. One element type per DecodeXMLFrom is the design anyway, so the rule costs nothing to follow.

func (c *Cell) DecodeXMLFrom(r *xml.Reader) error {
	ref, _ := r.Attr("r")
	c.Col, c.Row = parseRef(ref)
	if s, ok := r.Attr("s"); ok {
		c.Style, _ = s.Int()
	}
	el := r.Element()
	for {
		ok, err := r.NextChild(el)
		if err != nil || !ok {
			return err
		}
		if r.NameIs("v") {
			v, err := r.ElementText()
			if err != nil {
				return err
			}
			c.Value, err = v.Float()
		}
	}
}

Values

Attr, Text and ElementText return a Value, the bytes as written with entities still encoded. Its methods decode on demand: Int, Float, Bool parse the raw bytes, Equal compares decoded content, String and AppendTo decode into new or caller-owned storage. Content with no ampersand, which is nearly all of it, is never copied.

Float converts a plain decimal of at most 15 significant digits by one exact division, which is bit-identical to strconv.ParseFloat and twice as fast; anything else, an exponent, more digits, inf, goes to strconv.

Lifetime

Every slice the reader returns aliases its buffer and is valid until the next call that advances the reader: Next, Skip, NextChild, ElementText, RawElement, Decode. Keep a name or a value by copying it. The reader is not safe for concurrent use; reuse one across parts with Reset.

Names and namespaces

Names are matched as written, prefix included. Office XML generators emit canonical prefixes (w:, a:, r:, mc:), so NameIs("w:p") is enough, and it costs a byte comparison.

The reader also tracks xmlns declarations, so a caller that must tell one vocabulary from another can ask. Namespace returns the namespace of the current element's name, through its prefix or the default namespace; LookupNamespace resolves any prefix at the current position. Declarations are copied once, when they are read, which for an Office part is a handful of strings at the root; every start tag after that pays one byte comparison per attribute to notice there are none. The xmlns attributes stay visible through NextAttr as well.

On TinyGo

Everything above holds under TinyGo 0.42, with two differences a caller should know.

string(b) == "lit" allocates. The Go compiler elides the conversion in a comparison, a switch string(b) and a map index; TinyGo does not, and copies the bytes every time. Compare names with NameIs, values with Value.Equal, and anything else with xml.Equal(b, s), which allocates on neither compiler. This package uses nothing else internally, and the allocation tests run under tinygo test to hold it there.

The iterators allocate per loop. TinyGo puts the closure contexts of a range-over-func loop on the heap, the iterator's and the loop body's, so the cost grows with what the body captures: 32 bytes for a Tokens loop with an empty body, 80 for a Children loop with one, about 200 for a loop whose body decodes into a struct. It is paid once per loop rather than per iteration, and the tests pin the numbers. The explicit Next and NextChild calls the iterators wrap allocate nothing, so a decoder that must not allocate on TinyGo writes the explicit loop.

testing.AllocsPerRun returns zero under TinyGo whatever the code does, so the allocation tests here measure runtime.MemStats.TotalAlloc instead, and testing.B.Loop is unimplemented there, so the benchmarks use b.N.

Bounds and refusals

  • The buffer grows only to hold one token or one capture, never past Options.MaxBufferBytes (1 MiB by default). A larger token is ErrTooLarge.
  • Nesting stops at Options.MaxDepth (1024 by default).
  • A DOCTYPE is refused unless Options.AllowDoctype is set; an internal subset is refused always. There are no external entities to expand.
  • An XML declaration naming an encoding other than UTF-8, or a UTF-16 byte order mark, is ErrEncoding.
  • An end tag that does not match its start tag is a SyntaxError, checked by a 32-bit hash of the name; a malformed attribute is one too. That is the whole of the well-formedness checking done: duplicate attributes, the characters a name or text may contain, and UTF-8 validity are not checked. This is a reader for documents a writer produced, not a validator.
  • Decoding normalizes line ends in text (\r\n and \r to \n) as XML requires; it does not apply the further whitespace normalization XML specifies for attribute values.

Three fuzz targets check that no input panics the reader and that a document reads identically as a byte slice and as a one-byte-at-a-time stream through a buffer too small for any token.

Not in scope

Writing: this reader serves a viewer, which produces nothing. Also reflection-based mapping, DTDs, and encodings other than UTF-8.

Documentation

Overview

Package xml is a pull reader for XML documents shaped like the parts of an Office Open XML package: one encoding, no DTD, deep repetition of a few element shapes, and a caller that wants a handful of the elements and none of the rest.

It is not a replacement for encoding/xml. It does not resolve namespaces, does not decode into arbitrary Go values, and does not accept a DOCTYPE. In exchange the reader allocates nothing in steady state: every name, attribute and text it returns is a slice into its own buffer, valid until the next call that advances the reader, and a caller copies only what it keeps.

Reading

Next advances one token at a time. On a StartElement the caller reads the name and the attributes it wants, then either descends into the children, reads the text content with ElementText, captures the whole subtree with RawElement, or discards it with Skip. NextChild walks the direct children of an element and skips whatever the caller did not consume, so a decoder for one element is a loop over its children with a switch on the name:

el := r.Element()
for {
	ok, err := r.NextChild(el)
	if err != nil || !ok {
		return err
	}
	switch {
	case r.NameIs("v"):
		text, err := r.ElementText()
		...
	case r.NameIs("is"):
		err = inline.DecodeXMLFrom(r)
	}
}

Names and namespaces

Names are matched as written, prefix included: "w:p" is "w:p". Office XML generators emit canonical prefixes, so matching the qualified name is enough for that input and costs nothing. The xmlns attributes are visible through NextAttr for a caller that needs to resolve them.

Text

Text and Attr return the bytes as written, entities included. A Value knows whether it holds any, and unescapes into a caller-owned buffer, so the common case of text with no ampersand is a slice and no copy.

Bounds

The buffer grows only to hold one token, or one capture, and never past Options.MaxBufferBytes; nesting stops at Options.MaxDepth. A document that needs more is refused, not accommodated.

Example
package main

import (
	"fmt"
	"strings"

	"github.com/shibukawa/tinygodriver/encoding/xml"
)

// A cell of a worksheet, read by hand: the attributes it wants, the child
// it wants, and nothing kept as a string.
type cell struct {
	Ref    string
	Style  int64
	Shared bool
	Value  float64
	Index  int64
}

func (c *cell) DecodeXMLFrom(r *xml.Reader) error {
	*c = cell{} // the caller reuses one cell, and an absent attribute must not keep the last value
	ref, _ := r.Attr("r")
	c.Ref = ref.String()
	if s, ok := r.Attr("s"); ok {
		c.Style, _ = s.Int()
	}
	if t, ok := r.Attr("t"); ok {
		c.Shared = t.Equal("s")
	}
	for name := range r.Children(r.Element()) {
		if !xml.Equal(name, "v") {
			continue
		}
		v, err := r.ElementText()
		if err != nil {
			return err
		}
		if c.Shared {
			c.Index, err = v.Int()
		} else {
			c.Value, err = v.Float()
		}
		if err != nil {
			return err
		}
	}
	return r.Err()
}

func main() {
	const sheet = `<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<worksheet xmlns="http://schemas.openxmlformats.org/spreadsheetml/2006/main">
  <dimension ref="A1:B2"/>
  <sheetData>
    <row r="1"><c r="A1" t="s"><v>0</v></c><c r="B1" s="2"><v>12.5</v></c></row>
    <row r="2"><c r="A2"><f>B1*2</f><v>25</v></c></row>
  </sheetData>
  <mergeCells count="1"><mergeCell ref="A1:B1"/></mergeCells>
</worksheet>`

	r := xml.NewReader(strings.NewReader(sheet), xml.Options{})
	var c cell
	for k := range r.Tokens() {
		if k != xml.StartElement {
			continue
		}
		switch {
		case r.NameIs("sheetData"):
			for range r.Children(r.Element()) { // each <row>
				for range r.Children(r.Element()) { // each <c>
					if err := r.Decode(&c); err != nil {
						panic(err)
					}
					if c.Shared {
						fmt.Printf("%s style=%d shared string #%d\n", c.Ref, c.Style, c.Index)
					} else {
						fmt.Printf("%s style=%d value=%g\n", c.Ref, c.Style, c.Value)
					}
				}
			}
		case r.NameIs("mergeCell"):
			ref, _ := r.Attr("ref")
			fmt.Printf("merged %s\n", ref)
		}
	}
	if err := r.Err(); err != nil {
		panic(err)
	}
}
Output:
A1 style=0 shared string #0
B1 style=2 value=12.5
A2 style=0 value=25
merged A1:B1

Index

Examples

Constants

View Source
const XMLNamespace = "http://www.w3.org/XML/1998/namespace"

XMLNamespace is the namespace the xml prefix is bound to without a declaration.

Variables

View Source
var (
	// ErrTruncated is returned when the input ends inside a token or inside an
	// open element.
	ErrTruncated = errors.New("xml: unexpected end of input")
	// ErrTooLarge is returned when a token or a capture does not fit in
	// Options.MaxBufferBytes.
	ErrTooLarge = errors.New("xml: token exceeds MaxBufferBytes")
	// ErrTooDeep is returned when nesting exceeds Options.MaxDepth.
	ErrTooDeep = errors.New("xml: nesting exceeds MaxDepth")
	// ErrDoctype is returned for a DOCTYPE unless Options.AllowDoctype is set.
	ErrDoctype = errors.New("xml: DOCTYPE refused")
	// ErrEncoding is returned when the XML declaration names an encoding other
	// than UTF-8.
	ErrEncoding = errors.New("xml: declared encoding is not UTF-8")
	// ErrNotStart is returned by the element-level calls when the reader is
	// not positioned on a StartElement.
	ErrNotStart = errors.New("xml: not positioned on a start element")
)

Functions

func Equal

func Equal(b []byte, s string) bool

Equal reports whether b holds exactly the bytes of s. It allocates on neither compiler: the Go compiler elides the conversion in string(b) == s, but TinyGo does not, and copies b for every comparison, every switch on string(b) and every map index by string(b). Compare names through this, NameIs or Value.Equal when the binary is a TinyGo one.

func Unescape

func Unescape(dst, src []byte) []byte

Unescape appends src to dst with the five predefined entities and numeric character references decoded and line ends normalized to "\n". A reference it does not recognize is copied as written.

Types

type Decodable

type Decodable interface {
	DecodeXMLFrom(r *Reader) error
}

Decodable is implemented by a type that reads itself from an element. The reader is positioned on the element's StartElement when the method is called, and the method returns with the reader on the matching EndElement; NextChild and ElementText both end there.

type Element

type Element struct {
	// contains filtered or unexported fields
}

Element identifies an element the reader has entered, for NextChild. It is a value, held on the caller's stack.

type Kind

type Kind uint8

Kind is the kind of token the reader is positioned on.

const (
	// None is the kind before the first Next and after an error.
	None Kind = iota
	// StartElement is an opening tag. A self-closing tag is a StartElement
	// followed by an EndElement, as encoding/xml reports it.
	StartElement
	// EndElement is a closing tag.
	EndElement
	// Text is character data between tags, as written: entities are not
	// decoded and whitespace is not trimmed.
	Text
	// CData is the content of a CDATA section. It contains no entities.
	CData
	// Comment is the body of a comment.
	Comment
	// ProcInst is a processing instruction, including the XML declaration,
	// whose target is "xml".
	ProcInst
	// Directive is a DOCTYPE, reported only when Options.AllowDoctype is set.
	Directive
	// EOF is reported once at the end of a well-formed document.
	EOF
)

func (Kind) String

func (k Kind) String() string

type Options

type Options struct {
	// BufferSize is the initial buffer, 64 KiB by default. It grows only when
	// one token does not fit, and only up to MaxBufferBytes.
	BufferSize int
	// MaxBufferBytes bounds the buffer, and so the largest single token, the
	// largest ElementText and the largest RawElement. 1 MiB by default.
	MaxBufferBytes int
	// MaxDepth bounds element nesting, 1024 by default.
	MaxDepth int
	// AllowDoctype reports a DOCTYPE as a Directive instead of refusing it.
	// An internal subset is refused either way.
	AllowDoctype bool
}

Options bounds a Reader. The zero value selects the defaults noted on each field.

type Reader

type Reader struct {
	// contains filtered or unexported fields
}

Reader reads XML tokens from an io.Reader or from a byte slice.

Every slice a Reader returns aliases its buffer and is valid until the next call that advances the reader: Next, Skip, NextChild, ElementText, RawElement and Decode. Callers that keep a name or a value copy it.

A Reader is not safe for concurrent use. Reuse one across documents with Reset rather than allocating one per document.

func NewBytesReader

func NewBytesReader(data []byte, opts Options) *Reader

NewBytesReader returns a Reader over data, which it neither copies nor modifies. BufferSize and MaxBufferBytes do not apply: the buffer is data.

func NewReader

func NewReader(src io.Reader, opts Options) *Reader

NewReader returns a Reader over src.

func (*Reader) Attr

func (r *Reader) Attr(name string) (Value, bool)

Attr returns the value of the named attribute of the current StartElement, as written. The start tag was indexed as it was scanned, so this is a comparison against each of the element's few names, not a rescan.

func (*Reader) Children

func (r *Reader) Children(e Element) iter.Seq[[]byte]

Children iterates the direct child elements of e, yielding each one's qualified name with the reader positioned on its StartElement. It is NextChild as a range loop, with the same skipping of whatever the body does not consume, and it ends silently on an error:

for name := range r.Children(r.Element()) {
	switch {
	case xml.Equal(name, "v"):
		text, err := r.ElementText()
		...
	}
}

The name is compared through Equal rather than string(name) because the conversion allocates under TinyGo; see Equal.

if err := r.Err(); err != nil {
	return err
}

A break leaves the reader on the child's StartElement; the element e is then still open, and NextChild or Skip finish it.

func (*Reader) Decode

func (r *Reader) Decode(d Decodable) error

Decode reads the current element into d and checks that d consumed exactly that element.

func (*Reader) Depth

func (r *Reader) Depth() int

Depth reports how many elements are open. On a StartElement it counts that element; on its EndElement it no longer does.

func (*Reader) Element

func (r *Reader) Element() Element

Element returns a handle to the element the reader is in, for NextChild. On a StartElement that is the element itself; anywhere else it is the innermost open element.

func (*Reader) ElementText

func (r *Reader) ElementText() (Value, error)

ElementText advances from a StartElement to its EndElement and returns the element's own character data with entities decoded. Text in child elements is not included; the children are skipped. When the text is one run with no entity the result aliases the buffer; otherwise it is assembled in a scratch buffer the next call reuses.

func (*Reader) Err

func (r *Reader) Err() error

Err returns the error that stopped the reader, or nil. The iterators end silently on an error, so a loop over Tokens or Children checks Err after the loop, as a bufio.Scanner caller checks Scanner.Err.

func (*Reader) Kind

func (r *Reader) Kind() Kind

Kind reports the kind of the current token.

func (*Reader) LocalName

func (r *Reader) LocalName() []byte

LocalName returns the name after the prefix, or the whole name when there is none.

func (*Reader) LookupNamespace

func (r *Reader) LookupNamespace(prefix []byte) (string, bool)

LookupNamespace returns the namespace bound to prefix at the current position. An empty prefix looks up the default namespace. The xml prefix is always bound.

func (*Reader) Name

func (r *Reader) Name() []byte

Name returns the qualified name of a StartElement or EndElement, or the target of a ProcInst, as written.

func (*Reader) NameIs

func (r *Reader) NameIs(s string) bool

NameIs reports whether the qualified name equals s. It allocates nothing.

func (*Reader) Namespace

func (r *Reader) Namespace() string

Namespace returns the namespace the current element's name is in: the one its prefix is bound to, or the default namespace when it has none. It is empty when nothing binds the prefix. Names are still matched as written; this is for the caller that must tell one vocabulary from another under the same prefix, or an unprefixed name under a default namespace from one without.

func (*Reader) Next

func (r *Reader) Next() (Kind, error)

Next advances to the next token and reports its kind. After an error every later call returns the same error.

func (*Reader) NextAttr

func (r *Reader) NextAttr() (name []byte, value Value, ok bool)

NextAttr returns the attributes of the current StartElement in document order, one per call, and reports false after the last. It restarts on each new token.

func (*Reader) NextChild

func (r *Reader) NextChild(e Element) (bool, error)

NextChild advances to the next direct child element of e and reports whether there was one. Anything between children is passed over, and a child the caller did not consume is skipped, so a loop over NextChild ends on the EndElement of e however much of each child it read.

func (*Reader) Offset

func (r *Reader) Offset() int64

Offset reports how many bytes of the document precede the next token.

func (*Reader) Prefix

func (r *Reader) Prefix() []byte

Prefix returns the namespace prefix of the name, or nil when there is none.

func (*Reader) RawElement

func (r *Reader) RawElement() ([]byte, error)

RawElement advances from a StartElement to its EndElement and returns the bytes of the whole element, tags included, as written. The result aliases the buffer, so the element must fit in MaxBufferBytes.

func (*Reader) Reset

func (r *Reader) Reset(src io.Reader)

Reset points the Reader at a new source, keeping its options and its buffer.

func (*Reader) ResetBytes

func (r *Reader) ResetBytes(data []byte)

ResetBytes points the Reader at data, keeping its options. Slices returned before the call still alias the old input.

func (*Reader) Skip

func (r *Reader) Skip() error

Skip advances from a StartElement to its matching EndElement, reading nothing in between. It scans the subtree raw, tracking only tags, quotes and depth, so it neither indexes attributes nor checks that end tags match inside what it skips; the reader is left on the end tag with Name set, exactly as Next would leave it.

func (*Reader) Text

func (r *Reader) Text() Value

Text returns the body of a Text, CData, Comment, ProcInst or Directive token as written. For Text the entities are still encoded; see Value.

func (*Reader) Tokens

func (r *Reader) Tokens() iter.Seq[Kind]

Tokens iterates the remaining tokens, ending at EOF or at an error. The reader is the loop variable's context: inside the body Name, Attr, Text and the element-level calls all refer to the token just yielded.

for k := range r.Tokens() {
	if k == xml.StartElement && r.NameIs("sheetData") {
		r.Decode(&sheet)
	}
}
if err := r.Err(); err != nil {
	return err
}

type SyntaxError

type SyntaxError struct {
	Offset int64
	Msg    string
}

SyntaxError reports malformed input at a byte offset from the start of the document.

func (*SyntaxError) Error

func (e *SyntaxError) Error() string

type Value

type Value []byte

Value is attribute or text content as it appears in the document, entities included. Its methods decode on demand, so content with no ampersand, which is nearly all of it, is never copied. A Value aliases the reader's buffer and is valid until the reader advances.

Decoding replaces the five predefined entities and numeric character references, and normalizes line ends as XML requires of text: "\r\n" and a lone "\r" both become "\n". The further whitespace normalization XML applies to attribute values, tabs and newlines to spaces, is not done; Office writers escape such characters as references, which decoding leaves as the characters they name.

func (Value) AppendTo

func (v Value) AppendTo(dst []byte) []byte

AppendTo appends the decoded content to dst.

func (Value) Bool

func (v Value) Bool() (bool, error)

Bool parses an xsd:boolean: "1", "0", "true" or "false". Office attributes use the digits.

func (Value) Equal

func (v Value) Equal(s string) bool

Equal reports whether the decoded content equals s.

func (Value) EqualFold

func (v Value) EqualFold(s string) bool

EqualFold reports whether the decoded content equals s under ASCII case folding.

func (Value) Float

func (v Value) Float() (float64, error)

Float parses a floating point number. A plain decimal of at most 15 significant digits and at most 22 fraction digits, which is every number a spreadsheet writer emits for a cell, is converted by one exact division and is bit-identical to strconv's answer; anything else goes to strconv.

func (Value) HasEntities

func (v Value) HasEntities() bool

HasEntities reports whether decoding would change the bytes: an entity or character reference, or a carriage return.

func (Value) Int

func (v Value) Int() (int64, error)

Int parses a decimal integer. Numbers carry no entities, so the raw bytes are parsed directly.

func (Value) String

func (v Value) String() string

String returns the decoded content as a new string.

func (Value) Uint

func (v Value) Uint() (uint64, error)

Uint parses a decimal unsigned integer.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL