Documentation
¶
Overview ¶
Package xdm implements the XQuery/XPath Data Model (XDM) that XPath 2.0 and XSLT 2.0 are defined over.
The central difference from XPath 1.0 is that every value is a *sequence* of items, and an item is either a node or a typed atomic value. XPath 1.0 had four types (node-set, string, number, boolean) with implicit coercion everywhere; 2.0 has the full XML Schema datatype hierarchy with explicit promotion rules. Modelling that faithfully here is what lets the rest of the engine avoid the 1.0-style "just call ToString" shortcuts that make 2.0 stylesheets silently produce wrong answers.
Index ¶
- Constants
- Variables
- func CompareDT(a, b *DateTime, implicitTZ int) int
- func ErrCast(format string, args ...any) error
- func ErrType(format string, args ...any) error
- func ErrorCode(err error) string
- func Errorf(code, format string, args ...any) error
- func IsGregorian(t TypeCode) bool
- func IsNCName(s string) bool
- func IsXMLWhitespace(s string) bool
- func LexicalGregorian(dt *DateTime, t TypeCode) string
- func SplitQName(s string) (prefix, local string)
- type Atomic
- func NewAnyURI(s string) *Atomic
- func NewBinary(s string, t TypeCode) *Atomic
- func NewBoolean(v bool) *Atomic
- func NewDateTime(dt *DateTime, t TypeCode) *Atomic
- func NewDecimal(r *big.Rat) *Atomic
- func NewDouble(v float64) *Atomic
- func NewDuration(d *Duration, t TypeCode) *Atomic
- func NewFloat(v float64) *Atomic
- func NewGregorian(dt *DateTime, t TypeCode) *Atomic
- func NewInteger(v int64) *Atomic
- func NewIntegerFromRat(r *big.Rat) *Atomic
- func NewQNameValue(q QName) *Atomic
- func NewString(s string) *Atomic
- func NewUntypedAtomic(s string) *Atomic
- func (a *Atomic) Bool() bool
- func (a *Atomic) DateTimeVal() *DateTime
- func (a *Atomic) Derived() string
- func (a *Atomic) DurationVal() *Duration
- func (a *Atomic) FitsInt64() bool
- func (a *Atomic) Float64() float64
- func (a *Atomic) Int64() int64
- func (a *Atomic) IsNaN() bool
- func (a *Atomic) QName() *QName
- func (a *Atomic) Rat() *big.Rat
- func (a *Atomic) Str() string
- func (a *Atomic) String() string
- func (a *Atomic) TypeName() string
- func (a *Atomic) WithDerived(name string) *Atomic
- type DateTime
- type Duration
- type Error
- type Item
- type Node
- func (n *Node) AddAttr(a *Node)
- func (n *Node) AddNamespace(prefix, uri string)
- func (n *Node) AppendChild(c *Node)
- func (n *Node) Atomize() *Atomic
- func (n *Node) Attr(uri, local string) *Node
- func (n *Node) AttrValue(local string) string
- func (n *Node) ChildElements() []*Node
- func (n *Node) Compare(o *Node) int
- func (n *Node) InScopeNamespaces() map[string]string
- func (n *Node) IsElement(uri, local string) bool
- func (n *Node) LookupPrefix(prefix string) (string, bool)
- func (n *Node) Order() int
- func (n *Node) Position() (line, col int, ok bool)
- func (n *Node) Root() *Node
- func (n *Node) StringValue() string
- func (n *Node) Tree() *Tree
- func (n *Node) TypeName() string
- type NodeKind
- type Opaque
- type ParseOptions
- type QName
- type Sequence
- type Tree
- type TypeCode
Constants ¶
const ( // DefaultMaxDepth is the nesting limit. DefaultMaxDepth = 1000 // DefaultMaxBytes is the source-size limit: 64 MB, far above any // schema or stylesheet and above most real instance documents, while // still bounding what a single parse can be asked to read. DefaultMaxBytes int64 = 64 << 20 // DefaultMaxNodes is the node-count limit. At roughly 200 bytes a node // this bounds a tree to about 2 GB, which is the point of it: the // number is chosen to bound *memory*, and it is the limit that actually // binds on the documents designed to be expensive. DefaultMaxNodes = 10_000_000 )
Limits applied when the corresponding ParseOptions field is zero.
const ( NSXSL = "http://www.w3.org/1999/XSL/Transform" NSXML = "http://www.w3.org/XML/1998/namespace" NSXMLNS = "http://www.w3.org/2000/xmlns/" NSXS = "http://www.w3.org/2001/XMLSchema" NSXSI = "http://www.w3.org/2001/XMLSchema-instance" NSFN = "http://www.w3.org/2005/xpath-functions" NSErr = "http://www.w3.org/2005/xqt-errors" NSSVRL = "http://purl.oclc.org/dsdl/svrl" NSSchema = "http://purl.oclc.org/dsdl/schematron" // NSGoxslt is this engine's extension namespace. Extensions live outside // the fn: namespace so that a stylesheet written for another processor // cannot silently pick one up in place of a standard function, and so // that a stylesheet using them is visibly engine-specific. NSGoxslt = "https://github.com/knroy/go-xml" )
Well-known namespace URIs used throughout the engine.
Variables ¶
var Empty = Sequence(nil)
Empty is the canonical empty sequence.
Functions ¶
func ErrType ¶
ErrType is the XPath type error, XPTY0004. It is returned rather than panicked so that a stylesheet error degrades one transform.
func ErrorCode ¶
ErrorCode returns the spec error code carried by err, or "" if it has none.
It unwraps, so a code survives being wrapped with fmt.Errorf("%w"). Errors produced before this type existed still carry their code as a message prefix, so those are recognised too rather than silently reporting "".
func IsGregorian ¶
IsGregorian reports whether t is one of the five Gregorian types.
func IsNCName ¶
IsNCName reports whether s is an XML non-colonised name: a name with no prefix, which is what an element, attribute or processing-instruction name must be once the prefix has been split off.
It lives here rather than in a consumer because more than one caller needs it, and because the cost of *not* checking is that a computed name reaches the serialiser unvalidated. A name is written to output as-is, so a name holding "><script>" produces markup rather than a name — output that is either malformed or, in HTML, an injected element.
func IsXMLWhitespace ¶
IsXMLWhitespace reports whether s consists entirely of XML whitespace.
XML defines whitespace as exactly four characters: space, tab, carriage return and line feed. Go's strings.TrimSpace uses unicode.IsSpace, which additionally matches U+00A0 (no-break space) and other Unicode separators — so using it to decide whether a text node is "just whitespace" silently deletes a that the author put there deliberately.
func LexicalGregorian ¶
LexicalGregorian returns the canonical lexical form.
func SplitQName ¶
SplitQName splits a lexical QName into prefix and local part. It does not resolve the prefix; resolution needs a namespace context and is done by the caller that has one.
Types ¶
type Atomic ¶
type Atomic struct {
Type TypeCode
// contains filtered or unexported fields
}
Atomic is a typed atomic value.
The representation is a tagged union rather than an interface per type. The evaluator switches on Type constantly — every arithmetic op, comparison and function call — and a type switch across seventeen concrete types in those hot paths costs more than a single integer compare. It also keeps the numeric tower in one place, where the promotion rules are easy to audit.
func NewBinary ¶
NewBinary returns an xs:hexBinary or xs:base64Binary holding the given lexical form.
The value keeps its own type rather than collapsing to xs:string, because the two binary types are inter-convertible: casting hexBinary to base64Binary has to re-encode the underlying octets, and a value that has forgotten which encoding its lexical form uses cannot be decoded.
func NewDateTime ¶
NewDateTime returns a date, time or dateTime atomic value.
func NewDecimal ¶
NewDecimal returns an xs:decimal holding an exact value.
func NewDuration ¶
NewDuration returns a duration atomic value of the given duration type.
func NewFloat ¶
NewFloat returns an xs:float. The value is rounded to float32 precision on construction, because xs:float operations must produce float32 results.
func NewGregorian ¶
NewGregorian returns one of the five Gregorian atomic values.
func NewInteger ¶
NewInteger returns an xs:integer. Integers are held as exact rationals so that they participate in decimal arithmetic without precision loss.
func NewIntegerFromRat ¶
NewIntegerFromRat returns an xs:integer from an exact rational, which must have denominator 1. Used by arithmetic that has already established integrality (idiv, string-length, count).
func NewUntypedAtomic ¶
NewUntypedAtomic returns an xs:untypedAtomic, the type produced by atomising a node in a document that has not been schema-validated.
func (*Atomic) DateTimeVal ¶
DateTimeVal returns the date/time value, or nil.
func (*Atomic) Derived ¶
Derived returns the narrower XML Schema type this value was constructed as, or "" if it was not built by a derived-type constructor.
func (*Atomic) DurationVal ¶
Duration returns the duration value, or nil.
func (*Atomic) FitsInt64 ¶
FitsInt64 reports whether the value can be represented as an int64 without wrapping.
xs:integer is arbitrary-precision, so this is a real question: Int64() truncates the big.Int and silently returns a different number, which is worse than refusing.
func (*Atomic) Float64 ¶
Float64 returns the value as a float64 for any numeric type. Decimal and integer values are converted, which may lose precision; callers doing exact arithmetic must use Rat instead.
func (*Atomic) IsNaN ¶
IsNaN reports whether a is a double or float NaN. NaN needs its own check throughout comparison, because it is the one value where the general "compare and negate" shortcut produces wrong answers.
func (*Atomic) String ¶
String returns the XPath 2.0 canonical lexical representation, which is what fn:string and every implicit string conversion must produce. It is not a debug format: the exact spelling of doubles and decimals here is observable in stylesheet output.
func (*Atomic) WithDerived ¶
WithDerived returns a copy of a annotated as the named derived type.
type DateTime ¶
type DateTime struct {
Year int // proleptic Gregorian; negative for BCE. No year zero.
Month int // 1-12
Day int // 1-31
Hour int // 0-24 (24 only as the lexical form 24:00:00)
Minute int // 0-59
Second *big.Rat // seconds including fraction, [0,60)
// TZOffset is the timezone offset in minutes east of UTC.
// HasTZ distinguishes "no timezone" from "+00:00", which are different
// values under XML Schema equality.
TZOffset int
HasTZ bool
}
DateTime represents xs:date, xs:time and xs:dateTime.
It is not time.Time. XML Schema dates carry three properties that time.Time cannot express: an optional timezone (distinct from UTC — an unzoned date is a different value from a UTC one), a year range that exceeds int64 nanoseconds, and second values with arbitrary fractional precision. Comparison of unzoned values against zoned ones is defined against an implicit timezone supplied by the dynamic context, which only works if "absent" is representable.
func ParseDateTime ¶
ParseDateTime parses the lexical form of xs:date, xs:time or xs:dateTime according to the requested type.
func ParseGregorian ¶
ParseGregorian parses the lexical form of one of the five Gregorian types.
Each has its own leading-hyphen convention — "--01" is a month, "---15" a day — which exists so that the forms cannot be confused with a truncated date. Getting the hyphen count wrong silently reinterprets the value, so each form is matched exactly rather than by a permissive scan.
func (*DateTime) ToSeconds ¶
ToSeconds returns the value as seconds since 1972-12-31T00:00:00Z, adjusted to UTC using implicitTZ (in minutes) when the value carries no timezone.
Comparison and subtraction are defined on this normalised timeline, so having one conversion point means the timezone rules are applied uniformly rather than re-derived at each comparison site.
type Duration ¶
type Duration struct {
Negative bool
Months int // years*12 + months
Seconds *big.Rat // days*86400 + hours*3600 + minutes*60 + seconds
}
Duration represents xs:duration and its two subtypes.
XML Schema durations have two independent components — months and seconds — that cannot be converted into one another, because the number of days in a month is not fixed. That is why xs:duration is only partially ordered and why the two totally-ordered subtypes (xs:yearMonthDuration and xs:dayTimeDuration) exist. Keeping the components separate rather than normalising to a single scalar is what makes the ordering rules implementable at all.
func ParseDuration ¶
ParseDuration parses the lexical form of xs:duration, xs:yearMonthDuration or xs:dayTimeDuration, rejecting components the requested subtype does not permit.
func (*Duration) SignedMonths ¶
SignedMonths returns the month component with the sign applied.
func (*Duration) SignedSeconds ¶
SignedSeconds returns the second component with the sign applied.
type Error ¶
type Error struct {
// Code is the spec error code, such as "XPTY0004". Codes live in the
// http://www.w3.org/2005/xqt-errors namespace; the local name alone is
// carried here because it is unique across the specs and is how the
// documents themselves refer to them.
Code string
// Message is the human-readable detail, without the code prefix.
Message string
// Err is an underlying cause, if any.
Err error
}
Error is an XPath, XQuery or XSLT error carrying its specification error code.
The specs define a code for every error condition — XPTY0004 for a type error, FORG0001 for a failed cast, FODC0002 for an unretrievable document — and those codes are the stable, translatable part of an error. A message is prose that may be reworded; a code is what a caller can branch on and what a conformance suite compares.
The codes were already present as string prefixes on every error this engine produces, which reads correctly but cannot be inspected: a caller wanting to distinguish "the document was malformed" from "the stylesheet is wrong" had to match on substrings. This type makes the code a field while keeping the rendered message byte-identical, so nothing that reads error text changes.
type Item ¶
type Item interface {
// TypeName returns the QName of the item's type, for error messages and
// instance-of tests.
TypeName() string
// contains filtered or unexported methods
}
Item is a single member of a sequence: either a Node or an atomic value.
The interface is closed to outside implementations (unexported marker method). XDM defines exactly these two kinds of item in XPath 2.0; function items arrive in 3.0 and would be added here.
type Node ¶
type Node struct {
Kind NodeKind
Name QName
// Value is the text content for text, comment, PI and attribute nodes,
// and the namespace URI for namespace nodes. Element and document nodes
// derive their string value from descendants; see StringValue.
Value string
Parent *Node
Children []*Node
// Attrs and Namespaces hold attribute and namespace nodes for elements.
// They are kept out of Children because the child axis must not return
// them — a fact that a single mixed slice makes easy to get wrong.
Attrs []*Node
Namespaces []*Node
// BaseURI is the resolved base URI, used by fn:document and fn:doc.
BaseURI string
// TypeAnnotation records a schema type when the document has been
// validated. Untyped documents leave this empty, and atomisation then
// yields xs:untypedAtomic, which is the schemaless default.
TypeAnnotation string
// contains filtered or unexported fields
}
Node is a node in an XDM tree.
This is a concrete struct rather than an interface. Every node kind shares most of its fields, the evaluator switches on Kind rather than dispatching, and the axes need to walk parent/sibling links tens of thousands of times per document — an interface would add a pointer chase and an indirect call to each step for no expressiveness gained.
Trees are built by the parser in this package and are immutable afterwards. That immutability is what makes it safe to share one compiled stylesheet tree across concurrent transforms.
func (*Node) AddNamespace ¶
AddNamespace links a namespace node to n.
func (*Node) AppendChild ¶
AppendChild links c as the last child of n, setting the parent link. It does not assign document order; call Finalize once the tree is complete.
func (*Node) Atomize ¶
Atomize returns the typed value of a node. Without schema validation every node atomises to xs:untypedAtomic, which is what makes untyped comparison rules apply throughout a schemaless transform.
func (*Node) AttrValue ¶
AttrValue returns the value of a no-namespace attribute, or "". Most attributes the stylesheet compiler reads (match, select, name, test) are unprefixed, so this is the common case worth a helper.
func (*Node) ChildElements ¶
ChildElements returns the element children, which is what almost every stylesheet-compilation walk wants.
func (*Node) Compare ¶
Compare orders two nodes in document order, returning -1, 0 or 1. Nodes in different trees are ordered by tree id, which is stable within a transform.
func (*Node) InScopeNamespaces ¶
InScopeNamespaces returns every prefix-to-URI binding visible at n, with inner declarations shadowing outer ones. Used when copying elements and when resolving QNames in stylesheet attribute values.
func (*Node) LookupPrefix ¶
LookupPrefix resolves a namespace prefix against the in-scope namespaces of n, walking up the tree. Returns the URI and whether the prefix was bound.
func (*Node) Order ¶
Order returns the document-order index. Only meaningful against nodes from the same tree; use Compare for the general case.
func (*Node) Position ¶
Position returns the 1-based line and column where the node starts, and false if the position is unknown — the node was built by a transform rather than parsed, or the source text was not retained.
func (*Node) Root ¶
Root returns the root of the containing tree, walking parent links. For a well-formed parsed document this is the document node.
func (*Node) StringValue ¶
StringValue returns the node's string value per XDM: the concatenation of all descendant text for document and element nodes, and the value itself for the leaf kinds.
type Opaque ¶
type Opaque struct {
// Label names the kind of value, for error messages.
Label string
// Value is the wrapped payload.
Value any
}
Opaque wraps an arbitrary Go value as an Item.
It exists so that layers above this package can thread their own state through an evaluation context, which binds sequences rather than typed fields. The XSLT engine uses it for the transform runtime and grouping state, which the xpath package cannot name without an import cycle.
An Opaque is not a legal XDM value: it has no string value, does not atomise, and must never reach a stylesheet. Every producer binds it under a reserved namespace that no stylesheet can spell.
type ParseOptions ¶
type ParseOptions struct {
// BaseURI is recorded on the document node and used to resolve relative
// references in fn:document and xsl:include.
BaseURI string
// StripSpace removes whitespace-only text nodes. XSLT applies this per
// element name via xsl:strip-space, so the transform layer passes a
// predicate; a plain bool here would not express "strip in these elements
// only".
StripSpace func(elem QName) bool
// AllowDOCTYPE permits a DOCTYPE declaration. It defaults to false: a
// DOCTYPE is the entry point for both XXE (parser-executed file:// and
// http:// reads) and entity-expansion blowup, and a validator that
// happily expands entities from untrusted input is a liability. Callers
// that genuinely need DTD-declared entities opt in explicitly.
AllowDOCTYPE bool
// TrackPositions records where each element starts, so that a validator
// can report the line a failure occurred on. It retains the source text
// for the life of the tree, which measures at about 10% more memory on a
// typical invoice and no extra parse time. It is opt-in because that cost
// buys nothing for a caller that never asks for a position.
TrackPositions bool
// MaxDepth bounds nesting. Deeply nested input is the cheapest way to
// drive a recursive descent into stack exhaustion, so the limit is
// enforced during construction rather than left to the runtime.
MaxDepth int
// MaxBytes bounds the source document. Zero means DefaultMaxBytes;
// a negative value means no limit, for a caller reading input it
// produced itself.
MaxBytes int64
// MaxNodes bounds the tree. Zero means DefaultMaxNodes; a negative
// value means no limit.
//
// Both limits exist because neither alone is a memory bound. A node
// costs a fixed ~200 bytes whatever it contains, so the heap a document
// needs depends on how many nodes it has rather than how long it is:
// a megabyte of "<a/>" is fifty times the memory of a megabyte of text.
// MaxBytes bounds the read; MaxNodes bounds what the read can allocate.
MaxNodes int
// contains filtered or unexported fields
}
ParseOptions controls document construction.
type QName ¶
QName is an expanded name: namespace URI plus local part, with the prefix retained only for serialisation.
Equality in XPath is defined on (URI, Local) alone — the prefix is not part of the value — so Equal deliberately ignores Prefix. Keeping the prefix around anyway matters because a literal result element must be serialised with the prefix the stylesheet author wrote, not one we invent.
func (QName) Clark ¶
Clark returns the {uri}local form, which is unambiguous without a namespace context and is therefore what error messages and map keys use.
type Sequence ¶
type Sequence []Item
Sequence is an ordered list of items. The empty sequence is a nil or zero-length slice; both are treated identically by every operation, so callers never have to normalise before comparing.
A sequence is flat: XDM has no nested sequences, and every constructor in this package maintains that invariant.
func Atomize ¶
Atomize converts a sequence to atomic values, replacing each node with its typed value. This is the fn:data() operation, applied implicitly wherever XPath 2.0 requires atomic operands.
Every item in the result is an *Atomic. Callers rely on that: two dozen of them assert the type without checking, because within the data model there is nothing else atomisation can produce.
Opaque items are the exception, and they are dropped here rather than passed through. They carry engine-internal state — the transform runtime, grouping bookkeeping — through the closed Item interface, and a stylesheet that names the internal namespace could reach one:
xmlns:gi="urn:goxslt:internal" ... distinct-values($gi:runtime)
Passing it through made that expression panic with an interface-conversion error, which in a server embedding this engine is a denial of service triggered by stylesheet text. An Opaque has no typed value, so dropping it is also what the data model implies: it is not a node and not an atomic value, so fn:data has nothing to return for it.
func One ¶
One wraps a single item as a sequence. Named for how often it is needed: most XPath operations produce exactly one item and must still return a sequence.
func SortDocumentOrder ¶
SortDocumentOrder sorts a sequence of nodes into document order and removes duplicates.
Every path expression in XPath 2.0 returns nodes in document order with duplicates removed, and so do the union, intersect and except operators. Doing it in one place means the axis implementations can emit nodes in whatever order is natural for them (reverse axes emit backwards) without each having to re-sort.
Items that are not nodes are an error at the call sites that use this, so they are passed through unsorted rather than silently dropped; the caller type-checks first.
func (Sequence) First ¶
First returns the first item, or nil if the sequence is empty. Callers that require exactly one item should use Single instead so that a length > 1 is reported rather than silently truncated.
func (Sequence) Single ¶
Single returns the sole item of a one-item sequence. It reports an error for any other length, because the places that call it (operands of arithmetic, the argument of a function declared to take exactly one item) are precisely the places where XPath 2.0 raises XPTY0004 rather than coercing.
type Tree ¶
type Tree struct {
Root *Node
// DocType is the DOCTYPE declaration's text, when the document had one
// and AllowDOCTYPE permitted it. Empty otherwise.
//
// It is retained because the internal subset is the only place a
// document's own DTD lives, and validating against it needs the text —
// encoding/xml hands the declaration over as one opaque token and keeps
// nothing. The dtd package parses it; this package applies only the two
// declarations whose absence is visible in the data model.
DocType string
// contains filtered or unexported fields
}
Tree owns a document and the counter used to assign document order.
func Parse ¶
func Parse(r io.Reader, opts ParseOptions) (*Tree, error)
Parse builds an XDM tree from an XML document.
It uses encoding/xml as a tokeniser only. The Go decoder's own namespace handling is not usable here: it resolves prefixes into Name.Space but discards the prefix and the declarations themselves, and XSLT needs both — namespace nodes are addressable on the namespace axis, and a literal result element must be serialised with the prefix the author wrote.
func ParseString ¶
func ParseString(s string, opts ParseOptions) (*Tree, error)
ParseString is Parse over a string, which is what most tests and the stylesheet compiler want.
type TypeCode ¶
type TypeCode int
TypeCode identifies an atomic type from the XML Schema built-in hierarchy.
Only the types XPath 2.0 gives special treatment are enumerated. The rest of the schema hierarchy (xs:token, xs:NMTOKEN, and the other string subtypes) behaves identically to its base type for every operation this engine performs, so carrying them as distinct codes would add branches with no behavioural difference.
const ( // TypeUntypedAtomic is the type of atomised nodes in a schemaless // document. It is the reason XPath 2.0 needs so few explicit casts: an // untypedAtomic operand is converted to the required type at the point of // use, but *only* in the specific contexts the spec lists. TypeUntypedAtomic TypeCode = iota TypeString TypeBoolean TypeDecimal TypeInteger TypeDouble TypeFloat TypeQName TypeAnyURI TypeDate TypeTime TypeDateTime TypeDuration TypeYearMonthDuration TypeDayTimeDuration TypeHexBinary TypeBase64Binary // The five Gregorian types denote a recurring or partial calendar point: // a year, a year and month, a month, a month and day, or a day. TypeGYear TypeGYearMonth TypeGMonth TypeGMonthDay TypeGDay )
func NumericPromote ¶
NumericPromote returns the common type for a binary numeric operation, per the XPath 2.0 promotion lattice: integer -> decimal -> float -> double. Both operands are converted to that type before the operation runs.