Documentation
¶
Overview ¶
Package xpath implements XPath 2.0: lexing, parsing to an AST, static analysis, and evaluation against an XDM tree.
Index ¶
- Constants
- func CastAtomic(a *xdm.Atomic, target xdm.TypeCode) (*xdm.Atomic, error)
- func CastToDerived(a *xdm.Atomic, target xdm.TypeCode, facet string) (*xdm.Atomic, error)
- func CompileRegexp(pattern, flags string) (*regexp.Regexp, error)
- func EffectiveBooleanValue(seq xdm.Sequence) (bool, error)
- func Eval(src string, ctx *Context, ns NamespaceResolver) (xdm.Sequence, error)
- func RegisterCollation(uri string, c Collation)
- func RegisterXSLTFuncs(l *Library)
- func TranslateSchemaRegexp(pattern string) (string, error)
- func TranslateSchemaRegexpVersion(pattern string, xsd11 bool) (string, error)
- type Axis
- type BinaryOp
- type Binding
- type CastExpr
- type Collation
- type CollectionResolver
- type Compiled
- type Context
- func (c *Context) ContextNode() (*xdm.Node, error)
- func (c *Context) Descend() (*Context, error)
- func (c *Context) Err() error
- func (c *Context) LookupVar(name xdm.QName) (xdm.Sequence, bool)
- func (c *Context) WithFocus(item xdm.Item, pos, size int) *Context
- func (c *Context) WithNow(t time.Time) *Context
- func (c *Context) WithVar(name xdm.QName, val xdm.Sequence) *Context
- type ContextItem
- type DocumentResolver
- type Expr
- type FilterExpr
- type ForExpr
- type FuncCall
- type Function
- type FunctionLibrary
- type IfExpr
- type InstanceOfExpr
- type KindTest
- type Lexer
- type Library
- type Literal
- type NameTest
- type NamespaceResolver
- type NodeTest
- type Parser
- type PathExpr
- type QuantifiedExpr
- type SequenceExpr
- type SequenceType
- type Step
- type Token
- type TokenKind
- type TreatExpr
- type UnaryOp
- type VarRef
Constants ¶
const CodepointCollation = "http://www.w3.org/2005/xpath-functions/collation/codepoint"
CodepointCollation is the one collation this implementation provides. It is the only collation the spec requires every processor to support, and it compares strings by Unicode codepoint.
const HTMLASCIICaseInsensitive = "http://www.w3.org/2005/xpath-functions/collation/html-ascii-case-insensitive"
HTMLASCIICaseInsensitive is the collation that compares ASCII letters without regard to case. It is defined by the spec and is the one non-codepoint collation that needs no locale data.
const MaxDepth = 500
MaxDepth bounds recursive evaluation.
const MaxItems = 5_000_000
MaxItems bounds the number of items an evaluation may materialise.
Depth bounds the stack and Ctx bounds the wall clock, but neither bounds memory: "count(1 to 9999999)" is one shallow, fast expression that allocates nine million *Atomic values and peaked at 1.8 GB of resident memory. The range operator had its own limit, but "for $a in 1 to 3000, $b in 1 to 3000" walked straight past it, because the sequence is built by the for-expression rather than by the range.
The bound is deliberately generous: a real stylesheet over a large document works in thousands of nodes, not tens of millions, so this only fires on input designed to exhaust memory or on a genuine runaway.
Variables ¶
This section is empty.
Functions ¶
func CastAtomic ¶
CastAtomic converts an atomic value to a target type, per the XPath 2.0 casting table.
Casting is stricter than the string conversions of XPath 1.0: "abc" cast to xs:double is an error (FORG0001), not NaN. Silent NaN is how a validator ends up reporting a document as valid because a numeric comparison quietly became false.
func CastToDerived ¶
CastToDerived casts to a derived type named by its local name, applying the facet that the type code cannot carry. An unknown name falls back to a plain cast to the primitive.
func CompileRegexp ¶
CompileRegexp exposes the XPath-to-Go regular expression translation for the XSLT layer, which needs it for xsl:analyze-string. The compiled result is cached exactly as it is for fn:matches.
func EffectiveBooleanValue ¶
EffectiveBooleanValue computes fn:boolean over a sequence.
XPath 2.0 restricts this compared with 1.0: a sequence of two or more atomic values raises FORG0006 rather than being truthy. That strictness is deliberate — it catches "if ($seq)" where the author meant "if (exists($seq))" — so it is enforced rather than relaxed.
func Eval ¶
Eval is a one-shot compile-and-evaluate, for callers that will not reuse the expression.
func RegisterCollation ¶
RegisterCollation makes a collation available under a URI.
The spec requires exactly two collations — codepoint and the HTML ASCII case-insensitive one — and leaves the rest implementation-defined. This is how a host application supplies its own: a locale-aware comparison, or the case-blind collation the W3C test catalog defines for its own use.
Registering a URI that is already known replaces it, which is what makes a host override sensible. Registration is expected during setup, before any evaluation; it is guarded so that a late one cannot race a lookup, but it is not a way to change collations while expressions are running.
func RegisterXSLTFuncs ¶
func RegisterXSLTFuncs(l *Library)
RegisterXSLTFuncs adds the functions that XSLT 2.0 defines but XPath 2.0 does not.
fn:unparsed-text, fn:format-date and friends are in the XSLT specification, not the XPath one — a bare XPath 2.0 processor is required to report XPST0017 for them. This engine is an XSLT engine, so they must exist when a stylesheet is running, and must not when a plain XPath expression is being evaluated against the XPath library alone. Keeping them out of Builtins and adding them here is what makes both true.
func TranslateSchemaRegexp ¶
TranslateSchemaRegexp rewrites an XML Schema regular expression into RE2 syntax, without anchoring it.
The XML Schema flavour and the XPath flavour share a grammar — Part 2 Appendix F defines the one, and XPath's fn:matches extends it — so the translation is the same in both directions: the multi-character escapes \i and \c, the block and category escapes, and character class subtraction, none of which RE2 accepts as written.
The result is deliberately unanchored, because the two flavours differ exactly there. fn:matches is a containment test, while a pattern facet must span the whole value. A caller using this for a pattern facet has to wrap the result — see the xsd package, which does so with \A(?:...)\z.
func TranslateSchemaRegexpVersion ¶
TranslateSchemaRegexpVersion is TranslateSchemaRegexp with the one grammar rule that XSD 1.1 changed made selectable.
1.1 stopped treating an unrecognised \p{Is...} block name as an error and began reading it as a class that matches every character, so the same pattern is invalid under 1.0 and valid under 1.1. reK88 asserts exactly that pair, which is why the version has to reach the grammar check.
Types ¶
type Axis ¶
type Axis int
Axis identifies one of the thirteen XPath axes.
func (Axis) IsReverse ¶
IsReverse reports whether the axis is a reverse axis. Reverse axes number their positions backwards, which changes what position() means inside a predicate — the one place the distinction is observable.
func (Axis) PrincipalKind ¶
PrincipalKind is the node kind an axis selects when the node test is a name or wildcard: attributes on the attribute axis, namespaces on the namespace axis, elements everywhere else.
type BinaryOp ¶
BinaryOp is any infix operator. Keeping them in one node with an Op field rather than one type per operator keeps the parser's precedence ladder short and puts all the operand-conversion rules in one evaluator function, which is where they are easiest to check against the spec's tables.
type CastExpr ¶
type CastExpr struct {
Operand Expr
Type SequenceType
Castable bool // true for "castable as", which yields a boolean
}
CastExpr is "expr cast as type" and "expr castable as type".
type Collation ¶
type Collation interface {
Compare(a, b string) int
Contains(s, sub string) bool
StartsWith(s, prefix string) bool
EndsWith(s, suffix string) bool
IndexOf(s, sub string) int
}
Collation compares two strings. Only the operations XPath actually needs are exposed, because a general Compare is not enough: fn:contains under a case-insensitive collation is not "compare the folded strings", it is "does the folded needle occur in the folded haystack".
func ResolveCollation ¶
ResolveCollation returns the collation a URI names.
A relative URI is accepted when its tail matches a known collation, because the QT3 suite and some stylesheets write "collation/codepoint" rather than the full URI. Resolving it against the static base URI would be more correct still, but the base is not threaded into every function that takes a collation argument, and matching the tail covers the forms that occur.
type CollectionResolver ¶
type CollectionResolver interface {
// ResolveCollection returns the documents in uri, resolved against base.
//
// The result is a sequence rather than a []*xdm.Tree because a collection
// is permitted to contain items that are not document nodes.
ResolveCollection(uri, base string) (xdm.Sequence, error)
}
CollectionResolver loads a named set of documents for fn:collection.
It is deliberately separate from DocumentResolver rather than an extra method on it. A caller who wants fn:doc for the code lists shipped beside a stylesheet does not thereby want fn:collection to enumerate a directory, and folding the two together would make enabling one enable the other.
The empty uri is the default collection — fn:collection() with no argument. A resolver that has no default should return an error for it rather than an empty sequence, for the reason given on fnCollection.
type Compiled ¶
type Compiled struct {
// contains filtered or unexported fields
}
Compiled is a parsed XPath expression, ready to evaluate against any context.
Compiling once and evaluating many times is the intended usage: parsing dominates the cost of a short expression, and a stylesheet evaluates the same expression once per node. A Compiled value is immutable and safe for concurrent use.
func Compile ¶
func Compile(src string, ns NamespaceResolver) (*Compiled, error)
Compile parses src, resolving namespace prefixes with ns.
func MustCompile ¶
func MustCompile(src string, ns NamespaceResolver) *Compiled
MustCompile is Compile, panicking on error. For tests and for expressions that are literals in this package's own source.
func (*Compiled) EvalString ¶
EvalString evaluates and returns the string value of the result, which is the concatenation rule of fn:string applied to the first item, or "" for the empty sequence.
type Context ¶
type Context struct {
// Item is the context item. It is nil where there is no context item,
// which is an error to reference rather than an empty sequence.
Item xdm.Item
// Position is the context position, 1-based. Zero means "no focus".
Position int
// Size is the context size.
Size int
// Vars holds in-scope variable bindings, keyed by expanded name.
// Lookups walk to Parent, so a nested scope does not copy the map.
Vars map[string]xdm.Sequence
Parent *Context
// Funcs resolves function calls. Supplied by the caller so that XSLT can
// add xsl:function declarations and extension functions without this
// package knowing about them.
Funcs FunctionLibrary
// StaticBaseURI is the base URI of the expression itself — the stylesheet
// or query it was written in — which is what fn:static-base-uri returns
// and what fn:resolve-uri resolves against by default.
//
// It is distinct from a *node's* base URI, which comes from the document
// the node was parsed from. Returning the context node's was the nearest
// thing available before this existed, and it is a different value: a
// stylesheet in one place can perfectly well be applied to a document
// from another.
StaticBaseURI string
// ImplicitTimezone is the offset in minutes applied to date/time values
// that carry no timezone. The spec requires the dynamic context to supply
// one; defaulting to UTC keeps results reproducible across machines,
// which matters more for a validator than matching local time.
ImplicitTimezone int
// Ctx carries cancellation. A stylesheet can loop for a long time on
// pathological input, and the caller needs a way out that does not
// involve killing the process.
Ctx context.Context
// Docs resolves fn:doc and fn:document URIs. Nil disables them, which is
// the safe default: a stylesheet that can open arbitrary URIs is an SSRF
// and file-disclosure vector.
Docs DocumentResolver
// Collections resolves fn:collection URIs. Nil disables it, for the same
// reason nil disables Docs, and setting Docs does not set this: see
// CollectionResolver.
Collections CollectionResolver
// Depth guards against unbounded recursion in user-defined functions and
// named templates, which the spec does not bound.
Depth int
// Now is the value fn:current-dateTime and its siblings return.
//
// The spec requires these to be stable for the whole of one evaluation:
// calling current-dateTime() twice must give the same answer, or a
// stylesheet that stamps a document and then checks the stamp against
// "now" can disagree with itself. Reading the clock here once, rather
// than per call, is what guarantees that. A zero value means the caller
// did not set one and the functions are unavailable.
Now time.Time
// HasNow distinguishes an unset clock from a legitimately zero time.
HasNow bool
// contains filtered or unexported fields
}
Context is the XPath dynamic context: everything an expression can observe beyond its own AST.
The focus (item, position, size) changes on every step and predicate, while the rest (variables, functions, the implicit timezone) changes rarely. They are kept in one struct anyway, copied cheaply by value in the hot paths, because splitting them means every evaluator function takes two parameters and the copy is a handful of words either way.
func NewContext ¶
func NewContext(item xdm.Item, funcs FunctionLibrary) *Context
NewContext returns a context with the given focus and library.
func (*Context) ContextNode ¶
ContextNode returns the context item as a node, or an error when there is no context item or it is an atomic value.
Steps require a node context; the distinct error codes matter because XPDY0002 (absent) and XPTY0020 (present but not a node) mean different things to a stylesheet author.
func (*Context) Descend ¶
Descend returns a copy with the recursion depth incremented, erroring past the limit.
func (*Context) LookupVar ¶
LookupVar resolves a variable by expanded name, walking enclosing scopes.
func (*Context) WithFocus ¶
WithFocus returns a copy of ctx with a new context item, position and size, sharing the variable scope.
This is the operation performed once per node per step. It copies the struct rather than allocating a child scope, so variable lookups still resolve through the same maps without a new one being built.
The copy itself does allocate — it is the largest single allocation site in the engine, around a quarter of what a stylesheet render allocates. Reusing one context across a step loop was measured and made no difference at all (4,963,596 vs 4,964,187 bytes per render), so it was reverted: WithVar builds children holding a pointer back to this context, and the aliasing risk that reuse introduces buys nothing. Anyone tempted to try it again should measure first.
func (*Context) WithVar ¶
WithVar returns a child context binding name to val.
A child scope with its own one-entry map is used rather than mutating the parent's, because a for-expression binds a fresh value per iteration while the body may capture it; mutation would make all iterations observe the last value.
type ContextItem ¶
type ContextItem struct{}
ContextItem is the "." expression.
func (*ContextItem) Eval ¶
func (e *ContextItem) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for the context item.
func (*ContextItem) String ¶
func (e *ContextItem) String() string
type DocumentResolver ¶
type DocumentResolver interface {
// ResolveDocument returns the tree for uri, resolved against base.
ResolveDocument(uri, base string) (*xdm.Tree, error)
}
DocumentResolver loads a document by URI for fn:doc and fn:document.
type Expr ¶
type Expr interface {
// Eval evaluates the expression in ctx and returns a sequence.
Eval(ctx *Context) (xdm.Sequence, error)
// String returns a source-like rendering, used in error messages and to
// make test failures readable.
String() string
}
Expr is a node in the XPath abstract syntax tree.
Evaluation is a method on the AST rather than a separate visitor. XPath evaluation is a simple recursive walk with no multi-pass analysis, so a visitor would add an indirection layer without buying anything; the one place a second pass would help (static typing) is not implemented, and the spec permits a dynamically-typed implementation.
type FilterExpr ¶
FilterExpr applies predicates to an arbitrary expression, as in "(1 to 10)[. mod 2 = 0]".
func (*FilterExpr) Eval ¶
func (e *FilterExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for a filtered expression.
func (*FilterExpr) String ¶
func (e *FilterExpr) String() string
type ForExpr ¶
ForExpr is "for $x in seq return expr".
type FuncCall ¶
FuncCall is a function call. Resolution happens at evaluation time against the context's function library, so that a stylesheet's own xsl:function declarations are visible without a separate binding pass.
type Function ¶
type Function struct {
Name xdm.QName
Arity int
// Call receives the already-evaluated arguments. Functions that need the
// context item (fn:string with no argument, fn:position) read it from ctx.
Call func(ctx *Context, args []xdm.Sequence) (xdm.Sequence, error)
}
Function is a callable XPath function.
type FunctionLibrary ¶
type FunctionLibrary interface {
// Lookup returns the function with the given name and arity.
Lookup(name xdm.QName, arity int) (Function, bool)
}
FunctionLibrary resolves and calls functions.
func Builtins ¶
func Builtins() FunctionLibrary
Builtins returns the standard fn: function library.
The library is built once and shared: Function values hold no mutable state (everything they need comes from the Context passed at call time), so a single instance is safe for concurrent transforms. Rebuilding it per transform would cost several hundred map inserts for no benefit.
type IfExpr ¶
type IfExpr struct {
Cond, Then, Else Expr
}
IfExpr is "if (cond) then a else b". Both branches are required by the grammar; there is no one-armed form.
type InstanceOfExpr ¶
type InstanceOfExpr struct {
Operand Expr
Type SequenceType
}
InstanceOfExpr is "expr instance of type".
func (*InstanceOfExpr) Eval ¶
func (e *InstanceOfExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for "instance of".
func (*InstanceOfExpr) String ¶
func (e *InstanceOfExpr) String() string
type KindTest ¶
type KindTest struct {
Kind xdm.NodeKind
// Any matches every kind: the node() test.
Any bool
// Name constrains element()/attribute()/processing-instruction() tests
// that name a target.
Name *xdm.QName
HasName bool
// Content constrains the root element of a document-node() test:
// document-node(element(invoice)) matches only a document whose element
// child satisfies the inner test. Nil means the document's content is
// unconstrained.
Content NodeTest
}
KindTest matches by node kind: text(), comment(), node(), element(name), and so on.
type Lexer ¶
type Lexer struct {
// contains filtered or unexported fields
}
Lexer turns XPath source into tokens.
XPath 2.0's grammar is not context-free at the lexical level: whether `*` means multiplication or "any element", and whether `div`, `and`, `is` and friends are operators or element names, depends on what preceded them. The spec resolves this with a rule stated in terms of the previous token, and that is what prevOperand tracks. Trying to decide these in the parser instead means the lexer must emit ambiguous tokens and the parser must re-lex, which is worse.
type Library ¶
type Library struct {
// Parent is consulted when a name is not found locally, so a stylesheet's
// own functions can shadow and extend the builtins without copying them.
Parent FunctionLibrary
// contains filtered or unexported fields
}
Library is a mutable function library keyed by expanded name and arity.
Arity is part of the key because XPath overloads on it: fn:string() and fn:string($arg) are different functions, and fn:substring has both a two- and a three-argument form with different behaviour.
func NewLibrary ¶
func NewLibrary(parent FunctionLibrary) *Library
NewLibrary returns an empty library chained to parent.
type Literal ¶
Literal is a constant atomic value.
type NameTest ¶
type NameTest struct {
// Name is the expanded name to match. Wildcards leave one or both parts
// unconstrained; see AnyURI and AnyLocal.
Name xdm.QName
// AnyURI matches any namespace ("*" and "*:local").
AnyURI bool
// AnyLocal matches any local name ("*" and "prefix:*").
AnyLocal bool
}
NameTest matches by expanded name.
type NamespaceResolver ¶
type NamespaceResolver interface {
// ResolvePrefix returns the URI bound to prefix, or false if unbound.
ResolvePrefix(prefix string) (string, bool)
// DefaultElementNamespace returns the namespace applied to unprefixed
// element name tests. XSLT sets this from xpath-default-namespace; it is
// empty by default, and it never applies to attribute names or function
// names.
DefaultElementNamespace() string
// DefaultFunctionNamespace returns the namespace for unprefixed function
// names, which is the fn: namespace in XPath and XSLT.
DefaultFunctionNamespace() string
}
NamespaceResolver resolves a namespace prefix to a URI at parse time.
Prefixes must be resolved when the expression is compiled, not when it runs: an XPath expression in a stylesheet is bound to the namespace declarations in scope at the point it appears, and by evaluation time the relevant element is long out of view.
type NodeTest ¶
type NodeTest interface {
// Matches reports whether n is selected on an axis whose principal node
// kind is principal.
Matches(n *xdm.Node, principal xdm.NodeKind) bool
String() string
}
NodeTest decides whether a node on an axis is selected.
type Parser ¶
type Parser struct {
// contains filtered or unexported fields
}
Parser builds an AST from tokens.
type PathExpr ¶
type PathExpr struct {
// Root marks a path that starts at the document root ("/foo" rather
// than "foo").
Root bool
Steps []Expr // Step, or an arbitrary expression in "(...)/foo" form
}
PathExpr is a sequence of steps evaluated left to right, each against the nodes produced by the previous one.
type QuantifiedExpr ¶
QuantifiedExpr is "some $x in seq satisfies test" or the "every" form.
func (*QuantifiedExpr) Eval ¶
func (e *QuantifiedExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for quantified expressions.
func (*QuantifiedExpr) String ¶
func (e *QuantifiedExpr) String() string
type SequenceExpr ¶
type SequenceExpr struct{ Items []Expr }
SequenceExpr is a comma-separated sequence constructor.
func (*SequenceExpr) Eval ¶
func (e *SequenceExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for a sequence constructor.
func (*SequenceExpr) String ¶
func (e *SequenceExpr) String() string
type SequenceType ¶
type SequenceType struct {
// Empty is the empty-sequence() type.
Empty bool
// ItemType is nil for item(), which matches anything.
ItemType NodeTest
// AtomicType names an atomic type when the item type is one.
AtomicType xdm.TypeCode
HasAtomicType bool
// FacetName is the derived type actually written, when it differs from
// AtomicType — "byte" for xs:byte, which is an xs:integer with a range.
// The code alone cannot express the bound, and dropping it made
// "128 castable as xs:byte" answer true.
FacetName string
// Occurrence is "", "?", "*" or "+".
Occurrence string
}
SequenceType is a type annotation: an item type plus an occurrence indicator.
func (SequenceType) Matches ¶
func (t SequenceType) Matches(seq xdm.Sequence) bool
Matches reports whether seq conforms to the sequence type.
func (SequenceType) String ¶
func (t SequenceType) String() string
type Step ¶
Step is one step of a path: an axis, a node test, and zero or more predicates.
func (*Step) Eval ¶
Eval implements Expr for a single axis step.
A step is evaluated against the context item alone; iterating a step over many context nodes is PathExpr's job. Splitting it this way means the predicate's context size is the number of nodes selected by *this* step from *this* node, which is what the spec requires and what a combined implementation typically gets wrong.
type Token ¶
type Token struct {
Kind TokenKind
Val string
Pos int
// Num holds the parsed numeric value and NumType its XPath type, so the
// parser does not re-parse the literal. A numeric literal's type is fixed
// by its lexical form: no dot or E means integer, a dot means decimal, an
// E means double.
Num float64
NumType numLiteralKind
}
Token is a lexical token with its source offset, which error messages use to point at the offending construct.
type TreatExpr ¶
type TreatExpr struct {
Operand Expr
Type SequenceType
}
TreatExpr is "expr treat as type": a static assertion that does not convert.
Source Files
¶
- ast.go
- ast_string.go
- axes.go
- builtins.go
- cast.go
- classdiff.go
- collation.go
- context.go
- eval.go
- fn_date.go
- fn_misc.go
- fn_node.go
- fn_qname.go
- fn_regex.go
- fn_seq.go
- fn_string.go
- functions.go
- lexer.go
- operators.go
- optimize.go
- parser.go
- parser_path.go
- rangeprops.go
- regex_backref.go
- schema_anchors.go
- schema_grammar.go
- typeexpr.go
- xpath.go