Documentation
¶
Overview ¶
Package xpath implements XPath 2.0: lexing, parsing to an AST, static analysis, and evaluation against an XDM tree.
Index ¶
- Constants
- func BacktrackingRegexEnabled() bool
- func BuiltinAtomicTypeCode(local string) (xdm.TypeCode, bool)
- func CastAtomic(a *xdm.Atomic, target xdm.TypeCode) (*xdm.Atomic, error)
- func CastToDerived(a *xdm.Atomic, target xdm.TypeCode, facet string) (*xdm.Atomic, error)
- func EffectiveBooleanValue(seq xdm.Sequence) (bool, error)
- func Eval(src string, ctx *Context, ns NamespaceResolver) (xdm.Sequence, error)
- func FragmentIsValidXMLName(uri string) bool
- func GroupingEqual(a, b *xdm.Atomic, coll Collation, implicitTZ int) bool
- func GroupingKey(a *xdm.Atomic, coll Collation, implicitTZ int) (string, error)
- func RegexpErr(re Regexp) error
- func RegisterCollation(uri string, c Collation)
- func RegisterEXSLTFuncs(l *Library)
- func RegisterHarnessFuncs(l *Library)
- func RegisterXSLTFuncs(l *Library)
- func SetBacktrackingRegex(on bool)
- func TranslateSchemaRegexp(pattern string) (string, error)
- func TranslateSchemaRegexpVersion(pattern string, xsd11 bool) (string, error)
- type Axis
- type BinaryOp
- type Binding
- type CastExpr
- type Collation
- type CollectionResolver
- type Compiled
- func (c *Compiled) CompatMode() bool
- func (c *Compiled) Eval(ctx *Context) (xdm.Sequence, error)
- func (c *Compiled) EvalBool(ctx *Context) (bool, error)
- func (c *Compiled) EvalString(ctx *Context) (string, error)
- func (c *Compiled) Expr() Expr
- func (c *Compiled) Source() string
- func (c *Compiled) WithCompatMode(on bool) *Compiled
- func (c *Compiled) WithDefaultCollation(coll Collation) *Compiled
- func (c *Compiled) WithStaticBaseURI(base string) *Compiled
- type Context
- func (c *Context) ContextNode() (*xdm.Node, error)
- func (c *Context) Descend() (*Context, error)
- func (c *Context) Err() error
- func (c *Context) LookupVar(name xdm.QName) (xdm.Sequence, bool)
- func (c *Context) WithFocus(item xdm.Item, pos, size int) *Context
- func (c *Context) WithNow(t time.Time) *Context
- func (c *Context) WithVar(name xdm.QName, val xdm.Sequence) *Context
- type ContextItem
- type DocumentResolver
- type Expr
- type FilterExpr
- type ForExpr
- type FuncCall
- type Function
- type FunctionLibrary
- type IfExpr
- type InstanceOfExpr
- type KindTest
- type Lexer
- type Library
- type Literal
- type NameTest
- type NamespaceResolver
- type NodeTest
- type Parser
- type PathExpr
- type QuantifiedExpr
- type Regexp
- type SchemaTypes
- type SequenceExpr
- type SequenceType
- type SimpleMap
- type Step
- type StringConcat
- type TextResolver
- type Token
- type TokenKind
- type TreatExpr
- type UnaryOp
- type VarRef
Constants ¶
const CodepointCollation = "http://www.w3.org/2005/xpath-functions/collation/codepoint"
CodepointCollation is the one collation this implementation provides. It is the only collation the spec requires every processor to support, and it compares strings by Unicode codepoint.
const HTMLASCIICaseInsensitive = "http://www.w3.org/2005/xpath-functions/collation/html-ascii-case-insensitive"
HTMLASCIICaseInsensitive is the collation that compares ASCII letters without regard to case. It is defined by the spec and is the one non-codepoint collation that needs no locale data.
const MaxDepth = 500
MaxDepth bounds recursive evaluation.
const MaxItems = 5_000_000
MaxItems bounds the number of items an evaluation may materialise.
Depth bounds the stack and Ctx bounds the wall clock, but neither bounds memory: "count(1 to 9999999)" is one shallow, fast expression that allocates nine million *Atomic values and peaked at 1.8 GB of resident memory. The range operator had its own limit, but "for $a in 1 to 3000, $b in 1 to 3000" walked straight past it, because the sequence is built by the for-expression rather than by the range.
The bound is deliberately generous: a real stylesheet over a large document works in thousands of nodes, not tens of millions, so this only fires on input designed to exhaust memory or on a genuine runaway.
const NSEXSLTCommon = "http://exslt.org/common"
NSEXSLTCommon is the namespace of the EXSLT "common" module.
EXSLT is a community extension library from the XSLT 1.0 era, not a W3C specification. It is named here rather than in xdm because nothing in the data model depends on it.
const UCACollation = "http://www.w3.org/2013/collation/UCA"
UCACollation is the URI family defined by XSLT 3.0 section 5.3.3 and F&O 5.3.4 for the Unicode Collation Algorithm. The base URI selects the root (DUCET) collation; query parameters tailor it.
Variables ¶
This section is empty.
Functions ¶
func BacktrackingRegexEnabled ¶ added in v1.0.0
func BacktrackingRegexEnabled() bool
BacktrackingRegexEnabled reports the current setting.
func BuiltinAtomicTypeCode ¶ added in v1.0.0
BuiltinAtomicTypeCode returns the type code for a built-in xs: type, given its local name.
It exists so that a caller holding a schema's primitive type — the xsd package, which cannot be imported from here — can map it to the code this engine's values carry, without a second copy of the table in parser_path.go drifting away from the first.
func CastAtomic ¶
CastAtomic converts an atomic value to a target type, per the XPath 2.0 casting table.
Casting is stricter than the string conversions of XPath 1.0: "abc" cast to xs:double is an error (FORG0001), not NaN. Silent NaN is how a validator ends up reporting a document as valid because a numeric comparison quietly became false.
func CastToDerived ¶
CastToDerived casts to a derived type named by its local name, applying the facet that the type code cannot carry. An unknown name falls back to a plain cast to the primitive.
func EffectiveBooleanValue ¶
EffectiveBooleanValue computes fn:boolean over a sequence.
XPath 2.0 restricts this compared with 1.0: a sequence of two or more atomic values raises FORG0006 rather than being truthy. That strictness is deliberate — it catches "if ($seq)" where the author meant "if (exists($seq))" — so it is enforced rather than relaxed.
func Eval ¶
Eval is a one-shot compile-and-evaluate, for callers that will not reuse the expression.
func FragmentIsValidXMLName ¶ added in v1.0.0
FragmentIsValidXMLName reports whether the fragment identifier in uri, if there is one, conforms to the rules for the XML media types.
XSLT 2.0 16.1 makes it a recoverable dynamic error (XTRE1160) if "the fragment identifier does not conform to the rules for fragment identifiers for that media type". For text/xml and application/xml those rules (RFC 7303) admit a bare name, which must be an XML Name, or an XPointer -- and an XPointer scheme part is itself a QName. So a fragment of "123456789" is not a legal one for XML however the resource is fetched, which is what lets the error be raised without retrieving anything: error-1160a names a w3.org URL this engine will not fetch at all, and diagnosing the fragment is the only way to reach the right answer offline.
It is exported because xslt/rtfuncs.go registers its own fn:document#1, which shadows the one here for arity 1 and needs the same rule.
A uri with no fragment, or an empty one, is not an error here: only a present and malformed fragment is.
func GroupingEqual ¶ added in v1.0.0
GroupingEqual reports whether two grouping key values are the same value under the "eq" operator, with the given collation comparing strings.
GroupingKey is the fast path for grouping: equal values hash alike, so a map finds the group in one lookup. It is not a complete answer, though, because the value comparison XSLT grouping uses is not transitive across the numeric types — erratum E25 spells this out. xs:float(1.0) equals xs:decimal(1.0000000000100000000001) (the decimal is promoted to float), and that decimal equals xs:double(1.00000000001) (promoted to double), yet the float and the double are different values. No single hash can express that, so a caller that missed in the map falls back to comparing against the key each existing group was opened with, in order.
A pair with no ordering at all — a string against a number — is simply not equal rather than an error, because grouping puts such values in separate groups rather than failing.
func GroupingKey ¶ added in v1.0.0
GroupingKey returns a string that is identical for two atomic values that compare equal, and different otherwise.
XSLT needs this for xsl:for-each-group, whose grouping keys are compared by value rather than by lexical form: xs:dateTime("2000-01-01T00:00:00Z") and xs:dateTime("2000-01-01T01:00:00+01:00") name the same instant and belong in one group, but their string forms differ. Keying on the string put them in two.
coll may be nil, in which case strings key on themselves.
func RegexpErr ¶ added in v1.0.0
RegexpErr reports a match-time failure from the most recent operation on re.
It is nil for an RE2 pattern, which cannot fail at match time, and it is the budget error for a backtracking pattern that ran out of steps.
func RegisterCollation ¶
RegisterCollation makes a collation available under a URI.
The spec requires exactly two collations — codepoint and the HTML ASCII case-insensitive one — and leaves the rest implementation-defined. This is how a host application supplies its own: a locale-aware comparison, or the case-blind collation the W3C test catalog defines for its own use.
Registering a URI that is already known replaces it, which is what makes a host override sensible. Registration is expected during setup, before any evaluation; it is guarded so that a late one cannot race a lookup, but it is not a way to change collations while expressions are running.
func RegisterEXSLTFuncs ¶ added in v1.0.0
func RegisterEXSLTFuncs(l *Library)
RegisterEXSLTFuncs adds the EXSLT common-module functions this processor implements to l.
Like RegisterXSLTFuncs and RegisterHarnessFuncs, these are deliberately NOT in Builtins(): EXSLT is an extension, and an XPath processor is required to report XPST0017 for a function it does not have. Registering into a chained Library keeps the populations separate, and — because function-available answers from the same library — it also makes function-available('exsl: node-set') report true exactly when the function is really callable.
Only node-set is provided. Every other EXSLT function is surface area that would then have to be supported, so they are added when something actually calls them, not speculatively.
func RegisterHarnessFuncs ¶ added in v1.0.0
func RegisterHarnessFuncs(l *Library)
RegisterHarnessFuncs adds the XPath 3.0 functions that the W3C conformance harness needs in order to SET UP a test, as distinct from the functions a stylesheet under test may call.
It exists for the same reason ParseExtended does, and observes the same boundary. Several XSLT test-set environments describe their initial context with an XPath 3.0 expression even when the stylesheet they then run is an XSLT 2.0 one — id-043's environment is
<source select="parse-xml('<root/>')" role="."/>
while id-043.xsl itself declares version="2.0". The setup expression is the harness's own XPath, not the test subject's, so it is written in the 3.0 language by design.
These functions are deliberately NOT in Builtins(). An XPath 2.0 processor is required to report XPST0017 for fn:parse-xml, and a stylesheet compiled against the builtin library still does, which is what the suite asserts elsewhere. Registering into a chained Library keeps the two populations separate in the same way RegisterXSLTFuncs does for the XSLT-only functions.
func RegisterXSLTFuncs ¶
func RegisterXSLTFuncs(l *Library)
RegisterXSLTFuncs adds the functions that XSLT 2.0 defines but XPath 2.0 does not.
fn:unparsed-text, fn:format-date and friends are in the XSLT specification, not the XPath one — a bare XPath 2.0 processor is required to report XPST0017 for them. This engine is an XSLT engine, so they must exist when a stylesheet is running, and must not when a plain XPath expression is being evaluated against the XPath library alone. Keeping them out of Builtins and adding them here is what makes both true.
func SetBacktrackingRegex ¶ added in v1.0.0
func SetBacktrackingRegex(on bool)
SetBacktrackingRegex turns the backtracking matcher on or off. See BacktrackingRegex.
It is safe to call while other goroutines are evaluating patterns: the setting is read atomically, and the compiled-pattern cache is keyed on it, so neither a stale compilation nor a torn read is possible. What is not guaranteed is which setting a call already in flight observes.
func TranslateSchemaRegexp ¶
TranslateSchemaRegexp rewrites an XML Schema regular expression into RE2 syntax, without anchoring it.
The XML Schema flavour and the XPath flavour share a grammar — Part 2 Appendix F defines the one, and XPath's fn:matches extends it — so the translation is the same in both directions: the multi-character escapes \i and \c, the block and category escapes, and character class subtraction, none of which RE2 accepts as written.
The result is deliberately unanchored, because the two flavours differ exactly there. fn:matches is a containment test, while a pattern facet must span the whole value. A caller using this for a pattern facet has to wrap the result — see the xsd package, which does so with \A(?:...)\z.
func TranslateSchemaRegexpVersion ¶
TranslateSchemaRegexpVersion is TranslateSchemaRegexp with the one grammar rule that XSD 1.1 changed made selectable.
1.1 stopped treating an unrecognised \p{Is...} block name as an error and began reading it as a class that matches every character, so the same pattern is invalid under 1.0 and valid under 1.1. reK88 asserts exactly that pair, which is why the version has to reach the grammar check.
Types ¶
type Axis ¶
type Axis int
Axis identifies one of the thirteen XPath axes.
func (Axis) IsReverse ¶
IsReverse reports whether the axis is a reverse axis. Reverse axes number their positions backwards, which changes what position() means inside a predicate — the one place the distinction is observable.
func (Axis) PrincipalKind ¶
PrincipalKind is the node kind an axis selects when the node test is a name or wildcard: attributes on the attribute axis, namespaces on the namespace axis, elements everywhere else.
type BinaryOp ¶
type BinaryOp struct {
Op string
Left, Right Expr
// ResolveQName binds a prefix in the static context of this operator, for
// the one conversion that needs it. A general comparison casts an
// untypedAtomic operand to the *other* operand's type, and when that type
// is xs:QName the lexical form carries a prefix whose namespace lives in
// the static context -- which the runtime Context deliberately does not
// carry, since namespaces are a static property. Capturing the resolver
// on the node is what makes the binding available where the cast happens.
// Nil for every operator other than a comparison, and for a comparison
// parsed without a namespace resolver.
ResolveQName func(prefix string) (string, bool)
}
BinaryOp is any infix operator. Keeping them in one node with an Op field rather than one type per operator keeps the parser's precedence ladder short and puts all the operand-conversion rules in one evaluator function, which is where they are easiest to check against the spec's tables.
type CastExpr ¶
type CastExpr struct {
Operand Expr
Type SequenceType
Castable bool // true for "castable as", which yields a boolean
}
CastExpr is "expr cast as type" and "expr castable as type".
type Collation ¶
type Collation interface {
Compare(a, b string) int
Contains(s, sub string) bool
StartsWith(s, prefix string) bool
EndsWith(s, suffix string) bool
IndexOf(s, sub string) int
}
Collation compares two strings. Only the operations XPath actually needs are exposed, because a general Compare is not enough: fn:contains under a case-insensitive collation is not "compare the folded strings", it is "does the folded needle occur in the folded haystack".
func ResolveCollation ¶
ResolveCollation returns the collation a URI names.
A relative URI is accepted when its tail matches a known collation, because the QT3 suite and some stylesheets write "collation/codepoint" rather than the full URI. Resolving it against the static base URI would be more correct still, but the base is not threaded into every function that takes a collation argument, and matching the tail covers the forms that occur.
type CollectionResolver ¶
type CollectionResolver interface {
// ResolveCollection returns the documents in uri, resolved against base.
//
// The result is a sequence rather than a []*xdm.Tree because a collection
// is permitted to contain items that are not document nodes.
ResolveCollection(uri, base string) (xdm.Sequence, error)
}
CollectionResolver loads a named set of documents for fn:collection.
It is deliberately separate from DocumentResolver rather than an extra method on it. A caller who wants fn:doc for the code lists shipped beside a stylesheet does not thereby want fn:collection to enumerate a directory, and folding the two together would make enabling one enable the other.
The empty uri is the default collection — fn:collection() with no argument. A resolver that has no default should return an error for it rather than an empty sequence, for the reason given on fnCollection.
type Compiled ¶
type Compiled struct {
// contains filtered or unexported fields
}
Compiled is a parsed XPath expression, ready to evaluate against any context.
Compiling once and evaluating many times is the intended usage: parsing dominates the cost of a short expression, and a stylesheet evaluates the same expression once per node. A Compiled value is immutable and safe for concurrent use.
func Compile ¶
func Compile(src string, ns NamespaceResolver) (*Compiled, error)
Compile parses src, resolving namespace prefixes with ns.
func MustCompile ¶
func MustCompile(src string, ns NamespaceResolver) *Compiled
MustCompile is Compile, panicking on error. For tests and for expressions that are literals in this package's own source.
func (*Compiled) CompatMode ¶ added in v1.0.0
CompatMode reports whether c evaluates under XPath 1.0 compatibility mode.
func (*Compiled) EvalString ¶
EvalString evaluates and returns the string value of the result, which is the concatenation rule of fn:string applied to the first item, or "" for the empty sequence.
func (*Compiled) Expr ¶
Expr returns the root of the AST, for callers that need to inspect or rewrite it (the XSLT layer analyses patterns this way).
func (*Compiled) WithCompatMode ¶ added in v1.0.0
WithCompatMode returns a copy of c evaluated under XPath 1.0 compatibility mode.
The mode is static, exactly as the base URI and the default collation are: XSLT 3.8 fixes it from the [xsl:]version attribute of the nearest ancestor-or-self of the element the expression is written on, which cannot change between evaluations. Binding it to the compiled expression rather than threading it through the dynamic context is therefore both correct and what keeps an ordinary 2.0 expression byte-identical to what it was: a Compiled that was never given the flag never sets it on the context, so no evaluation outside a 1.0 scope can observe it.
func (*Compiled) WithDefaultCollation ¶ added in v1.0.0
WithDefaultCollation returns a copy of c whose functions use coll when no collation argument is given.
func (*Compiled) WithStaticBaseURI ¶ added in v1.0.0
WithStaticBaseURI returns a copy of c whose expressions resolve relative references against base.
It exists because the static base URI really is static: xml:base is written in the stylesheet and cannot change between evaluations, so binding it to the compiled expression is both correct and cheaper than threading it through the dynamic context.
type Context ¶
type Context struct {
// Item is the context item. It is nil where there is no context item,
// which is an error to reference rather than an empty sequence.
Item xdm.Item
// Position is the context position, 1-based. Zero means "no focus".
Position int
// Size is the context size.
Size int
// Vars holds in-scope variable bindings, keyed by expanded name.
// Lookups walk to Parent, so a nested scope does not copy the map.
Vars map[string]xdm.Sequence
Parent *Context
// Funcs resolves function calls. Supplied by the caller so that XSLT can
// add xsl:function declarations and extension functions without this
// package knowing about them.
Funcs FunctionLibrary
// StaticBaseURI is the base URI of the expression itself — the stylesheet
// or query it was written in — which is what fn:static-base-uri returns
// and what fn:resolve-uri resolves against by default.
//
// It is distinct from a *node's* base URI, which comes from the document
// the node was parsed from. Returning the context node's was the nearest
// thing available before this existed, and it is a different value: a
// stylesheet in one place can perfectly well be applied to a document
// from another.
StaticBaseURI string
// ImplicitTimezone is the offset in minutes applied to date/time values
// that carry no timezone. The spec requires the dynamic context to supply
// one; defaulting to UTC keeps results reproducible across machines,
// which matters more for a validator than matching local time.
ImplicitTimezone int
// Ctx carries cancellation. A stylesheet can loop for a long time on
// pathological input, and the caller needs a way out that does not
// involve killing the process.
Ctx context.Context
// Docs resolves fn:doc and fn:document URIs. Nil disables them, which is
// the safe default: a stylesheet that can open arbitrary URIs is an SSRF
// and file-disclosure vector.
Docs DocumentResolver
// Collections resolves fn:collection URIs. Nil disables it, for the same
// reason nil disables Docs, and setting Docs does not set this: see
// CollectionResolver.
Collections CollectionResolver
// Texts resolves fn:unparsed-text URIs. Nil disables it, and setting
// Docs does not set this: reading a file as raw text is a wider grant
// than reading it as a parsed document. See TextResolver.
Texts TextResolver
// Compat is XPath 1.0 compatibility mode, which XSLT 3.8 puts in force for
// expressions written on an element whose effective [xsl:]version is below
// 2.0. Under it the coercion rules of XPath 2.0 appendix B.1 apply: a
// multi-item argument to a parameter expecting a string, a number or a
// node is truncated to its first item instead of raising XPTY0004,
// arithmetic on a non-numeric operand yields NaN rather than a type error,
// and a general comparison converts its operands the way XPath 1.0 did.
//
// It defaults to false and is set only by a Compiled that was given it, so
// ordinary 2.0 evaluation never sees it.
Compat bool
// Depth guards against unbounded recursion in user-defined functions and
// named templates, which the spec does not bound.
Depth int
// Now is the value fn:current-dateTime and its siblings return.
//
// The spec requires these to be stable for the whole of one evaluation:
// calling current-dateTime() twice must give the same answer, or a
// stylesheet that stamps a document and then checks the stamp against
// "now" can disagree with itself. Reading the clock here once, rather
// than per call, is what guarantees that. A zero value means the caller
// did not set one and the functions are unavailable.
Now time.Time
// HasNow distinguishes an unset clock from a legitimately zero time.
HasNow bool
// contains filtered or unexported fields
}
Context is the XPath dynamic context: everything an expression can observe beyond its own AST.
The focus (item, position, size) changes on every step and predicate, while the rest (variables, functions, the implicit timezone) changes rarely. They are kept in one struct anyway, copied cheaply by value in the hot paths, because splitting them means every evaluator function takes two parameters and the copy is a handful of words either way.
func NewContext ¶
func NewContext(item xdm.Item, funcs FunctionLibrary) *Context
NewContext returns a context with the given focus and library.
func (*Context) ContextNode ¶
ContextNode returns the context item as a node, or an error when there is no context item or it is an atomic value.
Steps require a node context; the distinct error codes matter because XPDY0002 (absent) and XPTY0020 (present but not a node) mean different things to a stylesheet author.
func (*Context) Descend ¶
Descend returns a copy with the recursion depth incremented, erroring past the limit.
func (*Context) LookupVar ¶
LookupVar resolves a variable by expanded name, walking enclosing scopes.
func (*Context) WithFocus ¶
WithFocus returns a copy of ctx with a new context item, position and size, sharing the variable scope.
This is the operation performed once per node per step. It copies the struct rather than allocating a child scope, so variable lookups still resolve through the same maps without a new one being built.
The copy itself does allocate — it is the largest single allocation site in the engine, around a quarter of what a stylesheet render allocates. Reusing one context across a step loop was measured and made no difference at all (4,963,596 vs 4,964,187 bytes per render), so it was reverted: WithVar builds children holding a pointer back to this context, and the aliasing risk that reuse introduces buys nothing. Anyone tempted to try it again should measure first.
func (*Context) WithVar ¶
WithVar returns a child context binding name to val.
A child scope with its own one-entry map is used rather than mutating the parent's, because a for-expression binds a fresh value per iteration while the body may capture it; mutation would make all iterations observe the last value.
type ContextItem ¶
type ContextItem struct{}
ContextItem is the "." expression.
func (*ContextItem) Eval ¶
func (e *ContextItem) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for the context item.
func (*ContextItem) String ¶
func (e *ContextItem) String() string
type DocumentResolver ¶
type DocumentResolver interface {
// ResolveDocument returns the tree for uri, resolved against base.
ResolveDocument(uri, base string) (*xdm.Tree, error)
}
DocumentResolver loads a document by URI for fn:doc and fn:document.
type Expr ¶
type Expr interface {
// Eval evaluates the expression in ctx and returns a sequence.
Eval(ctx *Context) (xdm.Sequence, error)
// String returns a source-like rendering, used in error messages and to
// make test failures readable.
String() string
}
Expr is a node in the XPath abstract syntax tree.
Evaluation is a method on the AST rather than a separate visitor. XPath evaluation is a simple recursive walk with no multi-pass analysis, so a visitor would add an indirection layer without buying anything; the one place a second pass would help (static typing) is not implemented, and the spec permits a dynamically-typed implementation.
func Parse ¶
func Parse(src string, ns NamespaceResolver) (Expr, error)
Parse compiles an XPath 2.0 expression.
func ParseExtended ¶ added in v1.0.0
func ParseExtended(src string, ns NamespaceResolver) (Expr, error)
ParseExtended compiles an expression in which the XPath 3.0 braced URI literal Q{uri}local is also accepted.
A stylesheet compiled with Parse still rejects it, which is what a 2.0 processor must do. It exists for a caller that is itself writing XPath rather than running someone else's — specifically the conformance harness, whose assertion expressions are written in the 3.0 language even for tests whose stylesheets are 2.0.
The simple map operator "!" used to be gated here too. It is now accepted unconditionally, along with "||" and "=>": see Lexer.extended.
type FilterExpr ¶
FilterExpr applies predicates to an arbitrary expression, as in "(1 to 10)[. mod 2 = 0]".
func (*FilterExpr) Eval ¶
func (e *FilterExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for a filtered expression.
func (*FilterExpr) String ¶
func (e *FilterExpr) String() string
type ForExpr ¶
ForExpr is "for $x in seq return expr".
type FuncCall ¶
FuncCall is a function call. Resolution happens at evaluation time against the context's function library, so that a stylesheet's own xsl:function declarations are visible without a separate binding pass.
type Function ¶
type Function struct {
Name xdm.QName
Arity int
// Call receives the already-evaluated arguments. Functions that need the
// context item (fn:string with no argument, fn:position) read it from ctx.
Call func(ctx *Context, args []xdm.Sequence) (xdm.Sequence, error)
}
Function is a callable XPath function.
type FunctionLibrary ¶
type FunctionLibrary interface {
// Lookup returns the function with the given name and arity.
Lookup(name xdm.QName, arity int) (Function, bool)
}
FunctionLibrary resolves and calls functions.
func Builtins ¶
func Builtins() FunctionLibrary
Builtins returns the standard fn: function library.
The library is built once and shared: Function values hold no mutable state (everything they need comes from the Context passed at call time), so a single instance is safe for concurrent transforms. Rebuilding it per transform would cost several hundred map inserts for no benefit.
type IfExpr ¶
type IfExpr struct {
Cond, Then, Else Expr
}
IfExpr is "if (cond) then a else b". Both branches are required by the grammar; there is no one-armed form.
type InstanceOfExpr ¶
type InstanceOfExpr struct {
Operand Expr
Type SequenceType
}
InstanceOfExpr is "expr instance of type".
func (*InstanceOfExpr) Eval ¶
func (e *InstanceOfExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for "instance of".
func (*InstanceOfExpr) String ¶
func (e *InstanceOfExpr) String() string
type KindTest ¶
type KindTest struct {
Kind xdm.NodeKind
// Any matches every kind: the node() test.
Any bool
// Name constrains element()/attribute()/processing-instruction() tests
// that name a target.
Name *xdm.QName
HasName bool
// Content constrains the root element of a document-node() test:
// document-node(element(invoice)) matches only a document whose element
// child satisfies the inner test. Nil means the document's content is
// unconstrained.
Content NodeTest
// TypeName is the second argument of element(name, type) and its
// attribute() counterpart, resolved to the key the data model records
// type annotations under: a namespace-qualified {uri}local for a schema
// type, the bare local name for a built-in. The empty string means the
// test carried no type argument and constrains only the name.
//
// It is resolved at parse time because the prefix binding lives in the
// static context, which is gone by the time the test runs. Comparing the
// lexical form instead forced the comparison down to local parts, and two
// types sharing a local name in different namespaces then matched each
// other.
TypeName string
// TypeNameLexical is TypeName as the author wrote it, kept only so that
// String() renders the test back in the syntax it was parsed from.
TypeNameLexical string
// TypeNillable records the "?" of element(name, type?), which lets a
// nilled element match even though its content is absent.
TypeNillable bool
// SchemaDeclared marks a schema-element() or schema-attribute() test,
// which names a global declaration rather than an element name. It
// matches the named declaration and, for an element, the members of its
// substitution group.
SchemaDeclared bool
// SubstitutionGroup holds the other names schema-element(E) admits: the
// members of E's substitution group, resolved from the imported schema
// when the test was parsed.
//
// It is resolved at parse time because nothing carries a schema into the
// evaluator, and the group is fixed once the schema is imported. Nil for
// every test that is not a schema-element(), and for a declaration that
// heads no group.
SubstitutionGroup []xdm.QName
// DeclaredType is the local name of the type the schema-element() or
// schema-attribute() declaration names, resolved at parse time for the
// same reason SubstitutionGroup is. A node whose annotation is neither
// that type nor derived from it was validated against some *other*
// declaration of the same name — a local one — and does not match.
//
// Empty when the declaration's type is anonymous, in which case there is
// no name to compare and the test checks only that the node was
// validated.
DeclaredType string
}
KindTest matches by node kind: text(), comment(), node(), element(name), and so on.
type Lexer ¶
type Lexer struct {
// contains filtered or unexported fields
}
Lexer turns XPath source into tokens.
XPath 2.0's grammar is not context-free at the lexical level: whether `*` means multiplication or "any element", and whether `div`, `and`, `is` and friends are operators or element names, depends on what preceded them. The spec resolves this with a rule stated in terms of the previous token, and that is what prevOperand tracks. Trying to decide these in the parser instead means the lexer must emit ambiguous tokens and the parser must re-lex, which is worse.
type Library ¶
type Library struct {
// Parent is consulted when a name is not found locally, so a stylesheet's
// own functions can shadow and extend the builtins without copying them.
Parent FunctionLibrary
// contains filtered or unexported fields
}
Library is a mutable function library keyed by expanded name and arity.
Arity is part of the key because XPath overloads on it: fn:string() and fn:string($arg) are different functions, and fn:substring has both a two- and a three-argument form with different behaviour.
func NewLibrary ¶
func NewLibrary(parent FunctionLibrary) *Library
NewLibrary returns an empty library chained to parent.
type Literal ¶
Literal is a constant atomic value.
type NameTest ¶
type NameTest struct {
// Name is the expanded name to match. Wildcards leave one or both parts
// unconstrained; see AnyURI and AnyLocal.
Name xdm.QName
// AnyURI matches any namespace ("*" and "*:local").
AnyURI bool
// AnyLocal matches any local name ("*" and "prefix:*").
AnyLocal bool
}
NameTest matches by expanded name.
type NamespaceResolver ¶
type NamespaceResolver interface {
// ResolvePrefix returns the URI bound to prefix, or false if unbound.
ResolvePrefix(prefix string) (string, bool)
// DefaultElementNamespace returns the namespace applied to unprefixed
// element name tests. XSLT sets this from xpath-default-namespace; it is
// empty by default, and it never applies to attribute names or function
// names.
DefaultElementNamespace() string
// DefaultFunctionNamespace returns the namespace for unprefixed function
// names, which is the fn: namespace in XPath and XSLT.
DefaultFunctionNamespace() string
}
NamespaceResolver resolves a namespace prefix to a URI at parse time.
Prefixes must be resolved when the expression is compiled, not when it runs: an XPath expression in a stylesheet is bound to the namespace declarations in scope at the point it appears, and by evaluation time the relevant element is long out of view.
type NodeTest ¶
type NodeTest interface {
// Matches reports whether n is selected on an axis whose principal node
// kind is principal.
Matches(n *xdm.Node, principal xdm.NodeKind) bool
String() string
}
NodeTest decides whether a node on an axis is selected.
type Parser ¶
type Parser struct {
// contains filtered or unexported fields
}
Parser builds an AST from tokens.
type PathExpr ¶
type PathExpr struct {
// Root marks a path that starts at the document root ("/foo" rather
// than "foo").
Root bool
Steps []Expr // Step, or an arbitrary expression in "(...)/foo" form
}
PathExpr is a sequence of steps evaluated left to right, each against the nodes produced by the previous one.
type QuantifiedExpr ¶
QuantifiedExpr is "some $x in seq satisfies test" or the "every" form.
func (*QuantifiedExpr) Eval ¶
func (e *QuantifiedExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for quantified expressions.
func (*QuantifiedExpr) String ¶
func (e *QuantifiedExpr) String() string
type Regexp ¶ added in v1.0.0
type Regexp interface {
MatchString(s string) bool
FindAllStringSubmatchIndex(s string, n int) [][]int
NumSubexp() int
}
Regexp is what CompileRegexp hands back: the subset of *regexp.Regexp that the XSLT layer uses, so that a pattern needing the backtracking engine can be returned in its place without the caller knowing which it got.
The two implementations differ in one way callers must respect. RE2 cannot fail at match time, so *regexp.Regexp's methods have nowhere to report an error and need none. The backtracking engine *can* fail at match time, by exhausting its step budget, and it reports that through Err() rather than by answering false — answering false would be a guess, and precisely on the inputs where the answer was hardest to get. So a caller that may be holding a backtracking pattern must check Err() after any operation whose result it intends to use. RegexpErr does that check for both implementations.
func CompileRegexp ¶
CompileRegexp exposes the XPath-to-Go regular expression translation for the XSLT layer, which needs it for xsl:analyze-string. The compiled result is cached exactly as it is for fn:matches.
A pattern with a backreference RE2 cannot express is compiled by the backtracking engine instead, but only when that engine is enabled; when it is not, the pattern is refused exactly as before.
type SchemaTypes ¶ added in v1.0.0
type SchemaTypes interface {
// LookupSchemaType reports whether name is a type in the static context,
// and which primitive an atomic value of it erases to.
//
// The primitive is what the type *system* needs: "instance of" and "treat
// as" compare against the type hierarchy, and a value of a derived atomic
// type is a value of its primitive with facets applied. A complex type,
// or a list or union with no single primitive, returns ok with a zero
// code and false for atomic — enough to stop XPST0051 without claiming
// the value is comparable as an atomic.
LookupSchemaType(name xdm.QName) (prim xdm.TypeCode, atomic, ok bool)
// LookupSchemaDeclaration reports whether name is a global element or
// attribute declaration in the static context.
//
// It is what schema-element() and schema-attribute() need: both name a
// *declaration* rather than a type, and both are XPST0008 when no schema
// declares the name.
LookupSchemaDeclaration(name xdm.QName, attribute bool) bool
// SubstitutionGroupMembers returns the global element declarations that
// may substitute for name, transitively and not including name itself.
//
// schema-element(E) matches E and every member of E's substitution
// group, so a schema that declares "surname" as substitutable for "last"
// makes schema-element(z:last) match a z:surname element. Resolving the
// members here rather than at match time is what keeps the node test
// self-contained: nothing carries a schema into the evaluator, and the
// group is fixed once the schema is imported.
//
// An implementation with no schema, or a name with no members, returns
// nil.
SubstitutionGroupMembers(name xdm.QName) []xdm.QName
// SchemaDeclarationType returns the local name of the type a global
// element or attribute declaration names, and whether there is one.
//
// It is a LOCAL name, unlike the type name in an element() test, which is
// resolved to a namespace-qualified annotation key at parse time. The
// difference is deliberate: this string is produced by an implementation
// of this interface, which has no obligation to know about annotation
// keys, so the comparison against it (in nodeTypeMatches, reached through
// KindTest.DeclaredType) accepts a bare name as a local-part match. That
// is narrower ground than it sounds: the check exists only to tell a
// global declaration from a LOCAL declaration of the same name in the
// same schema, where the namespace is not in question.
//
// schema-element(E) matches a node only when it was validated against
// E's *declaration*, and a node may carry E's name while having been
// validated against a local declaration of a different type — which is
// the case the suite draws the line on. The evaluator sees only the
// compiled test, so the declared type is resolved here, while the schema
// is still reachable, and compared against the node's annotation at
// match time.
//
// An anonymous type has no name to return, so a declaration using one
// returns false and the test falls back to checking only that the node
// was validated at all.
SchemaDeclarationType(name xdm.QName, attribute bool) (string, bool)
// ValidateSchemaValue checks a lexical value against a named simple type
// in the static context, reporting whether the name is a simple type at
// all and, if so, whether the value is in its value space.
//
// "castable as my:hatsize" is that question. The engine can cast to the
// built-in the type derives from, but the facets the schema author wrote
// live only in the schema — so without asking, a cast to a restriction of
// xs:integer accepted every integer and the restriction meant nothing.
ValidateSchemaValue(name xdm.QName, value string) (known bool, err error)
}
SchemaTypes reports the types an imported schema contributes to the static context.
XPath 2.0 has an *in-scope schema definitions* component that this engine otherwise leaves empty: without it, only the built-in xs: types exist, and "instance of my:partNumberType" is XPST0051 no matter what the stylesheet imported. A stylesheet with xsl:import-schema is precisely the case where that component is not empty.
It is an interface here rather than an *xsd.Schema because xsd imports xpath — schema documents contain XPath expressions in their assertions and selectors — so the dependency cannot run the other way. The xslt package supplies the implementation, which is a few lines over the schema's own type table.
A resolver that also implements this is asked about a name only after the built-in table has declined it, so a schema cannot redefine xs:integer.
type SequenceExpr ¶
type SequenceExpr struct{ Items []Expr }
SequenceExpr is a comma-separated sequence constructor.
func (*SequenceExpr) Eval ¶
func (e *SequenceExpr) Eval(ctx *Context) (xdm.Sequence, error)
Eval implements Expr for a sequence constructor.
func (*SequenceExpr) String ¶
func (e *SequenceExpr) String() string
type SequenceType ¶
type SequenceType struct {
// Empty is the empty-sequence() type.
Empty bool
// ItemType is nil for item(), which matches anything.
ItemType NodeTest
// AtomicType names an atomic type when the item type is one.
AtomicType xdm.TypeCode
HasAtomicType bool
// FacetName is the derived type actually written, when it differs from
// AtomicType — "byte" for xs:byte, which is an xs:integer with a range.
// The code alone cannot express the bound, and dropping it made
// "128 castable as xs:byte" answer true.
FacetName string
// SchemaType is the lexical name of a type that came from an imported
// schema rather than from the built-in table.
//
// It is kept as written because that is what the schema's own type table
// is keyed by: matching a value against it means asking the schema, not
// the type codes here, and a derived type's identity is exactly its name.
SchemaType string
// SchemaValueValid checks a lexical value against the imported schema
// type SchemaType names, when that name is a simple type. It is captured
// at parse time, while the schema is still reachable through the
// resolver, for the same reason a schema-element() test's substitution
// group is: nothing carries a schema into the evaluator.
//
// nil when the type is not an imported simple type, in which case a cast
// is decided entirely by the built-in the type derives from.
SchemaValueValid func(value string) error
// Occurrence is "", "?", "*" or "+".
Occurrence string
}
SequenceType is a type annotation: an item type plus an occurrence indicator.
func (SequenceType) Matches ¶
func (t SequenceType) Matches(seq xdm.Sequence) bool
Matches reports whether seq conforms to the sequence type.
func (SequenceType) MatchesItem ¶ added in v1.0.0
func (t SequenceType) MatchesItem(it xdm.Item) bool
MatchesItem reports whether a single item conforms to the sequence type's item type, ignoring the occurrence indicator.
The function conversion rules need this separately from Matches: subtype substitution says an item that already conforms is passed through untouched, and only an item that does not conform is a candidate for atomisation, casting or promotion.
func (SequenceType) String ¶
func (t SequenceType) String() string
type SimpleMap ¶ added in v1.0.0
SimpleMap is the XPath 3.0 "!" operator: the right operand is evaluated once per item of the left, with that item as the context item, and the results are concatenated.
Unlike "/" it neither requires nodes nor sorts, which is the whole reason the suite's assertions use it — "string-to-codepoints(...)!string()" maps over integers, where "/" would raise XPTY0019.
type Step ¶
type Step struct {
Axis Axis
Test NodeTest
Predicates []Expr
// Explicit records that the axis was written out ("child::x") rather
// than abbreviated ("x"). Section 5.5.3 gives the abbreviated child axis
// in a pattern a wider reach than the written one — it is evaluated on
// the child-or-top axis, so "document-node()" alone matches the document
// node while "child::document-node()" is legal but matches nothing,
// since a document node is never a child.
Explicit bool
}
Step is one step of a path: an axis, a node test, and zero or more predicates.
func (*Step) Eval ¶
Eval implements Expr for a single axis step.
A step is evaluated against the context item alone; iterating a step over many context nodes is PathExpr's job. Splitting it this way means the predicate's context size is the number of nodes selected by *this* step from *this* node, which is what the spec requires and what a combined implementation typically gets wrong.
type StringConcat ¶ added in v1.0.0
StringConcat is the XPath 3.0 "||" operator.
It is defined as fn:concat($a, $b), which means it atomizes each operand, requires at most one item from each, and treats the empty sequence as the zero-length string rather than propagating it — so "() || 'x'" is "x" and not the empty sequence.
func (*StringConcat) Eval ¶ added in v1.0.0
func (e *StringConcat) Eval(ctx *Context) (xdm.Sequence, error)
func (*StringConcat) String ¶ added in v1.0.0
func (e *StringConcat) String() string
type TextResolver ¶ added in v1.0.0
type TextResolver interface {
// ResolveText returns the text of uri, resolved against base, decoded
// using encoding when one is named and as UTF-8 when it is empty.
ResolveText(uri, base, encoding string) (string, error)
}
TextResolver reads a resource as text for fn:unparsed-text.
It is deliberately separate from DocumentResolver rather than reusing it. fn:doc parses what it reads as XML, so a resolver granting it hands out well-formed documents; fn:unparsed-text hands the stylesheet the raw bytes of any file the resolver will open, which is a strictly larger disclosure and a different decision for the caller to make. Nil disables the function, which is the default and the safe one.
type Token ¶
type Token struct {
Kind TokenKind
Val string
Pos int
// Num holds the parsed numeric value and NumType its XPath type, so the
// parser does not re-parse the literal. A numeric literal's type is fixed
// by its lexical form: no dot or E means integer, a dot means decimal, an
// E means double.
Num float64
// contains filtered or unexported fields
}
Token is a lexical token with its source offset, which error messages use to point at the offending construct.
type TreatExpr ¶
type TreatExpr struct {
Operand Expr
Type SequenceType
}
TreatExpr is "expr treat as type": a static assertion that does not convert.
Source Files
¶
- ast.go
- ast_string.go
- axes.go
- builtins.go
- cast.go
- classdiff.go
- collation.go
- compat.go
- context.go
- eval.go
- fn_date.go
- fn_document.go
- fn_misc.go
- fn_node.go
- fn_qname.go
- fn_regex.go
- fn_seq.go
- fn_string.go
- functions.go
- lexer.go
- operators.go
- optimize.go
- parser.go
- parser_path.go
- rangeprops.go
- regex_backref.go
- regex_backtrack.go
- regex_grammar.go
- schema_anchors.go
- schema_grammar.go
- schema_types.go
- typeexpr.go
- xpath.go