xquery

package
v1.3.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 15, 2026 License: MIT Imports: 13 Imported by: 0

Documentation

Overview

Package xquery implements XQuery 3.1.

XQuery and XPath share an expression language, and this package does not reimplement it: expressions are compiled by the xpath package, which is at 100% of the QT3 suite for XPath 2.0, 3.0 and 3.1. What lives here is what XQuery has and XPath does not — constructors, FLWOR, the prolog, and the handful of expressions that are XQuery's alone.

Why the parser reads the source rather than a token stream

A direct element constructor puts XML syntax inside expression syntax, and an enclosed expression puts expression syntax back inside XML:

<a>{ $x + 1 }</a>

Whether "a" is a tag name or a name test, and whether "+" is an operator or literal text, is not decidable without knowing where in the nesting the reader is. XPath's lexer runs to completion before its parser starts, and its parser backtracks by rewinding an index into the finished token slice — a design that is correct for XPath and cannot answer the question above.

Rather than make that lexer re-entrant, which would put a conformant component at risk for a language it does not implement, this parser reads constructor syntax directly from the source and hands each enclosed expression to xpath as a substring. XML syntax stays here; expression syntax stays there. BaseX takes the same approach; Saxon instead uses a mode-switching tokeniser, which suits a codebase whose parser was written for both languages from the start.

Index

Examples

Constants

View Source
const DefaultMaxModuleBytes = 16 << 20

DefaultMaxModuleBytes bounds the total source text one compilation may read from its resolver when Options.MaxModuleBytes is zero.

The module count alone does not bound the work: 512 modules of a gigabyte each is within DefaultMaxModules and is not within anything worth calling a budget. So the bytes are counted too, cumulatively across the whole compilation rather than per module, because a budget spent one module at a time is not spent at all.

16 MB is far above any hand-written module — the largest library module in this tree's corpora is under 100 KB — and far below what would matter to the process.

View Source
const DefaultMaxModules = 512

DefaultMaxModules bounds how many library modules one compilation may load when Options.MaxModules is zero.

A module may import modules, so the graph is attacker-shaped in the same way a schema's include graph is: one small module that imports two others, each importing two more, costs nothing to write and everything to load. This is the same bound and the same reasoning as xsd's DefaultMaxDocuments, at the same value, because it guards the same shape of fan-out. Anything a real query imports is two orders of magnitude below it.

View Source
const DefaultMaxSchemaBytes = 16 << 20

DefaultMaxSchemaBytes bounds the total schema source text one compilation may read when Options.MaxSchemaBytes is zero.

A schema document may include and import further schema documents, so the graph is attacker-shaped in exactly the way a module graph and an xs:include graph are, and the bound is the same one and for the same reason: DefaultMaxModuleBytes, at the same value. The count of documents reached through a schema's OWN references is bounded separately, by xsd's own limits, which the assembler applies as it follows them -- so this budget bounds the bytes the QUERY's own imports pull in, cumulatively across the compilation rather than per import, because a budget spent one import at a time is not spent at all.

Variables

This section is empty.

Functions

func Eval

func Eval(src string, ctx *xpath.Context, opts Options) (xdm.Sequence, error)

Eval compiles and runs a query in one step.

Example

Eval compiles and runs a query in one step.

package main

import (
	"fmt"

	"github.com/knroy/go-xml/xpath"
	"github.com/knroy/go-xml/xquery"
)

func main() {
	seq, err := xquery.Eval(`for $i in 1 to 3 return $i * $i`,
		xpath.NewContext(nil, xpath.Builtins()), xquery.Options{})
	if err != nil {
		panic(err)
	}
	for _, item := range seq {
		fmt.Print(item, " ")
	}
}
Output:
1 4 9
Example (Document)

Binding the parsed document as the context item makes path expressions resolve against it, exactly as they do in XPath.

package main

import (
	"fmt"

	"github.com/knroy/go-xml/xdm"
	"github.com/knroy/go-xml/xpath"
	"github.com/knroy/go-xml/xquery"
)

func main() {
	doc, err := xdm.ParseString(
		`<books><book price="30"><t>A</t></book><book price="10"><t>B</t></book></books>`,
		xdm.ParseOptions{})
	if err != nil {
		panic(err)
	}
	ctx := xpath.NewContext(doc.Root, xpath.Builtins())

	seq, err := xquery.Eval(
		`for $b in //book order by xs:decimal($b/@price) return string($b/t)`,
		ctx, xquery.Options{})
	if err != nil {
		panic(err)
	}
	fmt.Println(seq)
}
Output:
[B A]
Example (Errors)

Errors carry the code the specification gives them, so a caller can match on the code rather than on the prose.

package main

import (
	"fmt"

	"github.com/knroy/go-xml/xpath"
	"github.com/knroy/go-xml/xquery"
)

func main() {
	_, err := xquery.Eval(`1 div 0`,
		xpath.NewContext(nil, xpath.Builtins()), xquery.Options{})
	fmt.Println(err)
}
Output:
FOAR0001: division by zero

Types

type BoundarySpace

type BoundarySpace int

BoundarySpace is the boundary-space policy of §4.3: whether whitespace that only separates markup survives into the constructed element.

const (
	// StripSpace discards boundary whitespace. It is the default, and the
	// value a query gets when it says nothing.
	StripSpace BoundarySpace = iota
	// PreserveSpace keeps it, as "declare boundary-space preserve" asks.
	PreserveSpace
)

type Construction

type Construction int

Construction is the construction mode of §4.10: whether a copied node keeps the type annotation it was validated with.

const (
	// PreserveTypes keeps annotations. It is the default.
	PreserveTypes Construction = iota
	// StripTypes replaces them with xs:untyped and xs:untypedAtomic.
	StripTypes
)

type EmptyOrder

type EmptyOrder int

EmptyOrder decides where the empty sequence sorts in an "order by" that does not say (§4.7).

const (
	// EmptyGreatest sorts the empty sequence last ascending. It is this
	// implementation's default, which the specification leaves open.
	EmptyGreatest EmptyOrder = iota
	// EmptyLeast sorts it first.
	EmptyLeast
)

type MapModuleResolver added in v1.3.0

type MapModuleResolver struct {
	// Modules maps a target namespace to the module's source text.
	Modules map[string]string
}

MapModuleResolver answers module imports from an in-memory table keyed by target namespace.

It exists so that a host can supply modules without granting any reach: nothing is opened and nothing is fetched, and a namespace not in the table is XQST0059. Options.Modules is the shorter way to say the same thing for a fixed set; this is for a host that builds the set at run time.

func (MapModuleResolver) Resolve added in v1.3.0

func (r MapModuleResolver) Resolve(namespace string, hints []string, base string) (
	io.ReadCloser, string, error)

Resolve implements ModuleResolver.

The location hints are ignored, deliberately. §4.12 permits it — they are hints — and honouring them here would mean the table's keys were not the only thing this resolver answers for, which is the property that makes it safe to hand an untrusted query.

type Module added in v1.3.0

type Module struct {
	// Namespace is the module's target namespace. It must match the "module
	// namespace" declaration in Source, which is checked on load: a store
	// keyed by a namespace the module does not declare would answer imports
	// with the wrong module. XQST0059 is the error, since from the import's
	// point of view no module with that namespace was found.
	Namespace string

	// Source is the library module's text.
	Source string

	// BaseURI is the static base URI of this module, used for the module's
	// own relative resolution. It may be empty.
	BaseURI string
}

A Module is a library module registered with the compilation directly, rather than fetched through a resolver.

This is the store §4.12 leaves to the implementation, and it is the only way to import a module without granting the query any reach at all: the caller supplies the source, so nothing is opened and nothing is fetched. It is also what an import with no "at" clause resolves against, since such an import names a namespace and nothing else.

type ModuleResolver added in v1.3.0

type ModuleResolver interface {
	// Resolve returns the source of the library module whose target
	// namespace is namespace. The hints are the import's "at" locations,
	// resolved relative to base when they are not absolute.
	//
	// Returning a nil reader and a nil error means "no such module", which
	// the caller reports as XQST0059. It is not a way to succeed quietly:
	// a query whose import found nothing has an incomplete static context,
	// and evaluating against that is exactly what this package refuses to
	// do.
	Resolve(namespace string, hints []string, base string) (io.ReadCloser, string, error)
}

A ModuleResolver locates the source of a library module.

The target namespace identifies the module; the location hints are the "at" clause of the import, in the order written, and may be empty. §4.12 lets a processor use a catalogue, a preloaded module, or nothing at all — which is why this is an interface rather than a built-in fetch. Following a location means reading whatever the query names, and only the caller can say whether that is allowed.

type Options

type Options struct {
	// BaseURI is the static base URI. It is stamped on constructed elements
	// and is what a relative reference resolves against.
	BaseURI string

	// DeclarationBaseURI is the URI of the resource the query text was read
	// from. It is used for one thing only: resolving a relative
	// "declare base-uri" against it, as §4.5 requires.
	//
	// It is deliberately separate from BaseURI. The two answer different
	// questions -- "what does a prolog declaration resolve against" and "what
	// does the query run under" -- and conflating them is what made every
	// earlier attempt at K2-BaseURIProlog-4 a net loss: one value for both
	// also stamped that value on constructed elements and on
	// fn:static-base-uri when the query declared nothing, which the suite's
	// base-URI-12/14/23/24 and K2-BaseURIFunc-30 all detect. Kept apart,
	// setting this one changes nothing except the resolution a declaration
	// asks for: a query with no "declare base-uri" is unaffected by it.
	DeclarationBaseURI string

	// BoundarySpace decides whether whitespace that only separates markup
	// survives into a constructed element. The zero value strips it, which is
	// what a query with no "declare boundary-space" gets.
	BoundarySpace BoundarySpace

	// Construction decides whether a copied node keeps its type annotation.
	// The zero value preserves it.
	Construction Construction

	// DefaultElementNamespace is applied to an unprefixed element name. It is
	// never applied to an attribute name.
	DefaultElementNamespace string

	// Namespaces are prefix bindings available to the query, as though its
	// prolog had declared them. The eight predeclared prefixes of §4.1 —
	// xml, xs, xsi, fn, local, math, map and array — are bound already, as is
	// err, which §3.16 binds so that "catch err:FODC0002" works with no
	// declaration. None of them need to appear here.
	Namespaces map[string]string

	// Modules are library modules available to "import module" (§4.12),
	// registered by target namespace with their source text.
	//
	// This is the module store the specification leaves to the
	// implementation, and it is the only way to import a module that grants
	// the query no reach whatever: the caller supplies the text, so nothing
	// is opened and nothing is fetched. An import with no "at" clause names
	// a namespace and nothing else, and resolves against this.
	//
	// The store is consulted before ModuleResolver, so a registered module
	// shadows any location a query might name for that namespace.
	Modules []Module

	// ModuleResolver locates a library module that Modules does not have.
	// When nil, NOTHING IS FETCHED: an "at" location is never opened, and an
	// import that the store cannot answer raises XQST0059.
	//
	// It is off by default for the reason xsd.Options.Resolver is — it hands
	// control of what this process reads to whoever wrote the query — and the
	// exposure here is worse than a schema's, because an "at" location is a
	// string chosen by the query's author and a query is the more commonly
	// untrusted input of the two. A host that sets this is granting the
	// queries it compiles the reach the resolver has, and should give one
	// that is rooted or table-driven rather than one that will open anything.
	//
	// MapModuleResolver answers from memory and follows no location at all.
	ModuleResolver ModuleResolver

	// MaxModules bounds how many library modules one compilation may load,
	// counting those reached transitively. A module that imports two modules
	// that each import two more is a fan-out with no natural bound, in the
	// same shape as a schema's include graph. Zero means DefaultMaxModules.
	//
	// Exceeding it FAILS the compilation with an error wrapping
	// xdm.ErrResourceLimit. It never yields a query compiled against the
	// modules that fitted: a partial static context is how an import comes to
	// look successful while half a library is missing.
	MaxModules int

	// MaxModuleBytes bounds the total source text one compilation may read
	// through ModuleResolver and Modules, cumulatively rather than per
	// module — a budget spent one module at a time is not spent at all. Zero
	// means DefaultMaxModuleBytes.
	//
	// Exceeding it fails the compilation with an error wrapping
	// xdm.ErrResourceLimit, on the same reasoning as MaxModules: a truncated
	// module is a module whose declarations are partly missing.
	MaxModuleBytes int64

	// Schemas are schemas available to "import schema" (§4.11), registered
	// by target namespace.
	//
	// This is the schema store the specification leaves to the
	// implementation, and it is the only way to import a schema that grants
	// the query no reach whatever: the caller supplies the components or the
	// text, so nothing is opened and nothing is fetched. An import with no
	// "at" clause names a namespace and nothing else, and resolves against
	// this.
	//
	// The store is consulted before SchemaResolver, so a registered schema
	// shadows any location a query might name for that namespace.
	Schemas []Schema

	// SchemaResolver locates a schema document that Schemas does not have.
	// When nil, NOTHING IS FETCHED: an "at" location is never opened, and an
	// import that the store cannot answer raises XQST0059.
	//
	// It is off by default for the same reason ModuleResolver is: an "at"
	// location is a string chosen by the query's author, and a query is
	// untrusted input. The same resolver is handed to xsd for the imported
	// schema's own xs:include and xs:import, so an imported schema can reach
	// no further than the query's import was granted.
	SchemaResolver SchemaResolver

	// MaxSchemaBytes bounds the total schema source text one compilation may
	// read through SchemaResolver and Schemas, cumulatively rather than per
	// import. Zero means DefaultMaxSchemaBytes.
	//
	// Exceeding it fails the compilation with an error wrapping
	// xdm.ErrResourceLimit rather than XQST0059: a truncated schema is a
	// schema whose components are partly missing, and compiling against a
	// partial static context is what this package refuses to do.
	MaxSchemaBytes int64
}

Options configure how a query is compiled.

The zero value is the specification's defaults: boundary whitespace is stripped, construction preserves types, an unprefixed element name is in no namespace, and an unprefixed function name is in the fn: namespace.

type Ordering

type Ordering int

Ordering is the ordering mode of §4.6.

const (
	// Ordered requires the document order a path expression would give. It
	// is the default.
	Ordered Ordering = iota
	// Unordered permits any order.
	Unordered
)

type Query

type Query struct {
	// contains filtered or unexported fields
}

A Query is a compiled query, safe for concurrent use.

Compiling separates what can be decided from the text — namespaces, the shape of every constructor, and every expression in it — from what depends on the input. Nothing about a Query changes when it runs, so one may be evaluated from several goroutines at once.

func Compile

func Compile(src string, opts Options) (*Query, error)

Compile compiles a query.

What is implemented is the prolog, the constructors, FLWOR and the XQuery-only expression forms: direct and computed constructors, every FLWOR clause including group by, order by and the two window clauses, try/catch, switch, typeswitch, quantified and ordered/unordered expressions, the extension expression and the string constructor. A query that is only an expression compiles too, since every XPath 3.1 expression is an XQuery expression.

"import module" is implemented (§4.12): a library module is found in Options.Modules or through Options.ModuleResolver, and contributes its public functions and variables. Nothing is fetched unless a resolver was configured -- see Options.ModuleResolver.

"import schema" is implemented (§4.11): a schema is found in Options.Schemas or through Options.SchemaResolver, and its type and declaration names reach the static context before the query body is parsed, so "cast as my:t", "instance of my:t", "element(*, my:t)", "schema-element(my:e)" and "validate" are all judged against it. Nothing is fetched unless a resolver was configured -- see Options.SchemaResolver.

Example

A compiled Query is immutable and safe to evaluate concurrently, so the cost of parsing is paid once however many times it runs.

package main

import (
	"os"

	"github.com/knroy/go-xml/xpath"
	"github.com/knroy/go-xml/xquery"
	"github.com/knroy/go-xml/xslt"
)

func main() {
	q, err := xquery.Compile(`<sum>{ 1 + 2 }</sum>`, xquery.Options{})
	if err != nil {
		panic(err)
	}
	seq, err := q.Eval(xpath.NewContext(nil, xpath.Builtins()))
	if err != nil {
		panic(err)
	}
	// A query returns a sequence; serialising it is a separate step.
	if err := xslt.Serialize(os.Stdout, seq,
		xslt.OutputSettings{OmitXMLDecl: true}, nil); err != nil {
		panic(err)
	}
}
Output:
<sum>3</sum>

func (*Query) Eval

func (q *Query) Eval(ctx *xpath.Context) (xdm.Sequence, error)

Eval runs the query and returns its result sequence.

ctx supplies the context item, the variable bindings and the function library, exactly as it does for an XPath expression. A nil function library is legal and means the query may not call anything.

Example (ExternalVariable)

An external variable is declared by the query and bound by the caller. The binding lives on the context rather than on the Query, which is what lets one compiled query serve many different bindings at once.

package main

import (
	"fmt"

	"github.com/knroy/go-xml/xdm"
	"github.com/knroy/go-xml/xpath"
	"github.com/knroy/go-xml/xquery"
)

func main() {
	q, err := xquery.Compile(
		`declare variable $who external; concat("hello ", $who)`,
		xquery.Options{})
	if err != nil {
		panic(err)
	}
	ctx := xpath.NewContext(nil, xpath.Builtins())
	ctx.Vars = map[string]xdm.Sequence{"who": {xdm.NewString("world")}}

	seq, err := q.Eval(ctx)
	if err != nil {
		panic(err)
	}
	fmt.Println(seq)
}
Output:
[hello world]

func (*Query) SerializationOptions

func (q *Query) SerializationOptions() map[string]string

SerializationOptions returns the serialization parameters the prolog declared, keyed by the parameter's local name and carrying its lexical value.

Evaluating a query and serialising its result are two steps, and only the first belongs here: Eval hands back a sequence, and what a caller does with it — write it as XML, as JSON, as nothing at all — is the caller's decision. But the *parameters* for that second step are stated in the query, by "declare option output:method" and its siblings (XQuery 3.1 §2.2.4), and a caller that never sees them would have to re-parse the prolog to find out what the query asked for. This is how it asks instead.

The values are unvalidated lexical forms. Whether "indent" says "yes" or something meaningless is the serialiser's judgement to make, since only it knows which parameters it honours; the names, however, are checked at compile time, an unknown one being XQST0109.

The returned map is a copy, so a caller may keep or modify it without disturbing the Query, which is otherwise immutable and safe for concurrent use. Nil is returned when the prolog declared none.

Example

The prolog can state how the result should be serialised. Evaluating and serialising are separate steps, so those parameters are handed back rather than acted on -- the caller decides whether to honour them.

package main

import (
	"fmt"

	"github.com/knroy/go-xml/xquery"
)

func main() {
	// "output" is not a predeclared prefix, so it must be bound first.
	q, err := xquery.Compile(`
		declare namespace output =
			"http://www.w3.org/2010/xslt-xquery-serialization";
		declare option output:method "json";
		1`, xquery.Options{})
	if err != nil {
		panic(err)
	}
	fmt.Println(q.SerializationOptions()["method"])
}
Output:
json

func (*Query) String

func (q *Query) String() string

String returns the query's source.

type Schema added in v1.3.0

type Schema struct {
	// Namespace is the target namespace this schema supplies. An import of
	// this namespace resolves to this entry.
	//
	// It is matched against the import's target namespace rather than against
	// the schema's own targetNamespace attribute: a caller registering a
	// Schema has said which namespace it answers for, and Components is
	// already-assembled rather than a document to re-read.
	Namespace string

	// Components is the assembled schema. A caller that has one from
	// xsd.Load, xsd.LoadFile or xslt's Stylesheet.Schema can register it
	// directly, which is what makes a host able to share one schema between a
	// stylesheet and a query without loading it twice and risking the two
	// disagreeing.
	//
	// It may be nil, which registers the namespace as known and empty. That
	// is not a useless entry: §4.11 makes importing a namespace legal whether
	// or not the processor has components for it, and a nil entry is how a
	// host says "this namespace is expected and I have nothing for it"
	// without the import failing.
	Components *xsd.Schema

	// Source is schema document text to load, used when Components is nil.
	// It is the convenient half of the store: a caller with a schema document
	// in a string need not call xsd.Load itself.
	Source string

	// BaseURI is what Source's own relative references resolve against. It
	// may be empty.
	BaseURI string
}

A Schema is a schema registered with the compilation directly, rather than fetched through a resolver.

This is the store §4.11 leaves to the implementation, and it is the only way to import a schema without granting the query any reach at all: the caller supplies the components, so nothing is opened and nothing is fetched. It is also what an import with no "at" clause resolves against, since such an import names a namespace and nothing else.

type SchemaResolver added in v1.3.0

type SchemaResolver = xsd.Resolver

A SchemaResolver locates the source of a schema document.

It is xsd.Resolver rather than a new interface of this package's own. The two would have had identical shapes, and a caller that already holds a resolver for xsd.Load or for xslt's SchemaResolver would have had to wrap it to pass it here. Sharing the type also means an imported schema's own xs:include and xs:import are followed by the SAME resolver the query's import was granted, rather than by a second one that could disagree about what this process may read.

type XQVersion added in v1.2.1

type XQVersion int

XQVersion is the version of the XQuery language a module is written in, as named by its version declaration (§4.1 VersionDecl).

This is not the same thing as xpath.Version, and the two cannot be collapsed into one. xpath.Version selects a version of the *expression* language, and XQuery's expression language is XPath's: XQuery 1.0's is XPath 2.0's, 3.0's is XPath 3.0's, 3.1's is XPath 3.1's. But XQuery has rules of its own that XPath has no counterpart for -- the prolog, the module system, node construction -- and those changed on XQuery's own schedule. An unprefixed "declare option" name is the clearest case: it is XPST0081 in XQuery 1.0 and legal from 3.0, and XPath has no option declaration at all, so no value of xpath.Version can carry that distinction. Keeping the two types separate is what lets a decision point say which language's rule it is applying.

The constants are in version order, so that "later than" is a comparison and the predicates below are the only place that knows which comparison. The zero value is therefore XQuery10 and is never what a module gets: newStaticContext sets the default explicitly, because the default is a policy decision about undeclared modules and not an accident of ordering.

const (
	// XQuery10 is XQuery 1.0, the 2007 Recommendation.
	XQuery10 XQVersion = iota
	// XQuery30 is XQuery 3.0, the 2014 Recommendation. It admits everything
	// in 1.0, and changes the answer to a handful of questions 1.0 had
	// already answered -- see the callers of atLeast30.
	XQuery30
	// XQuery31 is XQuery 3.1, the 2017 Recommendation. It is what this engine
	// implements, and what a module with no version declaration is compiled
	// as: §4.1 leaves the version of such a module implementation-defined.
	XQuery31
)

func (XQVersion) String added in v1.2.1

func (v XQVersion) String() string

String gives the version back in the spelling a version declaration uses, so an error message can quote what the module asked for.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL