Documentation
¶
Overview ¶
Package xdm implements the XQuery/XPath Data Model (XDM) that XPath 2.0 and XSLT 2.0 are defined over.
The central difference from XPath 1.0 is that every value is a *sequence* of items, and an item is either a node or a typed atomic value. XPath 1.0 had four types (node-set, string, number, boolean) with implicit coercion everywhere; 2.0 has the full XML Schema datatype hierarchy with explicit promotion rules. Modelling that faithfully here is what lets the rest of the engine avoid the 1.0-style "just call ToString" shortcuts that make 2.0 stylesheets silently produce wrong answers.
Index ¶
- Constants
- Variables
- func AnnotationLocal(annotation string) string
- func AnnotationName(uri, local string) string
- func CompareDT(a, b *DateTime, implicitTZ int) int
- func DerivedBase(name string) string
- func ErrCast(format string, args ...any) error
- func ErrType(format string, args ...any) error
- func ErrorCode(err error) string
- func Errorf(code, format string, args ...any) error
- func HasSimpleTypeAnnotation(annotation string) bool
- func IsGregorian(t TypeCode) bool
- func IsNCName(s string) bool
- func IsNameChar(r rune) bool
- func IsNameStartChar(r rune) bool
- func IsQualifiedAnnotation(annotation string) bool
- func IsXMLWhitespace(s string) bool
- func LexicalGregorian(dt *DateTime, t TypeCode) string
- func ListItemOf(name string) string
- func MapKeyOf(a *Atomic) (string, error)
- func ProcessXInclude(tree *Tree, opts XIncludeOptions) error
- func RegisterDerivedType(name, primitive string)
- func RegisterListType(name, itemType string)
- func RegisterUnionType(name string, members []string)
- func SplitAnnotationName(annotation string) (uri, local string)
- func SplitQName(s string) (prefix, local string)
- func SplitXMLSpace(s string) []string
- func TrimXMLSpace(s string) string
- func UnionMembersOf(name string) []string
- func ValidateXIncludeHref(href string) error
- type ArrayItem
- type Atomic
- func NewAnyURI(s string) *Atomic
- func NewBinary(s string, t TypeCode) *Atomic
- func NewBoolean(v bool) *Atomic
- func NewDateTime(dt *DateTime, t TypeCode) *Atomic
- func NewDecimal(r *big.Rat) *Atomic
- func NewDouble(v float64) *Atomic
- func NewDuration(d *Duration, t TypeCode) *Atomic
- func NewFloat(v float64) *Atomic
- func NewGregorian(dt *DateTime, t TypeCode) *Atomic
- func NewInteger(v int64) *Atomic
- func NewIntegerFromRat(r *big.Rat) *Atomic
- func NewQNameValue(q QName) *Atomic
- func NewString(s string) *Atomic
- func NewUntypedAtomic(s string) *Atomic
- func (a *Atomic) Bool() bool
- func (a *Atomic) DateTimeVal() *DateTime
- func (a *Atomic) Derived() string
- func (a *Atomic) DerivedMember() string
- func (a *Atomic) DurationVal() *Duration
- func (a *Atomic) FitsInt64() bool
- func (a *Atomic) Float64() float64
- func (a *Atomic) Int64() int64
- func (a *Atomic) IsNaN() bool
- func (a *Atomic) QName() *QName
- func (a *Atomic) Rat() *big.Rat
- func (a *Atomic) Str() string
- func (a *Atomic) String() string
- func (a *Atomic) TypeEnv() *TypeEnvironment
- func (a *Atomic) TypeName() string
- func (a *Atomic) WithDerived(name string) *Atomic
- func (a *Atomic) WithDerivedUnion(name, member string) *Atomic
- func (a *Atomic) WithTypeEnv(e *TypeEnvironment) *Atomic
- type DateTime
- type DecimalMagnitude
- type Duration
- type EntityBudget
- type EntityResolver
- type Error
- type FunctionItem
- type IdentityKey
- type IncludeResolver
- type Item
- type MapBuilder
- type MapItem
- func (m *MapItem) Entries(f func(key *Atomic, value Sequence) error) error
- func (m *MapItem) Get(key *Atomic) (Sequence, bool, error)
- func (m *MapItem) Keys() []*Atomic
- func (m *MapItem) Len() int
- func (m *MapItem) Put(key *Atomic, value Sequence) (*MapItem, error)
- func (m *MapItem) Remove(key *Atomic) (*MapItem, error)
- func (m *MapItem) RemoveAll(keys []*Atomic) (*MapItem, error)
- func (m *MapItem) TypeName() string
- type Node
- func (n *Node) AddAttr(a *Node)
- func (n *Node) AddNamespace(prefix, uri string)
- func (n *Node) AppendChild(c *Node)
- func (n *Node) ApplyTyping(t Typing)
- func (n *Node) Atomize() *Atomic
- func (n *Node) AtomizeList() (Sequence, bool)
- func (n *Node) Attr(uri, local string) *Node
- func (n *Node) AttrValue(local string) string
- func (n *Node) ChildElements() []*Node
- func (n *Node) Compare(o *Node) int
- func (n *Node) CopyTypingFrom(src *Node)
- func (n *Node) CopyTypingStrippedFrom(src *Node)
- func (n *Node) Identity() IdentityKey
- func (n *Node) InScopeNamespaces() map[string]string
- func (n *Node) Is(o *Node) bool
- func (n *Node) IsElement(uri, local string) bool
- func (n *Node) LookupPrefix(prefix string) (string, bool)
- func (n *Node) Order() int
- func (n *Node) Position() (line, col int, ok bool)
- func (n *Node) Root() *Node
- func (n *Node) SetSynthesizedOrder(owner *Node, offset int)
- func (n *Node) SetTypeAnnotation(annotation string)
- func (n *Node) SetTypeAnnotationResolved(annotation, derivedPrimitive, listItem string)
- func (n *Node) SetTypeEnv(e *TypeEnvironment)
- func (n *Node) StringValue() string
- func (n *Node) StripTyping()
- func (n *Node) Tree() *Tree
- func (n *Node) TypeEnv() *TypeEnvironment
- func (n *Node) TypeName() string
- type NodeKind
- type Opaque
- type ParseOptions
- type QName
- type Sequence
- func Atomize(seq Sequence) Sequence
- func AtomizeChecked(seq Sequence) (Sequence, error)
- func Concat(seqs ...Sequence) Sequence
- func Empty() Sequence
- func Except(a, b Sequence) Sequence
- func Flatten(seq Sequence) Sequence
- func Intersect(a, b Sequence) Sequence
- func One(it Item) Sequence
- func SortDocumentOrder(seq Sequence) Sequence
- func Union(a, b Sequence) Sequence
- type Tree
- type TypeCode
- type TypeEnvironment
- func (e *TypeEnvironment) DerivedBase(name string) string
- func (e *TypeEnvironment) Len() (derived, lists, unions int)
- func (e *TypeEnvironment) ListItemOf(name string) string
- func (e *TypeEnvironment) Merge(src *TypeEnvironment)
- func (e *TypeEnvironment) RegisterDerived(name, primitive string)
- func (e *TypeEnvironment) RegisterList(name, itemType string)
- func (e *TypeEnvironment) RegisterUnion(name string, members []string)
- func (e *TypeEnvironment) UnionMembersOf(name string) []string
- type Typing
- type VariadicSignature
- type XIncludeOptions
Constants ¶
const ( // DefaultMaxDepth is the nesting limit. DefaultMaxDepth = 1000 // DefaultMaxBytes is the source-size limit: 64 MB, far above any // schema or stylesheet and above most real instance documents, while // still bounding what a single parse can be asked to read. DefaultMaxBytes int64 = 64 << 20 // DefaultMaxNodes is the node-count limit. At roughly 200 bytes a node // this bounds a tree to about 2 GB, which is the point of it: the // number is chosen to bound *memory*, and it is the limit that actually // binds on the documents designed to be expensive. DefaultMaxNodes = 10_000_000 )
Limits applied when the corresponding ParseOptions field is zero.
const ( NSXSL = "http://www.w3.org/1999/XSL/Transform" NSXML = "http://www.w3.org/XML/1998/namespace" NSXMLNS = "http://www.w3.org/2000/xmlns/" NSXS = "http://www.w3.org/2001/XMLSchema" NSXSI = "http://www.w3.org/2001/XMLSchema-instance" NSFN = "http://www.w3.org/2005/xpath-functions" NSMath = "http://www.w3.org/2005/xpath-functions/math" NSArray = "http://www.w3.org/2005/xpath-functions/array" NSMap = "http://www.w3.org/2005/xpath-functions/map" NSErr = "http://www.w3.org/2005/xqt-errors" NSSVRL = "http://purl.oclc.org/dsdl/svrl" NSSchema = "http://purl.oclc.org/dsdl/schematron" // NSGoxslt is this engine's extension namespace. Extensions live outside // the fn: namespace so that a stylesheet written for another processor // cannot silently pick one up in place of a standard function, and so // that a stylesheet using them is visibly engine-specific. NSGoxslt = "https://github.com/knroy/go-xml" )
Well-known namespace URIs used throughout the engine.
const NSXInclude = "http://www.w3.org/2001/XInclude"
NSXInclude is the XInclude namespace. XInclude 1.0 section 3: "elements in the XInclude namespace ... http://www.w3.org/2001/XInclude".
Variables ¶
var ErrResourceLimit = errors.New("processor resource limit exceeded")
ErrResourceLimit marks an error raised because the processor refused to do the work, not because the input was wrong.
"This XPath is malformed" and "this document would have cost more than the processor is willing to spend" are different conditions with different remedies, but they arrive looking alike: most limits have no code of their own and borrow a semantic one, so an embedding caller reading the code alone concludes the input was wrong when in fact it was merely expensive. XPath's nesting guards report XPDY0130, which §2.3.1 does define for a limit, but that code also covers evaluation budgets, so even there the code alone does not say which limit; a sentinel can.
Wrap with %w to add it, never to replace the code:
fmt.Errorf("XPDY0130: expression nesting exceeds %d levels: %w",
maxParseDepth, ErrResourceLimit)
The rendered message still begins with the code, so ErrorCode and the conformance suites, which read the code out of the message, are unaffected, while a caller can now ask errors.Is(err, ErrResourceLimit) and back off rather than reporting a syntax error to its user.
Functions ¶
func AnnotationLocal ¶ added in v1.0.0
AnnotationLocal returns just the local part of an annotation key.
func AnnotationName ¶ added in v1.0.0
AnnotationName builds the key a type annotation is recorded and compared under.
The data model keys annotations by a single string, and that string used to be the type's bare local name. That conflated every type sharing a local part across namespaces, and the conflation was not theoretical: the W3C's own schema-for-xslt20.xsd deliberately declares an xsl:QName of its own, as a restriction of xs:Name, and says in its text why ("This schema does not use the built-in type xs:QName... a schema processor would expand unprefixed QNames incorrectly"). Keyed by "QName", loading that schema overwrote the built-in's entry in a package-level, process-global map, so every later schema in the same process saw xs:QName deriving from xs:Name.
Both meanings have to coexist rather than one displacing the other: import-schema-029 asserts that the SHADOWING xsl:QName does erase to a string, while type-functions-0501 asserts that the built-in xs:QName still atomises to a QName value. A registration that refuses to shadow breaks the first; a global "something shadowed a built-in" flag breaks the second, because one flag cannot hold two answers at once. Only qualifying the key separates them.
The encoding is Clark notation, {uri}local, which the codebase already spells through QName.Clark, with one deliberate exception: a type in the XML Schema namespace keys under its BARE local name. That exception is what keeps the change tractable. Built-in annotations are compared against bare literals — "QName", "NOTATION", "ID", "string" — at roughly a hundred sites across four packages, in switch statements, map lookups and equality tests. Qualifying them would have required rewriting every one of those, whereas leaving them bare means only names that were previously AMBIGUOUS change spelling, and a built-in's key is the same string it has always been.
The empty URI also keys bare, which is the no-namespace case and is already unambiguous.
func DerivedBase ¶ added in v1.0.0
DerivedBase returns the type a schema type derives from, or "" if the name is not a registered schema type. The name is an annotation name, and so is the result, so a chain can be walked by feeding one back in.
It is what makes the subtype relation work for schema types: a value annotated as a restriction of xs:NOTATION is an instance of xs:NOTATION as well as of its own type, and answering that means walking the chain the schema recorded.
func ErrType ¶
ErrType is the XPath type error, XPTY0004. It is returned rather than panicked so that a stylesheet error degrades one transform.
func ErrorCode ¶
ErrorCode returns the spec error code carried by err, or "" if it has none.
It unwraps, so a code survives being wrapped with fmt.Errorf("%w"). Errors produced before this type existed still carry their code as a message prefix, so those are recognised too rather than silently reporting "".
func HasSimpleTypeAnnotation ¶ added in v1.0.0
HasSimpleTypeAnnotation reports whether an annotation names a simple type, or a complex type with simple content.
XSLT 2.0 section 4.4 preserves whitespace-only text in such an element *regardless* of xsl:strip-space: that text is the element's entire typed value, which the schema validated, and stripping it would leave a node whose annotation describes a value it no longer holds. An element with element-only or mixed content has no such value and is stripped normally.
The registration table is the oracle rather than a list of names, because the annotation on such an element is the built-in its content type erases to — "string" for both an element of type xs:string and one whose anonymous complex type extends xs:string, which is exactly the pair section 4.4 groups together. A complex type with element-only content registers no derivation to a built-in, so it answers false, which is the distinction being drawn.
func IsGregorian ¶
IsGregorian reports whether t is one of the five Gregorian types.
func IsNCName ¶
IsNCName reports whether s is an XML non-colonised name: a name with no prefix, which is what an element, attribute or processing-instruction name must be once the prefix has been split off.
It lives here rather than in a consumer because more than one caller needs it, and because the cost of *not* checking is that a computed name reaches the serialiser unvalidated. A name is written to output as-is, so a name holding "><script>" produces markup rather than a name — output that is either malformed or, in HTML, an injected element.
func IsNameChar ¶ added in v1.2.0
IsNameChar reports whether r may appear in an XML Name after the first character. See IsNameStartChar for why this is exported.
func IsNameStartChar ¶ added in v1.2.0
IsNameStartChar reports whether r may begin an XML Name.
Exported for the sibling packages that parse names out of source rather than out of markup: an XQuery direct element constructor spells a QName in the query text, and needs the same production the XML parser uses rather than an approximation of it.
func IsQualifiedAnnotation ¶ added in v1.0.0
IsQualifiedAnnotation reports whether an annotation key carries a namespace, which is to say it names a type that is neither a built-in nor in no namespace.
func IsXMLWhitespace ¶
IsXMLWhitespace reports whether s consists entirely of XML whitespace.
XML defines whitespace as exactly four characters: space, tab, carriage return and line feed. Go's strings.TrimSpace uses unicode.IsSpace, which additionally matches U+00A0 (no-break space) and other Unicode separators — so using it to decide whether a text node is "just whitespace" silently deletes a that the author put there deliberately.
func LexicalGregorian ¶
LexicalGregorian returns the canonical lexical form.
func ListItemOf ¶ added in v1.0.0
ListItemOf returns the item type registered for a list type, or "" when the name is not a registered list.
func MapKeyOf ¶ added in v1.1.0
MapKeyOf returns the canonical form under which a key is compared.
Two keys are the same key when they are equal under the "eq" operator with no type promotion beyond the numeric hierarchy, so 1 and 1.0 collide while "1" stands apart. Encoding that as a string keeps the lookup a plain map access rather than a scan with a comparison function.
This is NOT the same relation as xpath.GroupingKey, and the difference is deliberate rather than an omission. Grouping substitutes the implicit timezone into an unzoned value, so xs:date("2015-04-08") and xs:date("2015-04-08Z") fall in one group. A map key must not: same-key-013, -014 and -015 build a three-entry map from exactly that pair and require all three entries to survive, because a key that depended on the implicit timezone would give one map different sizes in different dynamic contexts. The two keys answer different questions; neither is the other lagging behind.
The relation this encodes is stated directly as SameKey in samekey_oracle_test.go, and TestMapKeyOfMatchesSameKey asserts that the encoding and the relation agree in both directions. Change one and that test tells you whether the other has to follow.
func ProcessXInclude ¶ added in v1.2.1
func ProcessXInclude(tree *Tree, opts XIncludeOptions) error
ProcessXInclude performs XInclude processing on tree in place.
The tree is modified rather than copied. XInclude 1.0 section 4 is written as a transformation from one infoset to another, and building a second tree would be the more literal reading — but every node in this package carries a pointer to its Tree and its document order, so a copy would have to rebuild both anyway, and the caller's other references to the tree would then point at the *unincluded* document. Modifying in place and re-finalising is the behaviour a caller loading a document actually wants.
The tree is re-finalised before returning, so document order is correct over the merged content. Callers must not hold node identities across this call.
func RegisterDerivedType ¶ added in v1.0.0
func RegisterDerivedType(name, primitive string)
RegisterDerivedType records that a schema type erases to a built-in one.
Both arguments are annotation names, which AnnotationName builds; passing a bare local name for a type that has a namespace re-creates the conflation this keying exists to prevent.
The xsd package calls this as it loads a schema, so that a node annotated with a user-defined type still atomises to a typed value rather than to untypedAtomic. Without it, "instance of my:partNumberType" could never be true for a value read out of a validated document, because the value would have discarded the annotation on the way out of the tree.
func RegisterListType ¶ added in v1.0.0
func RegisterListType(name, itemType string)
RegisterListType records that a schema type is a list, and what its items are.
The xsd package calls this as it loads a schema, for the same reason it calls RegisterDerivedType: the typed value of a list-typed node is a SEQUENCE of one atomic per token, and nothing in the data model can work out from a bare type name that "numbers" is a list of xs:decimal. Without it a list-typed node atomises to one untypedAtomic holding the whole literal, so count(data(@list)) answers 1 and "data(@list) instance of xs:untypedAtomic" answers true for a node the schema plainly gave a typed value.
Both arguments are annotation names, as RegisterDerivedType's are.
itemType is the item type's own name, which may itself be a registered schema type; atomicForAnnotation and the derivation walk resolve it.
func RegisterUnionType ¶ added in v1.0.0
RegisterUnionType records that a schema type is a union, and what its member types are.
The xsd package calls this as it loads a schema, for the reason it calls RegisterListType: a union's base is always xs:anySimpleType, so the derivation chain RegisterDerivedType records dead-ends immediately and carries no information about what the value actually is. Without the member list a union-typed node atomises to xs:untypedAtomic — the walk finds anySimpleType, cannot build a value for it, and gives up — which makes "data(u) instance of xs:untypedAtomic" true for a node the schema plainly gave a typed value, and makes every question about the member it validated as answer false.
The name and every member are annotation names, as RegisterDerivedType's arguments are. This registry is keyed by the same strings as derivedPrimitives and listItems, so qualifying one of the three and not the others would leave unions silently unresolvable.
The members are the *declared* members, in declaration order; which of them a given value belongs to is a per-value fact recorded on the node, because XSD 1.0 §3.14.4 chooses the member by trying each one's lexical space against the value in turn.
func SplitAnnotationName ¶ added in v1.0.0
SplitAnnotationName is the inverse of AnnotationName: it returns the namespace URI and local part of an annotation key.
A bare key is a built-in or a no-namespace type, and the two are told apart by nothing here — the URI comes back empty for both, because the callers that care (the built-in switches) match on the local part they already expect. What this function exists for is the comparison path, which needs the local part of a qualified key without mistaking "{uri}local" for a prefixed lexical QName.
It must be used in place of SplitQName wherever the input is an annotation. SplitQName cuts at the first colon, so handed "{http://x}foo" it returns the prefix "{http" and the local part "//x}foo" — nonsense rather than an error, and silently wrong.
func SplitQName ¶
SplitQName splits a lexical QName into prefix and local part. It does not resolve the prefix; resolution needs a namespace context and is done by the caller that has one.
func SplitXMLSpace ¶ added in v1.3.0
SplitXMLSpace splits s on runs of XML whitespace, dropping empty tokens.
This is the tokenization XML Schema's whiteSpace="collapse" and every whitespace-separated list type (IDREFS, NMTOKENS, ENTITIES and any xs:list) are defined in terms of. It is deliberately not strings.Fields, which splits on the whole Unicode White_Space set: a no-break space inside a list value is data, not a separator, so Fields turned the one token "a<NBSP>b" into the two tokens "a" and "b".
func TrimXMLSpace ¶ added in v1.3.0
TrimXMLSpace removes leading and trailing XML whitespace, and nothing wider.
func UnionMembersOf ¶ added in v1.0.0
UnionMembersOf returns the member types of a registered union type, or nil when the name does not denote one.
The result must not be modified: it is the stored slice, shared with every other caller.
func ValidateXIncludeHref ¶ added in v1.2.1
ValidateXIncludeHref reports whether an href is one this package will accept before a resolver is consulted.
XInclude 1.0 section 4.1.1 forbids a fragment identifier in href: "the value of the href attribute must not contain a fragment identifier", because the xpointer attribute is where a subresource is named. It is a fatal error rather than a fallback condition, since it is a defect in the including document rather than a property of the resource.
Types ¶
type ArrayItem ¶ added in v1.1.0
type ArrayItem struct {
// contains filtered or unexported fields
}
ArrayItem is the fifth kind of XDM item, added in XPath 3.1.
An array holds a sequence of *members*, each of which is itself a sequence. That is the difference from a sequence, which is flat: [(1,2),(3)] has two members where (1,2,3) has three items, and the distinction survives every operation until fn:data or array:flatten deliberately removes it.
Like a map, an array is also a function item: "$a(1)" is the first member, which is what the lookup operator "$a?1" lowers to.
func (*ArrayItem) Member ¶ added in v1.1.0
Member returns the i'th member, counting from 1 as the data model does.
An index outside the array is FOAY0001, which is a different error from asking a map for a key it does not have: an array's positions are its whole domain, so a position outside them is a mistake rather than an absence.
type Atomic ¶
type Atomic struct {
Type TypeCode
// contains filtered or unexported fields
}
Atomic is a typed atomic value.
The representation is a tagged union rather than an interface per type. The evaluator switches on Type constantly — every arithmetic op, comparison and function call — and a type switch across seventeen concrete types in those hot paths costs more than a single integer compare. It also keeps the numeric tower in one place, where the promotion rules are easy to audit.
func NewBinary ¶
NewBinary returns an xs:hexBinary or xs:base64Binary holding the given lexical form.
The value keeps its own type rather than collapsing to xs:string, because the two binary types are inter-convertible: casting hexBinary to base64Binary has to re-encode the underlying octets, and a value that has forgotten which encoding its lexical form uses cannot be decoded.
func NewDateTime ¶
NewDateTime returns a date, time or dateTime atomic value.
func NewDecimal ¶
NewDecimal returns an xs:decimal holding an exact value.
func NewDuration ¶
NewDuration returns a duration atomic value of the given duration type.
func NewFloat ¶
NewFloat returns an xs:float. The value is rounded to float32 precision on construction, because xs:float operations must produce float32 results.
func NewGregorian ¶
NewGregorian returns one of the five Gregorian atomic values.
func NewInteger ¶
NewInteger returns an xs:integer. Integers are held as exact rationals so that they participate in decimal arithmetic without precision loss.
func NewIntegerFromRat ¶
NewIntegerFromRat returns an xs:integer from an exact rational, which must have denominator 1. Used by arithmetic that has already established integrality (idiv, string-length, count).
func NewUntypedAtomic ¶
NewUntypedAtomic returns an xs:untypedAtomic, the type produced by atomising a node in a document that has not been schema-validated.
func (*Atomic) DateTimeVal ¶
DateTimeVal returns the date/time value, or nil.
func (*Atomic) Derived ¶
Derived returns the narrower XML Schema type this value was constructed as, or "" if it was not built by a derived-type constructor.
func (*Atomic) DerivedMember ¶ added in v1.0.0
DerivedMember returns the union member type this value was validated as, or "" when the value's type is not a union.
It is a second answer alongside Derived, not a replacement for it: a value of a union type is an instance of both the union and the selected member.
func (*Atomic) DurationVal ¶
Duration returns the duration value, or nil.
func (*Atomic) FitsInt64 ¶
FitsInt64 reports whether the value can be represented as an int64 without wrapping.
xs:integer is arbitrary-precision, so this is a real question: Int64() truncates the big.Int and silently returns a different number, which is worse than refusing.
func (*Atomic) Float64 ¶
Float64 returns the value as a float64 for any numeric type. Decimal and integer values are converted, which may lose precision; callers doing exact arithmetic must use Rat instead.
func (*Atomic) IsNaN ¶
IsNaN reports whether a is a double or float NaN. NaN needs its own check throughout comparison, because it is the one value where the general "compare and negate" shortcut produces wrong answers.
func (*Atomic) String ¶
String returns the XPath 2.0 canonical lexical representation, which is what fn:string and every implicit string conversion must produce. It is not a debug format: the exact spelling of doubles and decimals here is observable in stylesheet output.
func (*Atomic) TypeEnv ¶ added in v1.3.0
func (a *Atomic) TypeEnv() *TypeEnvironment
TypeEnv returns the type environment of the schema that issued this value's derived annotation, or nil when no schema did.
func (*Atomic) WithDerived ¶
WithDerived returns a copy of a annotated as the named derived type.
The union member is cleared: re-annotating the value as a different type makes any member recorded for the previous one meaningless, and carrying it forward would let a value claim membership in a union it no longer has.
func (*Atomic) WithDerivedUnion ¶ added in v1.0.0
WithDerivedUnion returns a copy of a annotated as the named union type with the named member recorded as the one that accepted it.
func (*Atomic) WithTypeEnv ¶ added in v1.3.0
func (a *Atomic) WithTypeEnv(e *TypeEnvironment) *Atomic
WithTypeEnv returns a copy of a carrying the given type environment, so that questions about its derived type name are answered by the schema that issued the name rather than by the process-global table.
type DateTime ¶
type DateTime struct {
Year int // proleptic Gregorian; negative for BCE. No year zero.
Month int // 1-12
Day int // 1-31
Hour int // 0-24 (24 only as the lexical form 24:00:00)
Minute int // 0-59
Second *big.Rat // seconds including fraction, [0,60)
// TZOffset is the timezone offset in minutes east of UTC.
// HasTZ distinguishes "no timezone" from "+00:00", which are different
// values under XML Schema equality.
TZOffset int
HasTZ bool
}
DateTime represents xs:date, xs:time and xs:dateTime.
It is not time.Time. XML Schema dates carry three properties that time.Time cannot express: an optional timezone (distinct from UTC — an unzoned date is a different value from a UTC one), a year range that exceeds int64 nanoseconds, and second values with arbitrary fractional precision. Comparison of unzoned values against zoned ones is defined against an implicit timezone supplied by the dynamic context, which only works if "absent" is representable.
func ParseDateTime ¶
ParseDateTime parses the lexical form of xs:date, xs:time or xs:dateTime according to the requested type.
func ParseGregorian ¶
ParseGregorian parses the lexical form of one of the five Gregorian types.
Each has its own leading-hyphen convention — "--01" is a month, "---15" a day — which exists so that the forms cannot be confused with a truncated date. Getting the hyphen count wrong silently reinterprets the value, so each form is matched exactly rather than by a permissive scan.
func (*DateTime) ToSeconds ¶
ToSeconds returns the value as seconds since 1972-12-31T00:00:00Z, adjusted to UTC using implicitTZ (in minutes) when the value carries no timezone.
Comparison and subtraction are defined on this normalised timeline, so having one conversion point means the timezone rules are applied uniformly rather than re-derived at each comparison site.
type DecimalMagnitude ¶ added in v1.2.2
type DecimalMagnitude struct {
Coefficient *big.Int // absolute value; never negative
Scale int64 // fraction digits, exactly; never negative
}
DecimalMagnitude is a terminating decimal split into an integer coefficient and a decimal scale: the value is Coefficient / 10^Scale, with Coefficient the absolute value and Scale the exact number of fraction digits.
func DecimalMagnitudeOf ¶ added in v1.2.2
func DecimalMagnitudeOf(r *big.Rat) (DecimalMagnitude, bool)
DecimalMagnitudeOf converts a rational to its exact decimal form, reporting false if the value has no terminating decimal expansion.
The contract is strict on purpose. This primitive serves both a renderer and a validity decision, and a validity decision must never be handed a guess: a digit count invented for 1/3 would let it pass a fractionDigits facet it violates. So the only two answers are the exact one and an explicit false. Callers that want a fallback for a non-terminating value apply it themselves, visibly, at their own call site — see nonTerminatingScale in atomic.go.
There is deliberately no ceiling on the scale either. Capping it makes the lexical form disagree with the value: a cap of 18 printed a literal with 360 fraction digits as "0" while it compared unequal to zero, and raising the cap to 1024 moved the same contradiction to 10^-1025. The value is the thing that must not move, so the scale follows it.
big.Rat is kept in lowest terms, so a value is an exact decimal precisely when its denominator is 2^a*5^b with nothing left over, and then the scale is max(a, b) — no more digits are needed and no fewer will do. The factor 2^a comes off in a single shift. The factor 5^b is stripped by repeated squaring rather than one division per power: dividing off 5 at a time is quadratic in b, which costs 907ms for a 50000-digit value against 415µs here.
type Duration ¶
type Duration struct {
Negative bool
Months int // years*12 + months
Seconds *big.Rat // days*86400 + hours*3600 + minutes*60 + seconds
}
Duration represents xs:duration and its two subtypes.
XML Schema durations have two independent components — months and seconds — that cannot be converted into one another, because the number of days in a month is not fixed. That is why xs:duration is only partially ordered and why the two totally-ordered subtypes (xs:yearMonthDuration and xs:dayTimeDuration) exist. Keeping the components separate rather than normalising to a single scalar is what makes the ordering rules implementable at all.
func ParseDuration ¶
ParseDuration parses the lexical form of xs:duration, xs:yearMonthDuration or xs:dayTimeDuration, rejecting components the requested subtype does not permit.
func (*Duration) SignedMonths ¶
SignedMonths returns the month component with the sign applied.
func (*Duration) SignedSeconds ¶
SignedSeconds returns the second component with the sign applied.
type EntityBudget ¶ added in v1.3.0
type EntityBudget struct {
// contains filtered or unexported fields
}
EntityBudget is an entity-expansion allowance that spans more than one parse, handed to ParseOptions.WithEntityBudget.
It exists for a caller OUTSIDE this package that parses repeatedly on behalf of one larger operation, and for which "one document" is therefore the wrong boundary. fn:parse-xml is the case that forced it: it is an ordinary function in the default library, so an expression calls it once per node, and each call built its own entityTable and so its own fresh allowance. Sixty calls to a bomb that each stayed under the 1 MB ceiling therefore expanded 47,185,920 bytes from 1,328 bytes of XPath and were accepted; three hundred allocated 898 MB. Neither xpath's item budget nor its byte budget can see it — an expansion is a tree, not an intermediate sequence and not built string content.
The zero value is a fresh allowance. One value shared by many parses is what makes maxTotalEntityBytes bound the whole operation, and it is deliberately opaque: a caller may pass the same budget to several parses and may not read or reset the count, which is what stops the bound from being negotiable.
It is NOT safe for concurrent use. One belongs to one evaluation, the same way one includeProc belongs to one XInclude pass.
func NewEntityBudget ¶ added in v1.3.0
func NewEntityBudget() *EntityBudget
NewEntityBudget returns a fresh expansion allowance of maxTotalEntityBytes, to be shared across every parse that belongs to one operation.
type EntityResolver ¶ added in v1.0.0
type EntityResolver interface {
ResolveEntity(systemID, publicID, base string) (io.ReadCloser, string, error)
}
EntityResolver fetches the resource an external entity or an external DTD subset names.
It is the caller's, deliberately: xdm has no filesystem and no network, so every decision about what may be read — which schemes, which directories, how symlinks resolve — is made in code the caller owns and can audit. A resolver MUST refuse anything it is not certain of; returning an error makes the reference fail, which is the safe outcome.
systemID is the system identifier exactly as the document wrote it, which is usually relative. base is the absolute URI of the entity that contains the reference, against which systemID is to be resolved — note that for an entity declared inside an external DTD subset this is the SUBSET's URI, not the document's, as XML requires.
It returns the resource's content and the absolute URI it resolved to. That URI becomes the base for anything the fetched text itself references, so a resolver must return the URI it actually read, not the one it was asked for.
type Error ¶
type Error struct {
// Code is the spec error code, such as "XPTY0004". Codes live in the
// http://www.w3.org/2005/xqt-errors namespace; the local name alone is
// carried here because it is unique across the specs and is how the
// documents themselves refer to them.
Code string
// Message is the human-readable detail, without the code prefix.
Message string
// Err is an underlying cause, if any.
Err error
// CodeName is the full QName of the code, when the error came from
// fn:error with a QName naming a namespace other than the standard error
// one. Code alone cannot carry it, and XSLT 3.0's $err:code is a QName.
CodeName *QName
// Raised marks an error that fn:error produced rather than one the
// engine detected.
//
// It matters to try/catch. §3.16 catches dynamic errors only, and this
// engine tells a static error from a dynamic one by its code, because it
// resolves names at evaluation time and a static fault therefore arrives
// looking dynamic. fn:error breaks that inference: it raises a *dynamic*
// error whatever QName it is handed, so fn:error(xs:QName("err:XPST0008"))
// is catchable even though XPST names the static family. Without this
// flag the code heuristic refused to catch it.
Raised bool
// Value is the error object fn:error was given as its third argument.
//
// It exists for XSLT 3.0's xsl:catch, which exposes it as $err:value. Only
// fn:error can supply one, so it is nil on every error the engine raises
// itself, which is exactly the empty sequence the spec requires there.
Value Sequence
// Line and Module say where in a stylesheet the error was raised: the
// line number within the module, and the module's URI. XSLT 3.0 section
// 8.3 publishes them to an xsl:catch clause as $err:line-number and
// $err:module, and both are optional there — a processor that does not
// record the position reports the empty sequence. Line is 0 and Module
// is "" when nothing stamped them, which is how "not recorded" is
// spelled.
//
// They are stamped by the XSLT engine as the error passes out of the
// instruction that raised it, not at the point of construction: an error
// value is built in dozens of places across xdm, xpath and xsd, none of
// which knows anything about a stylesheet. See xslt/execSequence.
Line int
Module string
}
Error is an XPath, XQuery or XSLT error carrying its specification error code.
The specs define a code for every error condition — XPTY0004 for a type error, FORG0001 for a failed cast, FODC0002 for an unretrievable document — and those codes are the stable, translatable part of an error. A message is prose that may be reworded; a code is what a caller can branch on and what a conformance suite compares.
The codes were already present as string prefixes on every error this engine produces, which reads correctly but cannot be inspected: a caller wanting to distinguish "the document was malformed" from "the stylesheet is wrong" had to match on substrings. This type makes the code a field while keeping the rendered message byte-identical, so nothing that reads error text changes.
type FunctionItem ¶ added in v1.1.0
type FunctionItem struct {
// Name is the function's name, or the zero QName for an anonymous
// (inline) function. fn:function-name returns the empty sequence for the
// latter, which is why this is a value rather than a pointer: the zero
// QName is the "no name" case.
Name QName
// Arity is the number of parameters the function declares. It is part of
// a function's identity: fn:concat#2 and fn:concat#3 are different
// function items.
Arity int
// Signature is the function's declared parameter and return types, in
// source spelling ("xs:string", "item()*", "node()?"), with Signature[0]
// the return type and the rest the parameters in order.
//
// It is nil for a function whose signature was never recorded, which a
// typed function test then cannot judge: such an item matches on arity
// alone, as it did before signatures existed. Strings rather than a
// parsed type because the parsed form lives in the xpath package, which
// this one cannot import.
Signature []string
// VariadicSignature is the declared type of a VARIADIC function, whose
// parameters are all one type however many it is called with. It is nil
// for every fixed-arity function, which keeps Signature the ordinary
// path and this a compatible extension.
//
// It exists so that such a signature need not be materialised. fn:concat
// is declared at every arity from 2 upwards, and writing Signature for a
// call at arity N means an N+1 element slice of one repeated string --
// 16 bytes per argument, so 16MB at arity 2^20, measured. The arity is
// supplied by the CALLER, so that slice is an allocation an untrusted
// expression sizes: this is why the ceiling in xpath.synthesizeVariadic
// exists, and why raising it without this field is a memory-exhaustion
// hole rather than the conformance fix it looks like.
//
// MinArity is carried because it is part of the declared type rather
// than a fact about construction: F&O 3.1 declares fn:concat as two
// arguments or more, so an item claiming concat#1 is not merely
// unbuildable, it fails to match a function test of that arity.
VariadicSignature *VariadicSignature
// Invoke calls the function with the given arguments.
//
// The context is passed as an any because the type that carries it lives
// in the xpath package. The closure the xpath package installs here knows
// the concrete type and asserts it; no other package calls this directly.
Invoke func(ctx any, args []Sequence) (Sequence, error)
}
FunctionItem is the third kind of XDM item, introduced in XPath 3.0.
XPath 2.0 has exactly two kinds of item, a node and an atomic value, and the Item interface is closed around them. 3.0 adds this one: a function is a value, so it can be bound to a variable, passed to another function, and returned from one.
The body is deliberately not modelled here. Calling a function needs an evaluation context, a function library and the expression tree — all of which live in the xpath package, which imports this one. So this type carries only what the data model itself defines about a function item — its name, its arity, and an opaque payload — and the xpath package supplies the Invoke closure that knows how to run it.
The consequence worth stating: a caller doing an exhaustive type switch over Item now has a third case to consider. That is unavoidable — a new kind of item is exactly what 3.0 adds — but it is contained, because an XPath 2.0 expression can never produce one. Nothing in the 2.0 language constructs a function item, so a 2.0 caller's switch is never reached by one.
func (*FunctionItem) IsAnonymous ¶ added in v1.1.0
func (f *FunctionItem) IsAnonymous() bool
IsAnonymous reports whether the function item came from an inline function expression rather than a named function.
func (*FunctionItem) String ¶ added in v1.1.0
func (f *FunctionItem) String() string
String renders the function item for an error message.
A function item has no string value — fn:string of one is FOTY0014 — so this is not that, and is never the result of atomising one.
func (*FunctionItem) TypeName ¶ added in v1.1.0
func (f *FunctionItem) TypeName() string
TypeName implements Item.
The data model calls this type "function(*)". A more precise signature is not available: this value records the arity but not the declared parameter and return types, which the static type system would need.
type IdentityKey ¶ added in v1.3.0
type IdentityKey struct {
// contains filtered or unexported fields
}
IdentityKey is a comparable value equal for two node references exactly when Is reports them the same node, for use as a map key where a bare *Node pointer would split one synthesized namespace node into two entries.
Set membership — fn:intersect, fn:except, fn:innermost, fn:outermost — is keyed on this rather than on the pointer so that the set operators agree with "is". Every kind but namespace keys on the pointer itself, so this costs nothing and changes nothing for them.
type IncludeResolver ¶ added in v1.2.1
type IncludeResolver interface {
ResolveInclude(href, base, encoding string) (data []byte, uri string, err error)
}
IncludeResolver reads the resource an xi:include names.
It is separate from EntityResolver even though both read a URI, because the two answer different questions and a caller must be able to permit one without the other. An external entity is named by a document's own DOCTYPE and is refused by default as the XXE surface; an inclusion is named by an element the caller can see in the document it handed over. Folding them into one interface would mean enabling inclusions silently enabled entity reads.
The returned uri is the URI of the resource actually read, which is what anything *inside* the included resource resolves against — a resolver that follows a redirect, or that canonicalises a path, must report where it landed rather than where it was asked to look. XInclude 1.0 section 4.5.1 makes that base the one the included subtree carries.
The encoding argument carries the xi:include encoding attribute, which section 3.1 says "specifies the encoding of the resource" and applies only to parse="text". It is empty for an XML inclusion, where the encoding is the resource's own business and is discovered by the XML parser from a BOM or a declaration — section 4.4 is explicit that encoding "is ignored" there.
type Item ¶
type Item interface {
// TypeName returns the QName of the item's type, for error messages and
// instance-of tests.
TypeName() string
// contains filtered or unexported methods
}
Item is a single member of a sequence: either a Node or an atomic value.
The interface is closed to outside implementations (unexported marker method). XDM defines exactly these two kinds of item in XPath 2.0; function items arrive in 3.0 and would be added here.
type MapBuilder ¶ added in v1.1.0
type MapBuilder struct {
// contains filtered or unexported fields
}
MapBuilder accumulates entries into a map in one pass.
map:merge is handed half a million singleton maps by the suite (map-keys-014), and building the result with Put would allocate a trie path per entry. The builder owns the nodes it creates until Build hands the map over, so it writes through them in place; nothing else has a reference to observe the intermediate states, and a node that belonged to an earlier map is copied before it is written.
func NewMapBuilder ¶ added in v1.1.0
func NewMapBuilder() *MapBuilder
NewMapBuilder returns a builder over an empty map.
func NewMapBuilderFrom ¶ added in v1.1.0
func NewMapBuilderFrom(m *MapItem) *MapBuilder
NewMapBuilderFrom returns a builder seeded with m's entries, for the operations that start from an existing map. m itself is never written.
func (*MapBuilder) Build ¶ added in v1.1.0
func (b *MapBuilder) Build() *MapItem
Build returns the finished map. The builder must not be used afterwards, so that the map it hands out really is immutable: the edit token its nodes carry is dropped here, and no later builder will match it.
type MapItem ¶ added in v1.1.0
type MapItem struct {
// contains filtered or unexported fields
}
MapItem is the fourth kind of XDM item, added in XPath 3.1.
A map associates atomic keys with arbitrary sequences. It is a *function* item as well: calling a map with one argument looks a key up, which is what makes "$m('k')" and the lookup operator "$m?k" the same operation. That dual nature is in the specification rather than a convenience here — map:merge and fn:for-each can be handed a map wherever a function of arity one is expected.
Keys are compared by value, not by type identity, but xs:untypedAtomic is not admitted: a key arrives already atomized, and leaving an untyped one comparable to both a string and a number would make lookup depend on which happened to be asked for.
func (*MapItem) Entries ¶ added in v1.1.0
Entries calls f for each entry in insertion order, stopping on the first error.
func (*MapItem) Get ¶ added in v1.1.0
Get returns the value a key maps to, and whether the key is present.
An absent key is the empty sequence rather than an error, which is what makes "$m?missing" usable in a predicate.
func (*MapItem) Put ¶ added in v1.1.0
Put adds or replaces an entry, returning a new map and leaving the receiver untouched.
Maps are immutable in the data model: map:put returns a map rather than changing one, and a caller holding the original must still see it. The trie shares every node but the ones on the changed path, so this is O(log n) rather than the whole-map copy it used to be -- which, at the 421,875 entries of op:same-key-023, was 3.66ms per call and hours per case.
func (*MapItem) RemoveAll ¶ added in v1.1.0
RemoveAll returns a map without any of the given keys.
map:remove takes a *sequence* of keys, and each is removed along its own trie path. Absent keys are ignored rather than being an error, which is what makes map:remove($m, ("a", "nosuch")) legal, and a call that removes nothing is answered with the receiver itself: a map is immutable, so sharing it is safe.
type Node ¶
type Node struct {
Kind NodeKind
Name QName
// Value is the text content for text, comment, PI and attribute nodes,
// and the namespace URI for namespace nodes. Element and document nodes
// derive their string value from descendants; see StringValue.
Value string
Parent *Node
Children []*Node
// Attrs and Namespaces hold attribute and namespace nodes for elements.
// They are kept out of Children because the child axis must not return
// them — a fact that a single mixed slice makes easy to get wrong.
Attrs []*Node
Namespaces []*Node
// BaseURI is the resolved base URI, used by fn:document and fn:doc.
BaseURI string
// DocumentURI is the data model's dm:document-uri property, which
// fn:document-uri returns. It is meaningful only on a document node.
//
// It is deliberately NOT the same field as BaseURI, and not derived from
// it. dm:base-uri and dm:document-uri are separate accessors in the XDM,
// and the difference is observable: dm:document-uri is the absolute URI a
// document was RETRIEVED BY, so it is empty for any document that was not
// retrieved by URI at all — a temporary tree built by xsl:variable, a
// document node constructed by xsl:document, a tree parsed from a string.
// Those trees still need a base URI, for fn:base-uri and for resolving a
// relative reference written inside them, so BaseURI on a temporary tree
// is set on purpose (xslt/runtime.go does this) and cannot double as the
// document URI. XPath F&O fn:document-uri: "returns the empty sequence if
// $arg is not a document node, or if the document node was not retrieved
// via a URI".
//
// The invariant a caller must maintain: set this ONLY when the document
// was fetched by that URI and registered in the document pool, so that
// fn:doc of this value returns this same node. Setting it on a tree that
// fn:doc cannot retrieve would make "doc(document-uri($d)) is $d" false
// while claiming it should be true. Parse sets it from
// ParseOptions.DocumentURI, which defaults to empty.
DocumentURI string
// TypeAnnotation records a schema type when the document has been
// validated. Untyped documents leave this empty, and atomisation then
// yields xs:untypedAtomic, which is the schemaless default.
//
// It holds an ANNOTATION NAME, which AnnotationName builds and
// SplitAnnotationName takes apart: a type in the XML Schema namespace
// keys under its bare local name ("string", "QName"), and any other type
// under Clark notation, {uri}local. Producers must go through
// AnnotationName; consumers comparing against a qualified name must use
// SplitAnnotationName rather than SplitQName, which would cut a Clark key
// at the colon inside its URI and yield nonsense without an error.
//
// The namespace is load-bearing rather than decorative. This string is
// the key into a process-global derivation table, so a bare local name
// let one schema's type displace a built-in of the same name for every
// later schema in the process — see the commentary on derivedPrimitives.
TypeAnnotation string
// UnionMember records which member type of a union simple type actually
// accepted this node's value, when TypeAnnotation names (or has simple
// content of) a union.
//
// It is separate state from TypeAnnotation because the two facts are
// different and both are needed. XSD 1.0 §3.14.4 makes member selection a
// property of the *value*, not of the type: "100" validated against
// union(my:partNumberType, xs:integer) is an xs:integer while "123-AB" is
// a my:partNumberType, and the same annotation covers both. Folding the
// winner into TypeAnnotation would answer "instance of xs:integer" at the
// cost of "instance of my:partIntegerUnion", which the union's own
// identity requires; keeping only the union answers the second and loses
// the first. A node must answer both, so both are recorded.
//
// It is an annotation name, like TypeAnnotation, and for the same reason:
// it is compared against, and walked through, the same registries.
//
// Empty for every node whose type is not a union, which is almost all of
// them, so the common path pays only the field.
UnionMember string
// DerivedPrimitive and ListItem record what TypeAnnotation MEANS, as the
// schema that validated this node defined it: the built-in the annotated
// type erases to, and — when the annotated type is a list — the type of
// its items. Both are annotation names, like TypeAnnotation itself.
//
// They exist for the reason UnionMember and IsID do, and the three are
// one pattern: the property is fixed by the assessment that produced the
// node, so it is recorded when that assessment happens rather than
// recomputed afterwards from a name. Recomputing means asking the
// process-global derivation registries (derivedPrimitives, listItems),
// which are keyed by QName alone and hold whatever schema loaded LAST.
// Two schemas may legitimately define {urn:x}T differently — one deriving
// from xs:decimal, one from xs:string — and a node validated against the
// first then atomised as the second's type: same lexical form, a
// confidently wrong answer rather than a fallback, so "'10' lt '9'" came
// back true under string ordering where the numeric comparison is false.
// A cached stylesheet was corrupted permanently by any later load of a
// colliding name, because compile-time registration is not replayed per
// transform. Unions were immune precisely because UnionMember already
// recorded the answer on the node; this extends that to the other two.
//
// Empty means "not recorded", not "no derivation": a node annotated by
// something other than schema assessment — DTD attribute types, the XSLT
// validation instructions, a plain struct literal — leaves them empty and
// atomisation falls back to the registries, exactly as it did before.
// So these fields are a per-node OVERRIDE of a global answer, never a
// precondition for having one.
DerivedPrimitive string
ListItem string
// IsID and IsIDREFS are the data model's is-id and is-idrefs properties
// (XDM §5.2, §6.2). They are deliberately *separate* state from
// TypeAnnotation rather than being derived from it, because XSLT 2.0
// §3.5 requires them to survive input-type-annotations="strip": that
// setting turns every annotation into xs:untyped/xs:untypedAtomic while
// leaving is-id and is-idrefs exactly as they were. Deriving them from
// the annotation would lose them at precisely the point the
// specification says they must be kept, and fn:id/fn:idref — which are
// defined over these properties, not over the annotation — would then
// find nothing in a stripped document whose ID attributes happen not to
// be spelled "id".
//
// Two bools rather than one enum: an attribute of a union type can in
// principle be neither, and nothing in the model makes them exclusive.
// They are set wherever an annotation is assigned (schema assessment,
// DTD attribute types) by whoever knows the declared type; a node whose
// type was never determined leaves both false, which is the correct
// answer for an unvalidated document.
IsID bool
IsIDREFS bool
// IsNilled is the data model's dm:nilled property (XDM 5.10): true for an
// element that a schema assessment found nil, false for every other node.
//
// It is separate state from TypeAnnotation, and separate from the xsi:nil
// attribute, for the same reason IsID is. The property is fixed by the
// assessment that produced the node and does not follow from what the
// node looks like afterwards: xsi:nil on an element whose declaration is
// not nillable is an ERROR rather than a nilled element, and only the
// validator can tell those apart.
//
// Inferring it from "carries an annotation AND has xsi:nil='true'" gets
// xsl:copy validation="preserve" wrong, which is what validation-1204 is
// written to catch. That instruction CONSTRUCTS an element and preserves
// the annotation onto it without assessing anything, so the new element
// had both halves of that test while nothing had ever assessed it. The
// property is a fact about an event that either happened or did not, so
// it is recorded when it happens and copied when the node is.
//
// A node whose type was never determined leaves this false, which is the
// correct answer for an unvalidated document.
IsNilled bool
// NoTypedValue is the data model's "typed value is absent" property (XDM
// 3.1 §6.2.4): an element validated against a complex type with
// element-only or empty content HAS no typed value, and fn:data applied
// to one is FOTY0012.
//
// It is separate state from TypeAnnotation for the reason IsNilled is:
// the annotation cannot answer it. An anonymous complex type annotates
// with the nearest named base, ordinarily "anyType" -- which is also what
// a MIXED-content type and a genuine xs:anyType element carry, and both
// of those DO have a typed value (their string value). Only the validator
// knows which of the three it assessed, so only the validator can record
// it.
//
// A node nothing assessed leaves this false, which is correct: an
// unvalidated element is xs:untypedAtomic of its string value and
// atomizes without complaint.
NoTypedValue bool
// contains filtered or unexported fields
}
Node is a node in an XDM tree.
This is a concrete struct rather than an interface. Every node kind shares most of its fields, the evaluator switches on Kind rather than dispatching, and the axes need to walk parent/sibling links tens of thousands of times per document — an interface would add a pointer chase and an indirect call to each step for no expressiveness gained.
Trees are built by the parser in this package and are immutable afterwards. That immutability is what makes it safe to share one compiled stylesheet tree across concurrent transforms.
func ElementByID ¶ added in v1.2.1
ElementByID finds the element whose ID is id.
It is exported because a bare-name fragment identifier means the same thing wherever it appears: XPointer Framework section 3.2 defines the shorthand pointer as selecting the element with a matching ID, and xsl:source-document resolves its href fragment by that rule (see xslt/sourcedoc.go).
xml:id is honoured unconditionally, and a DTD-declared ID attribute is honoured through Node.IsID, which the DTD machinery sets. A plain attribute merely *named* "id" is deliberately not treated as one: without a DTD or a schema saying so it is an ordinary attribute, and guessing would make an inclusion resolve differently depending on data the document never declared.
func (*Node) AddNamespace ¶
AddNamespace links a namespace node to n.
func (*Node) AppendChild ¶
AppendChild links c as the last child of n, setting the parent link. It does not assign document order; call Finalize once the tree is complete.
func (*Node) ApplyTyping ¶ added in v1.3.0
ApplyTyping writes a Typing's properties onto n, replacing whatever was there.
The fields are ASSIGNED, not or-ed, and nothing is derived from the annotation name -- for the reason CopyTypingFrom assigns rather than going through SetTypeAnnotation. The caller already holds every answer, so re-deriving would be redundant where it agreed and wrong where it did not: SetTypeAnnotation only ever turns is-id ON, which would let a node inherit a marking its assessment did not give it.
func (*Node) Atomize ¶
Atomize returns the typed value of a node. Without schema validation every node atomises to xs:untypedAtomic, which is what makes untyped comparison rules apply throughout a schemaless transform.
func (*Node) AtomizeList ¶ added in v1.0.0
AtomizeList returns the typed value of a node whose annotation is a list type, as one atomic value per whitespace-separated token.
The second result reports whether the annotation is in fact a list type; a caller that gets false must fall back to Atomize, which yields the single value that every non-list node has.
Only the three built-in list types are recognised here. A user-defined list type is registered by the schema layer with its item type, and that derivation chain is what DerivedBase walks; a list type derived by restriction from one of these three therefore resolves to it and is expanded with its item type.
The empty string atomizes to the empty sequence rather than to one zero-length token, which is what "a list of no items" means and what strings.Fields already produces.
func (*Node) AttrValue ¶
AttrValue returns the value of a no-namespace attribute, or "". Most attributes the stylesheet compiler reads (match, select, name, test) are unprefixed, so this is the common case worth a helper.
func (*Node) ChildElements ¶
ChildElements returns the element children, which is what almost every stylesheet-compilation walk wants.
func (*Node) Compare ¶
Compare orders two nodes in document order, returning -1, 0 or 1. Nodes in different trees are ordered by tree id, which is stable within a transform.
func (*Node) CopyTypingFrom ¶ added in v1.2.2
CopyTypingFrom copies every PSVI property of src onto n, so that the copy answers each of them exactly as the original does.
It exists because there is no such thing as "the important half" of a node's typing. Eight properties record what an assessment concluded -- TypeAnnotation, UnionMember, DerivedPrimitive, ListItem, IsID, IsIDREFS, IsNilled, NoTypedValue -- and each one of them has, at some point in this repository, been dropped by a copy site that hand-picked the fields it thought mattered. Each omission was silent and each produced a confidently wrong answer rather than a missing one: a union-typed value atomising to xs:untypedAtomic, fn:id finding nothing, nilled() going false on a preserved copy, a list type erased to the wrong primitive by the process-global registries. The failure mode is always the same shape, so the fix is one operation rather than seven more careful field lists.
The fields are ASSIGNED rather than or-ed. The destination is a copy of the source and holds no assessment of its own; anything already on it is either identical or wrong.
It deliberately does NOT go through SetTypeAnnotation. That setter is for a PRODUCER of annotations -- schema assessment, DTD attribute types, the XSLT validation instructions -- which knows a name and must derive the rest from it. Here every property is already known, so deriving would be at best redundant and at worst wrong: SetTypeAnnotation only ever turns is-id and is-idrefs ON, which would make a copy of a non-ID node inherit a marking the original does not have. The invariant SetTypeAnnotation protects is upheld here by construction: the resolved fields cannot outlive their annotation, because src is a coherent node and all eight fields travel together.
func (*Node) CopyTypingStrippedFrom ¶ added in v1.2.2
CopyTypingStrippedFrom copies onto n the PSVI properties of src that survive input-type-annotations="strip" and validation="strip", and clears the rest.
The split between the two groups is not a judgement call; XSLT 2.0 sections 3.5 and 19.2 draw it explicitly, and it lands differently on each field:
TypeAnnotation, UnionMember, DerivedPrimitive and ListItem are CLEARED. They are one fact in four parts -- the name of the type, which member of a union accepted the value, the built-in the type erases to, and the item type of a list -- and stripping removes that fact. Clearing the name alone is precisely the bug SetTypeAnnotation guards against: atomisation gates on TypeAnnotation != "", so a surviving ListItem would go unread until something else re-annotated the node, and then describe a type the node no longer claims. The four go together in both directions.
IsID and IsIDREFS are KEPT. Section 3.5 says so in as many words: the setting "does not change the is-id and is-idrefs properties". They are separate state for exactly this reason (see Node.IsID), and fn:id and fn:idref are defined over them rather than over the annotation, so a stripped document whose ID attributes are not spelled "id" would otherwise become invisible to both.
IsNilled is CLEARED. XDM 5.10 makes dm:nilled a property of an element that a schema assessment found nil, and a stripped tree is one nothing assessed. Section 3.5 states the consequence directly: after stripping, the nilled property of every element is false. It parts company with is-id here because is-id survives by explicit exemption and this does not; the xsi:nil attribute, being an ordinary attribute once the type is gone, is a separate question that belongs to whoever is doing the copy.
Fields outside the PSVI set -- name, value, base URI, children -- are the caller's business, exactly as in CopyTypingFrom.
func (*Node) Identity ¶ added in v1.3.0
func (n *Node) Identity() IdentityKey
Identity returns the key that stands for this node's identity.
func (*Node) InScopeNamespaces ¶
InScopeNamespaces returns every prefix-to-URI binding visible at n, with inner declarations shadowing outer ones. Used when copying elements and when resolving QNames in stylesheet attribute values.
func (*Node) Is ¶ added in v1.3.0
Is reports whether n and o are the same node, which is what the "is" operator asks and what XDM means by node identity.
For every node the parser builds this is pointer equality, because each such node exists exactly once. Namespace nodes are the exception: they are not stored, they are synthesized on demand from the element's in-scope bindings, so a second walk over the same axis hands back a fresh pointer for a binding that is, by every other measure the engine applies, the same node. Comparing those pointers made "/*/namespace::xlink is /*/namespace::*[. = '...']" answer false where the spec requires true.
Order() is the identity the rest of the engine already uses — fn:generate-id is defined as "N" plus this number, and Compare reads the same field — and SetSynthesizedOrder derives it from the owning element and the binding's position in the sorted prefix list. So two synthesized nodes for one element and prefix already share it, and two for different elements, or different prefixes on one element, already do not. Deferring to it here makes "is" agree with generate-id, with "<<" and ">>", and with the document-order deduplication a path expression performs, rather than standing alone as the only operator that could see one node as two.
func (*Node) LookupPrefix ¶
LookupPrefix resolves a namespace prefix against the in-scope namespaces of n, walking up the tree. Returns the URI and whether the prefix was bound.
func (*Node) Order ¶
Order returns a number that identifies the node uniquely within the process.
It is the document-order index within the node's own tree, combined with the tree's identity so that nodes from two documents cannot collide. Callers wanting relative position must use Compare: this value orders nodes within one tree but says nothing across trees.
The combination is what fn:generate-id() needs. Returning the bare per-tree index gave the same answer to the first node of every document, so a stylesheet comparing generated identities across documents — the case key-042 in the XSLT suite exists to check — saw distinct nodes as identical.
A tree built by a sequence constructor is never finalized and has no tree of its own; those nodes take the identity assigned on demand to their root, the same one cross-tree comparison uses, so two parentless elements are also distinguished.
func (*Node) Position ¶
Position returns the 1-based line and column where the node starts, and false if the position is unknown — the node was built by a transform rather than parsed, or the source text was not retained.
func (*Node) Root ¶
Root returns the root of the containing tree, walking parent links. For a well-formed parsed document this is the document node.
func (*Node) SetSynthesizedOrder ¶ added in v1.0.0
SetSynthesizedOrder places a node the parser did not build into the document order of an existing tree, immediately after owner.
The namespace axis is the case this exists for: its nodes are synthesized on demand from the in-scope bindings, so they have no order of their own. Left at zero they sort before every real node, and — because generate-id() is derived from the order — every one of them answers "N0", colliding with each other and with the document node.
The offset separates the bindings of one element from each other while keeping them all adjacent to their owner. It is deliberately not an attempt at a spec-defined position: XPath leaves the relative order of namespace nodes implementation-dependent, and what a caller needs is that the order is stable and the identities distinct.
func (*Node) SetTypeAnnotation ¶ added in v1.0.0
SetTypeAnnotation records a type annotation and the is-id / is-idrefs properties that go with it.
It exists so that every producer of annotations — schema assessment, DTD attribute types, the XSLT validation instructions — sets the two properties the same way. Assigning TypeAnnotation directly is still legal but leaves is-id and is-idrefs at whatever they were, which is what a caller deliberately preserving them across a strip wants and what a caller annotating a fresh node does not.
The properties are only ever turned *on* here. A node that was already marked keeps its marking when re-annotated with a non-ID type, because the data model's properties describe how the node was validated originally and XSLT's stripping rules are the only thing entitled to change them — and those rules say the properties do not change at all.
func (*Node) SetTypeAnnotationResolved ¶ added in v1.2.2
SetTypeAnnotationResolved records an annotation together with what it means according to the schema doing the validating: the built-in the type erases to, and the item type when it is a list. Either may be empty when the caller has no answer, which leaves the corresponding field unset and lets atomisation fall back to the process-global registries.
It exists because SetTypeAnnotation cannot answer those questions. Resolving a name to its base means consulting the registries, and those are keyed by QName alone across the whole process, so the answer they give depends on which schema loaded most recently rather than on which schema validated this node. The validator holds the right schema at the right moment; this is how it hands the answer to the node instead of leaving it to be looked up again later against possibly different definitions. See the commentary on Node.DerivedPrimitive.
The fields are ASSIGNED rather than or-ed, unlike is-id and is-idrefs: they describe the assessment happening now, so re-annotating a node replaces them. Clearing them when the caller has no answer is deliberate — a stale value from a previous assessment would be a wrong answer rather than a missing one, and a missing one falls back correctly.
func (*Node) SetTypeEnv ¶ added in v1.3.0
func (n *Node) SetTypeEnv(e *TypeEnvironment)
SetTypeEnv records the type environment of the schema whose assessment produced this node's annotation.
The xsd package calls it as it annotates, so that later by-name questions about the node's type reach the definitions that schema actually made rather than whatever a later, unrelated schema registered under the same name.
func (*Node) StringValue ¶
StringValue returns the node's string value per XDM: the concatenation of all descendant text for document and element nodes, and the value itself for the leaf kinds.
func (*Node) StripTyping ¶ added in v1.2.2
func (n *Node) StripTyping()
StripTyping clears in place every PSVI property that stripping removes, keeping the two it preserves.
It is CopyTypingStrippedFrom applied to a node that is its own source, for the callers that strip a tree they already own rather than building a copy. Spelling it separately keeps those callers from writing n.CopyTypingStripped From(n), which reads as though it might do something else.
func (*Node) TypeEnv ¶ added in v1.3.0
func (n *Node) TypeEnv() *TypeEnvironment
TypeEnv returns the type environment of the schema that validated this node, or nil when no schema did.
Unlike TypeEnvOf this does NOT fall back to the global table: it reports what the node actually carries, which is what a test asserting that the stamping happened needs to see.
type Opaque ¶
type Opaque struct {
// Label names the kind of value, for error messages.
Label string
// Value is the wrapped payload.
Value any
}
Opaque wraps an arbitrary Go value as an Item.
It exists so that layers above this package can thread their own state through an evaluation context, which binds sequences rather than typed fields. The XSLT engine uses it for the transform runtime and grouping state, which the xpath package cannot name without an import cycle.
An Opaque is not a legal XDM value: it has no string value, does not atomise, and must never reach a stylesheet. Every producer binds it under a reserved namespace that no stylesheet can spell.
type ParseOptions ¶
type ParseOptions struct {
// BaseURI is recorded on the document node and used to resolve relative
// references in fn:document and xsl:include.
BaseURI string
// DocumentURI is recorded on the document node as its dm:document-uri
// property, which is what fn:document-uri returns. It is separate from
// BaseURI because the two accessors are separate in the data model: see
// Node.DocumentURI for why they cannot be the same field.
//
// It defaults to empty, which is the right answer for every caller that
// is parsing something it did not retrieve by URI — a stylesheet string,
// a re-parsed entity expansion, a test fixture. A caller that DID fetch
// the document from a URI, and that registers it in a document pool so
// that fn:doc of the same URI returns this same tree, sets it to that URI.
DocumentURI string
// StripSpace removes whitespace-only text nodes. XSLT applies this per
// element name via xsl:strip-space, so the transform layer passes a
// predicate; a plain bool here would not express "strip in these elements
// only".
StripSpace func(elem QName) bool
// AllowDOCTYPE permits a DOCTYPE declaration. It defaults to false: a
// DOCTYPE is the entry point for both XXE (parser-executed file:// and
// http:// reads) and entity-expansion blowup, and a validator that
// happily expands entities from untrusted input is a liability. Callers
// that genuinely need DTD-declared entities opt in explicitly.
AllowDOCTYPE bool
// ExternalEntities permits external entities — those declared SYSTEM or
// PUBLIC, and an external DTD subset — to be read, by supplying the
// resolver that reads them.
//
// It is nil by default, and nil means every external entity is refused
// exactly as before. It is deliberately SEPARATE from AllowDOCTYPE and
// is not implied by it: AllowDOCTYPE admits a DOCTYPE and its internal
// declarations, which cost nothing outside the document, while this
// admits reads of other resources — the XXE surface proper. A caller
// that wants entity declarations does not thereby want file reads.
//
// xdm has no filesystem and no network, so it can only read what a
// resolver hands it. Confinement — permitted schemes, permitted
// directories, symlink resolution — is entirely the resolver's, and
// xslt.FileResolver implements it. Expansion remains bounded by this
// package: fetched bytes are charged to the document's shared budget
// before they are expanded, and the number and nesting of fetches are
// capped. See xdm/dtd_external.go.
ExternalEntities EntityResolver
// TrackPositions records where each element starts, so that a validator
// can report the line a failure occurred on. It retains the source text
// for the life of the tree, which measures at about 10% more memory on a
// typical invoice and no extra parse time. It is opt-in because that cost
// buys nothing for a caller that never asks for a position.
TrackPositions bool
// MaxDepth bounds nesting. Deeply nested input is the cheapest way to
// drive a recursive descent into stack exhaustion, so the limit is
// enforced during construction rather than left to the runtime.
MaxDepth int
// MaxBytes bounds the source document. Zero means DefaultMaxBytes;
// a negative value means no limit, for a caller reading input it
// produced itself.
MaxBytes int64
// MaxNodes bounds the tree. Zero means DefaultMaxNodes; a negative
// value means no limit.
//
// Both limits exist because neither alone is a memory bound. A node
// costs a fixed ~200 bytes whatever it contains, so the heap a document
// needs depends on how many nodes it has rather than how long it is:
// a megabyte of "<a/>" is fifty times the memory of a megabyte of text.
// MaxBytes bounds the read; MaxNodes bounds what the read can allocate.
MaxNodes int
// contains filtered or unexported fields
}
ParseOptions controls document construction.
func (ParseOptions) WithEntityBudget ¶ added in v1.3.0
func (o ParseOptions) WithEntityBudget(b *EntityBudget) ParseOptions
WithEntityBudget returns a copy of opts whose parse charges its entity expansion against the shared allowance b, rather than minting a fresh one.
A nil b leaves opts alone, so a caller with no budget to share still gets the per-document allowance it had.
type QName ¶
QName is an expanded name: namespace URI plus local part, with the prefix retained only for serialisation.
Equality in XPath is defined on (URI, Local) alone — the prefix is not part of the value — so Equal deliberately ignores Prefix. Keeping the prefix around anyway matters because a literal result element must be serialised with the prefix the stylesheet author wrote, not one we invent.
func (QName) Clark ¶
Clark returns the {uri}local form, which is unambiguous without a namespace context and is therefore what error messages and map keys use.
type Sequence ¶
type Sequence []Item
Sequence is an ordered list of items. The empty sequence is a nil or zero-length slice; both are treated identically by every operation, so callers never have to normalise before comparing.
A sequence is flat: XDM has no nested sequences, and every constructor in this package maintains that invariant.
func Atomize ¶
Atomize converts a sequence to atomic values, replacing each node with its typed value. This is the fn:data() operation, applied implicitly wherever XPath 2.0 requires atomic operands.
Every item in the result is an *Atomic. Callers rely on that: two dozen of them assert the type without checking, because within the data model there is nothing else atomisation can produce.
Opaque items are the exception, and they are dropped here rather than passed through. They carry engine-internal state — the transform runtime, grouping bookkeeping — through the closed Item interface, and a stylesheet that names the internal namespace could reach one:
xmlns:gi="urn:goxslt:internal" ... distinct-values($gi:runtime)
Passing it through made that expression panic with an interface-conversion error, which in a server embedding this engine is a denial of service triggered by stylesheet text. An Opaque has no typed value, so dropping it is also what the data model implies: it is not a node and not an atomic value, so fn:data has nothing to return for it.
func AtomizeChecked ¶ added in v1.1.0
AtomizeChecked is Atomize for a caller that must report FOTY0013 rather than silently discard a function item.
XPath 3.0 makes atomising a function item an error, not a no-op: "data(f#1)" and "string(f#1)" both fail, and a function item reaching an arithmetic or comparison operator fails there too. Atomize cannot report it, so a caller that is about to demand a typed value uses this instead.
func Empty ¶
func Empty() Sequence
Empty is the canonical empty sequence.
A function rather than a variable. An exported package-level var of slice type is writable by anyone who imports the package, and a single stray assignment would corrupt the value every other caller reads -- a process-wide fault with no owner and no way to detect it. The value is nil, so this compiles to nothing.
func Flatten ¶ added in v1.1.0
Flatten replaces every array in a sequence with its members, recursively, which is what array:flatten and the function conversion rules do.
func One ¶
One wraps a single item as a sequence. Named for how often it is needed: most XPath operations produce exactly one item and must still return a sequence.
func SortDocumentOrder ¶
SortDocumentOrder sorts a sequence of nodes into document order and removes duplicates.
Every path expression in XPath 2.0 returns nodes in document order with duplicates removed, and so do the union, intersect and except operators. Doing it in one place means the axis implementations can emit nodes in whatever order is natural for them (reverse axes emit backwards) without each having to re-sort.
Items that are not nodes are an error at the call sites that use this, so they are passed through unsorted rather than silently dropped; the caller type-checks first.
func (Sequence) First ¶
First returns the first item, or nil if the sequence is empty. Callers that require exactly one item should use Single instead so that a length > 1 is reported rather than silently truncated.
func (Sequence) Single ¶
Single returns the sole item of a one-item sequence. It reports an error for any other length, because the places that call it (operands of arithmetic, the argument of a function declared to take exactly one item) are precisely the places where XPath 2.0 raises XPTY0004 rather than coercing.
type Tree ¶
type Tree struct {
Root *Node
// DocType is the DOCTYPE declaration's text, when the document had one
// and AllowDOCTYPE permitted it. Empty otherwise.
//
// It is retained because the internal subset is the only place a
// document's own DTD lives, and validating against it needs the text —
// encoding/xml hands the declaration over as one opaque token and keeps
// nothing. The dtd package parses it; this package applies only the two
// declarations whose absence is visible in the data model.
DocType string
// XMLVersion is the version the document's XML declaration names: "1.0"
// or "1.1", and "1.0" when there is no declaration, since that is the
// version such a document is read as. It is empty for a tree that was
// not parsed from text, which has no declaration to name one.
//
// It is recorded because some consumers are defined for one version
// only: Canonical XML is not defined for XML 1.1, and package c14n
// refuses a tree this reports as 1.1.
XMLVersion string
// contains filtered or unexported fields
}
Tree owns a document and the counter used to assign document order.
func Parse ¶
func Parse(r io.Reader, opts ParseOptions) (*Tree, error)
Parse builds an XDM tree from an XML document.
It uses encoding/xml as a tokeniser only. The Go decoder's own namespace handling is not usable here: it resolves prefixes into Name.Space but discards the prefix and the declarations themselves, and XSLT needs both — namespace nodes are addressable on the namespace axis, and a literal result element must be serialised with the prefix the author wrote.
func ParseString ¶
func ParseString(s string, opts ParseOptions) (*Tree, error)
ParseString is Parse over a string, which is what most tests and the stylesheet compiler want.
func (*Tree) CopyDTDFrom ¶ added in v1.1.0
CopyDTDFrom gives t the DTD context of src: the DOCTYPE as written and the external subset text, if one was read.
A tree built by copying nodes out of another -- xsl:copy-of over a document node, fn:snapshot over anything -- carries no DTD of its own, and so answers fn:unparsed-entity-uri with the empty string for entities the original declared. XSLT 3.0 27.2 says a snapshot's root "has the same unparsed entities as the tree from which it was taken", and the two-argument forms of those functions exist precisely so that such a copy can be asked.
Only the declarations matter, not the identity of the tree: both fields are text, and both are read-only after parsing.
func (*Tree) Finalize ¶
func (t *Tree) Finalize()
Finalize assigns document-order indices across the whole tree in a single pre-order walk. It must be called after the tree is fully built and before any node comparison; every parser entry point in this package does so.
func (*Tree) HasPositions ¶ added in v1.1.0
HasPositions reports whether the tree was parsed with TrackPositions, and so can answer Node.Position for the elements it holds.
A caller that caches trees needs it: one parsed without positions cannot serve a request that needs them, and the only way to tell is to ask.
func (*Tree) HasUnparsedEntities ¶ added in v1.3.0
HasUnparsedEntities reports whether the document declared any unparsed entity at all.
It exists for the one consumer that has to tell "this name is not among the declared unparsed entities" from "this document declares none, so there is nothing to check the name against". XML 1.0 section 3.3.1 makes only an unparsed entity's name a legal ENTITY value, but a schema-aware processor validating a document whose DTD declares no unparsed entity has no table to judge against, and Saxon accepts such a value rather than refusing it. The distinction is the difference between an invalid value and an unchecked one.
func (*Tree) UnparsedEntity ¶ added in v1.0.0
UnparsedEntity returns the system identifier and notation of an unparsed entity declared in a document's internal subset.
An unparsed entity is the one kind a processor never reads: it is declared SYSTEM or PUBLIC with an NDATA notation, referenced from an attribute of type ENTITY, and its identifier is data for the application rather than something to fetch. fn:unparsed-entity-uri and fn:unparsed-entity-public-id return exactly these.
The declarations are re-read from the retained DOCTYPE text rather than carried on every tree, since a document with unparsed entities is rare and the lookup happens at most once per call.
type TypeCode ¶
type TypeCode int
TypeCode identifies an atomic type from the XML Schema built-in hierarchy.
Only the types XPath 2.0 gives special treatment are enumerated. The rest of the schema hierarchy (xs:token, xs:NMTOKEN, and the other string subtypes) behaves identically to its base type for every operation this engine performs, so carrying them as distinct codes would add branches with no behavioural difference.
const ( // TypeUntypedAtomic is the type of atomised nodes in a schemaless // document. It is the reason XPath 2.0 needs so few explicit casts: an // untypedAtomic operand is converted to the required type at the point of // use, but *only* in the specific contexts the spec lists. TypeUntypedAtomic TypeCode = iota TypeString TypeBoolean TypeDecimal TypeInteger TypeDouble TypeFloat TypeQName TypeAnyURI TypeDate TypeTime TypeDateTime TypeDuration TypeYearMonthDuration TypeDayTimeDuration TypeHexBinary TypeBase64Binary // The five Gregorian types denote a recurring or partial calendar point: // a year, a year and month, a month, a month and day, or a day. TypeGYear TypeGYearMonth TypeGMonth TypeGMonthDay TypeGDay )
func NumericPromote ¶
NumericPromote returns the common type for a binary numeric operation, per the XPath 2.0 promotion lattice: integer -> decimal -> float -> double. Both operands are converted to that type before the operation runs.
type TypeEnvironment ¶ added in v1.3.0
type TypeEnvironment struct {
// contains filtered or unexported fields
}
TypeEnvironment owns the type derivation facts a single schema assembly established: which built-in each of its named types erases to, which of them are lists and what their items are, and which are unions and what their members are.
It exists because those three tables were process-global and keyed by type NAME alone. Two schemas may each legitimately define {urn:x}T -- one deriving from xs:decimal, one from xs:string -- and under one global table the second load silently rewrote the first's answer for every node already validated against it. Atomisation was fixed by recording the resolved metadata on the node (see Node.DerivedPrimitive), but every by-NAME consumer -- the subtype relation, "instance of", "castable as", xsl:validate -- has no node to read and so still asked the shared table.
The lifetime rule, which is not negotiable: an environment lives exactly as long as something can still reach it. A compiled schema holds one, so an environment stays reachable for as long as the schema is alive and becomes collectable only when it is. There is deliberately NO eviction -- no LRU, no TTL, no size cap. Evicting an entry a live consumer depends on would turn a correct answer into a wrong one at an arbitrary later moment, which is strictly worse than the memory it would save.
Environments are safe for concurrent use. Schemas load in parallel and atomisation reads on every typed value, so each table is guarded by its own RWMutex: a reader on the derivation table is never blocked by a writer on the union table.
func GlobalTypeEnvironment ¶ added in v1.3.0
func GlobalTypeEnvironment() *TypeEnvironment
GlobalTypeEnvironment returns the process-global environment the name-only registration functions conceptually correspond to.
It is exported so that a caller holding no schema can still state a derivation, and so that the fallback a by-name consumer uses is nameable in a test. Prefer a schema-owned environment: entries here are keyed by type name across every schema in the process and are subject to exactly the collision this type exists to prevent.
func NewTypeEnvironment ¶ added in v1.3.0
func NewTypeEnvironment() *TypeEnvironment
NewTypeEnvironment returns an empty environment.
The xsd package creates one per assembled schema and populates it as it walks the type definitions, which is the only place that knows a user-defined type's base, item type or member list.
func TypeEnvOf ¶ added in v1.3.0
func TypeEnvOf(n *Node) *TypeEnvironment
TypeEnvOf returns the environment to consult for questions about a node's type: the one the schema that validated it owns, or the process-global fallback when the node carries none.
Every by-NAME consumer of a schema's type facts that HOLDS a node should go through this rather than through the package-level DerivedBase, ListItemOf and UnionMembersOf. Those answer from the global table, which is keyed by type name across every schema in the process and so answers for whichever schema loaded last; this answers from the schema that actually produced the node's annotation, which is the only definition that can be correct for it.
func TypeEnvOfAtomic ¶ added in v1.3.0
func TypeEnvOfAtomic(a *Atomic) *TypeEnvironment
TypeEnvOfAtomic returns the environment to consult for questions about an atomic value's derived type: the one the schema that issued the annotation owns, or the process-global fallback when the value carries none.
func (*TypeEnvironment) DerivedBase ¶ added in v1.3.0
func (e *TypeEnvironment) DerivedBase(name string) string
DerivedBase returns the type the named schema type derives from, or "" when the name is not one this environment knows. The name is an annotation name and so is the result, so a chain is walked by feeding the result back in.
A nil environment answers "" for everything, which is the untyped answer and the one a caller with no schema should get.
func (*TypeEnvironment) Len ¶ added in v1.3.0
func (e *TypeEnvironment) Len() (derived, lists, unions int)
Len reports how many derivation, list and union entries the environment holds. It is here so a heap/lifetime test can state what it is measuring rather than reaching into unexported maps.
func (*TypeEnvironment) ListItemOf ¶ added in v1.3.0
func (e *TypeEnvironment) ListItemOf(name string) string
ListItemOf returns the item type recorded for a list type, or "" when the name does not denote one in this environment.
func (*TypeEnvironment) Merge ¶ added in v1.3.0
func (e *TypeEnvironment) Merge(src *TypeEnvironment)
Merge folds every fact src holds into e, without overwriting one e already has.
It exists for the schema assemblies that are themselves a merge of several loaded schemas: XSLT's xsl:import-schema and XQuery's "import schema" each fold every imported schema into one aggregate *xsd.Schema, and a merge that copied the type DEFINITIONS while leaving the derivation facts behind would produce a schema whose types no longer know what they derive from -- a NOTATION restriction that stopped being a NOTATION.
First writer wins, matching the component merge beside it: a name the destination already defines is the destination's, and an import does not redefine it.
func (*TypeEnvironment) RegisterDerived ¶ added in v1.3.0
func (e *TypeEnvironment) RegisterDerived(name, primitive string)
RegisterDerived records that a schema type erases to a built-in one.
Both arguments are annotation names, which AnnotationName builds; passing a bare local name for a type that has a namespace re-creates the conflation the qualified keying exists to prevent.
A nil environment is a no-op rather than a panic: a caller that has not been given one is saying it has no schema facts to record, which is the untyped case and not an error.
func (*TypeEnvironment) RegisterList ¶ added in v1.3.0
func (e *TypeEnvironment) RegisterList(name, itemType string)
RegisterList records that a schema type is a list, and what its items are.
Both arguments are annotation names. itemType is the item type's own name, which may itself be a type this environment knows; the derivation walk resolves it.
func (*TypeEnvironment) RegisterUnion ¶ added in v1.3.0
func (e *TypeEnvironment) RegisterUnion(name string, members []string)
RegisterUnion records that a schema type is a union, and what its declared member types are, in declaration order.
Which member a given value belongs to is a per-VALUE fact recorded on the node (Node.UnionMember), because XSD 1.0 §3.14.4 chooses the member by trying each one's lexical space against the value in turn. This records only the declaration.
func (*TypeEnvironment) UnionMembersOf ¶ added in v1.3.0
func (e *TypeEnvironment) UnionMembersOf(name string) []string
UnionMembersOf returns the declared member types of a union type, or nil when the name does not denote one in this environment.
The result must not be modified: it is the stored slice, shared with every other caller.
type Typing ¶ added in v1.3.0
type Typing struct {
TypeAnnotation string
UnionMember string
DerivedPrimitive string
ListItem string
IsID bool
IsIDREFS bool
IsNilled bool
NoTypedValue bool
}
Typing is the complete set of PSVI properties an assessment concludes about one node, detached from any node.
It exists so that a caller who ALREADY KNOWS these facts can hand them over as a unit instead of passing a type name and letting the receiver look the rest up. The lookup is the problem: resolving a name to its base, its item type or its ID kind means consulting derivedPrimitives, listItems and unionMembers, which are process-global and keyed by QName alone, so they answer for whichever schema loaded LAST rather than for the schema that validated this node. The validator holds the right schema at the right moment; Typing is the shape that lets it say so.
The field list is CopyTypingFrom's, and deliberately the same one: eight properties travel together or the copy is wrong, and every historical bug in this area was a hand-picked subset of them. A new PSVI property must be added here, to CopyTypingFrom and to CopyTypingStrippedFrom together.
The zero Typing means "nothing assessed this node", which is the correct state for an unvalidated node and is what the name-only convenience wrappers produce when given an empty annotation.
type VariadicSignature ¶ added in v1.3.0
type VariadicSignature struct {
// MinArity is the fewest arguments the function accepts; 2 for fn:concat.
MinArity int
// Result is the declared return type.
Result string
// Parameter is the declared type of EVERY parameter.
Parameter string
}
VariadicSignature is the declared type of a function that takes a minimum number of arguments and then any number more, all of one type.
Result and Parameter are source spellings, exactly as xdm.FunctionItem's Signature uses -- "xs:string", "item()*" -- so the two forms are read by the same subsumption code.
type XIncludeOptions ¶ added in v1.2.1
type XIncludeOptions struct {
// Resolver reads the resources. A nil Resolver makes every href fail,
// which is not the same as doing nothing: a failed inclusion still uses
// its xi:fallback, and is still a fatal error when it has none. That is
// the correct reading of section 4.3, and it means "no resolver" behaves
// as a resolver that refuses everything rather than as a silent no-op.
Resolver IncludeResolver
// Parse carries the options an *included* XML resource is parsed with.
// The including document's own limits are the natural choice and the
// caller supplies them: an inclusion is a part of the document as far as
// the data model is concerned, so it should not be able to sidestep a
// bound the including document was held to. BaseURI and DocumentURI are
// overwritten per resource and anything set here for them is ignored.
Parse ParseOptions
}
XIncludeOptions configures ProcessXInclude.