Documentation
¶
Overview ¶
Package xsd implements XML Schema 1.0 and 1.1 validation.
The package is organised around the spec's own vocabulary. XSD is defined in two layers: a *schema component model* (Part 1 §3), which is an abstract data structure, and a set of *validation rules* over that model. The XML syntax of a .xsd file is a third thing again — a concrete representation that maps onto components. Keeping the three apart is what makes the spec tractable, so the types here are named after components rather than after elements: a ComplexType is not an <xs:complexType>, it is what one maps to.
What is implemented ¶
XSD 1.0 and 1.1, both targeting the W3C xsdtests suite: the component model, schema assembly through include, import, redefine and override, content models, simple types and facets, xsi:type and xsi:nil, substitution groups, wildcards, identity constraints, and document-level ID/IDREF.
1.1 is opt-in through Options.Version, because it changes which documents are valid and a 1.0 schema must not acquire its behaviour by accident. It adds assertions, conditional type assignment with xs:alternative and inheritable attributes, open content, xs:override, the notNamespace and notQName wildcard forms, explicitTimezone, and conditional inclusion through the versioning attributes. The 1.1 constructs are always *parsed* — a schema that uses one is not made valid by pretending it is absent — but they are only honoured under Version11.
One constraint is deliberately not checked: Particle Valid (Restriction). See CheckConstraints.
Concurrency ¶
A Schema is immutable once loaded and safe to validate from any number of goroutines. Its one piece of lazily-built state, the content-model cache, is synchronised for that reason. A Schema still being assembled is not safe to share, and neither is a document tree being validated: parse one per goroutine.
Errata ¶
The 2nd Edition text folds in the 1st Edition errata, and this implementation follows the corrected text. Three corrections change behaviour enough to name here, because the uncorrected reading is the intuitive one:
- E1-26: an xs:all group may carry minOccurs="0". The original text required {min occurs}=1.
- E2-30: a pattern facet on a list type matches the whole space-separated literal, not each item separately.
- E1-51: the ur-type's attribute wildcard is processContents="lax", not strict.
Index ¶
- Constants
- Variables
- func CheckBuiltinValue(local, lexical string) (string, error)
- func FacetApplicable(t *SimpleType, f FacetKind) bool
- func ParseExpandedName(v string) (xdm.QName, bool)
- type Assertion
- type AttributeDecl
- type AttributeGroupDef
- type AttributeUse
- type CatalogEntry
- type CatalogResolver
- func (r *CatalogResolver) Add(namespace string, src []byte, aliases ...string)
- func (r *CatalogResolver) AddFromFS(fsys fs.FS, entries []CatalogEntry) error
- func (r *CatalogResolver) Resolve(namespace, location, base string) (io.ReadCloser, string, error)
- func (r *CatalogResolver) SetFallback(f Resolver)
- type CheckOptions
- type ComplexType
- type Component
- type Compositor
- type ContentKind
- type Derivation
- type DerivationSet
- type ElementDecl
- type FacetKind
- type FacetSet
- type FileResolver
- type HTTPResolver
- type ICPath
- type ICPathAlternative
- type ICStep
- type IdentityConstraint
- type IdentityConstraintKind
- type InstanceLocationPolicy
- type MapResolver
- type ModelGroup
- type ModelGroupDef
- type NamespaceConstraintKind
- type NotationDecl
- type OpenContent
- type OpenContentMode
- type Options
- type ParseError
- type Particle
- type Pattern
- type ProcessContents
- type Resolver
- type Schema
- func Load(root *xdm.Node, baseURI string, opts Options) (*Schema, error)
- func LoadFile(path string, opts Options) (*Schema, error)
- func LoadFiles(paths []string, opts Options) (*Schema, error)
- func NewSchema() *Schema
- func ParseSchema(root *xdm.Node) (*Schema, error)
- func SchemaForJSON() (*Schema, error)
- func (s *Schema) CanAssessStrictly(el *xdm.Node) bool
- func (s *Schema) CheckConstraints(opts CheckOptions) error
- func (s *Schema) HasAttributeDeclaration(name xdm.QName) bool
- func (s *Schema) HasElementDeclaration(name xdm.QName) bool
- func (s *Schema) HasSimpleType(typeName xdm.QName) bool
- func (s *Schema) IsListSimpleType(typeName xdm.QName) (itemType xdm.QName, ok bool)
- func (s *Schema) TypeEnv() *xdm.TypeEnvironment
- func (s *Schema) Validate(root *xdm.Node, opts ValidateOptions) error
- func (s *Schema) ValidateAgainstType(n *xdm.Node, typeName xdm.QName, opts ValidateOptions) error
- func (s *Schema) ValidateAttribute(at *xdm.Node, lax bool, opts ValidateOptions) error
- func (s *Schema) ValidateContext(ctx context.Context, root *xdm.Node, opts ValidateOptions) error
- func (s *Schema) ValidateElement(el *xdm.Node, opts ValidateOptions) error
- func (s *Schema) ValidateElementLax(el *xdm.Node, opts ValidateOptions) error
- func (s *Schema) ValidateExpandedQNameValue(typeName, value xdm.QName) (known bool, err error)
- func (s *Schema) ValidateValue(value string, typeName xdm.QName) error
- func (s *Schema) WithInstanceLocations(root *xdm.Node, policy InstanceLocationPolicy, opts Options) (*Schema, error)
- type SchemaErrors
- type Scope
- type SequenceMatcher
- type SimpleType
- type Term
- type Timezone
- type Type
- type TypeAlternative
- type ValidateOptions
- type ValidationError
- type ValidationErrors
- type ValueConstraint
- type Variety
- type Version
- type WhiteSpace
- type Wildcard
Constants ¶
const ( // NSSchema is the XML Schema namespace: the one <xs:schema> lives in. NSSchema = "http://www.w3.org/2001/XMLSchema" // NSInstance is the schema-instance namespace, holding xsi:type, // xsi:nil, xsi:schemaLocation and xsi:noNamespaceSchemaLocation. NSInstance = "http://www.w3.org/2001/XMLSchema-instance" // NSXML is the namespace of xml:lang, xml:space, xml:base and xml:id, // which a schema may reference without importing. NSXML = "http://www.w3.org/XML/1998/namespace" )
Namespaces used throughout the spec.
const ( // DefaultFetchTimeout bounds one network fetch. DefaultFetchTimeout = 30 * time.Second // DefaultMaxSchemaBytes bounds one fetched schema document. Real // schemas are far smaller; the W3C's own largest is under 200 kB. DefaultMaxSchemaBytes = 16 << 20 )
Defaults for HTTPResolver.
const AllDerivations = DerivationSet(0xff)
AllDerivations is the set of every derivation method, the meaning of final="#all" or block="#all".
Named for what it is a set *of*. A bare "All" at package scope reads as something universal when it means only this, and a package-level name in a surface this large cannot be corrected once it is frozen.
const DefaultMaxContentModelPositions = 8192
DefaultMaxContentModelPositions bounds a schema that does not set Options.MaxContentModelPositions.
This is a MEMORY bound, and it is the only thing standing between a schema of a few kilobytes and gigabytes of allocation. A group DAG in which each of n groups references the next twice is valid, acyclic and tiny, yet expands to 2^(n-1) positions. Measured through Load, cost is flat at ~400 bytes per position:
n=16 2^15 positions 10.7ms 12 MB n=20 2^19 positions 155ms 201 MB n=22 2^21 positions 577ms 786 MB n=24 2^23 positions 2.47s 3452 MB <- from a 2.7 KB schema
8192 positions is therefore about 3.3 MB for one model. The gate is INCREMENTAL — it fires having already allocated in proportion to the limit, so the limit is what is actually reserved, not merely what is refused:
limit 8192 refused in 2.4ms 2 MB limit 1048576 refused in 263ms 360 MB limit 4194304 refused in 1.08s 1521 MB
Raising it is a deliberate grant of memory to whoever wrote the schema. A host that does so is accepting that cost; it is exposed as an option because a trusted generated schema may legitimately need it, not because the default is conservative.
const DefaultMaxDepth = 1000
DefaultMaxDepth bounds validation recursion when MaxDepth is zero. It matches xdm.DefaultMaxDepth, so a document the parser accepts is one the validator will not refuse for depth alone.
const DefaultMaxDocuments = 512
DefaultMaxDocuments bounds an assembly that does not set MaxDocuments.
const DefaultMaxErrors = 100
DefaultMaxErrors bounds a run that does not set MaxErrors.
const DefaultMaxMatchStates = 4096
DefaultMaxMatchStates bounds the number of simultaneous content-model states the validator will carry while matching one element's children.
The set is the subset construction over the counter vectors, and its size is what makes the matcher exact rather than heuristic. On real schemas it stays at one or two: UBL 2.1, DocBook and both W3C suites never exceed a handful. A schema deep in nested repetitions with coinciding boundary positions can in principle make it grow, and a validator that grows without a ceiling is a way to be handed an unbounded allocation by a document being validated. Exceeding it is reported as an error rather than guessed at: a matcher that silently degraded to an approximation would do so precisely on the inputs where the answer was hardest to get.
const NSVersioning = "http://www.w3.org/2007/XMLSchema-versioning"
NSVersioning is the XSD 1.1 versioning namespace.
const Unbounded = -1
Unbounded is the {max occurs} of a particle written maxOccurs="unbounded".
Variables ¶
var ErrPrivateAddress = errors.New(
"address is in a private range; set AllowPrivateAddresses to permit it")
ErrPrivateAddress is returned when a fetch is refused because the host resolved to an address in a range HTTPResolver does not dial by default. It is wrapped by the dial error, so errors.Is finds it through the *url.Error and *net.OpError that net/http puts around it.
Functions ¶
func CheckBuiltinValue ¶
CheckBuiltinValue reports whether lexical is a legal value of the built-in simple type named local, and returns its canonical form.
It exists for callers outside schema validation that hold a value and a type name — the RELAX NG engine is the case it was added for, since that language names XSD types through its datatype library and would otherwise have to reimplement lexical checks that already exist here.
An unknown type name is an error rather than an acceptance: a validator that silently passes what it cannot check is the failure this exists to avoid.
func FacetApplicable ¶
func FacetApplicable(t *SimpleType, f FacetKind) bool
FacetApplicable reports whether a facet may be applied to a simple type.
Dispatch is on {variety} first, and the union case is the one that catches implementations out: a union admits only pattern and enumeration whatever its members admit. In particular it has no whiteSpace facet, because normalisation belongs to whichever member type validates the value.
func ParseExpandedName ¶ added in v1.0.0
ParseExpandedName reads the "{uri}local" spelling of an already-expanded name.
It is how a caller with no instance node states a QName value whose prefix it has already resolved. A lexical QName never has this shape — "{" is not an NCName character — so accepting it here cannot capture a real instance value, and no schema document can contain one.
Types ¶
type Assertion ¶
type Assertion struct {
// Test is the compiled XPath 2.0 expression.
Test *xpath.Compiled
// Source is the expression as written, for diagnostics.
Source string
// XPathDefaultNamespace is the namespace unprefixed element names in
// the test resolve to.
XPathDefaultNamespace string
}
Assertion is an <xs:assert> (XSD 1.1 §3.13).
type AttributeDecl ¶
type AttributeDecl struct {
Name xdm.QName
Type *SimpleType
Scope Scope
Constraint *ValueConstraint
// contains filtered or unexported fields
}
AttributeDecl is an Attribute Declaration (§3.2.1).
An attribute's type is always a simple type: XSD has no way to give an attribute element content.
func (*AttributeDecl) ComponentKind ¶
func (*AttributeDecl) ComponentKind() string
ComponentKind implements Component.
type AttributeGroupDef ¶
type AttributeGroupDef struct {
Name xdm.QName
AttributeUses []*AttributeUse
AttributeWildcard *Wildcard
// contains filtered or unexported fields
}
AttributeGroupDef is an Attribute Group Definition (§3.6.1).
func (*AttributeGroupDef) ComponentKind ¶
func (*AttributeGroupDef) ComponentKind() string
ComponentKind implements Component.
type AttributeUse ¶
type AttributeUse struct {
Required bool
Decl *AttributeDecl
Constraint *ValueConstraint
// Inheritable is the XSD 1.1 {inheritable}: the attribute is visible to
// conditional type assignment on descendant elements, not only on the
// element carrying it. It is how a schema lets an ancestor's xml:lang
// choose a descendant's type.
Inheritable bool
// Prohibited marks use="prohibited": the attribute is removed rather
// than declared. Such a use is not part of the type's {attribute uses}
// — it exists only so that inheritance can tell "this name was ruled
// out" from "this name was never mentioned", which is what stops the
// base's use being inherited straight back.
Prohibited bool
}
AttributeUse is an Attribute Use (§3.5.1).
This is the component that a <xs:attribute> inside a complex type maps to. It is distinct from the declaration because the same declaration can be used with different requiredness or a different value constraint in different types.
func (*AttributeUse) ComponentKind ¶
func (*AttributeUse) ComponentKind() string
ComponentKind implements Component.
type CatalogEntry ¶ added in v1.1.0
type CatalogEntry struct {
// Path locates the document within the fs.FS.
Path string
// Namespace is the schema's target namespace, or empty for a document
// only ever named by location.
Namespace string
// Aliases are the schemaLocation spellings this document answers.
Aliases []string
}
A CatalogEntry names one schema document to load into a catalog: where to read it from, what namespace it defines, and the locations it answers to.
func W3CEntries ¶ added in v1.1.0
func W3CEntries() []CatalogEntry
W3CEntries describes the schemas that the well-known W3C vocabularies are referred to by, with every spelling this package has seen in the wild.
It carries no schema content. A caller supplies the documents -- from a checkout, a go:embed, or the companion module -- and this says what to register them as, so that the aliasing is stated once here rather than rediscovered by everyone who needs it.
The paths are the conventional file names; a caller whose files are named differently can copy this and edit it, since it is data rather than behaviour.
type CatalogResolver ¶ added in v1.1.0
type CatalogResolver struct {
// contains filtered or unexported fields
}
A CatalogResolver answers schemaLocation from an in-memory table keyed by what a reference *means* rather than by how it is spelled.
MapResolver already resolves from memory, but it matches a location literally, and the well-known schemas are referred to by many spellings. The XSD 1.1 schema for schemas alone is written as http://www.w3.org/TR/xmlschema11-1/XMLSchema.xsd, as http://www.w3.org/2001/XMLSchema.xsd, as a bare relative XMLSchema.xsd, and as an xs:import carrying the namespace and no location at all. A table keyed on the literal string answers one of those and misses the rest.
So an entry is registered once, against a namespace and any number of aliases, and a lookup tries the location, then the location resolved against the base URI, then the namespace. Nothing is fetched and nothing is read from disk at resolve time: a reference to something not in the table is an error rather than a request, which is the property that makes this the resolver to reach for in a server.
The zero value is not usable; call NewCatalogResolver.
func NewCatalogResolver ¶ added in v1.1.0
func NewCatalogResolver() *CatalogResolver
NewCatalogResolver returns an empty catalog.
func (*CatalogResolver) Add ¶ added in v1.1.0
func (r *CatalogResolver) Add(namespace string, src []byte, aliases ...string)
Add registers one schema document under a target namespace and any number of location aliases.
The namespace may be empty, for a document that is only ever named by location. An alias may be an absolute URI or a bare relative reference; both are matched, and a relative one also matches wherever it lands after being resolved against the referring document's base URI.
Registering the same alias twice replaces it, so a caller may override a bundled entry with its own copy.
func (*CatalogResolver) AddFromFS ¶ added in v1.1.0
func (r *CatalogResolver) AddFromFS(fsys fs.FS, entries []CatalogEntry) error
AddFromFS registers every file named by aliases from an fs.FS, mapping each to the namespace given for it.
It is the bridge from a directory of schemas a caller already has -- checked in, embedded with go:embed, or unpacked at startup -- to a catalog, without this package needing to carry the documents itself.
entries maps a file path within fsys to the target namespace of the schema in it and the aliases it should answer to. A file that is missing is an error naming it, because a catalog that silently has fewer entries than the caller asked for fails later and somewhere less obvious.
func (*CatalogResolver) Resolve ¶ added in v1.1.0
func (r *CatalogResolver) Resolve(namespace, location, base string) (io.ReadCloser, string, error)
Resolve implements Resolver.
func (*CatalogResolver) SetFallback ¶ added in v1.1.0
func (r *CatalogResolver) SetFallback(f Resolver)
SetFallback names a resolver to consult when the catalog has no entry.
Nil, the default, makes a miss an error, which is what a server wants. A command-line tool that should still read a schema beside the one it was given sets a FileResolver here; one that may reach the network sets an HTTPResolver, and thereby says so deliberately.
type CheckOptions ¶
type CheckOptions struct {
// LaxUPA accepts a content model in which two competing particles are
// references to the same element declaration.
//
// The strict reading of §3.8.6 rejects those; Saxon and XSV accept
// them. Schemas written against either of those processors may rely on
// it, so this exists — but it is off by default, because the strict
// reading is the conforming one.
LaxUPA bool
// Version selects the UPA rule. XSD 1.1 relaxed the constraint so that
// an element particle competing with a wildcard is no longer an error;
// only element-against-element and wildcard-against-wildcard remain.
// The suite states it outright — the feature category is
// xsd1_1-Wildcards-RelaxationOfUPA, "wildcard/element competition no
// longer violates UPA" — and s3_10_1v04s through s3_10_1ii09s are
// version="1.1" schemaTests expected valid for exactly that shape.
//
// The zero value is 1.0, which keeps the stricter rule.
Version Version
}
CheckOptions configure the schema component constraint checks.
type ComplexType ¶
type ComplexType struct {
Name xdm.QName
Base Type
Abstract bool
FinalSet DerivationSet
Prohibits DerivationSet
// DerivationMethod is {derivation method}: extension or restriction.
DerivationMethod Derivation
// Content is the {content type} discriminator; SimpleContent and
// Particle carry the payload for the kinds that have one.
Content ContentKind
SimpleContent *SimpleType
Particle *Particle
AttributeUses []*AttributeUse
AttributeWildcard *Wildcard
// Assertions are the XSD 1.1 <xs:assert> co-constraints on this type
// (§3.4.1 {assertions}). They evaluate XPath 2.0 against the element
// being validated, which is the feature that makes XSD 1.1 unavailable
// to implementations without an XPath engine.
Assertions []*Assertion
// OpenContent is the XSD 1.1 {open content}, which permits elements the
// content model does not name. Nil means the type is closed.
OpenContent *OpenContent
// contains filtered or unexported fields
}
ComplexType is a Complex Type Definition (§3.4.1).
func (*ComplexType) ComponentKind ¶
func (*ComplexType) ComponentKind() string
ComponentKind implements Component.
type Component ¶
type Component interface {
// ComponentKind names the kind for error messages.
ComponentKind() string
}
A Component is any of the schema components of Part 1 §3.
The interface exists to give the component kinds a common type for diagnostics and for the schema-level tables; it deliberately says nothing about validation, because the components are data and the validation rules are separate.
type Compositor ¶
type Compositor uint8
Compositor is the kind of a model group (§3.8.1).
const ( // CompositorSequence requires its particles in order. CompositorSequence Compositor = iota // CompositorChoice requires exactly one of its particles. CompositorChoice // CompositorAll requires each particle at most once, in any order. // XSD 1.0 constrains where an all group may appear and what it may // contain; see All Group Limited (§3.8.6). CompositorAll )
The compositors.
func (Compositor) String ¶
func (c Compositor) String() string
String names the compositor for diagnostics.
type ContentKind ¶
type ContentKind uint8
ContentKind discriminates the {content type} of a complex type (§3.4.1).
The spec models {content type} as a three-way sum: empty, a simple type, or a particle paired with mixed/element-only. Go has no sum type, so the kind is explicit and the payload fields are only meaningful for the matching kind.
const ( // ContentEmpty permits no children at all. ContentEmpty ContentKind = iota // ContentSimple permits character data validated against a simple type, // and no element children. ContentSimple // ContentElementOnly permits element children matching a particle, and // no character data other than whitespace. ContentElementOnly // ContentMixed permits element children matching a particle, with // character data interleaved freely. ContentMixed )
The content kinds.
func (ContentKind) String ¶
func (k ContentKind) String() string
String names the kind for diagnostics.
type Derivation ¶
type Derivation uint8
Derivation methods, used by the {final}, {block}, {prohibited substitutions} and {disallowed substitutions} property sets.
These are sets rather than single values, and — a detail worth stating because conflating them is a classic bug — an element declaration carries *two* of them with different meanings. {disallowed substitutions} controls what may substitute for the element in an instance; {substitution group exclusions} controls what may join its substitution group at schema construction time. They are consulted at different moments.
const ( DerivationExtension Derivation = 1 << iota DerivationRestriction DerivationList DerivationUnion DerivationSubstitution )
The derivation methods. DerivationSubstitution is only meaningful in {disallowed substitutions}; the others apply to types as well.
func (Derivation) String ¶
func (d Derivation) String() string
String names a derivation method, so that a diagnostic reads "extension" rather than the bit value.
type DerivationSet ¶
type DerivationSet uint8
DerivationSet is a set of derivation methods. The zero value is the empty set, which is the correct default for every property that uses it.
func (DerivationSet) Has ¶
func (s DerivationSet) Has(d Derivation) bool
Has reports whether d is in the set.
func (DerivationSet) String ¶
func (s DerivationSet) String() string
String renders the set the way a schema author wrote it.
func (DerivationSet) With ¶
func (s DerivationSet) With(d Derivation) DerivationSet
With returns the set with d added.
type ElementDecl ¶
type ElementDecl struct {
Name xdm.QName
Type Type
Scope Scope
Nillable bool
Constraint *ValueConstraint
Abstract bool
// SubstitutionGroup is the {substitution group affiliation}: the head
// this declaration may substitute for. Only a global declaration may
// have one (erratum E1-36 requires the head be global too).
//
// It is the first of SubstitutionGroups, kept as its own field because
// almost every use has exactly one head.
SubstitutionGroup *ElementDecl
// SubstitutionGroups holds every head, which XSD 1.1 permits to be a
// list where 1.0 allowed only one.
SubstitutionGroups []*ElementDecl
// DisallowedSubstitutions is {disallowed substitutions}, from block=.
// It controls substitution in an instance.
DisallowedSubstitutions DerivationSet
// SubstitutionGroupExclusions is {substitution group exclusions}, from
// final=. It controls what may derive from this declaration's type and
// still substitute. Kept separate from DisallowedSubstitutions because
// the two are consulted at different times and conflating them is a
// classic source of wrong answers.
SubstitutionGroupExclusions DerivationSet
// IdentityConstraints holds the key, keyref and unique children.
IdentityConstraints []*IdentityConstraint
// Alternatives are the XSD 1.1 <xs:alternative> children, in order.
// The first whose test holds selects the type; conditional type
// assignment is the other half of what 1.1 needs XPath for.
Alternatives []*TypeAlternative
// contains filtered or unexported fields
}
ElementDecl is an Element Declaration (§3.3.1).
func (*ElementDecl) ComponentKind ¶
func (*ElementDecl) ComponentKind() string
ComponentKind implements Component.
func (*ElementDecl) SchemaElementMembers ¶ added in v1.3.0
func (d *ElementDecl) SchemaElementMembers() []*ElementDecl
SchemaElementMembers returns the members of this declaration's substitution group that a schema-element() test admits, transitively and not including the declaration itself.
It is NOT the same set as Substitutable(). That set answers a content-model question -- which declarations may appear where a particle names this one -- and a *validation episode* checks each candidate against the instance in front of it. A type test has no instance to check: schema-element(E) is asked whether some already-validated node could have been validated against E or something substitutable for it, so the members that could never yield such a node have to be dropped up front. Two do.
An ABSTRACT member is dropped because no element is ever validated against an abstract declaration (§3.3.6: an abstract declaration cannot itself validate an element, only its non-abstract substitutes can). It still has to stay in Substitutable(), which is why the filter lives here and not there: a content model naming the head admits the abstract member's own substitutes, so pruning the walk at the abstract declaration would lose them.
A member that is NOT nillable is dropped when the head IS. XPath 3.1 §2.5.5.3 makes the nilled property part of the element test: a test derived from a nillable declaration admits a nilled node, and a declaration that forbids nilling can never produce one, so it cannot stand in for the head across the whole of what the head's test accepts. The comparison is one-way -- a nillable member under a non-nillable head is fine, since it only ever yields nodes the head's test already admits.
substitution-020 through 025 are the four-and-two of exactly this: A is abstract and matches neither head, C is non-nillable under nillable heads and matches neither, while the plain member B matches both.
func (*ElementDecl) Substitutable ¶
func (d *ElementDecl) Substitutable() []*ElementDecl
Substitutable returns the element declarations that may substitute for this one, transitively, not including the declaration itself.
The list is empty until the schema is assembled, since a member may be declared in a document read after the head.
type FacetKind ¶
type FacetKind uint8
FacetKind identifies a constraining facet (Part 2 §4.3).
const ( FacetLength FacetKind = iota FacetMinLength FacetMaxLength FacetPattern FacetEnumeration FacetWhiteSpace FacetMaxInclusive FacetMaxExclusive FacetMinInclusive FacetMinExclusive FacetTotalDigits FacetFractionDigits // FacetExplicitTimezone is the XSD 1.1 facet constraining whether a // date or time value carries a timezone. FacetExplicitTimezone )
The twelve constraining facets. The fundamental facets (§4.2) are properties of a type rather than constraints an author writes, and are not modelled here.
type FacetSet ¶
type FacetSet struct {
Length *uint64
MinLength *uint64
MaxLength *uint64
TotalDigits *uint64
FractionDigits *uint64
// WhiteSpace is the whiteSpace facet, and Fixed records whether the
// schema wrote fixed="true" on it. The built-in types fix it, which is
// why a user type cannot loosen xs:token back to preserve.
WhiteSpace *WhiteSpace
WhiteSpaceFixed bool
// MinInclusive and friends hold bounds. They are stored as lexical
// strings because the type they must be parsed against is the type
// being defined, which is not complete while its own facets are read.
MinInclusive *string
MaxInclusive *string
MinExclusive *string
MaxExclusive *string
// Patterns from a single derivation step are alternatives: a value
// satisfying any one of them satisfies the step. Patterns from
// *different* steps are conjunctive. That asymmetry is why patterns are
// held per step rather than merged into one list.
Patterns []*Pattern
// Enumerations are the permitted values, lexical at this stage. An
// empty non-nil slice is meaningful — it permits nothing — so presence
// is tested with HasEnumeration rather than len.
Enumerations []string
HasEnumerations bool
// EnumerationQNames holds the expanded form of each enumeration value,
// parallel to Enumerations, for a type whose value space is QNames —
// xs:QName, xs:NOTATION and anything derived from them.
//
// Those two types compare by expanded name, not by spelling: Part 2
// §3.2.18 gives xs:NOTATION the value space of the QNames of the
// notations declared in the schema, so an instance writing "one:mp3"
// satisfies an enumeration written "smokey:mp3" whenever both prefixes
// bind the same URI. The prefix in the facet resolves against the
// *schema* document's namespaces and the prefix in the instance against
// the *instance* document's, so neither lexical form can be compared
// against the other; the expansion has to be recorded here, where the
// facet's own element is still in hand.
//
// An entry is the zero QName when the facet's prefix was not bound, and
// is only populated for a namespace-sensitive type, so an empty slice
// means "compare lexically" as before.
EnumerationQNames []xdm.QName
// Assertions are the XSD 1.1 <xs:assertion> facets. On a simple type an
// assertion is a facet rather than a component, though it compiles the
// same way.
Assertions []*Assertion
// ExplicitTimezone is the XSD 1.1 facet constraining whether a date or
// time value carries a timezone.
ExplicitTimezone *Timezone
// Fixed records which facets this step wrote fixed="true" on. Part 2
// §4.3 gives every constraining facet but pattern, enumeration and
// assertion a {fixed} property, and "Facet Valid (Restriction)" then
// forbids a derived type from stating a *different* value for one the
// base fixed.
//
// whiteSpace keeps its own flag rather than joining this map: the
// built-in types fix it, and they are constructed directly rather than
// parsed, so the boolean predates any schema document.
Fixed map[FacetKind]bool
}
FacetSet holds the constraining facets applied at one derivation step.
A facet is present only if the schema document set it at this step; the inherited value comes from walking the base chain. Keeping the steps separate rather than flattening them is what makes the schema-component constraints on restriction checkable, since those compare a derived facet against the value it narrows.
func (*FacetSet) IsEmpty ¶ added in v1.1.0
IsEmpty reports whether the set constrains nothing.
XPath 3.1 2.5 needs it: a *pure* union type is one whose {facets} property is empty, and only a pure union may be used as an item type. A union carrying facets is excluded because substituting a member for the union is unsafe there — the member is not constrained by the facets the union adds — which is the XSD 1.0 error XSD 1.1 corrected.
type FileResolver ¶
type FileResolver struct {
// Root, when set, confines resolution to a directory. A location that
// escapes it — through "..", a symlink, or an absolute path — is
// refused. Leaving it empty permits any readable path, which is the
// right default for a command-line tool and the wrong one for a server.
Root string
}
FileResolver resolves a schemaLocation against the filesystem.
It is the default because it is the case that cannot surprise anyone: a schema that includes a file beside it keeps working, and nothing leaves the machine.
func (*FileResolver) Resolve ¶
func (r *FileResolver) Resolve(namespace, location, base string) (io.ReadCloser, string, error)
Resolve implements Resolver.
type HTTPResolver ¶
type HTTPResolver struct {
// Client fetches remote documents. When nil, a client with Timeout is
// used. Supplying one is the hook for a caller that needs a proxy, a
// pinned CA set, or a transport that refuses private address ranges.
Client *http.Client
// Timeout bounds a single fetch. Zero means DefaultFetchTimeout.
Timeout time.Duration
// MaxBytes bounds a fetched document. Zero means DefaultMaxSchemaBytes.
// A schema is not a stream, so an unbounded read is a way to be handed
// an unbounded allocation.
MaxBytes int64
// AllowHost, when non-nil, reports whether a host may be fetched from.
// It runs before the request, and it is an allowlist of *names*.
//
// It is not an address check and must not be relied on as one. A name it
// admits may resolve to loopback, link-local or a private range, and a
// name checked here may resolve to something else by the time the
// connection is made — DNS rebinding defeats a name check by
// construction. Returning true for "schemas.example.com" says the name is
// permitted, not that the connection goes anywhere trustworthy.
//
// The addresses themselves are refused by the dialler, at the point they
// are known — see AllowPrivateAddresses. Use AllowHost to narrow the
// namespace and the dialler to enforce the boundary.
AllowHost func(host string) bool
// AllowPrivateAddresses re-permits the address ranges that are refused by
// default: loopback, link-local (including 169.254.169.254, the cloud
// instance metadata address), unique-local, and the RFC1918 private
// ranges. It is off by default, so the zero value refuses them.
//
// The check runs in the dialler, against the IP the connection is
// actually being made to rather than the name written in the document.
// That placement is what makes it a guarantee: a name resolves to an
// address only at dial time, so checking the name earlier leaves a
// rebinding window in which the name is re-resolved to a refused address
// after it was approved. Checking the resolved address closes it, and it
// covers redirects and every retry for free, because each connection is
// dialled through the same place.
//
// Turn it on for a caller that genuinely fetches schemas from a private
// network — an internal mirror, or a test server on loopback. It widens
// what the process can be made to reach by whoever writes the schema, so
// it is opt-in rather than a default.
AllowPrivateAddresses bool
// Files handles locations that are not remote. When nil, a FileResolver
// with no root is used.
Files Resolver
}
HTTPResolver resolves a schemaLocation over the network, falling back to the filesystem for locations that are not remote.
Network resolution is off by default and this type is how a caller turns it on, because it hands control of what the process fetches to whoever wrote the schema. That is a considered trade rather than a scary-sounding one: a schema naming http://internal/admin makes the validator fetch it, and a schemaLocation taken from an instance document lets whoever supplied the document choose the schema it is judged against.
The zero value is usable and applies the default limits.
func (*HTTPResolver) Resolve ¶
func (r *HTTPResolver) Resolve(namespace, location, base string) (io.ReadCloser, string, error)
Resolve implements Resolver.
type ICPath ¶
type ICPath struct {
// Source is the expression as written, kept for diagnostics.
Source string
// Alternatives are the "|"-separated branches.
Alternatives []ICPathAlternative
}
ICPath is a compiled selector or field XPath.
The subset the spec permits is far smaller than XPath 1.0, let alone the XPath 2.0 this repository implements. It is compiled by a dedicated parser rather than the general one, because accepting more than the subset would accept schemas that conforming processors reject.
type ICPathAlternative ¶
type ICPathAlternative struct {
// DescendantOrSelf records a leading ".//".
DescendantOrSelf bool
// Steps are the child-axis steps.
Steps []ICStep
// Attribute is the trailing "@name" step, which only a field may have
// and only in final position.
Attribute *xdm.QName
// AttributeWildcard records that the attribute step was written "@*"
// or "@prefix:*". Such a field is grammatical; whether it selects
// exactly one node is decided per instance document, not here.
AttributeWildcard bool
}
ICPathAlternative is one "|"-separated branch of a selector or field path.
type ICStep ¶
type ICStep struct {
// Wildcard records "*", which matches any element name.
Wildcard bool
// Name is the element name when Wildcard is false.
Name xdm.QName
}
ICStep is one step of an identity-constraint path.
type IdentityConstraint ¶
type IdentityConstraint struct {
Name xdm.QName
Kind IdentityConstraintKind
Selector *ICPath
Fields []*ICPath
// Refer is the key or unique that a keyref points at. It is nil for
// ICKey and ICUnique.
Refer *IdentityConstraint
// contains filtered or unexported fields
}
IdentityConstraint is an Identity-constraint Definition (§3.11.1).
func (*IdentityConstraint) ComponentKind ¶
func (*IdentityConstraint) ComponentKind() string
ComponentKind implements Component.
type IdentityConstraintKind ¶
type IdentityConstraintKind uint8
IdentityConstraintKind discriminates key, keyref and unique (§3.11.1).
const ( // ICKey requires the field values to be present and unique. ICKey IdentityConstraintKind = iota // ICUnique requires uniqueness but permits absence. ICUnique // ICKeyref requires each value to match one in a referenced key. ICKeyref )
The identity constraint categories.
func (IdentityConstraintKind) String ¶
func (k IdentityConstraintKind) String() string
String names the category for diagnostics.
type InstanceLocationPolicy ¶
type InstanceLocationPolicy struct {
// AllowNamespace reports whether the instance may supply a location for
// a namespace. It is consulted for every entry, and an entry it refuses
// is ignored rather than being an error: the location is a hint, and a
// hint this processor declines to take is not a fault in the document.
//
// A nil AllowNamespace allows nothing, so the zero value of this type
// is inert. That is deliberate — a policy that did nothing but exist
// should not be the one that opens the door.
AllowNamespace func(namespace string) bool
// AllowNoNamespace permits xsi:noNamespaceSchemaLocation, which names a
// document for the absent namespace. It is separate from
// AllowNamespace because "" is not a namespace a caller thinks about,
// and folding it in would grant it by accident.
AllowNoNamespace bool
// Resolver fetches the documents the instance names. When nil the
// schema's own resolver is used — but note that a FileResolver will
// read whatever path the instance gives it, relative to the base, so a
// caller following untrusted documents should supply a MapResolver or
// an HTTPResolver with AllowHost.
Resolver Resolver
// MaxDocuments bounds how many documents the instance may pull in.
// Zero means DefaultMaxDocuments.
MaxDocuments int
}
InstanceLocationPolicy decides whether an instance may name a schema document for a namespace, and is how a caller turns xsi:schemaLocation on.
It is a policy rather than a boolean because "follow what the document says" is not a safe thing to grant wholesale. The useful grant is narrow: a caller that already trusts a set of namespaces — a fixed vocabulary it ships, a catalogue it controls — and wants the instance to say which *version* of them applies.
type MapResolver ¶
type MapResolver struct {
// ByLocation maps a schemaLocation to schema source.
ByLocation map[string]string
// ByNamespace maps a target namespace to schema source, used when an
// import gives a namespace but no location.
ByNamespace map[string]string
// contains filtered or unexported fields
}
MapResolver resolves from an in-memory table, for callers that know every schema in advance.
It is the resolver to reach for in a server: nothing is fetched, nothing is read from disk, and a schema naming a location that is not in the table is an error rather than a request.
func (*MapResolver) Resolve ¶
func (r *MapResolver) Resolve(namespace, location, base string) (io.ReadCloser, string, error)
Resolve implements Resolver.
type ModelGroup ¶
type ModelGroup struct {
Compositor Compositor
Particles []*Particle
}
ModelGroup is a Model Group (§3.8.1).
func (*ModelGroup) ComponentKind ¶
func (*ModelGroup) ComponentKind() string
ComponentKind implements Component.
type ModelGroupDef ¶
type ModelGroupDef struct {
Name xdm.QName
Group *ModelGroup
}
ModelGroupDef is a Model Group Definition (§3.7.1) — a named <xs:group>.
It is a separate component from the group it names because a reference to it is a reference to the *definition*, and erratum E1-29 makes that distinction observable: particles reached through two references to the same definition are still distinct particles for the purpose of Unique Particle Attribution.
func (*ModelGroupDef) ComponentKind ¶
func (*ModelGroupDef) ComponentKind() string
ComponentKind implements Component.
type NamespaceConstraintKind ¶
type NamespaceConstraintKind uint8
NamespaceConstraintKind discriminates a wildcard's {namespace constraint}.
const ( // NSAny is ##any: every namespace, and unqualified names. NSAny NamespaceConstraintKind = iota // NSNot is ##other: every namespace except the named one — and *not* // unqualified names. The exclusion of the absent namespace is explicit // in Wildcard allows Namespace Name clause 2.3 and is easy to miss. NSNot // NSEnumerated is an explicit list, in which the empty string stands // for the absent namespace (##local). NSEnumerated )
The namespace constraint kinds.
type NotationDecl ¶
NotationDecl is a Notation Declaration (§3.12.1).
func (*NotationDecl) ComponentKind ¶
func (*NotationDecl) ComponentKind() string
ComponentKind implements Component.
type OpenContent ¶
type OpenContent struct {
Mode OpenContentMode
Wildcard *Wildcard
}
OpenContent is an <xs:openContent> or <xs:defaultOpenContent> (XSD 1.1).
It is how 1.1 lets a schema say "and anything else may appear here" without writing a wildcard into every content model, which is what makes a schema forward-compatible with documents produced against a later version of it.
type OpenContentMode ¶
type OpenContentMode uint8
OpenContentMode says where an open content wildcard may match (XSD 1.1 §3.4.1).
const ( // OpenNone is a closed content model: the default. OpenNone OpenContentMode = iota // OpenInterleave permits the wildcard to match anywhere among the // content model's own elements. OpenInterleave // OpenSuffix permits it only after everything the content model // requires. OpenSuffix )
The open content modes.
type Options ¶
type Options struct {
// Resolver locates the documents named by include, import and
// redefine. When nil, a FileResolver with no root is used, so a schema
// can include a file beside it but nothing is fetched over the network.
//
// To follow remote locations, set an HTTPResolver. That is off by
// default because it hands control of what this process fetches to
// whoever wrote the schema.
Resolver Resolver
// MaxDocuments bounds how many documents one assembly may read. A
// schema that includes a generator of schemas would otherwise be a way
// to spend the process. Zero means DefaultMaxDocuments.
MaxDocuments int
// MaxContentModelPositions bounds the number of positions in any one
// compiled content model. A model that exceeds it is not compiled, and
// the schema is refused with an error wrapping xdm.ErrResourceLimit —
// the constraints on that model are then undecided, and an undecided
// normative constraint is never reported as satisfied.
//
// It is the memory bound on one model: cost is flat at roughly 400
// bytes per position, so the default of 8192 caps one model at about
// 3.3 MB. Raising it AUTHORISES PROPORTIONAL MEMORY — a host that sets
// 2^23 is accepting that a 2.7 KB schema may allocate 3.4 GB. See
// DefaultMaxContentModelPositions for the measurements.
//
// Zero means DefaultMaxContentModelPositions.
MaxContentModelPositions int
// Version selects XSD 1.0 or 1.1. The zero value is 1.0, because a
// schema written for 1.0 must not acquire 1.1's relaxations by
// accident — 1.1 changes which schemas are legal, not only which
// documents are.
Version Version
// XPathVersion selects the version of XPath the 1.1 assertions and
// conditional type alternatives in this schema are written in.
//
// The zero value is XPath 2.0, which is what the specification requires:
// XSD 1.1 defines assertions against a subset of XPath 2.0, so a schema
// using a 3.0 construct is not portable and must not quietly work here.
//
// Raising it is for a host that controls its own schemas and wants the
// later function library in an assertion — fn:parse-json or a map, say —
// accepting that the schema is then this engine's rather than every
// engine's. It does not affect anything but assertions and alternatives:
// nothing else in a schema is XPath.
XPathVersion xpath.Version
// ParseOptions are passed to the XML parser for each schema document.
// The zero value refuses a DOCTYPE, which is the right default: a
// schema has no use for one and it is the entry point for entity
// expansion attacks.
ParseOptions xdm.ParseOptions
// LaxUPA relaxes Unique Particle Attribution, which loading enforces,
// to the reading Saxon and XSV take: two competing particles are
// tolerated when both are references to the same element declaration.
//
// Off by default because the strict reading is the conforming one —
// erratum E1-29 is explicit that particles at different points are
// distinct "even if they originated from the same named model group".
// It exists because schemas written against those processors do rely
// on the permissive rule, and such a schema would otherwise be
// unloadable rather than merely non-conforming.
LaxUPA bool
}
Options configure schema assembly.
type ParseError ¶
type ParseError struct {
// Code is the spec's error code, such as "src-element.1", or empty if
// the fault has no code in the spec.
Code string
// Message describes the fault.
Message string
// Line and Column locate it, when the node carried a position.
Line, Column int
// Document names the schema document the fault is in, when that is not
// the document the caller passed to Load. A fault in the caller's own
// document leaves this empty: they are already looking at it, and
// repeating its name on every line would be noise. It is set for a
// document reached through <xs:include>, <xs:import>, <xs:redefine> or
// <xs:override> — the case where the message otherwise describes a
// fault in a file the reader has no reason to suspect.
Document string
// contains filtered or unexported fields
}
ParseError reports a fault in a schema document.
The spec gives error codes for schema representation faults (the src-* and *-props-correct codes). Carrying the code rather than only a message lets a caller — and the conformance harness — distinguish "this schema is malformed" from "this schema is not the one you meant".
type Particle ¶
type Particle struct {
MinOccurs int
// MaxOccurs is Unbounded for maxOccurs="unbounded".
MaxOccurs int
Term Term
// contains filtered or unexported fields
}
Particle is a Particle (§3.9.1).
A particle is the occurrence-constrained use of a term. It has exactly three properties and, notably, no annotation — it is pure structure.
func (*Particle) ComponentKind ¶
ComponentKind implements Component.
type Pattern ¶
type Pattern struct {
// Source is the pattern as written, kept for diagnostics and for the
// schema-component constraints that compare patterns textually.
Source string
// contains filtered or unexported fields
}
A Pattern is a compiled pattern facet.
XSD patterns are anchored: the whole value must match. This differs from fn:matches in XPath, which is a containment test, so the regex translation this package shares with the xpath package must be wrapped before use. See compilePattern.
func CompilePattern ¶
CompilePattern compiles an XML Schema pattern facet.
The compiled form is anchored, matching the facet's semantics rather than fn:matches's containment test, and is safe to reuse across goroutines.
Exported for callers that need the XML Schema regular-expression flavour outside a schema — RELAX NG's <param name="pattern"> means exactly this, and reimplementing the translation would guarantee the two drift apart.
type ProcessContents ¶
type ProcessContents uint8
ProcessContents says how strictly a wildcard's matched content is validated (§3.10.1).
const ( // ProcessStrict requires a declaration to be found and the content to // be valid against it. ProcessStrict ProcessContents = iota // ProcessLax validates against a declaration if one is found, and // otherwise accepts. This is the ur-type's mode (erratum E1-51). ProcessLax // ProcessSkip accepts without looking for a declaration. ProcessSkip )
The processContents modes.
func (ProcessContents) String ¶
func (p ProcessContents) String() string
String names the mode for diagnostics.
type Resolver ¶
type Resolver interface {
// Resolve returns the contents of the schema document at location,
// which is resolved relative to base when it is not absolute. The
// namespace is the one the reference declared, or empty for an include.
//
// Returning a nil reader and a nil error means "no schema document" —
// which for an include is not an error (§4.2.1), and the caller
// distinguishes the cases.
Resolve(namespace, location, base string) (io.ReadCloser, string, error)
}
A Resolver turns a schemaLocation into schema source.
The location is a hint, not an identifier: §4.3.2 lets a processor use a catalogue, a preloaded schema, or nothing at all. Making resolution an interface rather than a built-in fetch is what lets a caller decide, since following a location means fetching whatever a document names.
type Schema ¶
type Schema struct {
// Elements and the maps beside it are the global components, keyed by
// expanded name. Local declarations are reachable only through the type
// that contains them and are not indexed here.
Elements map[xdm.QName]*ElementDecl
Attributes map[xdm.QName]*AttributeDecl
Types map[xdm.QName]Type
AttributeGroups map[xdm.QName]*AttributeGroupDef
ModelGroups map[xdm.QName]*ModelGroupDef
Notations map[xdm.QName]*NotationDecl
// Version selects XSD 1.0 or 1.1 behaviour. It governs whether the 1.1
// features a document may use — assertions, conditional type
// assignment, open content — are honoured; they are always *parsed*,
// because a schema that uses them is not made valid by pretending they
// are absent.
Version Version
// contains filtered or unexported fields
}
Schema is a set of schema components, assembled from one or more documents.
The spec is explicit that a schema is a set of components rather than a document or a collection of documents (§4.2). Several documents may contribute to one namespace, and once assembled there is no way to ask which document a component came from — nor any need to.
func Load ¶
Load assembles a schema from a document and everything it includes, imports or redefines.
The base URI locates the first document, so that relative locations in it resolve. It may be empty when the document names only absolute locations.
func LoadFiles ¶
LoadFiles assembles one schema from several documents.
A schema is a set of components, and nothing says they must come from a single file: a namespace is often split across documents that no one of them includes, with the caller naming them all. Loading each separately and merging afterwards would resolve each document's references against only what that document could see, so they are assembled together instead.
func NewSchema ¶
func NewSchema() *Schema
NewSchema returns an empty schema populated with the built-in types.
The built-ins are present in every schema "by definition" (§4.1.2) without being declared, so they are added here rather than by reading a document.
func ParseSchema ¶
ParseSchema reads one schema document into a new schema.
Include, import and redefine are not followed: this reads a single document. Assembling a schema from several documents is a separate concern, because it needs a resolver policy that a caller must supply — following a schemaLocation means fetching whatever the schema names, which is a decision about trust rather than about parsing.
func SchemaForJSON ¶ added in v1.1.0
SchemaForJSON returns the built-in schema for the XML representation of JSON, F&O 3.1 §C.2.
The schema is parsed once and shared: it is immutable after assembly, and parsing it on every xsl:import-schema would cost more than the whole of the json-to-xml test set. The caller must not modify the returned schema — mergeSchema in the xslt package copies components out of it rather than into it, which is the only use it has.
The error is a programming error rather than a runtime one: the source is a constant in this file, so a failure means this package can no longer parse a schema it ships with. It is returned rather than panicked because the callers are library entry points that already have an error to return.
func (*Schema) CanAssessStrictly ¶ added in v1.0.0
CanAssessStrictly reports whether strict assessment has something to assess an element against.
That is a weaker question than HasElementDeclaration, because §3.3.4 clause 1.2 also assesses an element carrying xsi:type against the type it names, declaration or no declaration. Validate already takes that path; without this the XSLT layer refused the element as undeclared before ever getting there, so <doc xsi:type="xs:anyType"> under validation="strict" reported no top-level declaration rather than validating.
func (*Schema) CheckConstraints ¶
func (s *Schema) CheckConstraints(opts CheckOptions) error
CheckConstraints applies the schema component constraints that are checked against a compiled content model: Unique Particle Attribution and Element Declarations Consistent.
Loading already applies both — see checkContentModelConstraints — so this re-runs work the schema has passed. It remains exported because it is the only way to ask for the *permissive* UPA reading after the fact: a caller that loaded with Options.LaxUPA set cannot otherwise re-check under the strict rule, and a caller holding a Schema from elsewhere may want the constraints stated as a list of errors rather than as a load failure.
func (*Schema) HasAttributeDeclaration ¶ added in v1.0.0
HasAttributeDeclaration reports whether the schema declares a global attribute of that name.
func (*Schema) HasElementDeclaration ¶ added in v1.0.0
HasElementDeclaration reports whether the schema declares a global element of that name, and HasAttributeDeclaration does the same for an attribute.
A caller that must distinguish "not declared" from "declared and invalid" — XSLT's XTTE1512 against XTTE1510 — needs to ask before validating, because once validation has run both look like a failure.
func (*Schema) HasSimpleType ¶ added in v1.0.0
HasSimpleType reports whether a name is a simple type in this schema, which is the precondition ValidateValue needs a caller to have checked.
func (*Schema) IsListSimpleType ¶ added in v1.1.0
IsListSimpleType reports whether a name is a simple type of variety list in this schema.
Casting to a list type is permitted by XPath 3.0 F&O 18.3 ("This section defines how other casts to non-primitive types operate, including casting to types derived by restriction, to union types, and to list types"), but a list value is a SEQUENCE of item-type values rather than one atomic item, so the caller has to know which it is dealing with before it can decide whether a cast target is legal. The engine's built-in list types are recognised by name; a schema-defined one can only be recognised by asking the schema.
func (*Schema) TypeEnv ¶ added in v1.3.0
func (s *Schema) TypeEnv() *xdm.TypeEnvironment
TypeEnv returns the type environment this schema owns: the derivation, list and union facts its own type definitions establish.
It is exported so that a caller holding a schema -- xslt's validation instructions are the case in this repository -- can answer a question about one of the schema's type NAMES without going through the process-global table, which is keyed by name across every schema in the process and so answers for whichever loaded last.
It is also what an aggregate schema merges: xsl:import-schema and XQuery's "import schema" fold several loaded schemas into one, and the derivation facts have to travel with the type definitions or the aggregate holds types that no longer know what they derive from.
func (*Schema) Validate ¶
func (s *Schema) Validate(root *xdm.Node, opts ValidateOptions) error
Validate checks a document against the schema.
It returns nil when the document is valid. The error, when there is one, is a *ValidationErrors holding every failure found up to the limit.
It cannot be cancelled. Use ValidateContext when the document comes from somewhere you do not control: identity-constraint checking is quadratic in nesting depth (see docs/security.md), and a deadline is what bounds it.
func (*Schema) ValidateAgainstType ¶ added in v1.0.0
ValidateAgainstType checks one element or attribute against a named type.
This is the [xsl:]type attribute, which names a type directly rather than letting the element's own name select a declaration — so an element called anything at all may be asked to match xs:integer.
func (*Schema) ValidateAttribute ¶ added in v1.0.0
ValidateAttribute checks one attribute against the global declaration for its name.
It is the attribute counterpart of ValidateElement, for a stylesheet that copies an attribute under validation="strict": the declaration selects the type, and the value has to satisfy it. lax passes an attribute the schema does not declare; strict rejects it.
func (*Schema) ValidateContext ¶ added in v1.2.2
ValidateContext is Validate with a cancellable context.
A cancelled or expired context stops the run and is returned as the context's own error — context.Canceled or context.DeadlineExceeded — rather than as a *ValidationErrors, so errors.Is tells "I gave up" apart from "the document is invalid". Whatever failures had been found are discarded: a partial list from a document that was never fully assessed would read as a verdict, and it is not one.
A nil ctx is treated as context.Background().
func (*Schema) ValidateElement ¶ added in v1.0.0
func (s *Schema) ValidateElement(el *xdm.Node, opts ValidateOptions) error
ValidateElement checks one element against the global declaration for its name.
This is xsl:element validation="strict": the element must have a global declaration and must be valid against it. An element with no declaration is an error rather than a pass, which is what distinguishes strict from lax.
func (*Schema) ValidateElementLax ¶ added in v1.0.0
func (s *Schema) ValidateElementLax(el *xdm.Node, opts ValidateOptions) error
ValidateElementLax checks one element against its global declaration when there is one, and passes when there is not.
That is what validation="lax" means: an element the schema does not describe is not thereby invalid. It is the mode that lets a stylesheet validate the parts of its output that are described without having to describe all of it.
"Does not describe" is not the same as "has no declaration", though. XSD 1.0 §3.3.4 clause 1.2.1.2 resolves an xsi:type attribute and assesses the element against the type it names whether or not a declaration exists, and lax assessment takes that path as readily as strict does -- CanAssessStrictly below states the same rule for the strict side. Skipping on the declaration alone let an element that named its own type escape the assessment it had asked for: "validate lax { <z:person xsi:type='xs:NCName'>abc 123</z:person> }" returned the element instead of the XQDY0027 the invalid NCName owes (qischema90621-err).
func (*Schema) ValidateExpandedQNameValue ¶ added in v1.0.0
ValidateExpandedQNameValue checks an already-expanded QName against a named simple type whose value space is the QName one.
The constructor of a type derived from xs:NOTATION or xs:QName has to resolve its argument's prefix while the static context of the expression still exists, so by the time the facets can be checked there is no node and no in-scope namespaces left — and comparing the raw lexical form against the schema's enumeration matched prefixes rather than namespaces, which made one:mp3 fail an enumeration written smokey:mp3 for the same URI. The expanded name is passed instead, in the "{uri}local" spelling that no lexical QName can have, and the lexical checks are skipped because the caller has already done the only one that applies to a QName literal.
known is false when the name is not a QName-valued simple type in this schema, in which case the caller keeps whatever answer it already had.
func (*Schema) ValidateValue ¶ added in v1.0.0
ValidateValue checks a lexical value against a named simple type, without a node to hang it on.
"castable as my:hatsize" asks exactly this question: whether a value is in the value space of a schema type, facets and all. The XPath engine holds only the type's name, so it has to be able to ask without constructing an element or attribute first — and without that, a cast to a user-defined type checked the built-in it derives from and ignored every facet the schema author wrote.
A name that is not a simple type in this schema is an error rather than a pass: a caller asking about a complex type has asked a question with no answer, and reporting success would make the cast succeed.
func (*Schema) WithInstanceLocations ¶
func (s *Schema) WithInstanceLocations(root *xdm.Node, policy InstanceLocationPolicy, opts Options) (*Schema, error)
WithInstanceLocations returns a schema extended by the documents an instance names in xsi:schemaLocation and xsi:noNamespaceSchemaLocation.
The receiver is not modified: a Schema is immutable once loaded and safe to share, so extending one produces another. The cost is a fresh assembly per instance, which is why this is a separate call rather than something Validate does — a caller validating many documents against one schema should not pay for it, and most callers should not use this at all.
Locations the policy refuses are ignored. §4.3.2 makes the attribute a hint, so declining to take it is not a fault in the document; a reference that really needed the components still fails, at the reference.
type SchemaErrors ¶
type SchemaErrors struct {
Errors []error
}
SchemaErrors collects every fault found in a schema.
func (*SchemaErrors) Error ¶
func (e *SchemaErrors) Error() string
Error implements error, listing every fault.
func (*SchemaErrors) Unwrap ¶
func (e *SchemaErrors) Unwrap() []error
Unwrap exposes the faults to errors.As and errors.Is.
type Scope ¶
type Scope uint8
Scope distinguishes a global declaration from a local one.
The distinction is not cosmetic: a global element declaration can be the head of a substitution group and can be the validation root, while a local one is reachable only through the type that contains it. Two local declarations with the same name in different types are different components.
type SequenceMatcher ¶
type SequenceMatcher struct {
// contains filtered or unexported fields
}
SequenceMatcher decides whether a sequence of element names satisfies a content model.
Schema validation does far more than this — it annotates types, collects identity-constraint tables, and reports where a failure happened — so it does not go through here. This exists for a caller that has a particle and a list of names and wants only the yes-or-no: the DTD validator is the case it was added for, since a DTD content model is a strict subset of what a particle expresses and rebuilding the automaton for it would be duplication.
A matcher is immutable once compiled and safe to share across goroutines.
func NewSequenceMatcher ¶
func NewSequenceMatcher(p *Particle) (*SequenceMatcher, error)
NewSequenceMatcher compiles p for repeated matching.
The model is bounded by DefaultMaxContentModelPositions. This entry takes a bare particle rather than a Schema — the DTD validator is its caller, and a DTD has no Options to carry — so there is no per-load budget to honour here and the default is the only bound available. A caller needing a different one goes through Load, where Options.MaxContentModelPositions applies.
func (*SequenceMatcher) Match ¶
func (s *SequenceMatcher) Match(names []xdm.QName) (bool, int)
Match reports whether names is admitted by the model, and when it is not, the index of the first name that could not be placed.
A rejection at index len(names) means the sequence ended early — every name was placed but the model required more.
type SimpleType ¶
type SimpleType struct {
Name xdm.QName
Base Type
FinalSet DerivationSet
Variety Variety
// Primitive is the primitive type this one erases to, for atomic
// varieties. A primitive type is its own primitive.
Primitive *SimpleType
// ItemType is {item type definition}, meaningful only for VarietyList.
ItemType *SimpleType
// MemberTypes is {member type definitions}, meaningful only for
// VarietyUnion. Order is significant: validation takes the *first*
// member that accepts the value, not the best or longest match.
MemberTypes []*SimpleType
// Facets are the constraining facets applied at this derivation step.
// A value must satisfy these and every ancestor's; see facet.go for how
// they combine, which is not simply union.
Facets *FacetSet
// contains filtered or unexported fields
}
SimpleType is a Simple Type Definition (Part 2 §4.1.1).
Part 1 §3.14.1 also defines this component, but that section is marked non-normative and disagrees with Part 2 about {final}. Part 2 governs here.
func BuiltinType ¶
func BuiltinType(local string) *SimpleType
BuiltinType returns a built-in simple type by local name, or nil.
It exists so that a caller validating a lone value against, say, xs:date does not have to construct a schema first.
func (*SimpleType) ComponentKind ¶
func (*SimpleType) ComponentKind() string
ComponentKind implements Component.
type Term ¶
type Term interface {
Component
// contains filtered or unexported methods
}
Term is what a particle constrains: an element declaration, a wildcard, or a model group.
type Timezone ¶
type Timezone uint8
Timezone is the value of the XSD 1.1 explicitTimezone facet.
type Type ¶
type Type interface {
Component
// TypeName is the {name} and {target namespace}. An anonymous type has
// the zero QName, which is why this cannot simply be a field access.
TypeName() xdm.QName
// BaseType is {base type definition}. For xs:anyType this is xs:anyType
// itself — a deliberate self-loop in the spec, so any walk up the base
// chain must test for it rather than for nil.
BaseType() Type
// Final is {final}: the derivations this type forbids.
Final() DerivationSet
// contains filtered or unexported methods
}
Type is a type definition: either a SimpleType or a ComplexType.
The interface is deliberately narrow. Almost every validation rule needs to know a type's name, its base, and how it may be derived from; the rules that need more do a type switch, because what they need differs entirely between the two kinds.
type TypeAlternative ¶
type TypeAlternative struct {
// Test is the compiled condition, or nil for the default alternative.
Test *xpath.Compiled
// Source is the condition as written.
Source string
// Type is the type to use when the test matches.
Type Type
// contains filtered or unexported fields
}
TypeAlternative is an <xs:alternative> (XSD 1.1 §3.3).
A declaration may carry several; the first whose test matches selects the type. An alternative with no test is the default, and may appear only last.
type ValidateOptions ¶
type ValidateOptions struct {
// MaxErrors stops validation once this many failures are found. Zero
// means DefaultMaxErrors; a negative value means no limit, as it does
// for dtd.Options.MaxErrors. A document that is wrong in every element
// would otherwise produce an error for each, which helps nobody and
// costs memory proportional to the document.
MaxErrors int
// Annotate writes the type of each validated node into its
// TypeAnnotation, producing the part of the PSVI that the XPath and
// XSLT layers consume. It is off by default because it mutates the
// tree the caller passed in.
Annotate bool
// SkipIDConstraints suppresses "Validation Root Valid (ID/IDREF)"
// (§3.3.4 clause 2) — the check that ID values are unique and that
// every IDREF resolves.
//
// That rule is scoped to the *document*, not to the subtree being
// assessed, so a caller assessing a bare element has no document for it
// to be true of. XSLT 2.0 §19.2.1.3 says so outright: when validating a
// constructed element, "the validation rule 'Validation Root Valid
// (ID/IDREF)' is not applied ... validation will not fail if there are
// non-unique ID values or dangling IDREF values in the subtree being
// validated". The same section applies it again for a document node.
//
// Off by default: validating a parsed document is the ordinary case and
// there the rule does apply.
SkipIDConstraints bool
// MaxDepth bounds how deep validation will recurse. Zero means
// DefaultMaxDepth; a negative value means no limit, which gives up the
// clean error for a fatal stack overflow that recover() cannot catch.
//
// This is not the parser's limit. A tree can be built by a transform
// rather than parsed, and a caller who raises xdm.ParseOptions.MaxDepth
// to accept a deep document has not thereby agreed to let the validator
// recurse that far — see validateElement for why exceeding it is fatal
// rather than recoverable.
MaxDepth int
}
ValidateOptions configure a validation run.
type ValidationError ¶
type ValidationError struct {
// Code is the spec's error code, such as "cvc-complex-type.2.4".
Code string
// Message describes the failure.
Message string
// Path is the location in the instance, as an element path.
Path string
// Line and Column locate it when the node carried a position.
Line, Column int
// Err is a sentinel this failure also is, for callers that need to
// branch on the KIND of failure rather than on its code.
//
// It exists for xdm.ErrResourceLimit. A depth refusal is reported with
// a cvc- code because that is the only vocabulary a validation result
// has, and cvc-elt.1 then says "this document is invalid" about a
// document that was never assessed. The code is kept -- callers and
// the conformance suites read it -- and the sentinel carried here
// alongside, so errors.Is can tell a refusal from a verdict. Nil on
// every ordinary failure, which is all of them but this one.
Err error
}
ValidationError reports one reason a document is not valid.
The spec gives each validation rule an error code — cvc-complex-type, cvc-datatype-valid and so on — and carrying it lets a caller distinguish the kinds of failure without matching on message text.
func (*ValidationError) Unwrap ¶ added in v1.2.2
func (e *ValidationError) Unwrap() error
Unwrap exposes Err, so errors.Is reaches a sentinel a failure carries.
type ValidationErrors ¶
type ValidationErrors struct {
Errors []*ValidationError
}
ValidationErrors is the set of failures found in one document.
func (*ValidationErrors) Unwrap ¶ added in v1.2.2
func (e *ValidationErrors) Unwrap() []error
Unwrap exposes the individual failures, so errors.Is over the set reaches a sentinel any one of them carries.
type ValueConstraint ¶
type ValueConstraint struct {
// Fixed distinguishes fixed="v" from default="v".
Fixed bool
// Lexical is the value as written in the schema document. It is stored
// unparsed because the type it must be parsed against is not always
// known when the schema is read.
Lexical string
// Value is the parsed form, filled in once the type is resolved.
Value *xdm.Atomic
}
ValueConstraint is the {value constraint} property: a default or fixed value.
The distinction matters at validation time, not just at defaulting time: a fixed value must equal the value in the instance, while a default supplies one when the instance has none.
type Variety ¶
type Variety uint8
Variety discriminates a simple type (Part 2 §4.1.1).
type Version ¶
type Version uint8
Version selects the XML Schema version a schema is interpreted under.
const ( // Version10 is XML Schema 1.0, the default. Version10 Version = iota // Version11 is XML Schema 1.1: assertions, conditional type // assignment, open content, and the relaxed rules that come with them. Version11 )
The supported versions.
func DocumentRequiresVersion ¶ added in v1.1.0
DocumentRequiresVersion reports the XSD version a schema document asks to be read under, and whether it asks for one at all.
A document whose <xs:schema> carries vc:minVersion="1.1" is telling a 1.0 processor to ignore it entirely -- includeElement drops the whole document, which is exactly what the attribute means. That is the right answer for a processor that only has 1.0, and the wrong one for a processor that has 1.1 and simply was not told to use it: the document then vanishes in silence, contributing no declaration at all.
A caller that supports both versions can use this to read the document under the version it names. Only minVersion is consulted: maxVersion says the document is for processors BELOW a version, which is not a request to upgrade.
type WhiteSpace ¶
type WhiteSpace uint8
WhiteSpace is the value of the whiteSpace facet (Part 2 §4.3.6).
The three modes are ordered by strength, and a derivation may only make the value stronger: preserve → replace → collapse. That ordering is what the comparison in checkWhiteSpaceRestriction relies on.
const ( // WhitePreserve leaves the value alone. It is the value for xs:string // and for xs:anySimpleType. WhitePreserve WhiteSpace = iota // WhiteReplace maps tab, newline and carriage return to a space. WhiteReplace // WhiteCollapse additionally merges runs of spaces and trims the ends. // Every built-in type except the string family collapses. WhiteCollapse )
The whiteSpace modes.
func EffectiveWhiteSpace ¶
func EffectiveWhiteSpace(t *SimpleType) WhiteSpace
EffectiveWhiteSpace returns the whiteSpace value in force for a type, walking up the base chain until a step sets one.
The walk terminates on xs:anySimpleType rather than on nil, and must also guard against xs:anyType, whose base type definition is *itself* — a deliberate self-loop in the spec that turns a naive walk into an infinite one.
func (WhiteSpace) Normalize ¶
func (w WhiteSpace) Normalize(s string) string
Normalize applies the whiteSpace facet to a lexical value.
Only the four XML whitespace characters take part. Using unicode.IsSpace here would fold characters such as U+00A0 that XML does not treat as whitespace, which would silently accept values the spec rejects.
func (WhiteSpace) String ¶
func (w WhiteSpace) String() string
String names the mode as it is spelled in a schema document.
type Wildcard ¶
type Wildcard struct {
Kind NamespaceConstraintKind
ProcessContents ProcessContents
// Namespace is the excluded namespace for NSNot, or the permitted set
// for NSEnumerated. It is unused for NSAny.
Namespace []string
// ExcludesAbsent records whether the absent namespace is excluded by an
// NSNot constraint.
//
// The two spellings differ here and the difference is easy to miss.
// XSD 1.0's ##other excludes unqualified names unconditionally —
// clause 2.3 of Wildcard allows Namespace Name — whereas XSD 1.1's
// notNamespace excludes only what it lists, so an unqualified name is
// permitted unless ##local appears. Applying ##other's rule to
// notNamespace rejects every unqualified attribute the wildcard was
// written to admit.
ExcludesAbsent bool
// DisallowedNames is XSD 1.1's {disallowed names} (§3.10.1): specific
// expanded names the wildcard refuses even though their namespace is
// admitted. It is the notQName attribute, and it is what lets a schema
// say "anything from this namespace except these".
//
// The namespace constraint and this set are independent tests: a name
// matches the wildcard only if the namespace admits it *and* it is not
// disallowed. That ordering matters because notQName may name a
// namespace the constraint would otherwise let through, which is the
// only reason to write it.
DisallowedNames []xdm.QName
// DisallowDefined is ##defined: refuse any name for which the schema
// has a global declaration of the matching kind. It is how a schema
// writes "anything the schema does not already know about", which is
// the useful form of an extension wildcard — one that cannot silently
// shadow a declared element.
DisallowDefined bool
// DisallowDefinedSibling is ##definedSibling: refuse any name declared
// by some other particle in the same content model. Unlike ##defined it
// is local, and it applies to elements only — an attribute wildcard has
// no siblings in the sense the keyword means, and the schema for
// schemas does not permit it there.
DisallowDefinedSibling bool
// contains filtered or unexported fields
}
Wildcard is a Wildcard (§3.10.1) — what <xs:any> and <xs:anyAttribute> map to.
func (*Wildcard) Allows ¶
Allows reports whether the wildcard permits a name in namespace ns, where the empty string means the absent namespace.
The NSNot case is the one worth reading twice: ##other excludes the absent namespace as well as the named one, so an unqualified element never matches a ##other wildcard.
func (*Wildcard) AllowsName ¶
AllowsName reports whether the wildcard admits an expanded name, applying both the namespace constraint and {disallowed names}.
func (*Wildcard) ComponentKind ¶
ComponentKind implements Component.
Source Files
¶
- assemble.go
- assert.go
- automaton.go
- builtin.go
- catalog.go
- component.go
- ctype_check.go
- edc_typetable.go
- facet.go
- facet_check.go
- icpath.go
- identity.go
- instloc.go
- jsonschema.go
- nfa.go
- parse.go
- parse_decl.go
- parse_type.go
- pattern.go
- redefine.go
- resolve.go
- restrict.go
- seqmatch.go
- srcmodel.go
- subsume.go
- temporal.go
- typevalidate.go
- upa.go
- validate.go
- validate_attr.go
- validate_simple.go
- versioning.go