Documentation
¶
Overview ¶
Package relaxng validates XML documents against RELAX NG schemas.
RELAX NG validates by a different model from XSD: a schema is a *pattern*, and validation computes the derivative of that pattern with respect to each item of input, accepting when what remains can match the empty sequence. That is why this is a separate engine rather than a use of the XSD automaton — there is no finite automaton to build, and interleave, which admits its branches in any order, is not something a Glushkov construction expresses.
The implementation follows James Clark's derivative algorithm, which is both the clearest description of the language and the one the conformance suite was written against.
Index ¶
Constants ¶
const DefaultMaxDepth = 1000
DefaultMaxDepth bounds validation recursion when MaxDepth is zero. It matches xdm.DefaultMaxDepth, so a document the parser accepts is one the validator will not refuse for depth alone.
const DefaultMaxPatternSize = 100_000
DefaultMaxPatternSize bounds the derivative pattern when MaxPatternSize is zero.
It is set high enough that no schema in the RELAX NG spec test suite comes near it — the whole suite passes unchanged — and low enough that the multiplicative blowup is refused in milliseconds rather than after a gigabyte of allocation.
const NS = "http://relaxng.org/ns/structure/1.0"
NS is the RELAX NG structure namespace.
Variables ¶
This section is empty.
Functions ¶
func ParseCompact ¶ added in v1.3.0
ParseCompact parses a RELAX NG compact syntax schema and returns the equivalent XML syntax document.
The result is exactly what an author would have written in the XML syntax, so it can be handed to Compile — and CompileCompact does just that. It is returned rather than kept private because the translation is useful on its own: it is how a caller converts a .rnc to a .rng, and how a reader checks what a compact schema actually means.
The compact syntax specification is not vendored in this repository; see the note at the top of compact_lex.go for what that means for the rules implemented here.
Types ¶
type FileResolver ¶ added in v1.3.0
type FileResolver struct {
// Root confines reads when non-empty.
Root string
// MaxBytes bounds one fetched schema. Zero uses defaultMaxSchemaBytes.
// A negative value refuses every read rather than disabling the bound.
MaxBytes int64
}
FileResolver is a filesystem-backed Resolver.
Root is a capability grant, not a cosmetic path prefix: with it set, every reference must remain below that directory even if a schema uses xml:base, .., an absolute file URL, or a symlink changed between validation and open. os.Root enforces the last property at open time. An empty Root intentionally remains unconfined for command-line callers that explicitly choose it. Network schemes are never fetched; a caller needing them must write a resolver with an explicit host and transport policy.
func (*FileResolver) ResolveSchema ¶ added in v1.3.0
func (r *FileResolver) ResolveSchema(href string) (*xdm.Node, error)
ResolveSchema implements Resolver.
type Options ¶
type Options struct {
// Resolver supplies the documents named by <externalRef> and <include>.
// When nil, both are refused.
Resolver Resolver
// BaseURI is the location the schema itself was read from, against which
// a relative href is resolved. It may be empty when the schema came from
// somewhere with no location, in which case only absolute hrefs work.
BaseURI string
}
Options configure compilation.
type Resolver ¶
Resolver fetches a schema document named by an <externalRef> or <include>.
It is an interface, and there is no default implementation, for the same reason DOCTYPE is refused by default and xsi:schemaLocation is ignored: an href in a schema is an instruction to go and read something, and where that read is allowed to reach is the caller's decision, not the schema author's. A caller that wants files supplies one that reads files; a caller that wants nothing supplies nothing, and every href is refused with an error that says so.
href is the value written in the schema, already resolved against the base URI in force — xml:base and the location the schema was loaded from — so an implementation receives one absolute reference rather than having to track the nesting itself.
type Schema ¶
type Schema struct {
// contains filtered or unexported fields
}
Schema is a compiled RELAX NG schema.
It is immutable once built and safe to share across goroutines, like a compiled XSD schema: validation takes derivatives of the pattern rather than mutating it.
func Compile ¶
Compile builds a schema from a parsed RELAX NG document in the XML syntax.
The grammar is checked as it is read rather than afterwards: RELAX NG's restrictions are mostly about what may appear where — an attribute inside an attribute, a text inside a list — and catching those at the point of use gives an error that names the construct.
func CompileCompact ¶ added in v1.3.0
CompileCompact compiles a schema written in the compact syntax.
It is ParseCompact followed by CompileWithOptions, and exists so that the ordinary case is one call. A Resolver supplied in opts serves "include" and "external" the same way it serves <include> and <externalRef>, since by the time compilation happens the two syntaxes are the same tree.
One asymmetry is worth naming: a Resolver returns an *xdm.Node, an XML syntax document. A compact schema that includes another compact schema therefore needs a Resolver that parses .rnc — ParseCompact is exported so that such a Resolver can be written in a few lines, and CompactResolver does it for the common case.
func CompileWithOptions ¶
CompileWithOptions builds a schema, with a Resolver for <externalRef> and <include>.
Compile is this with no options, which refuses both. Splitting them keeps the safe thing the short thing to write: reaching outside the schema document is something a caller opts into, not something a schema can decide for itself.
func (*Schema) Validate ¶
Validate checks a document against the schema.
The result is a single error rather than a list, which is the shape the derivative algorithm gives: a pattern that reaches notAllowedPat carries no record of the alternatives it tried, so there is one failure and it is the point at which every branch died. Reporting the *last* place the document was still viable is more useful than reporting the root.
func (*Schema) ValidateWithOptions ¶
func (s *Schema) ValidateWithOptions(doc *xdm.Node, opts ValidateOptions) error
ValidateWithOptions checks a document, with limits on the run.
type ValidateOptions ¶
type ValidateOptions struct {
// MaxDepth bounds how deep validation will recurse. Zero means
// DefaultMaxDepth; a negative value means no limit -- which ends in a
// fatal stack overflow rather than an error, since the runtime does not
// deliver stack exhaustion as a recoverable panic. Remove the bound only
// for input you produced yourself.
//
// This is not the parser's limit, and the distinction matters more here
// than elsewhere: taking derivatives over a nested document costs time and
// memory *quadratic* in the depth, since each level carries the pattern
// remaining at every level above it. A tree can also be built by a
// transform rather than parsed, and a caller who raises
// xdm.ParseOptions.MaxDepth to accept a deep document has not thereby
// agreed to let the validator spend a gigabyte on it.
MaxDepth int
// MaxPatternSize bounds the size of the derivative pattern carried
// during validation. Zero means DefaultMaxPatternSize; a negative value
// means no limit.
//
// It is a separate knob from MaxDepth because it bounds a different
// thing. MaxDepth bounds cost that grows with how deep the document is;
// this bounds cost that grows with how WIDE it is, which a schema
// nesting oneOrMore inside oneOrMore makes multiplicative — a 63-byte
// instance of fourteen children measured at 1.2 GB before this existed,
// at a depth of two, where no depth bound could reach it.
MaxPatternSize int
}
ValidateOptions bound one validation run.