xslt

package
v1.0.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 24, 2026 License: MIT Imports: 23 Imported by: 0

Documentation

Overview

Package xslt implements XSLT 2.0 transformation over the xdm data model, using the xpath package for expression evaluation.

Index

Constants

View Source
const DefaultMaxDepth = 1000

DefaultMaxDepth bounds template recursion when TransformOptions.MaxDepth is zero. It matches xdm.DefaultMaxDepth so that a document the parser accepts is one an identity transform can copy: the recursion counted here is the ordinary descent through the tree, not only a stylesheet calling itself.

Variables

This section is empty.

Functions

func SerializeAsXML added in v1.0.0

func SerializeAsXML(r *Result) string

SerializeAsXML renders a result with the xml output method, ignoring the stylesheet's own xsl:output.

It exists for a caller making a *tree* assertion about a result — the W3C conformance harness is the one in this repository. The stylesheet's method is part of what serialisation means, not part of what the tree is: the html method injects a content-type meta into <head> and writes void elements unclosed, so a result asserted as a tree would be compared against markup the stylesheet never produced, and would not parse back as XML at all.

Indentation and the other settings are deliberately left at their defaults rather than inherited, for the same reason.

Types

type CompileOptions

type CompileOptions struct {
	// Resolver loads xsl:include and xsl:import targets. Nil disables them,
	// which is the safe default for untrusted stylesheets.
	Resolver ModuleResolver
	// BaseURI of the stylesheet, for resolving relative include paths.
	BaseURI string
	// StaticParams supplies values for top-level xsl:param at compile time.
	StaticParams map[string]string

	// SchemaResolver loads the schemas named by xsl:import-schema. Nil
	// disables loading by location, for the same reason a nil Resolver
	// disables xsl:include: following a location means fetching whatever
	// the stylesheet names. An inline <xs:schema> child needs no resolver.
	SchemaResolver xsd.Resolver
}

CompileOptions configures compilation.

type DecimalFormat

type DecimalFormat struct {
	Name              xdm.QName
	DecimalSeparator  rune
	GroupingSeparator rune
	Percent           rune
	PerMille          rune
	ZeroDigit         rune
	Digit             rune
	PatternSeparator  rune
	MinusSign         rune
	Infinity          string
	NaN               string
}

DecimalFormat holds an xsl:decimal-format declaration.

Every symbol is configurable because the instruction exists to serve locales: a German invoice writes 1.234,56 where an English one writes 1,234.56, and the picture string is written once against whatever symbols the format declares.

type FileResolver

type FileResolver struct {
	// Roots are the directories a relative or absolute path may resolve
	// inside. A path escaping all of them is refused.
	Roots []string

	// AllowDOCTYPE permits a DOCTYPE declaration in the documents this
	// resolver parses. It is off by default, which is what keeps a
	// stylesheet from reaching a document that expands entities or names an
	// external one — the XXE entry point. A caller whose inputs are trusted,
	// a conformance suite among them, can turn it on.
	AllowDOCTYPE bool

	// ExternalEntities permits the documents this resolver parses to read
	// external entities and an external DTD subset, using this same resolver
	// — so they are confined to Roots on exactly the terms everything else
	// is, with the same scheme rejection and the same symlink handling.
	//
	// It is separate from AllowDOCTYPE and off by default. AllowDOCTYPE
	// admits declarations that cost nothing outside the document; this
	// admits reads of other files, which is the XXE surface proper. A caller
	// that wants DTD-declared entities does not thereby want file reads, and
	// making one imply the other would silently widen every existing caller.
	ExternalEntities bool

	// UnparsedText permits fn:unparsed-text to read files through this
	// resolver, confined to Roots on the same terms as everything else.
	//
	// It is separate from every other flag here and off by default, because
	// it is the widest of them. ResolveDocument hands the stylesheet a
	// parsed XML document, so a file that is not well-formed XML discloses
	// nothing; unparsed-text hands back the raw bytes of any file inside
	// Roots, so a root containing one XML data file and one private key
	// leaks the key. A caller who wants fn:doc does not thereby want that,
	// and folding the two together would silently widen every existing
	// caller of NewFileResolver.
	UnparsedText bool
	// contains filtered or unexported fields
}

FileResolver loads stylesheet modules and documents from the filesystem, confined to a set of allowed directories.

The confinement is the point. A stylesheet that can call document() on any path is a file-disclosure primitive, and one that can reach http:// is an SSRF primitive — but a blanket deny is not workable either, because real rule sets load code lists that ship beside the stylesheet. Naming the directories makes the trust boundary explicit and auditable.

func NewFileResolver

func NewFileResolver(roots ...string) (*FileResolver, error)

NewFileResolver returns a resolver confined to the given directories.

func (*FileResolver) Preload added in v1.0.0

func (r *FileResolver) Preload(uri string, tree *xdm.Tree)

Preload records a tree that has already been parsed as the answer for uri, so that a later fn:doc or fn:document naming the same resource hands back the very same nodes rather than a second parse of the same bytes.

Node identity is the point. XSLT 2.0 section 16.1 requires two retrievals of one absolute URI to return the same node, and the test the specification is written for is "fn:doc(fn:document-uri($arg)) is $arg". A caller that parses the principal source itself — every conformance harness does, because it has to annotate the tree before the transform sees it — supplies a document node the resolver has never heard of, and doc() of its own document-uri then parses the file again and answers a different node. Preloading closes that gap without weakening the containment check: the uri still has to resolve to a path inside a permitted root, and an unresolvable one is a no-op.

func (*FileResolver) ResolveDocument

func (r *FileResolver) ResolveDocument(uri, base string) (*xdm.Tree, error)

ResolveDocument implements xpath.DocumentResolver for fn:doc and fn:document.

func (*FileResolver) ResolveEntity added in v1.0.0

func (r *FileResolver) ResolveEntity(systemID, publicID, base string) (io.ReadCloser, string, error)

ResolveEntity implements xdm.EntityResolver, so that a document this resolver parses may read external entities — but only from inside Roots.

Every constraint the rest of this type enforces applies here unchanged, because the path goes through the same resolvePath: a non-file scheme is rejected before the filesystem is touched, symlinks are resolved before the containment check, and a path outside every root is refused. There is nothing entity-specific about the confinement, which is the point — an external entity is a file read like any other, and it gets the same gate rather than a second one written separately and drifting.

The base is the URI of the resource that made the reference, which for an entity declared in an external DTD subset is that subset rather than the document. Resolving against it is XML 1.0 section 4.4.3, and it is why a modular DTD in a subdirectory finds its siblings.

The returned URI is the file: URI of what was actually read, since that is what anything inside the fetched text resolves against.

func (*FileResolver) ResolveModule

func (r *FileResolver) ResolveModule(href, base string) (*xdm.Node, string, error)

ResolveModule implements ModuleResolver for xsl:include and xsl:import.

func (*FileResolver) ResolveText added in v1.0.0

func (r *FileResolver) ResolveText(uri, base, encoding string) (string, error)

ResolveText implements xpath.TextResolver for fn:unparsed-text.

The path goes through the same resolvePath as every other read this type performs, so the confinement is one implementation rather than two: a non-file scheme is rejected before the filesystem is touched, symlinks are resolved before the containment check, and a path outside every root is refused. UnparsedText only decides *whether* to ask; it does not relax where the answer may come from.

The encoding argument is honoured only for the encodings this package can decode without pulling in a converter. XSLT 2.0 section 16.2 requires an error for an encoding that is not supported, and reporting one is better than silently returning mojibake -- a stylesheet that reads a Shift-JIS file and gets bytes reinterpreted as UTF-8 produces wrong output with no indication anything went wrong.

type Instruction

type Instruction interface {
	// Execute runs the instruction, appending to out.
	Execute(rt *runtime, out *outputBuilder) error
}

Instruction is one compiled XSLT instruction.

Instructions write to an output builder rather than returning values, because an XSLT sequence constructor produces a *stream* of nodes and atomic values: xsl:element opens a node that subsequent instructions add children to. Returning trees from each instruction and concatenating them would mean building and copying the same subtree repeatedly.

type ModuleResolver

type ModuleResolver interface {
	ResolveModule(href, base string) (*xdm.Node, string, error)
}

ModuleResolver loads an included or imported stylesheet module.

type OutputSettings

type OutputSettings struct {
	Method        string // "xml", "html", "text"
	Indent        bool
	OmitXMLDecl   bool
	Encoding      string
	DocTypePublic string
	DocTypeSystem string
	CDataElements []xdm.QName
	Standalone    string
	// Version is xsl:output/@version; "5.0" selects the HTML5 doctype.
	Version string
	// UseCharacterMaps names the xsl:character-map declarations applied at
	// serialisation.
	UseCharacterMaps []xdm.QName
	// ByteOrderMark writes a BOM at the start of the output. It is "no" by
	// default for every method, which is what makes UTF-8 output usable by
	// readers that do not expect one.
	ByteOrderMark bool
	// IncludeContentType controls whether the HTML and XHTML methods insert
	// the content-type meta element. It defaults to true, which is why it is
	// stored as a pointer: an explicit "no" has to be distinguishable from
	// the attribute being absent.
	IncludeContentType *bool
	// EscapeURIAttributes controls percent-escaping of URI-valued attributes
	// in the HTML and XHTML methods. It defaults to true, and is a pointer
	// for the same reason.
	EscapeURIAttributes *bool
	// UndeclarePrefixes emits namespace undeclarations, which only XML 1.1
	// permits. This parser implements XML 1.0, so it is recorded and ignored.
	UndeclarePrefixes bool
	// ItemSeparator is xsl:output/@item-separator: the string inserted
	// between adjacent items of the result sequence during sequence
	// normalisation (XSLT 3.0 section 5.7.1 step 3).
	//
	// It is a pointer because an explicit zero-length separator is not the
	// same as the attribute being absent: absent means the default rule
	// (adjacent atomic values separated by a single space, nodes not
	// separated at all), while item-separator="" means every adjacency gets
	// nothing, including between two atomic values.
	ItemSeparator *string
	// MediaType is the media type of the output. It affects no serialised
	// character; it is metadata a caller passes on.
	MediaType string
	// NormalizationForm names a Unicode normalisation applied to the output.
	// Only "none" is implemented; any other value the serialiser does not
	// support is a serialization error rather than something to ignore,
	// because output that was silently left unnormalised would be accepted
	// by a consumer that then compares it against a normalised form and
	// finds a spurious difference.
	NormalizationForm string
	// Version10Implicit records that the principal module declares version
	// "1.0" and this is the implicitly-created final result tree.
	//
	// It changes only the *default* output method. Under backwards
	// compatibility an XSLT 1.0 stylesheet has no xhtml method to select —
	// the method did not exist — so a result whose document element is html
	// in the XHTML namespace serialises as xml, not xhtml: URI-valued
	// attributes are left unescaped and no content-type meta element is
	// added. An explicit xsl:output/@method overrides this like any other
	// default, and xsl:result-document clears the flag, because the tree it
	// creates is not the implicit one.
	Version10Implicit bool
}

OutputSettings holds the xsl:output declaration.

type Pattern

type Pattern struct {
	// contains filtered or unexported fields
}

Pattern is a compiled xsl:template match pattern.

Patterns look like path expressions but mean something different: a path says "navigate from here", a pattern says "does this node match". The natural implementation of matching — evaluate the path and check membership — is quadratic, because it would visit every node in the document for every node being matched.

Instead a pattern is matched right-to-left from the candidate node: check the last step's node test against the node itself, then walk *up* verifying each preceding step. That makes a match cost O(depth) rather than O(document size), which is the difference between a transform that finishes and one that does not on a large invoice.

func CompilePattern

func CompilePattern(src string, ns xpath.NamespaceResolver) (*Pattern, error)

CompilePattern compiles an XSLT match pattern.

func (*Pattern) Alternatives added in v1.0.0

func (p *Pattern) Alternatives() []*Pattern

Alternatives splits a union pattern into one Pattern per branch.

Section 6.4 says a template rule whose match pattern is a union behaves as if it were several template rules, one per branch, each with the default priority computed for that branch alone. Keeping them fused would give the whole rule the highest branch's priority, so a low-priority branch would outrank templates it should lose to; it would also make xsl:next-match skip the rule entirely after the first branch fired, when the spec has it reconsider the rule for each remaining branch.

func (*Pattern) Matches

func (p *Pattern) Matches(node *xdm.Node, ctx *xpath.Context) (bool, error)

Matches reports whether node matches the pattern.

ctx supplies the focus for predicate evaluation. Predicates in a pattern are evaluated with the candidate node as the context item, and with a context position derived from its position among its like-named siblings — which is why "para[1]" as a pattern means "a para that is the first para child of its parent" rather than "the first para in the document".

func (*Pattern) Priority

func (p *Pattern) Priority() float64

Priority returns the pattern's default priority, per the XSLT rules: a specific name test scores 0, a namespace wildcard -0.25, a full wildcard or bare kind test -0.5, and anything more complex 0.5.

These numbers exist so that a more specific template wins over a general one without the author having to say so. Getting them wrong makes template selection silently pick the wrong rule, which is far harder to debug than a crash.

func (*Pattern) String

func (p *Pattern) String() string

String returns the pattern source.

type Result

type Result struct {

	// Nodes is the result sequence, which for a typical stylesheet is a
	// single element.
	Nodes xdm.Sequence
	// Messages holds xsl:message output, in the order produced.
	Messages []string
	// Secondary holds the documents produced by xsl:result-document, in the
	// order produced. It is empty for the great majority of stylesheets,
	// which produce a single result.
	Secondary []SecondaryResult
	// BaseURI is the URI this result tree is identified by, which Tree()
	// puts on the document node it manufactures. It is empty for the
	// principal result, whose document node has no URI of its own; a caller
	// assembling a Result from a SecondaryResult sets it from that
	// document's BaseURI so that base-uri(/) answers inside it.
	BaseURI string
	// contains filtered or unexported fields
}

Result is the outcome of a transform.

func (*Result) Serialize

func (r *Result) Serialize(w io.Writer) error

Serialize writes the result using the stylesheet's xsl:output settings.

Deliberately not named WriteTo: that name implies io.WriterTo, whose contract returns a byte count this would have to fabricate.

func (*Result) String

func (r *Result) String() string

String renders the result using the stylesheet's output settings.

func (*Result) Tree

func (r *Result) Tree() *xdm.Node

Tree returns the result as a document node, for callers that want to keep navigating it rather than serialise it — which is what a Schematron driver does with an SVRL report.

type SecondaryResult

type SecondaryResult struct {
	// Href is the resolved @href value, as written by the stylesheet. It is
	// the caller's choice what to do with it — this engine never writes to
	// the filesystem on a stylesheet's behalf, since a transform that can
	// create files anywhere the process can write is a hazard the caller
	// should be the one to opt into.
	Href string
	// BaseURI is Href resolved against the base output URI, which is the
	// base URI of every node in this document that does not override it with
	// xml:base.
	//
	// Section 19.1 makes the base output URI implementation-defined when the
	// caller does not supply one, and the stylesheet's own location is the
	// only URI this engine has: an @href of "out/second.xml" written in a
	// stylesheet read from .../foo.xsl means .../out/second.xml, which is
	// also where a caller honouring the href would write it. Leaving it
	// empty made base-uri() answer "" for every node in a secondary result,
	// and made a relative xml:base inside one resolve against nothing.
	BaseURI string
	// Nodes is the result sequence for this document.
	Nodes xdm.Sequence
	// Output holds the serialisation settings that apply to this document,
	// taken from @format and any serialisation attributes on the instruction.
	Output OutputSettings
	// contains filtered or unexported fields
}

SecondaryResult is one document produced by xsl:result-document.

It is kept separate from the principal result rather than merged into it: the whole point of the instruction is that the stylesheet author wants two distinct documents, and folding them together would give a caller expecting several outputs a single plausible-looking wrong one.

func (*SecondaryResult) Serialize

func (sr *SecondaryResult) Serialize(w io.Writer, charMap map[rune]string) error

Serialize writes the secondary document using its own output settings.

A caller holding a SecondaryResult would otherwise have no way to render it: the serialiser is unexported, and re-deriving these settings from the stylesheet is exactly the duplication @format exists to avoid. The charMap argument overrides the table resolved from the document's own @use-character-maps; passing nil uses that table, which is what a caller almost always wants.

func (*SecondaryResult) String

func (sr *SecondaryResult) String() string

String renders the secondary document using its own output settings.

type Stylesheet

type Stylesheet struct {
	// contains filtered or unexported fields
}

Stylesheet is a compiled XSLT 2.0 stylesheet.

Compilation is separated from execution so that a stylesheet compiles once and transforms many documents concurrently. Everything reachable from here is immutable after Compile returns; all per-transform state lives in the runtime context. That is what makes a compiled EN 16931 rule set — tens of megabytes — shareable rather than per-worker.

func Compile

func Compile(doc *xdm.Node, opts CompileOptions) (*Stylesheet, error)

Compile compiles a stylesheet from a parsed XSLT document.

func (*Stylesheet) Output

func (s *Stylesheet) Output() OutputSettings

Output returns the stylesheet's output settings.

func (*Stylesheet) Schema

func (s *Stylesheet) Schema() *xsd.Schema

Schema returns the schema assembled from the stylesheet's xsl:import-schema declarations, or nil when it has none.

It is exposed so that a caller can validate a source document against the same schema the stylesheet declares, rather than having to load it twice and risk the two disagreeing.

func (*Stylesheet) Transform

func (s *Stylesheet) Transform(ctx context.Context, source *xdm.Node, opts TransformOptions) (*Result, error)

Transform applies the stylesheet to a source document.

The Stylesheet is not mutated, so one compiled stylesheet may be used from many goroutines concurrently.

type Template

type Template struct {
	Match    *Pattern
	Name     xdm.QName
	HasName  bool
	Mode     []string // empty means the default (unnamed) mode
	Priority float64
	Params   []*Variable
	Body     []Instruction
	// contains filtered or unexported fields
}

Template is a compiled xsl:template.

type TransformOptions

type TransformOptions struct {
	// Params supplies values for top-level xsl:param, keyed by Clark name
	// ("{uri}local", or just "local" for a no-namespace parameter).
	Params map[string]xdm.Sequence

	// Documents resolves fn:doc and fn:document. Nil disables them, which is
	// the default: a stylesheet that can open arbitrary URIs is an SSRF and
	// file-disclosure vector, and validation rule sets need at most the code
	// lists shipped beside them.
	Documents xpath.DocumentResolver

	// Collections resolves fn:collection. Nil disables it, which is the
	// default, and setting Documents does not set this: enabling fn:doc for
	// a known code list should not also let a stylesheet enumerate whatever
	// a collection URI happens to name.
	Collections xpath.CollectionResolver

	// Texts resolves fn:unparsed-text. Nil disables it, which is the
	// default, and setting Documents does not set this: fn:doc hands back a
	// parsed XML document, while fn:unparsed-text hands back the raw bytes
	// of whatever the resolver will open. See xpath.TextResolver.
	Texts xpath.TextResolver

	// MaxDepth bounds template recursion. Zero means DefaultMaxDepth; a
	// negative value means no limit.
	//
	// The bound catches a stylesheet that recurses without a base case,
	// which is the common authoring mistake. But it also counts the ordinary
	// descent of an identity transform, so a limit below the parser's left
	// this refusing documents it had just accepted: at the old fixed 300, a
	// legal 500-deep document could be parsed and not transformed.
	MaxDepth int

	// InitialMode names the mode for the initial apply-templates.
	InitialMode string

	// InitialTemplate names a template to invoke instead of matching the
	// document root, which is how a stylesheet with only named templates is
	// entered.
	InitialTemplate string
	// InitialTemplateURI is the namespace URI of InitialTemplate, for a
	// caller that has already resolved the prefix in its own namespace
	// context. Empty means resolve any prefix in InitialTemplate against the
	// stylesheet's own declarations instead.
	//
	// The two are not interchangeable. A caller naming the template from
	// outside the stylesheet — a test catalog, a command line with its own
	// bindings — binds the prefix itself, and resolving it a second time
	// against the stylesheet can silently select a DIFFERENT template that
	// happens to spell another namespace with the same prefix.
	InitialTemplateURI string

	// Now fixes the value fn:current-dateTime returns. Leave it zero to use
	// the wall clock; set it to make a transform reproducible, which is what
	// a golden-file test needs.
	Now time.Time

	// ImplicitTimezone is the offset in minutes for date values with no
	// timezone. Defaults to UTC so that results are reproducible across
	// machines.
	ImplicitTimezone int
}

TransformOptions configures one transform.

type Variable

type Variable struct {
	Name xdm.QName
	// Select is the value expression, or nil when the value comes from the
	// element's content.
	Select *xpath.Compiled
	// Body is the sequence constructor used when Select is absent. A variable
	// with content builds a temporary tree, which is why the two forms are
	// not interchangeable: "select" yields whatever the expression yields,
	// while content always yields a document node.
	Body []Instruction
	// Required marks a parameter that must be supplied.
	Required bool
	// IsParam distinguishes an xsl:param from an xsl:variable. The two are
	// compiled to the same structure because they evaluate identically, but
	// only a parameter can be supplied from outside, and section 10.1.1
	// gives the two different error codes when a declared "as" type rejects
	// the value: a parameter left unsupplied whose required type excludes
	// the empty sequence is XTDE0610, while a variable is only ever the
	// plain type error.
	IsParam bool
	// Tunnel marks a tunnel parameter, which passes through templates that do
	// not declare it.
	Tunnel bool
	// contains filtered or unexported fields
}

Variable is a compiled xsl:variable, xsl:param or xsl:with-param.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL