c14n

package
v1.4.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 26, 2026 License: MIT Imports: 11 Imported by: 0

Documentation

Overview

Package c14n implements Canonical XML 1.0, Exclusive Canonical XML 1.0 and Canonical XML 1.1.

Canonicalization produces a byte-exact serialization of a document or document subset, so that logically equivalent inputs yield identical octets. It is the foundation of XML Signature, and it is also useful for content-addressed storage, cache keys and semantic comparison.

The algorithms differ chiefly in how they treat namespace declarations and xml:* attributes inherited from ancestors that are not themselves part of the output. Exclusive canonicalization renders only the namespace declarations an element visibly uses, which makes a signed subtree relocatable into another document; inclusive canonicalization renders every declaration in scope, which does not. Choose by the specification you are implementing, never by preference.

Input

The input is an xdm tree, so everything the parser decides is decided once, there: line endings, character and entity references, CDATA sections, attribute value normalisation and DTD default attributes all arrive already resolved. The parser's defaults apply unchanged — in particular a DOCTYPE is refused unless the caller parsed with AllowDOCTYPE, and nothing is fetched from a network or a filesystem. C14N 1.1's xml:base fix-up is URI arithmetic only; it resolves nothing.

Two inputs the specifications leave undefined are refused: a tree parsed from an XML 1.1 document (ErrXML11), and a namespace declaration binding a relative URI (ErrRelativeNamespaceURI), which C14N 1.0 and 1.1 section 2.1 require an implementation to report as a failure.

Namespace nodes

The specifications define their input as an XPath 1.0 node-set, which can contain some of an element's namespace nodes and not others. By default an element's namespace declarations participate exactly when the element itself is in the set: that is what a same-document reference, the enveloped-signature transform and every constructor here except one produce. A node set that decides namespace-node membership itself implements NamespaceSet, and then canonicalization follows the specifications' namespace-node rules literally — a partial namespace axis, and the namespace nodes of an element outside the set. FromXPathFilter, the XML-DSig XPath Filter transform, is the constructor that does.

Index

Examples

Constants

View Source
const DefaultMaxDepth = 500

DefaultMaxDepth is MaxDepth's initial value, and what zero or a negative value means.

Variables

View Source
var (
	// ErrUnsupportedAlgorithm is returned for an Algorithm outside the
	// six constants.
	ErrUnsupportedAlgorithm = errors.New("c14n: unsupported algorithm")

	// ErrNoAlgorithm is returned when Options.Algorithm is empty.
	ErrNoAlgorithm = errors.New("c14n: no algorithm specified")

	// ErrUnsupportedNode is returned when the input starts at a node kind
	// that has no canonical form on its own, such as an attribute.
	ErrUnsupportedNode = errors.New("c14n: node kind has no canonical form")

	// ErrDepthExceeded is returned when element nesting exceeds MaxDepth.
	ErrDepthExceeded = errors.New("c14n: maximum depth exceeded")

	// ErrRelativeNamespaceURI is returned when a namespace declaration that
	// can reach the output binds a relative URI. C14N 1.0 and 1.1 section
	// 2.1: implementations "MUST report an operation failure on documents
	// containing relative namespace URIs".
	ErrRelativeNamespaceURI = errors.New("c14n: relative namespace URI")

	// ErrXML11 is returned for a tree parsed from an XML 1.1 document.
	// Canonical XML 1.1 "is applicable to XML 1.0 ... It is not defined for
	// XML 1.1", and 1.0 and Exclusive C14N share its XML 1.0 data model.
	ErrXML11 = errors.New("c14n: XML 1.1 input has no canonical form")
)
View Source
var MaxDepth = DefaultMaxDepth

MaxDepth bounds element nesting below the node a canonicalization starts from; deeper input fails with ErrDepthExceeded. It is a variable rather than a constant so a caller with unusual input can raise it knowingly. It is read at the start of each canonicalization; changing it concurrently with one is a data race.

Zero or negative means DefaultMaxDepth, as xdm.ParseOptions.MaxDepth does: a bound of zero or below would refuse every document, and there is no unlimited setting because the walk recurses. Raise it with a number.

Functions

func Bytes

func Bytes(n *xdm.Node, opts Options) ([]byte, error)

Bytes returns the canonical form of n.

n may be a document node, in which case the whole document is canonicalized, or an element node, in which case the element and its descendants are, with the namespace and xml:* axis treatment the chosen algorithm specifies for a subset apex.

func BytesNodeSet

func BytesNodeSet(ns NodeSet, opts Options) ([]byte, error)

BytesNodeSet mirrors Bytes for a subset.

func Digest

func Digest(h hash.Hash, n *xdm.Node, opts Options) ([]byte, error)

Digest canonicalizes n and writes the result into h, returning h's sum.

This is the form signature verification wants. It streams, so the canonical octets are never materialized.

sum, err := c14n.Digest(sha256.New(), elem, c14n.Options{
    Algorithm: c14n.Exclusive10,
})

func DigestNodeSet

func DigestNodeSet(h hash.Hash, ns NodeSet, opts Options) ([]byte, error)

DigestNodeSet mirrors Digest for a subset.

func Equal

func Equal(a, b *xdm.Node, opts Options) (bool, error)

Equal reports whether two nodes have identical canonical forms under opts. It stops at the first differing octet.

This is the non-security use: semantic equality of two XML documents, independent of insignificant serialization differences.

func FormatPrefixList

func FormatPrefixList(prefixes []string) string

FormatPrefixList is the inverse of ParsePrefixList.

func ParsePrefixList

func ParsePrefixList(s string) []string

ParsePrefixList parses the space-separated PrefixList attribute value of an ec:InclusiveNamespaces element, mapping "#default" to the empty string. Repeated and empty tokens are ignored.

func Write

func Write(w io.Writer, n *xdm.Node, opts Options) error

Write canonicalizes n and streams the result to w.

n must be a document or element node. Output is written incrementally: memory use is bounded by the document's element depth and the attribute count of one element, not by its size. A write error from w is returned unwrapped, so callers can use errors.Is against their own sentinel values.

func WriteNodeSet

func WriteNodeSet(w io.Writer, ns NodeSet, opts Options) error

WriteNodeSet canonicalizes an arbitrary document subset.

See NodeSet for how membership is determined, and the package documentation for the one limitation on namespace node membership.

Types

type Algorithm

type Algorithm string

Algorithm identifies a canonicalization algorithm by its W3C URI.

The URI is the identifier because that is how every consuming specification names the algorithm: a value read out of a ds:CanonicalizationMethod element can be passed through unchanged.

const (
	Inclusive10             Algorithm = "http://www.w3.org/TR/2001/REC-xml-c14n-20010315"
	Inclusive10WithComments Algorithm = "http://www.w3.org/TR/2001/REC-xml-c14n-20010315#WithComments"
	Exclusive10             Algorithm = "http://www.w3.org/2001/10/xml-exc-c14n#"
	Exclusive10WithComments Algorithm = "http://www.w3.org/2001/10/xml-exc-c14n#WithComments"
	Inclusive11             Algorithm = "http://www.w3.org/2006/12/xml-c14n11"
	Inclusive11WithComments Algorithm = "http://www.w3.org/2006/12/xml-c14n11#WithComments"
)

func (Algorithm) Exclusive

func (a Algorithm) Exclusive() bool

Exclusive reports whether a is one of the exclusive variants.

func (Algorithm) Valid

func (a Algorithm) Valid() bool

Valid reports whether a is one of the six supported algorithms.

func (Algorithm) WithComments

func (a Algorithm) WithComments() bool

WithComments reports whether a retains comments.

type NamespaceSet

type NamespaceSet interface {
	NodeSet

	// ContainsNamespace reports whether the namespace node binding prefix
	// on elem is a member. The empty prefix is the default namespace, which
	// in the XPath 1.0 data model has a namespace node only while a
	// non-empty default is in scope.
	ContainsNamespace(elem *xdm.Node, prefix string) bool
}

NamespaceSet is a NodeSet that decides namespace-node membership itself, rather than letting each element's namespace declarations follow the element. Canonicalization then applies C14N 1.0 section 2.3 and Exclusive C14N section 3 to namespace nodes as members in their own right: an element may keep some of its namespace nodes and not others, and an element outside the set still renders the namespace nodes of it that are in the set.

Most callers never need it; FromXPathFilter returns one.

type NodeSet

type NodeSet interface {
	// Root returns a node whose subtree contains every member, normally the
	// document node. Canonicalization walks the subtree in document order
	// from here; namespace and xml:* inheritance still consult ancestors
	// above it.
	Root() *xdm.Node

	// Contains reports whether n is a member.
	Contains(n *xdm.Node) bool
}

NodeSet identifies the nodes to canonicalize.

A NodeSet is a set of element, attribute, text, comment and processing instruction nodes, and the document node. Namespace declarations are not members in their own right: an element's declarations participate exactly when the element does. See the package documentation.

Implementations must be safe for concurrent use and must not change while a canonicalization is in progress.

func Document

func Document(d *xdm.Node) NodeSet

Document returns the node set consisting of the entire document that contains d, including the document node.

func ExcludeSubtree

func ExcludeSubtree(root, exclude *xdm.Node) NodeSet

ExcludeSubtree returns every node in root's subtree except exclude and its descendants.

This is the XML-DSig enveloped-signature transform: the whole document minus the ds:Signature element that contains the reference.

Example

The enveloped-signature transform: the document minus the ds:Signature element that holds the reference, then canonicalized and digested. Digest streams into the hash, so the canonical octets are never held in memory; Bytes is used here only to show them.

package main

import (
	"crypto/sha256"
	"fmt"

	"github.com/knroy/go-xml/c14n"
	"github.com/knroy/go-xml/xdm"
)

func main() {
	tr, err := xdm.ParseString(`<Order xmlns="urn:o" id="42"><Item>tea</Item>`+
		`<ds:Signature xmlns:ds="http://www.w3.org/2000/09/xmldsig#"><ds:SignedInfo/></ds:Signature></Order>`, xdm.ParseOptions{})
	if err != nil {
		panic(err)
	}
	order := tr.Root.Children[0]
	sig := order.Children[1]
	set := c14n.ExcludeSubtree(tr.Root, sig)
	opts := c14n.Options{Algorithm: c14n.Exclusive10}

	out, err := c14n.BytesNodeSet(set, opts)
	if err != nil {
		panic(err)
	}
	fmt.Println(string(out))

	sum, err := c14n.DigestNodeSet(sha256.New(), set, opts)
	if err != nil {
		panic(err)
	}
	fmt.Printf("%x\n", sum)
}
Output:
<Order xmlns="urn:o" id="42"><Item>tea</Item></Order>
f657247c609c986de332beaa737e807ea02422cdd907a50e5ca393ab38a19ed4

func FromXPath

func FromXPath(doc *xdm.Node, expr string, ns map[string]string) (NodeSet, error)

FromXPath evaluates expr against doc and returns the resulting nodes as a node set rooted at the top of doc's tree.

ns supplies prefix bindings for the expression, which is compiled as XPath 2.0 (an XPath 1.0 transform expression is almost always also valid 2.0). Namespace nodes the expression returns are ignored; see the package documentation. ExcludeSubtree is faster and should be preferred where it applies.

func FromXPathFilter

func FromXPathFilter(doc *xdm.Node, filter string, ns map[string]string) (NodeSet, error)

FromXPathFilter is the XML-DSig XPath Filter transform (http://www.w3.org/TR/1999/REC-xpath-19991116): filter is evaluated as a boolean with each node of doc's tree as the context node — element, attribute, text, comment, processing instruction and namespace nodes alike — and the node set holds the nodes for which it is true.

Because namespace nodes are filtered like any other, the result is a NamespaceSet and may keep part of an element's namespace axis. FromXPath, by contrast, takes the nodes an expression returns and lets namespace declarations follow their elements. The filter is compiled as XPath 2.0, with ns supplying its prefix bindings; the XML-DSig here() function is not provided.

func Func

func Func(root *xdm.Node, f func(*xdm.Node) bool) NodeSet

Func returns a node set defined by a predicate over root's subtree.

f is called at most once per node during a canonicalization, in document order. It must be deterministic and free of side effects.

func Subtree

func Subtree(n *xdm.Node) NodeSet

Subtree returns the node set consisting of n and all its descendants, with their attributes. This is the set a same-document reference of the form "#id" denotes.

type Options

type Options struct {
	// Algorithm is required.
	Algorithm Algorithm

	// InclusiveNamespacePrefixes is the PrefixList of an exclusive
	// canonicalization: the prefixes to render even when not visibly
	// utilised. The empty string denotes the default namespace, which
	// appears on the wire as "#default".
	//
	// Ignored by the inclusive algorithms. Passing it with an inclusive
	// algorithm is not an error, because a caller forwarding a parsed
	// transform's parameters should not have to branch.
	InclusiveNamespacePrefixes []string
}

Options configures a canonicalization.

The zero Options is not valid: Algorithm must be set. There is no default algorithm, deliberately. Specifications that use canonicalization always name the algorithm, and a library default invites the single most damaging error in this domain, which is silently canonicalizing with the wrong one and producing a signature that no peer accepts and no local test detects.

Example (PrefixList)

Exclusive canonicalization of a signed subtree, as a WS-Security Reference with an ec:InclusiveNamespaces PrefixList would request it. The Body renders only the prefixes it visibly uses (soap, wsu) plus the one the PrefixList names (ext); the unrelated declaration on the Envelope is dropped, which is what lets the subtree move between envelopes.

package main

import (
	"fmt"

	"github.com/knroy/go-xml/c14n"
	"github.com/knroy/go-xml/xdm"
)

func main() {
	tr, err := xdm.ParseString(`<soap:Envelope xmlns:soap="urn:soap" xmlns:wsu="urn:wsu" xmlns:ext="urn:ext" xmlns:other="urn:other">`+
		`<soap:Body wsu:Id="body"><m:Op xmlns:m="urn:m">hi</m:Op></soap:Body></soap:Envelope>`, xdm.ParseOptions{})
	if err != nil {
		panic(err)
	}
	body := tr.Root.Children[0].Children[0]
	out, err := c14n.Bytes(body, c14n.Options{
		Algorithm:                  c14n.Exclusive10,
		InclusiveNamespacePrefixes: c14n.ParsePrefixList("ext"),
	})
	if err != nil {
		panic(err)
	}
	fmt.Println(string(out))
}
Output:
<soap:Body xmlns:ext="urn:ext" xmlns:soap="urn:soap" xmlns:wsu="urn:wsu" wsu:Id="body"><m:Op xmlns:m="urn:m">hi</m:Op></soap:Body>

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL