Documentation
¶
Overview ¶
Package dtd validates an XML document against the DTD in its internal subset.
It lives outside xdm because it needs the content-model automaton in xsd, and xdm is what xsd is built on — putting it there would invert the dependency. The split also keeps the parser's job clear: xdm reads a document and applies the two declarations whose absence is visible in the data model (attribute defaults and internal entities), while deciding whether the document *satisfies* its DTD is validation and belongs here.
Scope ¶
A DTD is a smaller language than XSD, and almost all of it maps onto machinery that already exists:
- <!ELEMENT> content models are a strict subset of what xsd's Glushkov automaton compiles — DTD has sequence, choice, and the ?, * and + quantifiers, and no numeric occurrence bounds at all.
- <!ATTLIST> required/implied/fixed maps onto attribute use.
- ID, IDREF and IDREFS are the same document-scoped uniqueness and reference checks XSD defines.
Two rules have no XSD counterpart and are checked here directly, because both point from an attribute at a declaration somewhere else in the DTD rather than at the attribute's own value space (XML 1.0 §3.3.1): every name in a NOTATION attribute's enumeration must be declared by a <!NOTATION>, and an ENTITY or ENTITIES attribute must name entities declared with an NDATA notation. Both are skipped when only half the DTD was read, since neither can distinguish an undeclared name from an unread declaration.
What DTD has that XSD does not is the *external* subset, which is a file reference, along with the parameter entities and conditional sections that only exist there.
Two entry points ¶
Parse reads a DOCTYPE's internal subset and fetches nothing. It is what a caller wants for a document that arrived over the wire and whose DTD is wholly inline.
Load reads both subsets. Fetching the external one is the attack AllowDOCTYPE exists to gate, so it happens only through a caller-supplied LoadOptions.Resolver — nil in the zero value, following xsd.Options.Resolver. With none, a DOCTYPE naming an external subset is REFUSED rather than validated against half a DTD: the external subset routinely holds every element declaration in the language, so reporting a document valid against the internal half alone would turn "I could not read the constraints" into "the constraints hold". LoadOptions.InternalSubsetOnly is how a caller asks for that partial reading deliberately.
Index ¶
Constants ¶
const ( // DefaultMaxExternalDocuments bounds how many external resources one DTD // may pull in when LoadOptions.MaxExternalDocuments is zero. // // DTD modularisation in the wild runs to a handful of .ent and .mod // files per subset; DocBook, the largest in common use, is under fifty. // 64 leaves real DTDs room while bounding the fan-out of one written to // be expensive. It follows xsd.DefaultMaxDocuments in shape — a // resource bound named as an Options field, with zero meaning the // default — and differs in magnitude because a DTD is not a schema // assembly. DefaultMaxExternalDocuments = 64 // DefaultMaxExternalBytes bounds the total bytes read from external // resources when LoadOptions.MaxExternalBytes is zero: 4 MB. // // It is a bound on the whole load rather than on one file, because a // resolver handing back a hundred 512 KB files is the same memory // exhaustion as one handing back a 50 MB file, and only a total sees // both. TEI Lite's subset is 79 KB and DocBook's is around 400 KB, so // 4 MB carries real DTDs an order of magnitude over. DefaultMaxExternalBytes = 4 << 20 // DefaultMaxEntityBytes bounds the total EXPANDED size of parameter // entities when LoadOptions.MaxEntityBytes is zero: 1 MB. // // This is a different quantity from MaxExternalBytes and is measured for // a different reason. MaxExternalBytes bounds what was read; this bounds // what substitution produced, which is where the exponential of a // billion-laughs lives. Both are charged to the same counter — see // loader.charge — so a bomb split across the two subsets meets whichever // binds first rather than getting a fresh allowance per subset. DefaultMaxEntityBytes = 1 << 20 )
const DefaultMaxErrors = 100
DefaultMaxErrors bounds how many failures are reported.
const DefaultMaxResolverBytes = 4 << 20
DefaultMaxResolverBytes bounds one file a FileResolver reads when MaxBytes is zero: 4 MB, matching DefaultMaxExternalBytes so that neither limit is silently the tighter one.
Variables ¶
var ErrNoResolver = errors.New("no Resolver is configured")
ErrNoResolver reports that a DOCTYPE named an external subset and no Resolver was configured to read it.
It is an ERROR rather than "validate against the internal subset only", and the choice is deliberate. A DTD is a closed description: the external subset routinely holds every <!ELEMENT> in the language, with the internal one holding a handful of overrides. Validating against the internal half alone would report the document valid — or report every element undeclared, which a caller silences with AllowUndeclared — and in neither case has anything been proven about the document. That is the defect class this repository's governing invariant names: a budget or a policy may decline to answer, but must never turn "I could not prove the constraint" into "the constraint holds".
A caller who genuinely wants the internal-subset-only reading can still have it, and has to ask for it: Parse (which fetches nothing, and never did) is unchanged, and LoadOptions.InternalSubsetOnly makes Load say the same thing explicitly. What is refused is getting that reading by accident.
Functions ¶
func Validate ¶
Validate checks a document against the DTD in its own internal subset.
The DTD is read from the document rather than supplied separately, which is what a DOCTYPE means. A document with no DOCTYPE is valid trivially: there are no constraints to violate.
The document must have been parsed with xdm.ParseOptions.AllowDOCTYPE set, since without it the parse fails before this is reachable.
What is checked: element content models, attribute presence (#REQUIRED and #FIXED), enumerated attribute values, ID/IDREF, and the two §3.3.1 rules that tie an attribute to a declaration elsewhere in the DTD — every name in a NOTATION attribute's enumeration must be declared by a <!NOTATION>, and the value of an ENTITY or ENTITIES attribute must name an entity declared with an NDATA notation.
Those last two turn on a name being ABSENT from the DTD, so they are skipped when the DTD is only half of one: a DOCTYPE that named an external subset which was not read (HasExternalSubset with no ExternalSubset, which only LoadOptions.InternalSubsetOnly or Parse produces) may be missing the very declarations they look for, and reporting them would reject a valid document.
An attribute whose declared type or default declaration is outside the sets XML 1.0 §3.3 closes is reported here rather than skipped. Such a declaration constrains the attribute somehow and this package cannot say how, so leaving it silent would report an unexamined attribute as valid — the one thing this package must never do. Compare HasExternalSubset, which says the same about declarations that were never read.
Which declarations reach here is decided by how the DTD was read, not by this function. Parse reads the internal subset alone, and a DTD from it carries HasExternalSubset so a caller knows the check was partial; Load with a Resolver reads both subsets, and everything either one declares is applied on the same terms.
Types ¶
type AttrDefault ¶
type AttrDefault int
AttrDefault is how an attribute's presence is constrained.
const ( // AttrImplied is #IMPLIED: optional, no default. AttrImplied AttrDefault = iota // AttrRequired is #REQUIRED: must be present. AttrRequired // AttrFixed is #FIXED "v": if present the value must be v. AttrFixed // AttrDefaulted is a bare "v": supplied when absent. AttrDefaulted // AttrInvalid is a "#"-prefixed token that is none of the three // keywords — "#REQUIRE" for "#REQUIRED", say. // // XML 1.0 §3.3.2 closes this position to #REQUIRED, #IMPLIED, #FIXED and // a literal; a literal cannot begin with "#" unquoted. So the token // constrains the attribute somehow and this package cannot say how, // which is not the same as a bare default value that happens to look odd. // Validate reports it rather than assuming, because reading "#REQUIRE" as // the default string "#REQUIRE" turns a required attribute into an // optional one and reports the document clean. AttrInvalid )
type Attribute ¶
type Attribute struct {
Element string
Name string
// Type is the declared type: CDATA, ID, IDREF, IDREFS, NMTOKEN,
// NMTOKENS, ENTITY, ENTITIES, NOTATION, or ENUMERATION for a
// parenthesised list.
//
// XML 1.0 §3.3.1 closes the set to those; anything else is kept verbatim
// so Validate can report it rather than dropping it. A type it does not
// recognise is a type it cannot enforce, and saying so is the only way a
// caller can tell a checked attribute from an unchecked one.
Type string
// Enum holds the permitted values of an enumeration or NOTATION type.
Enum []string
Default AttrDefault
Value string
}
Attribute is one attribute definition within an <!ATTLIST>.
type ContentKind ¶
type ContentKind int
ContentKind is what an element's content model permits.
const ( // ContentEmpty is EMPTY: no child elements and no character data. ContentEmpty ContentKind = iota // ContentAny is ANY: anything, unchecked. ContentAny // ContentMixed is (#PCDATA | a | b)*: text interleaved with a set of // element names, in any order and any number. ContentMixed // ContentChildren is an element-only model such as (a, b*, (c|d)?). ContentChildren )
type DTD ¶
type DTD struct {
// Elements maps an element name to its content model.
Elements map[string]*Element
// Attributes maps an element name to its declared attributes.
Attributes map[string][]*Attribute
// Notations is the set of names declared by <!NOTATION>.
//
// It exists only so that Validate can apply XML 1.0 §3.3.1's rule that
// every name in a NOTATION attribute's enumeration be declared. The
// declaration's own system or public identifier is data for the
// application, not for validation, so nothing here keeps it.
Notations map[string]bool
// Unparsed is the set of entity names declared with an NDATA notation.
//
// §3.3.1 makes these, and only these, the permitted values of an ENTITY
// or ENTITIES attribute: an entity with replacement text is parsed and
// naming it there is a validity error just as naming an undeclared one
// is.
Unparsed map[string]bool
// HasExternalSubset records that the DOCTYPE named a SYSTEM or PUBLIC
// identifier.
//
// Parse fetches nothing, so with it this means validation is against the
// internal subset alone and callers are told rather than misled. Load
// with a Resolver does fetch, and then this says only that there was one
// — ExternalSubset holds what was read.
HasExternalSubset bool
// ExternalSubset is the external subset's text as it stood after
// parameter-entity substitution and conditional-section resolution, when
// Load read one. Empty otherwise.
//
// It is the text rather than the declarations because the declarations
// are already merged into Elements and Attributes; this exists so a
// caller can see what the DOCTYPE actually pulled in, which for a modular
// DTD is not something any single file contains.
ExternalSubset string
}
DTD is the subset of a document type declaration this package applies.
func Load ¶ added in v1.3.0
func Load(directive string, opts LoadOptions) (*DTD, error)
Load reads a DOCTYPE, fetches the external subset it names, and returns the declarations of both halves.
The argument is the directive text as encoding/xml hands it over, the same as Parse takes: "DOCTYPE name SYSTEM "..." [...]" including the brackets.
Precedence is XML 1.0 §2.8: the internal subset is read FIRST and its declarations bind. Where both subsets declare the same element, attribute or parameter entity, the internal one wins and the external one is IGNORED — not an error, which is what §2.8 says and what makes the "internal subset as a set of overrides" idiom work.
Everything this can decline to do, it declines loudly. There is no return path on which a subset that could not be read produces a DTD.
func Parse ¶
Parse reads the declarations out of a DOCTYPE's internal subset.
The argument is the directive text as encoding/xml hands it over — the whole "DOCTYPE name [...]" including the brackets. Anything the grammar here does not recognise is skipped rather than guessed at, so an unusual declaration leaves the document less constrained rather than wrongly rejected.
type Element ¶
type Element struct {
Name string
Kind ContentKind
// Mixed is the set of names a mixed model admits. Order and repetition
// are unconstrained there, so a set is the whole model.
Mixed map[string]bool
// Particle is the compiled model for ContentChildren, expressed in xsd's
// component model so that the existing automaton can run it.
Particle *xsd.Particle
}
Element is one <!ELEMENT> declaration.
type Error ¶
type Error struct {
// Path locates the element, as "/root/child".
Path string
// Message says what was wrong.
Message string
}
Error is one validity failure.
type Errors ¶
type Errors struct{ Errors []*Error }
Errors is what Validate returns when a document is not valid.
type FileResolver ¶ added in v1.3.0
type FileResolver struct {
// Root confines every read. A system identifier that resolves outside
// it, whether by "..", by an absolute path, or through a symlink, is
// refused before the file is opened.
//
// An empty Root means the process's working directory, which is almost
// never what a caller wants for an untrusted document — set it.
Root string
// MaxBytes bounds one file this resolver reads. Zero means
// DefaultMaxResolverBytes; a negative value means no limit, for a DTD
// the caller produced itself.
//
// It is a second bound rather than a duplicate of
// LoadOptions.MaxExternalBytes: that one bounds the whole load and is
// applied after the read, while this bounds what one call puts in memory
// at all. The loader's LimitReader already caps the read, so this exists
// for callers who use the resolver directly.
MaxBytes int64
}
FileResolver reads an external subset from the filesystem, confined to a directory.
It is off by default in the only sense that matters: LoadOptions.Resolver is nil unless a caller sets it, so nothing here runs for a document that merely arrived. Supplying one hands control of what this process reads to whoever wrote the DOCTYPE, which is why Root is not optional in practice — see the note on it. This mirrors xsd.FileResolver and xslt.FileResolver rather than reusing either, because dtd sits beneath both and importing one would invert the dependency.
func (*FileResolver) ResolveExternal ¶ added in v1.3.0
func (r *FileResolver) ResolveExternal(systemID, publicID, base string) (io.ReadCloser, string, error)
ResolveExternal implements Resolver.
The returned URI is the file: URI of what was actually read, because that is what a parameter-entity module inside the fetched text resolves against — XML 1.0 §4.4.3.
type LoadOptions ¶ added in v1.3.0
type LoadOptions struct {
// Resolver is how an external subset's system identifier becomes bytes.
//
// Nil — the zero value — means nothing is fetched, and a DOCTYPE naming
// an external subset is refused with ErrNoResolver rather than validated
// against half a DTD. It is off by default because a resolver hands
// control of what this process reads to whoever wrote the DOCTYPE, which
// for an untrusted document is the attacker. Pass a FileResolver rooted
// at the directory the DTD really lives in to say what may be read.
Resolver Resolver
// BaseURI is the URI the DOCTYPE's own system identifier resolves
// against, normally the document's. Empty is permitted; a resolver that
// needs one will say so.
BaseURI string
// InternalSubsetOnly validates against the internal subset alone,
// without fetching anything, and without the ErrNoResolver refusal.
//
// It exists so that the old behaviour is still reachable, but only by
// asking for it. The resulting DTD has HasExternalSubset set, so a
// caller can still see that it is partial, and AllowUndeclared is the
// companion that makes a partial subset usable.
InternalSubsetOnly bool
// MaxExternalDocuments bounds how many external resources one load may
// read — the subset itself and every parameter-entity module it pulls
// in. Zero means DefaultMaxExternalDocuments; a negative value means no
// limit, for a DTD the caller produced itself.
MaxExternalDocuments int
// MaxExternalBytes bounds the total bytes read from external resources.
// Zero means DefaultMaxExternalBytes; a negative value means no limit.
MaxExternalBytes int64
// MaxEntityBytes bounds the total expanded size of parameter entities,
// across both subsets. Zero means DefaultMaxEntityBytes; a negative
// value means no limit.
//
// This is the billion-laughs bound. It is charged on the same counter as
// the bytes read, so a bomb whose halves live in different subsets meets
// one budget rather than one per subset.
MaxEntityBytes int64
}
LoadOptions configures Load.
The zero value fetches nothing, which is the setting for a document that arrived over the wire.
type MapResolver ¶ added in v1.3.0
MapResolver resolves from an in-memory table, for callers that know every resource in advance — a build tool with its DTD modules embedded, or a test.
It reads nothing, so it is the one resolver that is safe to hand an untrusted document without further thought: a system identifier that is not a key is refused rather than searched for.
func (*MapResolver) ResolveExternal ¶ added in v1.3.0
func (r *MapResolver) ResolveExternal(systemID, publicID, base string) (io.ReadCloser, string, error)
ResolveExternal implements Resolver.
Lookup is by the identifier as written and then by its last path segment, so that a module referenced as "ent/iso-lat1.ent" from inside a subset found at "dtd/docbook.dtd" is found under either spelling. There is no filesystem here, so neither form can escape anywhere.
type Options ¶
type Options struct {
// MaxErrors stops after this many failures. Zero means
// DefaultMaxErrors; a negative value means no limit.
//
// A document wrong in every element would otherwise produce an error per
// element, which helps nobody and costs memory proportional to the input.
MaxErrors int
// AllowUndeclared skips elements the DTD says nothing about instead of
// reporting them.
//
// Strictly, an undeclared element is a validity error: a DTD is a closed
// description, unlike a schema where a wildcard may admit the unknown.
// But a document whose DOCTYPE names an *external* subset and declares
// only a few things internally is the common real-world shape — the
// W3C's own RFC 3986 type library declares one element and one attribute
// list, purely so that an external DTD's attributes work — and validating
// that against its internal subset alone reports every other element as
// undeclared, which is noise rather than a finding.
//
// Off by default, so the strict reading is what a caller gets unless they
// ask otherwise. Turn it on when the DTD is known to be partial;
// HasExternalSubset is how to detect that case.
AllowUndeclared bool
}
Options configures Validate.
type Resolver ¶ added in v1.3.0
type Resolver interface {
ResolveExternal(systemID, publicID, base string) (io.ReadCloser, string, error)
}
A Resolver turns the system identifier of an external subset, or of a parameter entity declared inside one, into its text.
It is the caller's, deliberately: this package has no filesystem and no network, so every decision about what may be read — which schemes, which directories, how symlinks resolve — is made in code the caller owns and can audit. A resolver MUST refuse anything it is not certain of; returning an error makes the reference fail, which is the safe outcome.
systemID is the identifier exactly as the DTD wrote it, usually relative. base is the absolute URI of the resource that contains the reference, which for a parameter entity declared in an external subset is that SUBSET's URI and not the document's — XML 1.0 §4.4.3.
It returns the resource's content and the absolute URI it resolved to. That URI becomes the base for anything the fetched text itself references, so a resolver must return the URI it actually read, not the one it was asked for.