regexpsyntax

package
v0.1.0-alpha.6 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 8, 2026 License: Apache-2.0 Imports: 6 Imported by: 0

Documentation

Overview

Package regexpsyntax parses ECMAScript regular expression patterns. It is shared by the regexp engine (package engine), which compiles the tree, and the syntax package, which reports invalid regexp literals as early errors. It depends on the standard library only.

Index

Constants

View Source
const MaxCodePoint = 0x10FFFF

MaxCodePoint is the largest Unicode code point.

View Source
const UnicodeVersion = unicodeVersion

UnicodeVersion is the UCD version of the generated tables.

Variables

This section is empty.

Functions

func IsIDContinue

func IsIDContinue(r rune) bool

func IsIDStart

func IsIDStart(r rune) bool

IsIDStart and IsIDContinue report the ID_Start and ID_Continue properties at unicodeVersion. The lexer uses them for non-ASCII identifier characters, so identifiers follow the same Unicode version as property escapes.

Types

type AST

type AST struct {
	Root     *Node
	Names    []string // capture names by group number (index 0 unused); "" for unnamed
	HasNames bool
	DupNames bool // a name is shared by groups in different alternatives
}

AST is a parsed pattern.

func Parse

func Parse(pattern []uint16, flags Flags) (*AST, error)

Parse parses a pattern (UTF-16 code units). Errors are *Error.

type CaseMap

type CaseMap struct {
	// contains filtered or unexported fields
}

CaseMap is a decoded code point mapping (identity outside keys).

func Canonicalizer

func Canonicalizer(unicodeMode bool) *CaseMap

Canonicalizer returns the Canonicalize map for the flags (u or v select simple case folding).

func (*CaseMap) Canonicalize

func (m *CaseMap) Canonicalize(c rune) rune

Canonicalize is Canonicalize(rer, ch) for an ignoreCase pattern.

type Checker

type Checker struct {
	// contains filtered or unexported fields
}

A Checker validates patterns without building their tree, keeping its buffers from one pattern to the next. The zero value is ready to use.

func (*Checker) Check

func (c *Checker) Check(pattern []uint16, flags Flags) error

Check returns the error Parse would return for pattern, or nil. It skips only the work that cannot fail: the tree and the character sets of plain atoms and non-v classes.

type Error

type Error struct{ Msg string }

Error reports a malformed pattern; Msg is the V8 wording without the "Invalid regular expression: /.../: " prefix.

func (*Error) Error

func (e *Error) Error() string

type Flags

type Flags struct {
	IgnoreCase, Multiline, DotAll, Unicode, UnicodeSets bool
}

Flags are the flags that affect parsing.

type Node

type Node struct {
	Op        Op
	Greedy    bool // OpRepeat
	Negate    bool // OpLook
	Behind    bool // OpLook
	Multiline bool // OpBegin, OpEnd: the m flag
	// Icase is the i flag in effect for OpBackref (Canonicalize compare)
	// and OpWordB/OpNotWordB (u-mode word characters include U+017F and
	// U+212A).
	Icase    bool
	Index    int // OpCapture: group number
	Min, Max int // OpRepeat
	// CapLo and CapHi delimit the groups [CapLo, CapHi) inside a repeated or
	// lookaround body.
	CapLo, CapHi int
	Set          Set    // OpChar
	Refs         []int  // OpBackref: group numbers
	Name         string // OpBackref by name, until resolved
	Subs         []*Node
}

Node is one node of a parsed pattern.

func (*Node) CanBeEmpty

func (n *Node) CanBeEmpty() bool

CanBeEmpty reports whether n can match the empty string.

func (*Node) Walk

func (n *Node) Walk(fn func(*Node))

Walk calls fn on n and its descendants in pre-order.

type Op

type Op uint8

Op is the kind of a Node.

const (
	OpEmpty    Op = iota
	OpChar        // one character in set
	OpSeq         // subs in order
	OpAlt         // subs in priority order
	OpCapture     // group index around subs[0]
	OpRepeat      // subs[0] repeated min..max times (max -1: unbounded)
	OpLook        // lookahead or lookbehind (behind, negate) of subs[0]
	OpBackref     // backreference to the groups in refs
	OpBegin       // ^
	OpEnd         // $
	OpWordB       // \b
	OpNotWordB    // \B
)

type Set

type Set []rune

Set is a set of characters (code units or code points) as sorted, non-overlapping, non-adjacent inclusive ranges [lo0, hi0, lo1, hi1, ...]. Every constructor below returns a normalized set; the set operations of the regexp parser (classes, v-flag difference and intersection, case closure) work on this form.

func DecodeRanges

func DecodeRanges(s string) Set

DecodeRanges decodes a table in the generator's range encoding, which engine/unicode_norm_tables.go shares.

func (Set) Complement

func (s Set) Complement() Set

Complement returns the characters in [0, MaxCodePoint] not in s.

func (Set) ComplementIn

func (s Set) ComplementIn(top rune) Set

ComplementIn returns the characters in [0, top] not in s.

func (Set) Has

func (s Set) Has(c rune) bool

Has reports whether c is in s.

func (Set) Intersect

func (s Set) Intersect(t Set) Set

Intersect returns s & t.

func (Set) Normalize

func (s Set) Normalize() Set

Normalize sorts the ranges and merges overlapping or adjacent ones.

func (Set) Single

func (s Set) Single() (rune, bool)

Single reports whether s is exactly one character.

func (Set) Subtract

func (s Set) Subtract(t Set) Set

Subtract returns s minus t.

func (Set) Union

func (s Set) Union(t Set) Set

Union returns s | t.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL