lex

package
v0.1.4 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 4, 2026 License: GPL-3.0 Imports: 4 Imported by: 0

Documentation

Overview

Package lex implements the TDL lexer, turning source text into a stream of tokens for the parser.

Index

Constants

View Source
const (
	IdentPattern  = `[_A-Za-z][_A-Za-z0-9]*`
	IntPattern    = `-?[0-9]+`
	FloatPattern  = `-?[0-9]+\.[0-9]+`
	StringPattern = `"([^"\\\n]|\\["\\nt])*"`
	DocPattern    = `///[^\n]*`
	RegexPattern  = `/([^/\\\n]|\\[^\n])*/`

	// LineCommentPattern is the shape scanComment consumes and discards.
	// It has no Kind, since the lexer never emits one, but it is a fact
	// about the language a second parser has to know: tree-sitter keeps
	// comments as extras where this one drops them.
	//
	// It also matches a doc comment, because `///` begins with `//`, so a
	// consumer tries DocPattern first. Three slashes or more is a doc
	// comment.
	LineCommentPattern = `//[^\n]*`
)

Patterns for the token classes scanned by shape. They are the shapes Lexer.Next and Lexer.RescanRegexAt accept, written as regular expressions because a generator consumes them; TestPatternsMatchTheLexer holds them to that.

Unanchored, and a caller matching at a position anchors them itself.

Variables

This section is empty.

Functions

func IsKeyword

func IsKeyword(text string) bool

IsKeyword reports whether text is a reserved keyword.

func Keywords added in v0.1.1

func Keywords() []string

Keywords returns every reserved keyword, sorted.

func Pattern added in v0.1.1

func Pattern(k Kind) string

Pattern returns a regular expression matching the source text of a token class scanned by shape, and "" for every other kind.

A kind with a fixed spelling reports "" rather than a pattern quoting itself: Spelling is the answer for those, and returning one from both would give a caller two ways to ask.

func Punctuation added in v0.1.1

func Punctuation() []string

Punctuation returns the spelling of every operator and delimiter, sorted.

func Spelling added in v0.1.1

func Spelling(k Kind) string

Spelling returns the source text of a kind that has exactly one, and "" for a class scanned by shape.

Types

type Kind

type Kind int

Kind identifies the lexical class of a Token.

const (
	ILLEGAL Kind = iota
	EOF

	IDENT  // identifiers and keywords share a scanning path; Kind distinguishes them
	DOC    // /// doc comment, text only
	STRING // "..."
	INT    // 123
	FLOAT  // 1.23
	REGEX  // /.../, scanned only on demand; see [Lexer.RescanRegexAt]

	// Reserved keywords. Declaration keywords are reserved; modifiers and
	// constraint names (key, owned, deprecated, min, length, ...) are
	// contextual and lex as IDENT.
	PACKAGE
	IMPORT
	AS
	PRIMITIVE
	UNIT
	ALIAS
	TYPE
	VALUE
	ENTITY
	ENUM
	CLASS
	MIXIN
	INSTANCE
	TARGET
	FOR
	REQUIRES
	WHERE
	INCLUDE
	NULL
	TRUE
	FALSE

	LBRACE   // {
	RBRACE   // }
	LPAREN   // (
	RPAREN   // )
	LBRACK   // [
	RBRACK   // ]
	LT       // <
	GT       // >
	COLON    // :
	COMMA    // ,
	EQUAL    // =
	QUESTION // ?
	DOT      // .
	PIPE     // |
	ARROW    // ->
	RANGE    // ..
	CARET    // ^
	STAR     // *
	SLASH    // /
	FATARROW // =>

)

func Lookup added in v0.1.1

func Lookup(text string) (Kind, bool)

Lookup returns the kind the lexer produces for a fixed spelling, whether keyword, operator, or delimiter. It reports false for anything scanned by shape, which has no single spelling to look up.

func LookupIdent

func LookupIdent(ident string) Kind

LookupIdent returns the keyword Kind for ident, or IDENT if ident is not a reserved keyword.

func (Kind) String

func (k Kind) String() string

type Lexer

type Lexer struct {
	// contains filtered or unexported fields
}

Lexer scans TDL source text into a stream of [Token]s.

Whitespace is insignificant in TDL: the lexer emits no newline tokens and the parser has no separator rules. An item ends where the next begins.

func New

func New(filename, src string) *Lexer

New returns a Lexer over src, reporting positions against filename.

func (*Lexer) Next

func (l *Lexer) Next() Token

Next scans and returns the next token. It returns an EOF token forever once the end of input is reached.

func (*Lexer) RescanRegexAt

func (l *Lexer) RescanRegexAt(pos Position) Token

RescanRegexAt rescans the input from pos as a regex literal and leaves the lexer positioned after it.

`/` is division in a unit expression and the delimiter of a regex literal, and nothing local to the token tells them apart: `matches` is a contextual keyword, so the preceding token is an ordinary identifier either way. The parser knows which it wants, so it asks. Every other token is scanned by Lexer.Next without context.

type Position

type Position struct {
	Filename string
	Line     int // 1-based
	Col      int // 1-based, in bytes
	Offset   int // 0-based byte offset
}

Position identifies a location in a source file.

func (Position) String

func (p Position) String() string

type Token

type Token struct {
	Kind Kind
	Text string // literal source text; decoded for STRING, body only for DOC and REGEX
	Pos  Position
}

Token is a single lexical token.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL