lex

package
v0.2.5 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 18, 2026 License: GPL-3.0 Imports: 5 Imported by: 0

Documentation

Overview

Package lex implements the TDL lexer, turning source text into a stream of tokens for the parser.

Index

Constants

View Source
const (
	IdentPattern  = `[_A-Za-z][_A-Za-z0-9]*`
	IntPattern    = `-?[0-9]+`
	FloatPattern  = `-?[0-9]+\.[0-9]+`
	StringPattern = `"([^"\\\n]|\\["\\nt])*"`
	DocPattern    = `///[^\n]*`
	RegexPattern  = `/([^/\\\n]|\\[^\n])*/`

	// LineCommentPattern is the shape scanComment consumes and discards.
	// It has no Kind, since the lexer never emits one, but it is a fact
	// about the language a second parser has to know: tree-sitter keeps
	// comments as extras where this one drops them.
	//
	// It also matches a doc comment, because `///` begins with `//`, so a
	// consumer tries DocPattern first. Three slashes or more is a doc
	// comment.
	LineCommentPattern = `//[^\n]*`
)

Patterns for the token classes scanned by shape. They are the shapes Lexer.Next and Lexer.RescanRegexAt accept, written as regular expressions because a generator consumes them; TestPatternsMatchTheLexer holds them to that, and internal/ebnf binds each to the grammar name a `/*@ token ... */` annotation gives it.

Unanchored, and a caller matching at a position anchors them itself.

Variables

This section is empty.

Functions

func IsKeyword

func IsKeyword(text string) bool

IsKeyword reports whether text is a reserved keyword.

func Keywords added in v0.1.1

func Keywords() []string

Keywords returns every reserved keyword, sorted.

func Punctuation added in v0.1.1

func Punctuation() []string

Punctuation returns the spelling of every operator and delimiter, sorted.

Types

type Comment added in v0.1.7

type Comment struct {
	Text string // the text after the slashes, with one leading space removed
	Pos  Position
	End  int // offset just past the comment's last character
}

Comment is an ordinary `//` comment, which Lexer.Next skips.

A doc comment is a DOC token instead, because it belongs to the declaration it precedes and the parser attaches it there. An ordinary comment belongs to nobody, so it is collected here and the formatter places it by position. Nothing between the two stages has to know it exists.

type Kind

type Kind int

Kind identifies the lexical class of a Token.

const (
	ILLEGAL Kind = iota
	EOF

	IDENT  // identifiers and keywords share a scanning path; Kind distinguishes them
	DOC    // /// doc comment, text only
	STRING // "..."
	INT    // 123
	FLOAT  // 1.23
	REGEX  // /.../, scanned only on demand; see [Lexer.RescanRegexAt]

	// Reserved keywords. Declaration keywords are reserved; modifiers and
	// constraint names (owned, deprecated, min, length, ...) are contextual
	// and lex as IDENT.
	PACKAGE
	IMPORT
	AS
	PRIMITIVE
	UNIT
	ALIAS
	TYPE
	ENUM
	CLASS
	MIXIN
	INSTANCE
	TARGET
	FOR
	REQUIRES
	WHERE
	INCLUDE
	NULL
	TRUE
	FALSE

	LBRACE   // {
	RBRACE   // }
	LPAREN   // (
	RPAREN   // )
	LBRACK   // [
	RBRACK   // ]
	LT       // <
	GT       // >
	COLON    // :
	COMMA    // ,
	EQUAL    // =
	QUESTION // ?
	DOT      // .
	PIPE     // |
	ARROW    // ->
	RANGE    // ..
	CARET    // ^
	STAR     // *
	SLASH    // /
	FATARROW // =>

)

func Lookup added in v0.1.1

func Lookup(text string) (Kind, bool)

Lookup returns the kind the lexer produces for a fixed spelling, whether keyword, operator, or delimiter. It reports false for anything scanned by shape, which has no single spelling to look up.

func LookupIdent

func LookupIdent(ident string) Kind

LookupIdent returns the keyword Kind for ident, or IDENT if ident is not a reserved keyword.

func (Kind) String

func (k Kind) String() string

type Lexer

type Lexer struct {
	// contains filtered or unexported fields
}

Lexer scans TDL source text into a stream of [Token]s.

Whitespace is insignificant in TDL: the lexer emits no newline tokens and the parser has no separator rules. An item ends where the next begins.

func New

func New(filename, src string) *Lexer

New returns a Lexer over src, reporting positions against filename.

func (*Lexer) Comments added in v0.1.7

func (l *Lexer) Comments() []Comment

Comments returns every ordinary comment scanned so far, in source order.

It is a copy, so a caller that edits what it gets does not edit what the next caller gets.

func (*Lexer) Next

func (l *Lexer) Next() Token

Next scans and returns the next token. It returns an EOF token forever once the end of input is reached.

func (*Lexer) RescanRegexAt

func (l *Lexer) RescanRegexAt(pos Position) Token

RescanRegexAt rescans the input from pos as a regex literal and leaves the lexer positioned after it.

`/` is division in a unit expression and the delimiter of a regex literal, and nothing local to the token tells them apart: `matches` is a contextual keyword, so the preceding token is an ordinary identifier either way. The parser knows which it wants, so it asks. Every other token is scanned by Lexer.Next without context.

type Position

type Position struct {
	Filename string
	Line     int // 1-based
	Col      int // 1-based, in bytes
	Offset   int // 0-based byte offset
}

Position identifies a location in a source file.

func (Position) String

func (p Position) String() string

type Token

type Token struct {
	Kind Kind
	Text string // literal source text; decoded for STRING, body only for DOC and REGEX
	Pos  Position
}

Token is a single lexical token.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL