Documentation
¶
Overview ¶
Package lex implements the TDL lexer, turning source text into a stream of tokens for the parser.
Index ¶
Constants ¶
const ( IdentPattern = `[_A-Za-z][_A-Za-z0-9]*` IntPattern = `-?[0-9]+` FloatPattern = `-?[0-9]+\.[0-9]+` StringPattern = `"([^"\\\n]|\\["\\nt])*"` DocPattern = `///[^\n]*` RegexPattern = `/([^/\\\n]|\\[^\n])*/` // LineCommentPattern is the shape scanComment consumes and discards. // It has no Kind, since the lexer never emits one, but it is a fact // about the language a second parser has to know: tree-sitter keeps // comments as extras where this one drops them. // // It also matches a doc comment, because `///` begins with `//`, so a // consumer tries DocPattern first. Three slashes or more is a doc // comment. LineCommentPattern = `//[^\n]*` )
Patterns for the token classes scanned by shape. They are the shapes Lexer.Next and Lexer.RescanRegexAt accept, written as regular expressions because a generator consumes them; TestPatternsMatchTheLexer holds them to that, and internal/ebnf binds each to the grammar name a `/*@ token ... */` annotation gives it.
Unanchored, and a caller matching at a position anchors them itself.
Variables ¶
This section is empty.
Functions ¶
func Keywords ¶ added in v0.1.1
func Keywords() []string
Keywords returns every reserved keyword, sorted.
func Punctuation ¶ added in v0.1.1
func Punctuation() []string
Punctuation returns the spelling of every operator and delimiter, sorted.
Types ¶
type Comment ¶ added in v0.1.7
type Comment struct {
Text string // the text after the slashes, with one leading space removed
Pos Position
End int // offset just past the comment's last character
}
Comment is an ordinary `//` comment, which Lexer.Next skips.
A doc comment is a DOC token instead, because it belongs to the declaration it precedes and the parser attaches it there. An ordinary comment belongs to nobody, so it is collected here and the formatter places it by position. Nothing between the two stages has to know it exists.
type Kind ¶
type Kind int
Kind identifies the lexical class of a Token.
const ( ILLEGAL Kind = iota EOF IDENT // identifiers and keywords share a scanning path; Kind distinguishes them DOC // /// doc comment, text only STRING // "..." INT // 123 FLOAT // 1.23 REGEX // /.../, scanned only on demand; see [Lexer.RescanRegexAt] // Reserved keywords. Declaration keywords are reserved; modifiers and // constraint names (key, owned, deprecated, min, length, ...) are // contextual and lex as IDENT. PACKAGE IMPORT AS PRIMITIVE UNIT ALIAS TYPE VALUE ENTITY ENUM CLASS MIXIN INSTANCE TARGET FOR REQUIRES WHERE INCLUDE NULL TRUE FALSE LBRACE // { RBRACE // } LPAREN // ( RPAREN // ) LBRACK // [ RBRACK // ] LT // < GT // > COLON // : COMMA // , EQUAL // = QUESTION // ? DOT // . PIPE // | ARROW // -> RANGE // .. CARET // ^ STAR // * SLASH // / FATARROW // => )
func Lookup ¶ added in v0.1.1
Lookup returns the kind the lexer produces for a fixed spelling, whether keyword, operator, or delimiter. It reports false for anything scanned by shape, which has no single spelling to look up.
func LookupIdent ¶
LookupIdent returns the keyword Kind for ident, or IDENT if ident is not a reserved keyword.
type Lexer ¶
type Lexer struct {
// contains filtered or unexported fields
}
Lexer scans TDL source text into a stream of [Token]s.
Whitespace is insignificant in TDL: the lexer emits no newline tokens and the parser has no separator rules. An item ends where the next begins.
func (*Lexer) Comments ¶ added in v0.1.7
Comments returns every ordinary comment scanned so far, in source order.
It is a copy, so a caller that edits what it gets does not edit what the next caller gets.
func (*Lexer) Next ¶
Next scans and returns the next token. It returns an EOF token forever once the end of input is reached.
func (*Lexer) RescanRegexAt ¶
RescanRegexAt rescans the input from pos as a regex literal and leaves the lexer positioned after it.
`/` is division in a unit expression and the delimiter of a regex literal, and nothing local to the token tells them apart: `matches` is a contextual keyword, so the preceding token is an ordinary identifier either way. The parser knows which it wants, so it asks. Every other token is scanned by Lexer.Next without context.