Documentation
¶
Overview ¶
Package syntax classifies source text one line at a time so a renderer can colour it.
The design constraint that shapes everything here is that an editor re-lexes on every keystroke. Lexing a whole buffer to redraw one line would make typing in a large file feel heavy, so Lex takes a single line plus the State left by the line above it and returns the State it leaves behind. A caller caches one State per line, and after an edit re-lexes downward only until the outgoing State matches what it had cached - at which point nothing below can have changed, and it can stop.
That convergence check is why State is a comparable value type rather than an interface or a pointer: == has to mean "the rest of the file is unaffected".
Every language is described the same way, by a .syntax file - nem's own embedded from languages/, and the user's read from beside init.lua - and compiled into a Language, which is the Lexer. tree-sitter is deliberately not used: it is a C library, and cgo would cost nem the single static binary Go was chosen for. A definition colours text rather than compiling it, so it is approximate at the edges, and the tests pin the cases that actually break highlighters rather than modelling each grammar completely.
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func BuiltinError ¶ added in v0.13.0
func BuiltinError() error
BuiltinError reports a problem with nem's own definitions, or nil.
func BuiltinSource ¶ added in v0.13.0
BuiltinSource returns the text of nem's own definition of a language, as it is written: what a user starts from to change it.
Types ¶
type Class ¶
type Class uint8
Class is what a span of text is, for colouring purposes.
The set is small and closed on purpose. A theme maps one style per class, so every class added is a colour a user has to choose and a decision a theme author has to make; a sprawling set makes a coherent theme impossible.
const ( // Plain is ordinary text. A region covered by no span is Plain, so a lexer // need not emit spans for the gaps between what it classifies. Plain Class = iota Keyword String Comment Number // Function is an identifier in a call or definition position. Function // Type is a type name where the grammar makes it unambiguous. Type // Constant is a literal like true, false, nil or iota. Constant Operator Punctuation )
type Def ¶ added in v0.13.0
type Def struct {
Name string
// Like names the language this one starts from, or "".
Like string
// Source is the file it was read from, for errors.
Source string
Files []string // globs, lower-cased
Shebangs []string
Headers []*regexp.Regexp
IgnoreCase bool
WordChars string
Words []WordDef
Declares []WordDef
Calls string // "", "paren", "lisp" or "none"
// Numbers is "" for the default, "none", or a pattern.
Numbers string
Operators *string // nil for the default
Punctuation *string
Regions []*RegionDef
Matches []*MatchDef
// contains filtered or unexported fields
}
Def is a language as a .syntax file describes it, before it is compiled. It is kept, rather than only the compiled Language, because another definition may start from it with like.
type Language ¶ added in v0.13.0
type Language struct {
// contains filtered or unexported fields
}
Language is one compiled definition, and is itself a Lexer.
A line is read left to right, one token at a time, and at each place the first of these that applies wins: a match rule, in the order written; the start of a comment, string or region, the longest if several start there; a number; a name, classified; an operator or punctuation rune. Reading tokens, rather than letting later rules paint over earlier ones as nano's files do, is what makes a // inside a string a string and a quote inside a comment a comment, with no ordering of rules to get right.
func Compile ¶ added in v0.13.0
Compile makes a Language of a definition, one already merged with what it is like.
func (*Language) Comment ¶ added in v0.13.0
Comment reports how the language writes a comment, for M-;: the first comment its definition lists, and that comment's end if it is a block comment. A language lists the one it writes first - Lisp's ;; before its ; - so a definition says which, with nothing else to learn. ok is false for a language with no comments at all.
type Lexer ¶
type Lexer interface {
// Lex classifies one line. Spans are ascending, non-overlapping, non-empty
// and within the line. Regions covered by no span are Plain.
//
// line is lent, not given: it is the buffer's own storage, and a lexer
// must neither change it nor keep it past the call. Copying it for every
// line lexed cost a copy of a 200KB line on every keystroke typed in it.
Lex(line []rune, in State) (spans []Span, out State)
// Name identifies the language, for the modeline and for tests.
Name() string
}
Lexer classifies one line at a time.
Implementations must be pure: the same line and incoming State must always produce the same spans and outgoing State, with nothing carried in the receiver. A caller re-lexing from the middle of a file relies on that, and a lexer holding hidden state between calls would give different colours depending on how the user happened to scroll.
func For ¶
For returns nem's lexer for a path, by file name. It never returns nil: a name nothing recognises gets the plain lexer.
func ForWithHeader ¶
ForWithHeader picks nem's lexer knowing the file's first line, which is what identifies a script with no extension by its #!.
type MatchDef ¶ added in v0.13.0
type MatchDef struct {
Class Class
Re *regexp.Regexp
// Named reports whether the groups name the classes.
Named bool
}
MatchDef is a pattern and what it colours: the whole match Class, or each of its named groups by its name.
type PlainLexer ¶
type PlainLexer struct{}
PlainLexer classifies nothing, for text nem has no grammar for.
func (PlainLexer) Lex ¶
func (PlainLexer) Lex(line []rune, _ State) ([]Span, State)
Lex returns the whole line as one Plain span, and never carries state.
func (PlainLexer) Name ¶
func (PlainLexer) Name() string
type RegionDef ¶ added in v0.13.0
type RegionDef struct {
Class Class
// Start is a literal delimiter, or StartRe a pattern.
Start string
StartRe *regexp.Regexp
// End is a literal delimiter, or EndRe a pattern, or EndTmpl a pattern
// that refers to what StartRe captured and is compiled once it has.
// ToEOL ends the region with the line: a line comment.
End string
EndRe *regexp.Regexp
EndTmpl string
ToEOL bool
// Escape is the rune that escapes the next one, or 0; EscapeRe a pattern
// that escapes what it matches. Doubled means the end written twice is
// the end escaped: "" inside "…".
Escape rune
EscapeRe *regexp.Regexp
Doubled bool
Nested bool
Multiline bool
}
RegionDef is a comment, a string or a region: text from a start to an end.
type Set ¶ added in v0.13.0
type Set struct {
// contains filtered or unexported fields
}
Set is the languages nem knows: its own, and the user's beside them or in their place.
func Load ¶ added in v0.13.0
Load returns nem's languages with the user's from dir: each *.syntax file there is a language, matched before nem's, and one named as one of nem's replaces it. A missing dir is no error; a bad file is reported and the rest still load.
func (*Set) Detect ¶ added in v0.13.0
Detect returns the language of a file, or nil: by its name, a whole name such as Makefile before any pattern, and then by its first line - the #! of a script with no extension, or a header such as <?xml.
func (*Set) For ¶ added in v0.13.0
For returns the lexer for a file: its language, or the plain lexer, so a caller never has to check before lexing.
type Span ¶
Span is a half-open run of one class within a line.
Start and End are RUNE indices, not byte offsets. The renderer converts them to display columns, and a byte offset would land in the wrong column the moment a line contains anything outside ASCII - which is exactly where a highlighter's mistakes are least forgivable, since the text still looks fine and only the colours are wrong.
type State ¶
type State uint32
State is what one line leaves open for the next: a block comment, a raw string, a fenced code block.
It is opaque, and its encoding belongs to the lexer that made it - a Language's is documented beside regionState - but the zero value means the same everywhere: nothing is open. It is therefore "start of file", which is what a caller passes for line 0 without asking the lexer for an initial value.