syntax

package
v0.14.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 29, 2026 License: MIT Imports: 16 Imported by: 0

Documentation

Overview

Package syntax classifies source text one line at a time so a renderer can colour it.

The design constraint that shapes everything here is that an editor re-lexes on every keystroke. Lexing a whole buffer to redraw one line would make typing in a large file feel heavy, so Lex takes a single line plus the State left by the line above it and returns the State it leaves behind. A caller caches one State per line, and after an edit re-lexes downward only until the outgoing State matches what it had cached - at which point nothing below can have changed, and it can stop.

That convergence check is why State is a comparable value type rather than an interface or a pointer: == has to mean "the rest of the file is unaffected".

Every language is described the same way, by a .syntax file - nem's own embedded from languages/, and the user's read from beside init.lua - and compiled into a Language, which is the Lexer. tree-sitter is deliberately not used: it is a C library, and cgo would cost nem the single static binary Go was chosen for. A definition colours text rather than compiling it, so it is approximate at the edges, and the tests pin the cases that actually break highlighters rather than modelling each grammar completely.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func BuiltinError added in v0.13.0

func BuiltinError() error

BuiltinError reports a problem with nem's own definitions, or nil.

func BuiltinSource added in v0.13.0

func BuiltinSource(name string) (string, bool)

BuiltinSource returns the text of nem's own definition of a language, as it is written: what a user starts from to change it.

func ValidName added in v0.13.0

func ValidName(name string) bool

ValidName reports whether name may name a language: lower-case letters, digits and + # . _ -.

Types

type Class

type Class uint8

Class is what a span of text is, for colouring purposes.

The set is small and closed on purpose. A theme maps one style per class, so every class added is a colour a user has to choose and a decision a theme author has to make; a sprawling set makes a coherent theme impossible.

const (
	// Plain is ordinary text. A region covered by no span is Plain, so a lexer
	// need not emit spans for the gaps between what it classifies.
	Plain Class = iota
	Keyword
	String
	Comment
	Number
	// Function is an identifier in a call or definition position.
	Function
	// Type is a type name where the grammar makes it unambiguous.
	Type
	// Constant is a literal like true, false, nil or iota.
	Constant
	Operator
	Punctuation
)

func (Class) String

func (c Class) String() string

String names the class, for test failures and debugging.

type Def added in v0.13.0

type Def struct {
	Name string
	// Like names the language this one starts from, or "".
	Like string
	// Source is the file it was read from, for errors.
	Source string

	Files    []string // globs, lower-cased
	Shebangs []string
	Headers  []*regexp.Regexp

	IgnoreCase bool
	WordChars  string
	Words      []WordDef
	Declares   []WordDef
	Calls      string // "", "paren", "lisp" or "none"

	// Numbers is "" for the default, "none", or a pattern.
	Numbers string

	Operators   *string // nil for the default
	Punctuation *string

	Regions []*RegionDef
	Matches []*MatchDef
	// contains filtered or unexported fields
}

Def is a language as a .syntax file describes it, before it is compiled. It is kept, rather than only the compiled Language, because another definition may start from it with like.

func ParseDef added in v0.13.0

func ParseDef(source string, r io.Reader) (*Def, error)

ParseDef reads one .syntax file. source names it in errors.

type Language added in v0.13.0

type Language struct {
	// contains filtered or unexported fields
}

Language is one compiled definition, and is itself a Lexer.

A line is read left to right, one token at a time, and at each place the first of these that applies wins: a match rule, in the order written; the start of a comment, string or region, the longest if several start there; a number; a name, classified; an operator or punctuation rune. Reading tokens, rather than letting later rules paint over earlier ones as nano's files do, is what makes a // inside a string a string and a quote inside a comment a comment, with no ordering of rules to get right.

func Compile added in v0.13.0

func Compile(d *Def) (*Language, error)

Compile makes a Language of a definition, one already merged with what it is like.

func (*Language) Comment added in v0.13.0

func (l *Language) Comment() (start, end string, ok bool)

Comment reports how the language writes a comment, for M-;: the first comment its definition lists, and that comment's end if it is a block comment. A language lists the one it writes first - Lisp's ;; before its ; - so a definition says which, with nothing else to learn. ok is false for a language with no comments at all.

func (*Language) Def added in v0.13.0

func (l *Language) Def() *Def

Def returns the definition the language was compiled from.

func (*Language) Lex added in v0.13.0

func (l *Language) Lex(line []rune, in State) ([]Span, State)

Lex classifies one line.

func (*Language) Name added in v0.13.0

func (l *Language) Name() string

Name reports the language.

func (*Language) Words added in v0.13.0

func (l *Language) Words(c Class) []string

Words lists a language's words of one class, sorted, for tests and for describing a language.

type Lexer

type Lexer interface {
	// Lex classifies one line. Spans are ascending, non-overlapping, non-empty
	// and within the line. Regions covered by no span are Plain.
	//
	// line is lent, not given: it is the buffer's own storage, and a lexer
	// must neither change it nor keep it past the call. Copying it for every
	// line lexed cost a copy of a 200KB line on every keystroke typed in it.
	Lex(line []rune, in State) (spans []Span, out State)
	// Name identifies the language, for the modeline and for tests.
	Name() string
}

Lexer classifies one line at a time.

Implementations must be pure: the same line and incoming State must always produce the same spans and outgoing State, with nothing carried in the receiver. A caller re-lexing from the middle of a file relies on that, and a lexer holding hidden state between calls would give different colours depending on how the user happened to scroll.

func For

func For(p string) Lexer

For returns nem's lexer for a path, by file name. It never returns nil: a name nothing recognises gets the plain lexer.

func ForWithHeader

func ForWithHeader(p, firstLine string) Lexer

ForWithHeader picks nem's lexer knowing the file's first line, which is what identifies a script with no extension by its #!.

type MatchDef added in v0.13.0

type MatchDef struct {
	Class Class
	Re    *regexp.Regexp
	// Named reports whether the groups name the classes.
	Named bool
}

MatchDef is a pattern and what it colours: the whole match Class, or each of its named groups by its name.

type PlainLexer

type PlainLexer struct{}

PlainLexer classifies nothing, for text nem has no grammar for.

func (PlainLexer) Lex

func (PlainLexer) Lex(line []rune, _ State) ([]Span, State)

Lex returns the whole line as one Plain span, and never carries state.

func (PlainLexer) Name

func (PlainLexer) Name() string

type RegionDef added in v0.13.0

type RegionDef struct {
	Class Class
	// Start is a literal delimiter, or StartRe a pattern.
	Start   string
	StartRe *regexp.Regexp
	// End is a literal delimiter, or EndRe a pattern, or EndTmpl a pattern
	// that refers to what StartRe captured and is compiled once it has.
	// ToEOL ends the region with the line: a line comment.
	End     string
	EndRe   *regexp.Regexp
	EndTmpl string
	ToEOL   bool

	// Escape is the rune that escapes the next one, or 0; EscapeRe a pattern
	// that escapes what it matches. Doubled means the end written twice is
	// the end escaped: "" inside "…".
	Escape   rune
	EscapeRe *regexp.Regexp
	Doubled  bool

	Nested    bool
	Multiline bool
}

RegionDef is a comment, a string or a region: text from a start to an end.

type Set added in v0.13.0

type Set struct {
	// contains filtered or unexported fields
}

Set is the languages nem knows: its own, and the user's beside them or in their place.

func Builtin added in v0.13.0

func Builtin() *Set

Builtin returns nem's own languages.

func Load added in v0.13.0

func Load(dir string) (*Set, []error)

Load returns nem's languages with the user's from dir: each *.syntax file there is a language, matched before nem's, and one named as one of nem's replaces it. A missing dir is no error; a bad file is reported and the rest still load.

func (*Set) Detect added in v0.13.0

func (s *Set) Detect(p, firstLine string) *Language

Detect returns the language of a file, or nil: by its name, a whole name such as Makefile before any pattern, and then by its first line - the #! of a script with no extension, or a header such as <?xml.

func (*Set) For added in v0.13.0

func (s *Set) For(p, firstLine string) Lexer

For returns the lexer for a file: its language, or the plain lexer, so a caller never has to check before lexing.

func (*Set) Language added in v0.13.0

func (s *Set) Language(name string) *Language

Language returns a language by name, or nil.

func (*Set) Languages added in v0.13.0

func (s *Set) Languages() []*Language

Languages lists the set's languages, by name.

type Span

type Span struct {
	Start, End int
	Class      Class
}

Span is a half-open run of one class within a line.

Start and End are RUNE indices, not byte offsets. The renderer converts them to display columns, and a byte offset would land in the wrong column the moment a line contains anything outside ASCII - which is exactly where a highlighter's mistakes are least forgivable, since the text still looks fine and only the colours are wrong.

type State

type State uint32

State is what one line leaves open for the next: a block comment, a raw string, a fenced code block.

It is opaque, and its encoding belongs to the lexer that made it - a Language's is documented beside regionState - but the zero value means the same everywhere: nothing is open. It is therefore "start of file", which is what a caller passes for line 0 without asking the lexer for an initial value.

type WordDef added in v0.13.0

type WordDef struct {
	Word  string
	Class Class
}

WordDef is a word and the class it is given.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL