Documentation
¶
Overview ¶
Package syntax classifies source text one line at a time so a renderer can colour it.
The design constraint that shapes everything here is that an editor re-lexes on every keystroke. Lexing a whole buffer to redraw one line would make typing in a large file feel heavy, so Lex takes a single line plus the State left by the line above it and returns the State it leaves behind. A caller caches one State per line, and after an edit re-lexes downward only until the outgoing State matches what it had cached - at which point nothing below can have changed, and it can stop.
That convergence check is why State is a comparable value type rather than an interface or a pointer: == has to mean "the rest of the file is unaffected".
tree-sitter is deliberately not used. It is a C library, and cgo would cost nem the single static binary that Go was chosen for. These are hand-written lexers, which are approximate at the edges - they are colouring text, not compiling it - and the tests pin the cases that actually break highlighters rather than trying to model each grammar completely.
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func EmbeddedError ¶
func EmbeddedError() error
EmbeddedError reports a problem with the bundled rule files, or nil.
Exported so a test can assert the binary's own definitions are sound; nothing at runtime can act on it.
func EmbeddedLanguages ¶
func EmbeddedLanguages() []string
EmbeddedLanguages lists the bundled languages, sorted.
Types ¶
type Class ¶
type Class uint8
Class is what a span of text is, for colouring purposes.
The set is small and closed on purpose. A theme maps one style per class, so every class added is a colour a user has to choose and a decision a theme author has to make; a sprawling set makes a coherent theme impossible.
const ( // Plain is ordinary text. A region covered by no span is Plain, so a lexer // need not emit spans for the gaps between what it classifies. Plain Class = iota Keyword String Comment Number // Function is an identifier in a call or definition position. Function // Type is a type name where the grammar makes it unambiguous. Type // Constant is a literal like true, false, nil or iota. Constant Operator Punctuation )
type Lexer ¶
type Lexer interface {
// Lex classifies one line. Spans are ascending, non-overlapping, non-empty
// and within the line. Regions covered by no span are Plain.
//
// line is lent, not given: it is the buffer's own storage, and a lexer
// must neither change it nor keep it past the call. Copying it for every
// line lexed cost a copy of a 200KB line on every keystroke typed in it.
Lex(line []rune, in State) (spans []Span, out State)
// Name identifies the language, for the modeline and for tests.
Name() string
}
Lexer classifies one line at a time.
Implementations must be pure: the same line and incoming State must always produce the same spans and outgoing State, with nothing carried in the receiver. A caller re-lexing from the middle of a file relies on that, and a lexer holding hidden state between calls would give different colours depending on how the user happened to scroll.
func For ¶
For returns the lexer for a path, by file name.
It never returns nil: a name nothing recognises gets the plain lexer, so a caller never has to check before lexing. Use ForWithHeader when the file's first line is available, which is what identifies a script with no extension.
func ForWithHeader ¶
ForWithHeader picks a lexer knowing the file's first line.
The first line is what identifies an extensionless script from its shebang - a file simply called `deploy` starting with #!/bin/sh. Pass "" when it is not known and extension matching still applies.
Precedence is hand-written lexer, then bundled rules, then plain. nano's definitions sit outside this package and are consulted by the caller only when this returns the plain lexer, so a language nem describes itself is never shadowed by nano's version of it.
type PlainLexer ¶
type PlainLexer struct{}
PlainLexer classifies nothing, for text nem has no grammar for.
func (PlainLexer) Lex ¶
func (PlainLexer) Lex(line []rune, _ State) ([]Span, State)
Lex returns the whole line as one Plain span, and never carries state.
func (PlainLexer) Name ¶
func (PlainLexer) Name() string
type Rule ¶
type Rule struct {
// contains filtered or unexported fields
}
Rule is one compiled pattern and the class it paints.
Either re is set, for a rule confined to one line, or startRe and endRe are, for a region that may span lines.
type RuleSet ¶
type RuleSet struct {
// contains filtered or unexported fields
}
RuleSet is one language, and is itself a Lexer.
func EmbeddedRuleSet ¶
EmbeddedRuleSet returns a bundled language by name, for tests.
func ParseRuleFile ¶
ParseRuleFile reads one .nemrc.
Unlike the nanorc reader, a problem here is fatal: these files are ours and ship in the binary, so a bad pattern is a build-time mistake rather than a third-party file to be tolerated. Failing loudly is what makes the "every embedded rule compiles" test meaningful.
func (*RuleSet) Lex ¶
Lex classifies one line against the set's rules.
Rules apply in file order and each paints over whatever came before. Painting into a per-rune array rather than collecting spans is what makes that exact: matches overlap constantly - a keyword inside a string, a number inside a comment - and resolving overlaps afterwards would need the same array anyway. The later rule wins because a .nemrc lists broad rules first and narrow ones after, which is how a string containing the word `if` stays a string.
type Span ¶
Span is a half-open run of one class within a line.
Start and End are RUNE indices, not byte offsets. The renderer converts them to display columns, and a byte offset would land in the wrong column the moment a line contains anything outside ASCII - which is exactly where a highlighter's mistakes are least forgivable, since the text still looks fine and only the colours are wrong.
type State ¶
type State uint32
State is what one line leaves open for the next: a block comment, a raw string, a fenced code block.
It is opaque and its encoding is private to each lexer, but its layout is fixed so that the zero value means the same thing everywhere:
bits 0..7 mode - lexer-specific; 0 always means "nothing is open"
bits 8..23 param - lexer-specific; a Lua long-bracket level, a Markdown
fence length and its delimiter
bits 24..31 unused, always zero
The zero State is therefore "start of file", which is what a caller passes for line 0 without needing to ask the lexer for an initial value.