syntax

package
v0.1.5 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 24, 2026 License: MIT Imports: 10 Imported by: 0

Documentation

Overview

Package syntax classifies source text one line at a time so a renderer can colour it.

The design constraint that shapes everything here is that an editor re-lexes on every keystroke. Lexing a whole buffer to redraw one line would make typing in a large file feel heavy, so Lex takes a single line plus the State left by the line above it and returns the State it leaves behind. A caller caches one State per line, and after an edit re-lexes downward only until the outgoing State matches what it had cached - at which point nothing below can have changed, and it can stop.

That convergence check is why State is a comparable value type rather than an interface or a pointer: == has to mean "the rest of the file is unaffected".

tree-sitter is deliberately not used. It is a C library, and cgo would cost nem the single static binary that Go was chosen for. These are hand-written lexers, which are approximate at the edges - they are colouring text, not compiling it - and the tests pin the cases that actually break highlighters rather than trying to model each grammar completely.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func EmbeddedError

func EmbeddedError() error

EmbeddedError reports a problem with the bundled rule files, or nil.

Exported so a test can assert the binary's own definitions are sound; nothing at runtime can act on it.

func EmbeddedLanguages

func EmbeddedLanguages() []string

EmbeddedLanguages lists the bundled languages, sorted.

Types

type Class

type Class uint8

Class is what a span of text is, for colouring purposes.

The set is small and closed on purpose. A theme maps one style per class, so every class added is a colour a user has to choose and a decision a theme author has to make; a sprawling set makes a coherent theme impossible.

const (
	// Plain is ordinary text. A region covered by no span is Plain, so a lexer
	// need not emit spans for the gaps between what it classifies.
	Plain Class = iota
	Keyword
	String
	Comment
	Number
	// Function is an identifier in a call or definition position.
	Function
	// Type is a type name where the grammar makes it unambiguous.
	Type
	// Constant is a literal like true, false, nil or iota.
	Constant
	Operator
	Punctuation
)

func (Class) String

func (c Class) String() string

String names the class, for test failures and debugging.

type Lexer

type Lexer interface {
	// Lex classifies one line. Spans are ascending, non-overlapping, non-empty
	// and within the line. Regions covered by no span are Plain.
	Lex(line []rune, in State) (spans []Span, out State)
	// Name identifies the language, for the modeline and for tests.
	Name() string
}

Lexer classifies one line at a time.

Implementations must be pure: the same line and incoming State must always produce the same spans and outgoing State, with nothing carried in the receiver. A caller re-lexing from the middle of a file relies on that, and a lexer holding hidden state between calls would give different colours depending on how the user happened to scroll.

func For

func For(path string) Lexer

For returns the lexer for a path, by file name.

It never returns nil: a name nothing recognises gets the plain lexer, so a caller never has to check before lexing. Use ForWithHeader when the file's first line is available, which is what identifies a script with no extension.

func ForWithHeader

func ForWithHeader(path, firstLine string) Lexer

ForWithHeader picks a lexer knowing the file's first line.

The first line is what identifies an extensionless script from its shebang - a file simply called `deploy` starting with #!/bin/sh. Pass "" when it is not known and extension matching still applies.

Precedence is hand-written lexer, then bundled rules, then plain. nano's definitions sit outside this package and are consulted by the caller only when this returns the plain lexer, so a language nem describes itself is never shadowed by nano's version of it.

type PlainLexer

type PlainLexer struct{}

PlainLexer classifies nothing, for text nem has no grammar for.

func (PlainLexer) Lex

func (PlainLexer) Lex(line []rune, _ State) ([]Span, State)

Lex returns the whole line as one Plain span, and never carries state.

func (PlainLexer) Name

func (PlainLexer) Name() string

type Rule

type Rule struct {
	// contains filtered or unexported fields
}

Rule is one compiled pattern and the class it paints.

Either re is set, for a rule confined to one line, or startRe and endRe are, for a region that may span lines.

type RuleSet

type RuleSet struct {
	// contains filtered or unexported fields
}

RuleSet is one language, and is itself a Lexer.

func EmbeddedRuleSet

func EmbeddedRuleSet(name string) *RuleSet

EmbeddedRuleSet returns a bundled language by name, for tests.

func ParseRuleFile

func ParseRuleFile(name string, r io.Reader) (*RuleSet, error)

ParseRuleFile reads one .nemrc.

Unlike the nanorc reader, a problem here is fatal: these files are ours and ship in the binary, so a bad pattern is a build-time mistake rather than a third-party file to be tolerated. Failing loudly is what makes the "every embedded rule compiles" test meaningful.

func (*RuleSet) Comment

func (rs *RuleSet) Comment() string

Comment reports the line-comment introducer, or "".

func (*RuleSet) Lex

func (rs *RuleSet) Lex(line []rune, in State) ([]Span, State)

Lex classifies one line against the set's rules.

Rules apply in file order and each paints over whatever came before. Painting into a per-rune array rather than collecting spans is what makes that exact: matches overlap constantly - a keyword inside a string, a number inside a comment - and resolving overlaps afterwards would need the same array anyway. The later rule wins because a .nemrc lists broad rules first and narrow ones after, which is how a string containing the word `if` stays a string.

func (*RuleSet) Name

func (rs *RuleSet) Name() string

Name reports the language.

func (*RuleSet) Rules

func (rs *RuleSet) Rules() int

Rules reports how many rules the set carries, for tests that assert coverage.

type Span

type Span struct {
	Start, End int
	Class      Class
}

Span is a half-open run of one class within a line.

Start and End are RUNE indices, not byte offsets. The renderer converts them to display columns, and a byte offset would land in the wrong column the moment a line contains anything outside ASCII - which is exactly where a highlighter's mistakes are least forgivable, since the text still looks fine and only the colours are wrong.

type State

type State uint32

State is what one line leaves open for the next: a block comment, a raw string, a fenced code block.

It is opaque and its encoding is private to each lexer, but its layout is fixed so that the zero value means the same thing everywhere:

bits 0..7    mode   - lexer-specific; 0 always means "nothing is open"
bits 8..23   param  - lexer-specific; a Lua long-bracket level, a Markdown
                      fence length and its delimiter
bits 24..31  unused, always zero

The zero State is therefore "start of file", which is what a caller passes for line 0 without needing to ask the lexer for an initial value.

Directories

Path Synopsis
Package nanorc reads GNU nano's syntax-highlighting definitions so nem can colour the languages it has no hand-written lexer for.
Package nanorc reads GNU nano's syntax-highlighting definitions so nem can colour the languages it has no hand-written lexer for.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL