lexical

package
v0.3.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 28, 2026 License: MIT Imports: 11 Imported by: 0

Documentation

Overview

Package lexical converts between Payload's Lexical rich-text JSON and Markdown, and extracts plain text from it.

Payload 3 stores a richText field as a Lexical editor state:

{"root": {"type": "root", "children": [ …block nodes… ], "direction": "ltr",
          "format": "", "indent": 0, "version": 1}}

Writing that shape by hand is the most error-prone thing an agent does against a Payload site: the node vocabulary is large, every element carries five bookkeeping keys, text formatting is a bitmask, and a link's target lives in a nested `fields` object whose shape depends on linkType. Agents measured doing a content migration wrote their own Markdown→Lexical builders and a Lexical→text parser before they could do anything else. This package is that code, written once, against the shapes Payload 3.87 and Lexical 0.41 store (see the grounding notes on the constructors in nodes.go).

Everything here is pure: no I/O, no clock, no environment. Inputs are the generic values encoding/json produces (map[string]any, []any, string, float64 or json.Number, bool, nil); outputs are the same kinds, so a result marshals straight into a request body.

API

  • FromMarkdown / FromMarkdownIssues: Markdown → editor state.
  • ToMarkdown / ToMarkdownIssues: editor state → Markdown.
  • PlainText: editor state → text.
  • IsEditorState, Walk, NodeTypes: recognise and traverse a state.
  • StripFrontmatter: drop a leading YAML block before converting (only a block whose lines look like YAML; a leading rule is content).

Dialect

CommonMark, restricted to what Lexical can hold, plus a few extensions:

# … ######            heading h1–h6 (ATX, closing #s allowed; setext = and - too)
paragraph             soft line breaks become a space
**b** __b__           bold (1)           *i* _i_      italic (2)
~~s~~                 strikethrough (4)  (exactly two tildes; "~5 min" stays text)
<u>u</u>              underline (8)      `c`          code (16)
<sub>x</sub>          subscript (32)     <sup>x</sup> superscript (64)
<mark>x</mark>        highlight (128)
<strong> <b> <em> <i> <s> <del> <ins> <code> are accepted as well
\ at end of line, two trailing spaces, <br>      hard break → linebreak node
&#9; or a tab         tab node           \* \_ \[ …  escapes; &amp; &#228; entities
[t](https://…)        link, linkType custom
[t](pay:pages/47)     link, linkType internal, doc {relationTo:"pages", value:47}
                      (all-digit ids are numbers, anything else a string id)
[t](pay:pages/slug=about)  the page whose slug is "about": doc.value is
                      {"$ref":"pages","slug":"about"}, which a write resolves to its id
[t](…){target=_blank} newTab; further `key=value` attributes become link fields,
                      `autolink` makes an autolink node
[t][ref] [ref][] [ref]  reference links, resolved against `[ref]: url "title"`
                      definitions anywhere in the document (the lines vanish)
<https://…> <a@b.c>   autolink node (mailto: for an address)
- * +  /  1. 1)       bullet / numbered list (start = first number), nested by
                      indentation into Lexical's wrapper items
- [ ] / - [x]         check list (checked false / true)
> quote               one paragraph → native quote (inline children);
                      several blocks → quote holding them
--- *** ___           horizontal rule (ToMarkdown writes ***, so its output
                      never opens with the front matter fence)
![media:40]()         upload node (Payload's own Markdown form)
<!-- lexical:TYPE {json} -->
                      any node, restored exactly: the JSON is the node minus
                      `type` and minus the keys at their export default

Code blocks become a paragraph of code-formatted lines, pipe tables stay paragraph text, other HTML stays literal text, other comments are dropped, and images other than the upload form stay text — each reported as an Issue. ToMarkdown writes exactly this dialect; FromMarkdown(ToMarkdown(x)) equals x for every node except the presentation attributes Markdown cannot spell (alignment, indent, inline CSS, three text-transform bits), which are reported, and derived bookkeeping (direction, paragraph textFormat, list item value/indent, link ids), which Lexical recomputes.

Index

Constants

View Source
const (
	FormatBold          = 1
	FormatItalic        = 2
	FormatStrikethrough = 4
	FormatUnderline     = 8
	FormatCode          = 16
	FormatSubscript     = 32
	FormatSuperscript   = 64
	FormatHighlight     = 128
)

Text format bits, exactly Lexical's TextNode format bitmask (lexical@0.41 IS_BOLD … IS_HIGHLIGHT). Bits 256 (lowercase), 512 (uppercase) and 1024 (capitalize) exist in Lexical but have no Payload feature and no Markdown spelling; ToMarkdown reports them as lossy.

View Source
const (
	// IssueHTMLAsText: raw HTML this dialect does not model was kept as
	// literal text (Lexical has no HTML node).
	IssueHTMLAsText = "html_as_text"
	// IssueCommentDropped: an HTML comment that is not a lexical placeholder
	// was removed.
	IssueCommentDropped = "comment_dropped"
	// IssueImageAsText: a Markdown image that is not the `![collection:id]()`
	// upload form was kept as literal text (an upload node needs a media id).
	IssueImageAsText = "image_as_text"
	// IssueCodeBlockFlattened: a fenced or indented code block became a
	// paragraph of inline-code text (Payload's default editor has no code
	// block node; projects model code as a block).
	IssueCodeBlockFlattened = "code_block_flattened"
	// IssueTableAsText: a GFM pipe table was kept as paragraph text.
	IssueTableAsText = "table_as_text"
	// IssueItemParagraphsJoined: a list item with several paragraphs was
	// joined with line breaks (a Lexical list item holds inline content only).
	IssueItemParagraphsJoined = "list_item_paragraphs_joined"
	// IssueListItemBlockFlattened: a heading or block quote inside a list
	// item lost its block type and became the item's inline text (a Lexical
	// list item holds inline content only).
	IssueListItemBlockFlattened = "list_item_block_flattened"
	// IssueListBecameChecklist: a bullet list mixing task items (`- [ ]`)
	// and plain items became a check list; the plain items are unchecked
	// checkboxes now (a Lexical list has one listType).
	IssueListBecameChecklist = "list_became_checklist"
	// IssueNestingCapped: quotes/lists nested deeper than the parser's bound;
	// the deeper lines were kept as paragraph text.
	IssueNestingCapped = "nesting_capped"
	// IssueLinkRefUndefined: a full or collapsed reference link
	// (`[text][label]`, `[label][]`) names a label no `[label]: url`
	// definition declares; it was kept as text.
	IssueLinkRefUndefined = "link_reference_undefined"
	// IssueLinkURLInvalid: a custom link whose URL is empty or contains a
	// space. Payload's default link validation (url required, and
	// validateUrlMinimal: no spaces) rejects the document on any write that
	// validates — a publish or a write without --draft. Verified live for the
	// empty URL: "link node failed to validate".
	IssueLinkURLInvalid = "link_url_invalid"
	// IssueLinkURLRelative: a link URL that is relative to the Markdown file
	// (`../brief.md`, `page.html`, `www.example.com`) rather than a site path
	// (`/…`), an anchor (`#…`), a query (`?…`) or an absolute URL. On the site
	// it resolves against the page's URL, which is almost never what the
	// source meant.
	IssueLinkURLRelative = "link_url_relative"

	// IssueLossyStateKeys: the editor state carries non-null keys next to
	// `root`; Markdown holds the root only.
	IssueLossyStateKeys = "lossy_state_keys"
	// IssueLossyAlignment: an element's format (alignment) or indent has no
	// Markdown spelling and is lost.
	IssueLossyAlignment = "lossy_alignment"
	// IssueLossyStyle: inline CSS (text `style`, element `textStyle`) is lost.
	IssueLossyStyle = "lossy_style"
	// IssueLossyFormat: a text format bit with no Markdown spelling
	// (lowercase/uppercase/capitalize) is lost.
	IssueLossyFormat = "lossy_format"
	// IssuePlaceholder: a node was written as a lossless
	// `<!-- lexical:TYPE {…} -->` placeholder because Markdown cannot express
	// it. Informational: FromMarkdown restores it exactly.
	IssuePlaceholder = "placeholder"
	// IssueNormalized: a shape Lexical's editor never produces (a quote
	// holding one paragraph, a list item holding a paragraph or text followed
	// by a nested list) was written as ordinary Markdown and reads back as
	// the editor's native shape. The content is unchanged.
	IssueNormalized = "normalized"
)

Issue codes. FromMarkdown raises the first group, ToMarkdown the second.

View Source
const RefKey = "$ref"

RefKey marks a document reference in a link's doc value — the "$ref" of PayCLI's write bodies (internal/manifest); lexical only produces and renders it.

Variables

View Source
var ErrNotEditorState = errors.New("not a Lexical editor state (expected an object with root.type \"root\" and root.children)")

ErrNotEditorState is returned by ToMarkdown for a value that is not a Lexical editor state.

Functions

func FromMarkdown

func FromMarkdown(md string) (map[string]any, error)

FromMarkdown converts Markdown to a Lexical editor state ({"root": {...}}) in the exact shape Payload 3.87's lexical editor stores. It fails only on a malformed `<!-- lexical:… -->` placeholder, because turning one into text would silently lose the node it carries; everything else converts, and what could not be represented is reported by FromMarkdownIssues.

func IsEditorState

func IsEditorState(v any) bool

IsEditorState reports whether v is a Lexical editor state: an object whose "root" is an object of type "root" with a children array. It is the test every "find the rich text in this document" walk uses, so it is strict about the root and says nothing about the children.

func NodeTypes

func NodeTypes(state any) map[string]int

NodeTypes counts the node types of an editor state, root included. It is what a caller prints to show what a conversion produced.

func PlainText

func PlainText(state any) string

PlainText extracts the readable text of an editor state: blocks separated by a blank line, list items and table rows by a newline, table cells by " | ", a linebreak as "\n" and a tab as "\t". Uploads, relationships, blocks, inline blocks and rules contribute nothing — they carry no text of their own. It returns "" for a value that is not an editor state.

The separators follow Payload's own convertLexicalToPlaintext heuristic (@payloadcms/richtext-lexical 3.87), except that a quote is a block here too, so quoted text does not run into the paragraph before it.

func PlainTextWithPlaceholders

func PlainTextWithPlaceholders(state any) string

PlainTextWithPlaceholders is PlainText for a reader who must SEE every node, text or not: a diff that has to show a swapped image, an outline that summarises a block holding only an upload. Each node without text of its own is rendered as a short bracketed placeholder where PlainText renders nothing:

upload          [upload media:12]       (relationTo:id; a populated value by its id)
relationship    [relationship pages:47]
block           [block cta]             (fields.blockType; [block] without one)
inlineBlock     [inlineBlock banner]
horizontalrule  ---
any other node type without children    [TYPE]

Separators are PlainText's. The placeholders are for reading only: they are not Markdown and FromMarkdown does not turn them back into nodes (the lossless form is ToMarkdown's <!-- lexical:… --> placeholder).

func StripFrontmatter

func StripFrontmatter(md string) (body, fm string, ok bool)

StripFrontmatter removes a leading YAML front matter block (`---` … `---`), which Markdown files written for static-site generators and knowledge bases carry and which CommonMark would otherwise read as a rule and a heading. ok reports whether one was found; fm is its body.

The block must look like YAML: its first line directly follows the opening `---` (no blank line) and every line is a `key:` line, a list item, a comment, a blank line or an indented continuation. Anything else is Markdown that happens to start with a thematic break — a rule, a paragraph, a second rule — and is left alone: stripping it would delete content.

func ToMarkdown

func ToMarkdown(state any) (string, error)

ToMarkdown renders a Lexical editor state as Markdown in this package's dialect. Every node survives FromMarkdown(ToMarkdown(x)): what Markdown can spell is spelled, and what it cannot is written as a lossless `<!-- lexical:TYPE {…} -->` placeholder. The only losses are presentation attributes Markdown has no syntax for (alignment, indent, inline CSS), which ToMarkdownIssues reports.

func Walk

func Walk(state any, fn func(n Node, path string) bool)

Walk visits every node of an editor state depth-first, parents before children. path is the node's address ("root", "root.children[2]", "root.children[2].children[0]"). Returning false from fn skips that node's children. Walk does nothing when state is not an editor state.

Types

type Issue

type Issue struct {
	// Code is a stable identifier (see the Issue* constants).
	Code string `json:"code"`
	// Path addresses the node in the editor state ("root.children[3]") for
	// ToMarkdown, and is empty for FromMarkdown.
	Path string `json:"path,omitempty"`
	// Line is the 1-based Markdown source line for FromMarkdown, 0 otherwise.
	Line int `json:"line,omitempty"`
	// Message says what happened in plain words.
	Message string `json:"message"`
}

Issue is a non-fatal finding of a conversion: something that was dropped, normalised or could not be expressed. Conversions never fail on content they can represent some other way; they report it here instead so the caller can decide whether the loss matters.

func FromMarkdownIssues

func FromMarkdownIssues(md string) (map[string]any, []Issue, error)

FromMarkdownIssues is FromMarkdown plus the list of non-fatal findings (dropped comments, HTML kept as text, flattened code blocks, link URLs Payload rejects or that are relative to the source file, …), in source order.

func ToMarkdownIssues

func ToMarkdownIssues(state any) (string, []Issue, error)

ToMarkdownIssues is ToMarkdown plus the findings: lossy attributes, placeholders written, and shapes normalised to Lexical's native form.

type Node

type Node = map[string]any

Node is one serialized Lexical node.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL