fastxml

package
v0.1.4 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 4, 2026 License: Apache-2.0 Imports: 9 Imported by: 0

Documentation

Overview

Package fastxml 提供 OFD 快速解析路径共用的零分配 XML lexer。它包装 github.com/tdewolff/parse/v2/xml,在原始字节上定位元素区间;低频、结构复杂 的子树用 DecodeSpan 截取字节后回退 encoding/xml,语义与 xml.Unmarshal 一致。

这里只放与具体模型无关的词法工具;页面内容、文档主体等模型重建仍留在 internal/models,避免为复用 lexer 而导出模型内部方法。

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func EndTagName

func EndTagName(b []byte) []byte

EndTagName 从 </Name> 结束标记中取出元素名。

func LocalName

func LocalName(b []byte) []byte

LocalName 去掉可能的命名空间前缀。

func StripBOM

func StripBOM(data []byte) ([]byte, bool)

StripBOM 去掉 UTF-8 BOM。若数据带 UTF-16 BOM,返回 utf16=true,调用方应回退 到 encoding/xml(UTF-16 编码的 OFD 极罕见)。

func TargetName

func TargetName(b []byte) []byte

TargetName 从 <Name ...> 开始标记中取出元素名。

func Unescape

func Unescape(b []byte) string

Unescape 反解码 XML 文本实体,无实体时零拷贝返回。

func Unquote

func Unquote(b []byte) []byte

Unquote 去掉属性值的引号。

Types

type Lexer

type Lexer struct {
	// contains filtered or unexported fields
}

Lexer 在原始字节上做 XML 词法分析,token 与区间偏移都指向原始切片,不拷贝。

func NewLexer

func NewLexer(data []byte) *Lexer

NewLexer 为 data 创建 lexer;data 需在解析期间保持有效。

func (*Lexer) AttrVal

func (l *Lexer) AttrVal() []byte

AttrVal 返回当前 AttributeToken 的属性值。

func (*Lexer) CData

func (l *Lexer) CData() []byte

CData 返回当前 CDATA token 的内容。tdewolff 的 xml lexer 在 Next() 中返回的是 含 "<![CDATA[" / "]]>" 包装的原始词素,去掉包装的正文在 Text() 里。

func (*Lexer) Data

func (l *Lexer) Data() []byte

Data 返回底层原始字节。

func (*Lexer) DecodeSpan

func (l *Lexer) DecodeSpan(start int, dst any) error

DecodeSpan 捕获 start 起的元素字节区间,交给 encoding/xml 完整解码。

func (*Lexer) Err

func (l *Lexer) Err() error

Err 返回词法错误;到达输入末尾时返回 io.EOF。

func (*Lexer) Next

func (l *Lexer) Next() (tt txml.TokenType, buf []byte, start int)

Next 返回下一个 token,start 是该 token 起始处的绝对字节偏移。绝对偏移直接 取自底层 parse.Input 的 Offset(),避免按 token 缓冲长度累加时被词法器内部 跳过的空白字节(如开始标签属性之间的空格)所扰动。

func (*Lexer) ReadAttrs

func (l *Lexer) ReadAttrs(h func(name string, value []byte) error) (void bool, err error)

ReadAttrs 消费开始标签后的属性序列,直到 '>' 或自闭合 '/>'。

func (*Lexer) SkipAttrs

func (l *Lexer) SkipAttrs() (bool, error)

SkipAttrs 跳过开始标签后的全部属性。

func (*Lexer) SpanElement

func (l *Lexer) SpanElement(start int) ([]byte, error)

SpanElement 从元素起始偏移 start 消费到匹配的结束标记,返回该元素的原始字节 区间。调用时该元素的 StartTag token 已被消费。

func (*Lexer) Text

func (l *Lexer) Text() []byte

Text 返回当前 token 的原始文本(属性名等)。

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL