Documentation
¶
Overview ¶
Package fastxml 提供 OFD 快速解析路径共用的零分配 XML lexer。它包装 github.com/tdewolff/parse/v2/xml,在原始字节上定位元素区间;低频、结构复杂 的子树用 DecodeSpan 截取字节后回退 encoding/xml,语义与 xml.Unmarshal 一致。
这里只放与具体模型无关的词法工具;页面内容、文档主体等模型重建仍留在 internal/models,避免为复用 lexer 而导出模型内部方法。
Index ¶
- func EndTagName(b []byte) []byte
- func LocalName(b []byte) []byte
- func StripBOM(data []byte) ([]byte, bool)
- func TargetName(b []byte) []byte
- func Unescape(b []byte) string
- func Unquote(b []byte) []byte
- type Lexer
- func (l *Lexer) AttrVal() []byte
- func (l *Lexer) CData() []byte
- func (l *Lexer) Data() []byte
- func (l *Lexer) DecodeSpan(start int, dst any) error
- func (l *Lexer) Err() error
- func (l *Lexer) Next() (tt txml.TokenType, buf []byte, start int)
- func (l *Lexer) ReadAttrs(h func(name string, value []byte) error) (void bool, err error)
- func (l *Lexer) SkipAttrs() (bool, error)
- func (l *Lexer) SpanElement(start int) ([]byte, error)
- func (l *Lexer) Text() []byte
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
Types ¶
type Lexer ¶
type Lexer struct {
// contains filtered or unexported fields
}
Lexer 在原始字节上做 XML 词法分析,token 与区间偏移都指向原始切片,不拷贝。
func (*Lexer) CData ¶
CData 返回当前 CDATA token 的内容。tdewolff 的 xml lexer 在 Next() 中返回的是 含 "<![CDATA[" / "]]>" 包装的原始词素,去掉包装的正文在 Text() 里。
func (*Lexer) DecodeSpan ¶
DecodeSpan 捕获 start 起的元素字节区间,交给 encoding/xml 完整解码。
func (*Lexer) Next ¶
Next 返回下一个 token,start 是该 token 起始处的绝对字节偏移。绝对偏移直接 取自底层 parse.Input 的 Offset(),避免按 token 缓冲长度累加时被词法器内部 跳过的空白字节(如开始标签属性之间的空格)所扰动。
func (*Lexer) SpanElement ¶
SpanElement 从元素起始偏移 start 消费到匹配的结束标记,返回该元素的原始字节 区间。调用时该元素的 StartTag token 已被消费。