gse

package
v1.23.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 13, 2026 License: Apache-2.0 Imports: 34 Imported by: 0

Documentation

Overview

Package gse provides a dictionary based tokenizer for Chinese, Japanese and other languages supported by github.com/go-ego/gse, plus a small gse-bleve style index API (New, Index, Search) built on riot.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func Langs

func Langs() []string

Langs lists the codes accepted by Option.Lang.

func NewAnalyzer

func NewAnalyzer(seg *gogse.Segmenter, opt Option) (*analysis.Analyzer, error)

NewAnalyzer builds an analyzer over an already loaded segmenter using opt.Opt as cut mode, so one segmenter can back both an index ("search-hmm") and a query ("hmm") analyzer. With a nil seg the analysis/lang analyzer named by opt.Lang is returned instead.

func NewLangAnalyzer

func NewLangAnalyzer(code string) (*analysis.Analyzer, error)

NewLangAnalyzer returns the analysis/lang analyzer for code.

func NewSegmenter

func NewSegmenter(opt Option) (*gogse.Segmenter, error)

NewSegmenter loads the dictionaries described by opt.

Types

type Hit

type Hit struct {
	ID    string
	Score float64
	// Fields holds the stored field values of the document.
	Fields map[string]string
	// Fragments holds highlighted snippets per field when requested.
	Fragments map[string][]string
}

Hit is one matched document.

type Index

type Index struct {
	// contains filtered or unexported fields
}

Index is a riot writer with gse (or analysis/lang) analyzers attached, in the spirit of gse-bleve: New, Index, Search, Close.

func New

func New(opt Option) (*Index, error)

New opens the index at opt.Index (in-memory when empty). With opt.Lang set the matching analysis/lang analyzer is used for both indexing and queries; otherwise the gse dictionaries are loaded, documents are cut with opt.Opt and queries with its non-search counterpart ("search-hmm" -> "hmm").

func (*Index) Close

func (x *Index) Close() error

Close closes the underlying writer.

func (*Index) Delete

func (x *Index) Delete(id string) error

Delete removes the document with id.

func (*Index) Field

func (x *Index) Field(name, text string) *riot.TermField

Field builds a stored, highlightable text field cut by the index analyzer; use it to add extra gse fields to documents written through Writer.

func (*Index) Index

func (x *Index) Index(id, text string) error

Index writes text under id into the configured field, replacing any existing document with that id.

func (*Index) Search

func (x *Index) Search(req *Request) (*Result, error)

Search runs req against the current index snapshot.

func (*Index) Segmenter

func (x *Index) Segmenter() *gogse.Segmenter

Segmenter returns the loaded gse segmenter, e.g. to add user words. It is nil when the index was opened with Option.Lang.

func (*Index) Writer

func (x *Index) Writer() *riot.Writer

Writer returns the underlying riot writer for batch or custom documents.

type Option

type Option struct {
	// Index is the on-disk index path; empty opens an in-memory index.
	Index string
	// Field is the document field written by Index.Index; default "text".
	Field string
	// Lang selects a riot analysis/lang analyzer ("en", "cjk", "de", ...;
	// see Langs) instead of gse. When set, no gse dictionary is loaded, the
	// Segmenter is nil and Dicts/Stop/Opt/Alpha are ignored.
	Lang string
	// Dicts selects the dictionaries: "zh", "zh_s", "zh_t", "ja"/"jp" or
	// comma separated dictionary file paths. Prefix with "embed, " to use the
	// dictionaries compiled into gse instead of reading files from the gse
	// module directory. Empty loads the embedded "zh" dictionary.
	Dicts string
	// Stop selects the stop word dictionary: "zh", "embed, zh" or file paths.
	// Empty loads none. Stop words are dropped from the token stream.
	Stop string
	// Opt is the cut mode: "" (shortest path), "hmm", "dag", "search",
	// "search-hmm" or "search-dag". The search modes additionally emit the
	// sub-words of every segment and are meant for indexing; the Index
	// built by New queries with the matching non-search mode.
	Opt string
	// Alpha makes gse emit every Latin letter/digit as its own token.
	Alpha bool
}

Option configures the segmenter, the cut mode and the index.

type Request

type Request struct {
	Query string
	// Field to match; empty uses the field configured in Option.
	Field string
	// Size is the number of hits to return (default 10) and From the number
	// of hits to skip.
	Size, From int
	// Highlight adds HTML <mark> fragments of the stored field to every hit.
	Highlight bool
}

Request is a query string search; build it with NewQueryString.

func NewQueryString

func NewQueryString(query string, enableHighlight ...bool) *Request

NewQueryString creates a request matching query against the index field. Pass true to also return highlighted fragments.

type Result

type Result struct {
	Total    uint64
	MaxScore float64
	Took     time.Duration
	Hits     []*Hit
}

Result is the outcome of Index.Search.

type Tokenizer

type Tokenizer struct {
	// contains filtered or unexported fields
}

Tokenizer adapts a loaded gse.Segmenter to the analysis.Tokenizer interface.

func NewTokenizer

func NewTokenizer(seg *gogse.Segmenter, search bool, hmm ...bool) *Tokenizer

NewTokenizer wraps seg. With search enabled every segment is additionally expanded into its sub-words (搜索引擎 -> 搜索 索引 引擎 搜索引擎) at the same position; use it for indexing so that shorter query terms still hit.

func (*Tokenizer) Tokenize

func (t *Tokenizer) Tokenize(input []byte) analysis.TokenStream

Tokenize cuts input and reports exact byte offsets for every token.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL