gse

package
v1.30.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 21, 2026 License: Apache-2.0 Imports: 39 Imported by: 0

Documentation

Overview

Package gse provides a dictionary based tokenizer for Chinese, Japanese and other languages supported by github.com/go-ego/gse, plus a small gse-bleve style index API (New, Index, Search) built on riot.

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

func Langs

func Langs() []string

Langs lists the codes accepted by Option.Lang.

func NewAnalyzer

func NewAnalyzer(seg *gogse.Segmenter, opt Option) (*analysis.Analyzer, error)

NewAnalyzer builds an analyzer over an already loaded segmenter using opt.Opt as cut mode, so one segmenter can back both an index ("search-hmm") and a query ("hmm") analyzer. With a nil seg the analysis/lang analyzer named by opt.Lang is returned instead.

func NewLangAnalyzer

func NewLangAnalyzer(code string) (*analysis.Analyzer, error)

NewLangAnalyzer returns the analysis/lang analyzer for code.

func NewSegmenter

func NewSegmenter(opt Option) (*gogse.Segmenter, error)

NewSegmenter loads the dictionaries described by opt.

func Query added in v1.30.0

func Query() *query.Builder

Query creates a fluent search builder. See package query for clause semantics.

Types

type Batch added in v1.30.0

type Batch struct {
	// contains filtered or unexported fields
}

Batch queues mapped updates and deletes for a single atomic writer batch. It is not safe for concurrent use. The last operation for an ID wins.

func (*Batch) Commit added in v1.30.0

func (b *Batch) Commit() error

Commit applies all queued operations together and resets the batch on success. On failure the operations remain queued for retry.

func (*Batch) Delete added in v1.30.0

func (b *Batch) Delete(id string)

Delete queues removal of id, replacing any queued update for that ID.

func (*Batch) Index added in v1.30.0

func (b *Batch) Index(id string, data any) error

Index queues a replacement using the same mapping and analyzer as Index.Index. Invalid data returns an error without changing the queued operations.

func (*Batch) Reset added in v1.30.0

func (b *Batch) Reset()

Reset discards all queued operations.

type Hit

type Hit struct {
	ID    string
	Score float64
	// Fields holds the stored field values of the document.
	Fields map[string]string
	// Fragments holds highlighted snippets per field when requested.
	Fragments map[string][]string
}

Hit is one matched document.

type Index

type Index struct {
	// contains filtered or unexported fields
}

Index is a riot writer with gse (or analysis/lang) analyzers attached, in the spirit of gse-bleve: New, Index, Search, Close.

func New

func New(opt Option) (*Index, error)

New opens the index at opt.Index (in-memory when empty). With opt.Lang set the matching analysis/lang analyzer is used for both indexing and queries; otherwise the gse dictionaries are loaded, documents are cut with opt.Opt and queries with its non-search counterpart ("search-hmm" -> "hmm").

func (*Index) Batch added in v1.30.0

func (x *Index) Batch() *Batch

Batch creates an empty batch. Changes are visible only after Commit succeeds.

func (*Index) Close

func (x *Index) Close() error

Close closes the underlying writer.

func (*Index) Delete

func (x *Index) Delete(id string) error

Delete removes the document with id.

func (*Index) Field

func (x *Index) Field(name, text string) *riot.TermField

Field builds a stored, highlightable text field cut by the index analyzer; use it to add extra gse fields to documents written through Writer.

func (*Index) Index

func (x *Index) Index(id string, data any) error

Index writes a string or a struct (or non-nil pointer to a struct) under id, replacing any existing document with that id. Strings use Option.Field. Structs follow encoding/json field names and tags, including "-" and omitempty. Nested fields use dotted names; arrays are joined with newlines; nulls are skipped. Mapped names must be nonempty, without dots or a leading underscore. Scalar values are stored and analyzed as text, not as numeric or date range fields. Use Request.Field to search a mapped field; the default search field is unchanged.

func (*Index) Search

func (x *Index) Search(req SearchRequest) (res *Result, err error)

Search runs req against the current index snapshot.

func (*Index) Segmenter

func (x *Index) Segmenter() *gogse.Segmenter

Segmenter returns the loaded gse segmenter, e.g. to add user words. It is nil when the index was opened with Option.Lang.

func (*Index) Writer

func (x *Index) Writer() *riot.Writer

Writer returns the underlying riot writer for batch or custom documents.

type Option

type Option struct {
	// Index is the on-disk index path; empty opens an in-memory index.
	Index string
	// Field is the field for string documents and default searches; default "text".
	// Struct documents use their JSON field names instead.
	Field string
	// Lang selects a riot analysis/lang analyzer ("en", "cjk", "de", ...;
	// see Langs) instead of gse. When set, no gse dictionary is loaded, the
	// Segmenter is nil and Dicts/Stop/Opt/Alpha are ignored.
	Lang string
	// Dicts selects the dictionaries: "zh", "zh_s", "zh_t", "ja"/"jp" or
	// comma separated dictionary file paths. Prefix with "embed, " to use the
	// dictionaries compiled into gse instead of reading files from the gse
	// module directory. Empty loads the embedded "zh" dictionary.
	Dicts string
	// Stop selects the stop word dictionary: "zh", "embed, zh" or file paths.
	// Empty loads none. Stop words are dropped from the token stream.
	Stop string
	// Opt is the cut mode: "" (shortest path), "hmm", "dag", "search",
	// "search-hmm" or "search-dag". The search modes additionally emit the
	// sub-words of every segment and are meant for indexing; the Index
	// built by New queries with the matching non-search mode.
	Opt string
	// Alpha makes gse emit every Latin letter/digit as its own token.
	Alpha bool
}

Option configures the segmenter, the cut mode and the index.

type Request

type Request struct {
	Query string
	// Field to match; empty uses the field configured in Option.
	Field string
	// Size is the number of hits to return (default 10) and From the number
	// of hits to skip.
	Size, From int
	// Highlight adds HTML <mark> fragments of the stored field to every hit.
	Highlight bool
}

Request is an analyzed text search; build it with QueryString.

func NewQueryString deprecated

func NewQueryString(text string, enableHighlight ...bool) *Request

NewQueryString creates an analyzed text request.

Deprecated: use QueryString instead.

func QueryString added in v1.30.0

func QueryString(text string, enableHighlight ...bool) *Request

QueryString creates an analyzed text request against the index field. It does not parse Elasticsearch query-string syntax. Pass true to also return highlighted fragments.

func (*Request) Build added in v1.30.0

func (r *Request) Build(field string, analyzer *analysis.Analyzer) (*riot.TopNSearch, error)

Build compiles a legacy request using the index's field and query analyzer.

func (*Request) HighlightEnabled added in v1.30.0

func (r *Request) HighlightEnabled() bool

HighlightEnabled reports whether the request includes highlighted fragments.

type Result

type Result struct {
	Total    uint64
	MaxScore float64
	Took     time.Duration
	Hits     []*Hit
}

Result is the outcome of Index.Search.

type SearchRequest added in v1.30.0

type SearchRequest interface {
	Build(string, *analysis.Analyzer) (*riot.TopNSearch, error)
	HighlightEnabled() bool
}

SearchRequest is implemented by Request and query.Builder.

type Tokenizer

type Tokenizer struct {
	// contains filtered or unexported fields
}

Tokenizer adapts a loaded gse.Segmenter to the analysis.Tokenizer interface.

func NewTokenizer

func NewTokenizer(seg *gogse.Segmenter, search bool, hmm ...bool) *Tokenizer

NewTokenizer wraps seg. With search enabled every segment is additionally expanded into its sub-words (搜索引擎 -> 搜索 索引 引擎 搜索引擎) at the same position; use it for indexing so that shorter query terms still hit.

func (*Tokenizer) Tokenize

func (t *Tokenizer) Tokenize(input []byte) analysis.TokenStream

Tokenize cuts input and reports exact byte offsets for every token.

Directories

Path Synopsis
Package query provides an immutable fluent builder for riot search requests.
Package query provides an immutable fluent builder for riot search requests.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL