Documentation
¶
Overview ¶
Package gse provides a dictionary based tokenizer for Chinese, Japanese and other languages supported by github.com/go-ego/gse, plus a small gse-bleve style index API (New, Index, Search) built on riot.
Index ¶
- func Langs() []string
- func NewAnalyzer(seg *gogse.Segmenter, opt Option) (*analysis.Analyzer, error)
- func NewLangAnalyzer(code string) (*analysis.Analyzer, error)
- func NewSegmenter(opt Option) (*gogse.Segmenter, error)
- func Query() *query.Builder
- type Batch
- type Hit
- type Index
- func (x *Index) Batch() *Batch
- func (x *Index) Close() error
- func (x *Index) Delete(id string) error
- func (x *Index) Field(name, text string) *riot.TermField
- func (x *Index) Index(id string, data any) error
- func (x *Index) Search(req SearchRequest) (res *Result, err error)
- func (x *Index) Segmenter() *gogse.Segmenter
- func (x *Index) Writer() *riot.Writer
- type Option
- type Request
- type Result
- type SearchRequest
- type Tokenizer
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func NewAnalyzer ¶
NewAnalyzer builds an analyzer over an already loaded segmenter using opt.Opt as cut mode, so one segmenter can back both an index ("search-hmm") and a query ("hmm") analyzer. With a nil seg the analysis/lang analyzer named by opt.Lang is returned instead.
func NewLangAnalyzer ¶
NewLangAnalyzer returns the analysis/lang analyzer for code.
func NewSegmenter ¶
NewSegmenter loads the dictionaries described by opt.
Types ¶
type Batch ¶ added in v1.30.0
type Batch struct {
// contains filtered or unexported fields
}
Batch queues mapped updates and deletes for a single atomic writer batch. It is not safe for concurrent use. The last operation for an ID wins.
func (*Batch) Commit ¶ added in v1.30.0
Commit applies all queued operations together and resets the batch on success. On failure the operations remain queued for retry.
func (*Batch) Delete ¶ added in v1.30.0
Delete queues removal of id, replacing any queued update for that ID.
type Hit ¶
type Hit struct {
ID string
Score float64
// Fields holds the stored field values of the document.
Fields map[string]string
// Fragments holds highlighted snippets per field when requested.
Fragments map[string][]string
}
Hit is one matched document.
type Index ¶
type Index struct {
// contains filtered or unexported fields
}
Index is a riot writer with gse (or analysis/lang) analyzers attached, in the spirit of gse-bleve: New, Index, Search, Close.
func New ¶
New opens the index at opt.Index (in-memory when empty). With opt.Lang set the matching analysis/lang analyzer is used for both indexing and queries; otherwise the gse dictionaries are loaded, documents are cut with opt.Opt and queries with its non-search counterpart ("search-hmm" -> "hmm").
func (*Index) Batch ¶ added in v1.30.0
Batch creates an empty batch. Changes are visible only after Commit succeeds.
func (*Index) Field ¶
Field builds a stored, highlightable text field cut by the index analyzer; use it to add extra gse fields to documents written through Writer.
func (*Index) Index ¶
Index writes a string or a struct (or non-nil pointer to a struct) under id, replacing any existing document with that id. Strings use Option.Field. Structs follow encoding/json field names and tags, including "-" and omitempty. Nested fields use dotted names; arrays are joined with newlines; nulls are skipped. Mapped names must be nonempty, without dots or a leading underscore. Scalar values are stored and analyzed as text, not as numeric or date range fields. Use Request.Field to search a mapped field; the default search field is unchanged.
func (*Index) Search ¶
func (x *Index) Search(req SearchRequest) (res *Result, err error)
Search runs req against the current index snapshot.
type Option ¶
type Option struct {
// Index is the on-disk index path; empty opens an in-memory index.
Index string
// Field is the field for string documents and default searches; default "text".
// Struct documents use their JSON field names instead.
Field string
// Lang selects a riot analysis/lang analyzer ("en", "cjk", "de", ...;
// see Langs) instead of gse. When set, no gse dictionary is loaded, the
// Segmenter is nil and Dicts/Stop/Opt/Alpha are ignored.
Lang string
// Dicts selects the dictionaries: "zh", "zh_s", "zh_t", "ja"/"jp" or
// comma separated dictionary file paths. Prefix with "embed, " to use the
// dictionaries compiled into gse instead of reading files from the gse
// module directory. Empty loads the embedded "zh" dictionary.
Dicts string
// Stop selects the stop word dictionary: "zh", "embed, zh" or file paths.
// Empty loads none. Stop words are dropped from the token stream.
Stop string
// Opt is the cut mode: "" (shortest path), "hmm", "dag", "search",
// "search-hmm" or "search-dag". The search modes additionally emit the
// sub-words of every segment and are meant for indexing; the Index
// built by New queries with the matching non-search mode.
Opt string
// Alpha makes gse emit every Latin letter/digit as its own token.
Alpha bool
}
Option configures the segmenter, the cut mode and the index.
type Request ¶
type Request struct {
Query string
// Field to match; empty uses the field configured in Option.
Field string
// Size is the number of hits to return (default 10) and From the number
// of hits to skip.
Size, From int
// Highlight adds HTML <mark> fragments of the stored field to every hit.
Highlight bool
}
Request is an analyzed text search; build it with QueryString.
func NewQueryString
deprecated
func QueryString ¶ added in v1.30.0
QueryString creates an analyzed text request against the index field. It does not parse Elasticsearch query-string syntax. Pass true to also return highlighted fragments.
func (*Request) Build ¶ added in v1.30.0
Build compiles a legacy request using the index's field and query analyzer.
func (*Request) HighlightEnabled ¶ added in v1.30.0
HighlightEnabled reports whether the request includes highlighted fragments.
type SearchRequest ¶ added in v1.30.0
type SearchRequest interface {
Build(string, *analysis.Analyzer) (*riot.TopNSearch, error)
HighlightEnabled() bool
}
SearchRequest is implemented by Request and query.Builder.
type Tokenizer ¶
type Tokenizer struct {
// contains filtered or unexported fields
}
Tokenizer adapts a loaded gse.Segmenter to the analysis.Tokenizer interface.
func NewTokenizer ¶
NewTokenizer wraps seg. With search enabled every segment is additionally expanded into its sub-words (搜索引擎 -> 搜索 索引 引擎 搜索引擎) at the same position; use it for indexing so that shorter query terms still hit.