Documentation
¶
Overview ¶
Package gse provides a dictionary based tokenizer for Chinese, Japanese and other languages supported by github.com/go-ego/gse, plus a small gse-bleve style index API (New, Index, Search) built on riot.
Index ¶
- func Langs() []string
- func NewAnalyzer(seg *gogse.Segmenter, opt Option) (*analysis.Analyzer, error)
- func NewLangAnalyzer(code string) (*analysis.Analyzer, error)
- func NewSegmenter(opt Option) (*gogse.Segmenter, error)
- type Hit
- type Index
- func (x *Index) Close() error
- func (x *Index) Delete(id string) error
- func (x *Index) Field(name, text string) *riot.TermField
- func (x *Index) Index(id, text string) error
- func (x *Index) Search(req *Request) (*Result, error)
- func (x *Index) Segmenter() *gogse.Segmenter
- func (x *Index) Writer() *riot.Writer
- type Option
- type Request
- type Result
- type Tokenizer
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
func NewAnalyzer ¶
NewAnalyzer builds an analyzer over an already loaded segmenter using opt.Opt as cut mode, so one segmenter can back both an index ("search-hmm") and a query ("hmm") analyzer. With a nil seg the analysis/lang analyzer named by opt.Lang is returned instead.
func NewLangAnalyzer ¶
NewLangAnalyzer returns the analysis/lang analyzer for code.
Types ¶
type Hit ¶
type Hit struct {
ID string
Score float64
// Fields holds the stored field values of the document.
Fields map[string]string
// Fragments holds highlighted snippets per field when requested.
Fragments map[string][]string
}
Hit is one matched document.
type Index ¶
type Index struct {
// contains filtered or unexported fields
}
Index is a riot writer with gse (or analysis/lang) analyzers attached, in the spirit of gse-bleve: New, Index, Search, Close.
func New ¶
New opens the index at opt.Index (in-memory when empty). With opt.Lang set the matching analysis/lang analyzer is used for both indexing and queries; otherwise the gse dictionaries are loaded, documents are cut with opt.Opt and queries with its non-search counterpart ("search-hmm" -> "hmm").
func (*Index) Field ¶
Field builds a stored, highlightable text field cut by the index analyzer; use it to add extra gse fields to documents written through Writer.
func (*Index) Index ¶
Index writes text under id into the configured field, replacing any existing document with that id.
type Option ¶
type Option struct {
// Index is the on-disk index path; empty opens an in-memory index.
Index string
// Field is the document field written by Index.Index; default "text".
Field string
// Lang selects a riot analysis/lang analyzer ("en", "cjk", "de", ...;
// see Langs) instead of gse. When set, no gse dictionary is loaded, the
// Segmenter is nil and Dicts/Stop/Opt/Alpha are ignored.
Lang string
// Dicts selects the dictionaries: "zh", "zh_s", "zh_t", "ja"/"jp" or
// comma separated dictionary file paths. Prefix with "embed, " to use the
// dictionaries compiled into gse instead of reading files from the gse
// module directory. Empty loads the embedded "zh" dictionary.
Dicts string
// Stop selects the stop word dictionary: "zh", "embed, zh" or file paths.
// Empty loads none. Stop words are dropped from the token stream.
Stop string
// Opt is the cut mode: "" (shortest path), "hmm", "dag", "search",
// "search-hmm" or "search-dag". The search modes additionally emit the
// sub-words of every segment and are meant for indexing; the Index
// built by New queries with the matching non-search mode.
Opt string
// Alpha makes gse emit every Latin letter/digit as its own token.
Alpha bool
}
Option configures the segmenter, the cut mode and the index.
type Request ¶
type Request struct {
Query string
// Field to match; empty uses the field configured in Option.
Field string
// Size is the number of hits to return (default 10) and From the number
// of hits to skip.
Size, From int
// Highlight adds HTML <mark> fragments of the stored field to every hit.
Highlight bool
}
Request is a query string search; build it with NewQueryString.
func NewQueryString ¶
NewQueryString creates a request matching query against the index field. Pass true to also return highlighted fragments.
type Tokenizer ¶
type Tokenizer struct {
// contains filtered or unexported fields
}
Tokenizer adapts a loaded gse.Segmenter to the analysis.Tokenizer interface.
func NewTokenizer ¶
NewTokenizer wraps seg. With search enabled every segment is additionally expanded into its sub-words (搜索引擎 -> 搜索 索引 引擎 搜索引擎) at the same position; use it for indexing so that shorter query terms still hit.