expr

package
v0.12.0-clickbench Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 16, 2026 License: AGPL-3.0 Imports: 28 Imported by: 0

Documentation

Overview

Package expr provides a typed expression engine for evaluating SQL expressions against record batches. It replaces the string-based expression parsing with a compiled expression tree built from the SQL parser AST.

Index

Constants

This section is empty.

Variables

View Source
var DefaultRegistry = NewFuncRegistry()

DefaultRegistry is the global function registry used by the expression engine.

View Source
var DefaultUDFs = NewUDFStore()

DefaultUDFs is the global UDF store.

Functions

func RegisterFunc

func RegisterFunc(name string, fn ScalarFunc)

RegisterFunc registers a custom scalar function in the default registry.

func ToFloat64

func ToFloat64(v any) float64

ToFloat64 converts any numeric value to float64.

func ToInt64

func ToInt64(v any) int64

ToInt64 converts any numeric value to int64.

Types

type And

type And struct {
	Left, Right Expr
}

And is a logical AND.

func (*And) Eval

func (e *And) Eval(b *batch.RecordBatch, row int) any

func (*And) EvalBool

func (e *And) EvalBool(b *batch.RecordBatch, row int) bool

type ArrayLitExpr

type ArrayLitExpr struct {
	Elements []Expr
}

ArrayLitExpr evaluates to a []any containing the evaluated elements.

func (*ArrayLitExpr) Eval

func (e *ArrayLitExpr) Eval(b *batch.RecordBatch, row int) any

type Between

type Between struct {
	Expr    Expr
	Low, Hi Expr
	Not     bool
}

Between checks if a value is between two bounds.

func (*Between) Eval

func (e *Between) Eval(b *batch.RecordBatch, row int) any

func (*Between) EvalBool

func (e *Between) EvalBool(b *batch.RecordBatch, row int) bool

type BinOp

type BinOp struct {
	Left, Right Expr
	Op          string // +, -, *, /, %
}

BinOp is a binary arithmetic expression (generic, uses ToFloat64).

func (*BinOp) Eval

func (e *BinOp) Eval(b *batch.RecordBatch, row int) any

type BinOpFloat64

type BinOpFloat64 struct {
	Left, Right Float64Expr
	Op          string
	// contains filtered or unexported fields
}

BinOpFloat64 is a typed binary op that operates on float64 without boxing. Uses a pre-resolved arithOp opcode for the hot EvalFloat64 path to avoid per-row string comparison on the Op field. The opcode is resolved lazily via opOnce so external callers can construct BinOpFloat64 directly with only Op populated; concurrent pipeline workers see the same opCode after the first call returns thanks to sync.Once's happens-before guarantee.

func (*BinOpFloat64) CloneVec

func (e *BinOpFloat64) CloneVec() *BinOpFloat64

CloneVec creates a deep copy of the BinOpFloat64 tree with fresh scratch buffers. Required for parallel pipeline execution where multiple workers must not share mutable vecBuf state. Stateless leaf nodes (ColRef, Literal) are shared; only BinOpFloat64 nodes (which own vecBuf) are cloned.

func (*BinOpFloat64) Eval

func (e *BinOpFloat64) Eval(b *batch.RecordBatch, row int) any

func (*BinOpFloat64) EvalFloat64

func (e *BinOpFloat64) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

func (*BinOpFloat64) EvalFloat64Vec

func (e *BinOpFloat64) EvalFloat64Vec(b *batch.RecordBatch, dst []float64, n int) bool

EvalFloat64Vec evaluates left and right operands in bulk, then applies the arithmetic op in a tight loop. Eliminates ~5 function calls per row.

type BinOpInt64

type BinOpInt64 struct {
	Left, Right Int64Expr
	Op          string
	// contains filtered or unexported fields
}

BinOpInt64 is a typed binary op that operates on int64 without boxing. opCode is resolved lazily via opOnce so external construction with only Op populated stays safe; see BinOpFloat64 for the same pattern.

func (*BinOpInt64) Eval

func (e *BinOpInt64) Eval(b *batch.RecordBatch, row int) any

func (*BinOpInt64) EvalFloat64

func (e *BinOpInt64) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

EvalFloat64 allows BinOpInt64 to be used as Float64Expr (int→float promotion).

func (*BinOpInt64) EvalInt64

func (e *BinOpInt64) EvalInt64(b *batch.RecordBatch, row int) (int64, bool)

type BoolExpr

type BoolExpr interface {
	EvalBool(b *batch.RecordBatch, row int) bool
}

BoolExpr evaluates a boolean expression (used for WHERE/HAVING/JOIN conditions).

type Case

type Case struct {
	Operand Expr       // optional: CASE <operand> WHEN ...
	Whens   []CaseWhen // WHEN condition THEN result
	Else    Expr       // optional ELSE clause
}

Case is a CASE WHEN ... THEN ... ELSE ... END expression.

func (*Case) Eval

func (e *Case) Eval(b *batch.RecordBatch, row int) any

type CaseWhen

type CaseWhen struct {
	Cond   Expr // the condition (or value to compare against operand)
	Result Expr
}

CaseWhen is a single WHEN clause in a CASE expression.

type Cast

type Cast struct {
	Operand  Expr
	DestType string // "int", "float", "string"
}

Cast wraps an expression with explicit type conversion.

func (*Cast) Eval

func (e *Cast) Eval(b *batch.RecordBatch, row int) any

type Cmp

type Cmp struct {
	Left, Right Expr
	Op          CmpOp
}

Cmp is a comparison expression.

func (*Cmp) Eval

func (e *Cmp) Eval(b *batch.RecordBatch, row int) any

func (*Cmp) EvalBool

func (e *Cmp) EvalBool(b *batch.RecordBatch, row int) bool

type CmpFloat64

type CmpFloat64 struct {
	Left, Right Float64Expr
	Op          CmpOp
}

CmpFloat64 is a typed comparison that operates on float64 without boxing.

func (*CmpFloat64) Eval

func (e *CmpFloat64) Eval(b *batch.RecordBatch, row int) any

func (*CmpFloat64) EvalBool

func (e *CmpFloat64) EvalBool(b *batch.RecordBatch, row int) bool

type CmpInt64

type CmpInt64 struct {
	Left, Right Int64Expr
	Op          CmpOp
}

CmpInt64 is a typed comparison that operates on int64 without boxing.

func (*CmpInt64) Eval

func (e *CmpInt64) Eval(b *batch.RecordBatch, row int) any

func (*CmpInt64) EvalBool

func (e *CmpInt64) EvalBool(b *batch.RecordBatch, row int) bool

type CmpOp

type CmpOp int

CmpOp represents a comparison operator.

const (
	CmpEq CmpOp = iota
	CmpNe
	CmpLt
	CmpLe
	CmpGt
	CmpGe
)

type CmpTemporalLit

type CmpTemporalLit struct {
	Col  *ColRef
	Lit  string // original literal text (generic-fallback operand)
	Op   CmpOp
	Flip bool // literal was the LEFT operand: evaluate as (lit OP col)
	// contains filtered or unexported fields
}

CmpTemporalLit compares a bare column against a string literal that parses as a date/timestamp, without per-row parsing, cache lookups, or boxing — the generic path spent 3.2% of SF100 worker CPU inside the date-parse memo's sync.Map.Load (interface-key hashing dominated; 2026-07-25 re-rank). The literal is parsed once into BOTH temporal units at compile time; the unit is chosen from the column's resolved type per batch. Every non-fast sub-case (non-temporal column, the epoch-zero literal guard) delegates to the generic compare() with the original operand order, keeping semantics bit-identical with Cmp.

func (*CmpTemporalLit) Eval

func (e *CmpTemporalLit) Eval(b *batch.RecordBatch, row int) any

func (*CmpTemporalLit) EvalBool

func (e *CmpTemporalLit) EvalBool(b *batch.RecordBatch, row int) bool

type Coalesce

type Coalesce struct {
	Args []Expr
}

Coalesce returns the first non-null argument.

func (*Coalesce) Eval

func (e *Coalesce) Eval(b *batch.RecordBatch, row int) any

type ColRef

type ColRef struct {
	Name string
	// contains filtered or unexported fields
}

ColRef reads a column value from the batch. Caches the column index and type after first resolution for zero-allocation reads on numeric types. The cache is guarded by sync.Once so concurrent callers (parallel pipeline workers sharing this *ColRef via captured expression closures) don't race on the resolution writes.

func (*ColRef) Eval

func (e *ColRef) Eval(b *batch.RecordBatch, row int) any

func (*ColRef) EvalFloat64

func (e *ColRef) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

EvalFloat64 reads the column value as float64 without any boxing. Returns (0, false) if null or column not found. Uses cached column type to dispatch directly to the typed data slice, avoiding the extra function call and redundant type switch in GetNumericFloat64.

func (*ColRef) EvalFloat64Vec

func (e *ColRef) EvalFloat64Vec(b *batch.RecordBatch, dst []float64, n int) bool

EvalFloat64Vec evaluates the column for all rows [0, n) into dst.

func (*ColRef) EvalInt64

func (e *ColRef) EvalInt64(b *batch.RecordBatch, row int) (int64, bool)

EvalInt64 reads the column value as int64 without boxing.

func (*ColRef) EvalString

func (e *ColRef) EvalString(b *batch.RecordBatch, row int) (string, bool)

EvalString reads the column value as string without boxing.

type CorrelatedExistsSubquery

type CorrelatedExistsSubquery struct {
	Runner          SubqueryRunner
	Not             bool
	OuterRefs       []plansql.OuterRef
	OuterTables     map[string]bool
	ParsedInfo      *plansql.SelectInfo
	UnqualOuterCols map[string]string
}

CorrelatedExistsSubquery evaluates a correlated EXISTS subquery per-row.

func (*CorrelatedExistsSubquery) Eval

func (*CorrelatedExistsSubquery) EvalBool

func (e *CorrelatedExistsSubquery) EvalBool(b *batch.RecordBatch, row int) bool

type CorrelatedInSubquery

type CorrelatedInSubquery struct {
	Expr            Expr
	Runner          SubqueryRunner
	Not             bool
	OuterRefs       []plansql.OuterRef
	OuterTables     map[string]bool
	ParsedInfo      *plansql.SelectInfo
	UnqualOuterCols map[string]string
}

CorrelatedInSubquery checks if a value is in the result set of a correlated subquery.

func (*CorrelatedInSubquery) Eval

func (e *CorrelatedInSubquery) Eval(b *batch.RecordBatch, row int) any

func (*CorrelatedInSubquery) EvalBool

func (e *CorrelatedInSubquery) EvalBool(b *batch.RecordBatch, row int) bool

type CorrelatedScalarSubquery

type CorrelatedScalarSubquery struct {
	Runner          SubqueryRunner
	OuterRefs       []plansql.OuterRef // correlated column references
	OuterTables     map[string]bool    // outer table aliases
	ParsedInfo      *plansql.SelectInfo
	UnqualOuterCols map[string]string // unqualified column → table mapping for outer refs
}

CorrelatedScalarSubquery evaluates a correlated scalar subquery per-row. Unlike ScalarSubquery, it cannot cache the result because the inner query depends on values from the outer row.

func (*CorrelatedScalarSubquery) Eval

type ExistsSubquery

type ExistsSubquery struct {
	SQL    string
	Runner SubqueryRunner
	Not    bool
	// contains filtered or unexported fields
}

ExistsSubquery evaluates to true if a subquery returns any rows. Example: WHERE EXISTS (SELECT 1 FROM orders WHERE orders.user_id = users.id) Uncorrelated: executed once and result cached.

func (*ExistsSubquery) Eval

func (e *ExistsSubquery) Eval(b *batch.RecordBatch, row int) any

func (*ExistsSubquery) EvalBool

func (e *ExistsSubquery) EvalBool(_ *batch.RecordBatch, _ int) bool

type Expr

type Expr interface {
	Eval(b *batch.RecordBatch, row int) any
}

Expr evaluates an expression against a record batch row, returning a typed value.

func Compile

func Compile(node plansql.Node) (Expr, error)

Compile converts our AST Node into an Expr tree.

func CompileSelectExpr

func CompileSelectExpr(expr plansql.Node, alias string) (Expr, string, error)

CompileSelectExpr compiles a SELECT column expression from our AST. Returns the compiled expression and the output column name.

func CompileWithFullScope

func CompileWithFullScope(node plansql.Node, runner SubqueryRunner, outerTables map[string]bool, outerCols map[string]string) (Expr, error)

CompileWithFullScope is like CompileWithScope but also accepts a column-to-table mapping for resolving unqualified column references in correlated subqueries.

func CompileWithRunner

func CompileWithRunner(node plansql.Node, runner SubqueryRunner) (Expr, error)

CompileWithRunner converts our AST Node into an Expr tree, with support for subquery expressions (scalar subqueries, IN subquery, EXISTS).

func CompileWithScope

func CompileWithScope(node plansql.Node, runner SubqueryRunner, outerTables map[string]bool) (Expr, error)

CompileWithScope converts our AST Node into an Expr tree with full scope information, enabling correlated subquery detection and per-row execution. outerTables contains the table names and aliases from the outer query.

type Float64Expr

type Float64Expr interface {
	EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)
}

Float64Expr evaluates to float64 without boxing.

type FuncCall

type FuncCall struct {
	Name string
	Args []Expr
	// contains filtered or unexported fields
}

FuncCall represents a scalar function call.

Note: this struct holds NO per-call mutable state. A previous version cached an args buffer on the receiver to avoid per-call allocation, but that was unsafe under parallel pipeline execution: aggPreProject closures (and other wrapped-expression paths) capture the same *FuncCall by pointer rather than cloning it per worker, so concurrent goroutines stomped on the shared args buffer and produced non-deterministic Q02 row counts at SF0.01 (and worse at SF100). The fn / vecFn lookup caches are guarded by sync.Once so concurrent first-time lookups don't race either.

func (*FuncCall) Eval

func (e *FuncCall) Eval(b *batch.RecordBatch, row int) any

func (*FuncCall) EvalVec

func (e *FuncCall) EvalVec(b *batch.RecordBatch, out *batch.Vector, n int)

EvalVec evaluates the function for an entire batch, writing results to out. Falls back to per-row Eval if no vectorized implementation exists or if argument types can't be resolved to column vectors.

type FuncRegistry

type FuncRegistry struct {
	// contains filtered or unexported fields
}

FuncRegistry is a concurrent-safe registry of scalar functions.

func NewFuncRegistry

func NewFuncRegistry() *FuncRegistry

NewFuncRegistry creates a new empty function registry.

func (*FuncRegistry) Has

func (r *FuncRegistry) Has(name string) bool

Has returns true if a function with the given name exists.

func (*FuncRegistry) Lookup

func (r *FuncRegistry) Lookup(name string) ScalarFunc

Lookup returns the function with the given name, or nil if not found.

func (*FuncRegistry) LookupVec

func (r *FuncRegistry) LookupVec(name string) VecScalarFunc

LookupVec returns the vectorized function with the given name, or nil if not found.

func (*FuncRegistry) Names

func (r *FuncRegistry) Names() []string

Names returns all registered function names.

func (*FuncRegistry) Register

func (r *FuncRegistry) Register(name string, fn ScalarFunc)

Register adds or replaces a scalar function.

func (*FuncRegistry) RegisterVec

func (r *FuncRegistry) RegisterVec(name string, fn VecScalarFunc)

RegisterVec adds a vectorized implementation for a scalar function.

func (*FuncRegistry) RegisterVecReturn

func (r *FuncRegistry) RegisterVecReturn(name string, dimFn func() int)

RegisterVecReturn marks a function as returning a VECTOR. dimFn is evaluated lazily (at plan time) to obtain the output dimensionality — embed(), for example, derives it from the configured embedding provider.

func (*FuncRegistry) Unregister

func (r *FuncRegistry) Unregister(name string) bool

Unregister removes a scalar function. Returns true if it existed.

func (*FuncRegistry) VecReturnDim

func (r *FuncRegistry) VecReturnDim(name string) (dim int, ok bool)

VecReturnDim reports whether the named function returns a VECTOR and, if so, its current output dimension. ok is false for non-vector-returning functions.

type In

type In struct {
	Expr   Expr
	Values []Expr
	Not    bool
}

In checks if a value is in a set.

func (*In) Eval

func (e *In) Eval(b *batch.RecordBatch, row int) any

func (*In) EvalBool

func (e *In) EvalBool(b *batch.RecordBatch, row int) bool

type InSubquery

type InSubquery struct {
	Expr   Expr
	SQL    string
	Runner SubqueryRunner
	Not    bool
	// contains filtered or unexported fields
}

InSubquery checks if a value is in the result set of a subquery. Example: WHERE user_id IN (SELECT user_id FROM active_users) Uncorrelated: executed once and result set cached in a hash set for O(1) lookup.

func (*InSubquery) Eval

func (e *InSubquery) Eval(b *batch.RecordBatch, row int) any

func (*InSubquery) EvalBool

func (e *InSubquery) EvalBool(b *batch.RecordBatch, row int) bool

type Int64Expr

type Int64Expr interface {
	EvalInt64(b *batch.RecordBatch, row int) (int64, bool)
}

Int64Expr evaluates to int64 without boxing.

type IntervalValue

type IntervalValue struct {
	Years   int
	Months  int
	Days    int
	Hours   int
	Minutes int
	Seconds int
}

IntervalValue represents a SQL INTERVAL (e.g., INTERVAL '30' DAY).

type IsNull

type IsNull struct {
	Operand Expr
	Not     bool // IS NOT NULL
}

IsNull checks if an expression is null.

func (*IsNull) Eval

func (e *IsNull) Eval(b *batch.RecordBatch, row int) any

func (*IsNull) EvalBool

func (e *IsNull) EvalBool(b *batch.RecordBatch, row int) bool

type Like

type Like struct {
	Expr    Expr
	Pattern Expr
	Not     bool
}

Like performs SQL LIKE pattern matching.

func (*Like) Eval

func (e *Like) Eval(b *batch.RecordBatch, row int) any

func (*Like) EvalBool

func (e *Like) EvalBool(b *batch.RecordBatch, row int) bool

type Lit

type Lit struct {
	Val any
}

Lit returns a constant value.

func (*Lit) Eval

func (e *Lit) Eval(_ *batch.RecordBatch, _ int) any

func (*Lit) EvalFloat64

func (e *Lit) EvalFloat64(_ *batch.RecordBatch, _ int) (float64, bool)

func (*Lit) EvalFloat64Vec

func (e *Lit) EvalFloat64Vec(_ *batch.RecordBatch, dst []float64, n int) bool

EvalFloat64Vec fills dst[0:n] with the literal value.

func (*Lit) EvalInt64

func (e *Lit) EvalInt64(_ *batch.RecordBatch, _ int) (int64, bool)

type Not

type Not struct {
	Operand Expr
}

Not is a logical NOT.

func (*Not) Eval

func (e *Not) Eval(b *batch.RecordBatch, row int) any

func (*Not) EvalBool

func (e *Not) EvalBool(b *batch.RecordBatch, row int) bool

type Or

type Or struct {
	Left, Right Expr
}

Or is a logical OR.

func (*Or) Eval

func (e *Or) Eval(b *batch.RecordBatch, row int) any

func (*Or) EvalBool

func (e *Or) EvalBool(b *batch.RecordBatch, row int) bool

type ParamRef

type ParamRef struct {
	Index int
	// contains filtered or unexported fields
}

ParamRef is an expression node that references a UDF parameter by index.

func (*ParamRef) Eval

func (e *ParamRef) Eval(_ *batch.RecordBatch, _ int) any

type ScalarFunc

type ScalarFunc func(args []any) any

ScalarFunc is a scalar function implementation.

type ScalarSubquery

type ScalarSubquery struct {
	SQL    string
	Runner SubqueryRunner
	// contains filtered or unexported fields
}

ScalarSubquery evaluates a subquery that returns a single scalar value. Example: WHERE price > (SELECT AVG(price) FROM products) Uncorrelated: executed once and result cached.

func (*ScalarSubquery) Eval

func (e *ScalarSubquery) Eval(_ *batch.RecordBatch, _ int) any

type SubqueryRunner

type SubqueryRunner func(sql string) ([]map[string]any, error)

SubqueryRunner executes a SQL subquery and returns its result rows. Each row is a map of column name to value.

type UDFCall

type UDFCall struct {
	Name     string
	ArgExprs []Expr // caller-supplied argument expressions
	Body     Expr   // compiled UDF body with ParamRef nodes
	// contains filtered or unexported fields
}

UDFCall evaluates a user-defined function by binding arguments, then evaluating the compiled body expression.

func (*UDFCall) Eval

func (e *UDFCall) Eval(b *batch.RecordBatch, row int) any

type UDFDef

type UDFDef struct {
	Name   string   // function name (lowercase)
	Params []string // parameter names (lowercase)
	Body   string   // SQL expression body (e.g. "param1 * 2 + param2")
	Owner  string   // who created this function (empty = system/unowned)
	Locked bool     // if true, only the owner (or admin) can modify/drop
}

UDFDef defines a user-defined function.

type UDFPersister

type UDFPersister func(udfs []UDFDef) error

UDFPersister is called after UDF register/unregister to persist the current state.

type UDFStore

type UDFStore struct {
	// contains filtered or unexported fields
}

UDFStore holds compiled UDF definitions for use by the expression engine. Thread-safe for concurrent reads and writes.

func NewUDFStore

func NewUDFStore() *UDFStore

NewUDFStore creates a new empty UDF store.

func (*UDFStore) CompileUDFCall

func (s *UDFStore) CompileUDFCall(name string, argExprs []Expr) (Expr, error)

CompileUDFCall creates a UDFCall expression node.

func (*UDFStore) Get

func (s *UDFStore) Get(name string) (UDFDef, bool)

Get returns a UDF definition by name.

func (*UDFStore) List

func (s *UDFStore) List() []UDFDef

List returns all registered UDF definitions.

func (*UDFStore) LoadDefs

func (s *UDFStore) LoadDefs(defs []UDFDef) int

LoadDefs registers pre-existing UDF definitions (e.g., from KV on startup). Skips compilation errors silently so one bad UDF doesn't block startup.

func (*UDFStore) Register

func (s *UDFStore) Register(def UDFDef, isAdmin bool) error

Register compiles and registers a UDF.

func (*UDFStore) SetPersister

func (s *UDFStore) SetPersister(p UDFPersister)

SetPersister sets the function called after UDF mutations to persist state.

func (*UDFStore) Unregister

func (s *UDFStore) Unregister(name, caller string, isAdmin bool) error

Unregister removes a UDF.

type UnaryOp

type UnaryOp struct {
	Operand Expr
	Op      string // -, +
}

UnaryOp is a unary arithmetic expression (negation).

func (*UnaryOp) Eval

func (e *UnaryOp) Eval(b *batch.RecordBatch, row int) any

type VecExpr

type VecExpr interface {
	EvalVec(b *batch.RecordBatch, out *batch.Vector, n int)
}

VecExpr evaluates an expression for an entire batch at once, writing results directly to the output vector. This avoids per-row interface dispatch and boxing.

type VecFloat64Expr

type VecFloat64Expr interface {
	EvalFloat64Vec(b *batch.RecordBatch, dst []float64, n int) bool
}

VecFloat64Expr evaluates an expression for all rows [0, n) at once, writing results to dst. Returns true if any output is null. Eliminates per-row function call overhead (~5 calls/row/expression).

type VecScalarFunc

type VecScalarFunc func(args []*batch.Vector, out *batch.Vector, n int)

VecScalarFunc is a vectorized scalar function that operates on entire columns at once, reading from input vectors and writing to an output vector. This avoids per-row interface dispatch and boxing overhead.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL