expr

package
v0.16.0-correctness Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 21, 2026 License: AGPL-3.0 Imports: 33 Imported by: 0

Documentation

Overview

Package expr provides a typed expression engine for evaluating SQL expressions against record batches. It replaces the string-based expression parsing with a compiled expression tree built from the SQL parser AST.

Index

Constants

View Source
const (
	// SessionUser is the user name reported by current_user / session_user /
	// user / current_role.
	SessionUser = "wadjet"
	// SessionCatalog is the database name reported by current_catalog /
	// current_database().
	SessionCatalog = "wadjet"
	// SessionSchema is the schema reported by current_schema.
	SessionSchema = "public"
	// ServerVersion is the answer to version(). PostgreSQL drivers parse the
	// leading "PostgreSQL <major>" to decide which protocol features and
	// catalog queries they may use, so the string keeps that prefix.
	ServerVersion = "PostgreSQL 15.0 (Wadjet analytical query engine)"
)

--- Session / catalog information functions ---

PostgreSQL clients (pgJDBC, DataGrip, psql, Superset) open a connection by asking who and where they are: current_user, current_schema, current_database. These are answered here rather than only in the pgwire introspection layer so that a query mixing them with real columns — or selecting three of them at once — executes as an ordinary query with an ordinary result shape.

The values are server constants. ScalarFunc is func([]any) any and DefaultRegistry is process-global, so a scalar function cannot see the calling connection's identity; a per-session answer would need a context-carrying evaluation path that does not exist. The constants match what pgwire reports for an unauthenticated session.

Variables

View Source
var (
	RetBool      = Ret{/* contains filtered or unexported fields */}
	RetInt64     = Ret{/* contains filtered or unexported fields */}
	RetFloat64   = Ret{/* contains filtered or unexported fields */}
	RetString    = Ret{/* contains filtered or unexported fields */}
	RetBytes     = Ret{/* contains filtered or unexported fields */}
	RetArray     = Ret{/* contains filtered or unexported fields */}
	RetMap       = Ret{/* contains filtered or unexported fields */}
	RetTimestamp = Ret{/* contains filtered or unexported fields */}

	// RetVector is embed()'s declaration. The output *dimension* is a
	// separate, deliberately dynamic answer the registry already carried
	// before this type existed — see RegisterVecReturn / VecReturnDim.
	RetVector = Ret{/* contains filtered or unexported fields */}

	// RetDynamic declares that only the value knows: element_at returns the
	// element type of its argument, json_extract whatever the document held.
	// The planner keeps its own fallback for these. It is an explicit
	// declaration, not an omission — a function whose vec kernel writes a
	// typed slice must never carry it.
	RetDynamic = Ret{/* contains filtered or unexported fields */}
)

The fixed declarations. These name the type the function's Go results are stored as, not the type SQL calls them: the date/time functions below return formatted strings, so they declare RetString.

View Source
var DefaultRegistry = NewFuncRegistry()

DefaultRegistry is the global function registry used by the expression engine.

View Source
var DefaultUDFs = NewUDFStore()

DefaultUDFs is the global UDF store.

Functions

func IntArithOn

func IntArithOn() bool

IntArithOn exposes the toggle to the planner: projection output types may only declare Int64 for arithmetic when the runtime will actually take the integer path (see inferProjectionTypeCols).

func IsUnknownFunc

func IsUnknownFunc(err error) bool

IsUnknownFunc reports whether err is, or wraps, an UnknownFuncError.

func RegisterFunc

func RegisterFunc(name string, fn ScalarFunc, ret Ret)

RegisterFunc registers a custom scalar function in the default registry. ret declares what the function returns; see Ret.

func ToFloat64

func ToFloat64(v any) float64

ToFloat64 converts any numeric value to float64.

func ToInt64

func ToInt64(v any) int64

ToInt64 converts any numeric value to int64.

Types

type And

type And struct {
	Left, Right Expr
}

And is a logical AND.

func (*And) Eval

func (e *And) Eval(b *batch.RecordBatch, row int) any

func (*And) EvalBool

func (e *And) EvalBool(b *batch.RecordBatch, row int) bool

func (*And) EvalBoolNull

func (e *And) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: FALSE AND anything is FALSE; otherwise a NULL operand makes it UNKNOWN. Short-circuits on a FALSE left operand.

type ArrayLitExpr

type ArrayLitExpr struct {
	Elements []Expr
}

ArrayLitExpr evaluates to a []any containing the evaluated elements.

func (*ArrayLitExpr) Eval

func (e *ArrayLitExpr) Eval(b *batch.RecordBatch, row int) any

type Between

type Between struct {
	Expr    Expr
	Low, Hi Expr
	Not     bool
}

Between checks if a value is between two bounds.

func (*Between) Eval

func (e *Between) Eval(b *batch.RecordBatch, row int) any

func (*Between) EvalBool

func (e *Between) EvalBool(b *batch.RecordBatch, row int) bool

func (*Between) EvalBoolNull

func (e *Between) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: BETWEEN is defined as (x >= lo AND x <= hi), so a NULL bound does not force UNKNOWN — the other half can still answer FALSE (`5 BETWEEN NULL AND 2` is false, and NOT BETWEEN flips it to true).

type BinOp

type BinOp struct {
	Left, Right Expr
	Op          string // +, -, *, /, %
}

BinOp is a binary arithmetic expression (generic, uses ToFloat64).

func (*BinOp) Eval

func (e *BinOp) Eval(b *batch.RecordBatch, row int) any

type BinOpFloat64

type BinOpFloat64 struct {
	Left, Right Float64Expr
	Op          string
	// contains filtered or unexported fields
}

BinOpFloat64 is a typed binary op that operates on float64 without boxing. Uses a pre-resolved arithOp opcode for the hot EvalFloat64 path to avoid per-row string comparison on the Op field. The opcode is resolved lazily via opOnce so external callers can construct BinOpFloat64 directly with only Op populated; concurrent pipeline workers see the same opCode after the first call returns thanks to sync.Once's happens-before guarantee.

func (*BinOpFloat64) CloneVec

func (e *BinOpFloat64) CloneVec() *BinOpFloat64

CloneVec creates a deep copy of the BinOpFloat64 tree with fresh scratch buffers. Required for parallel pipeline execution where multiple workers must not share mutable vecBuf state. Stateless leaf nodes (ColRef, Literal) are shared; only BinOpFloat64 nodes (which own vecBuf) are cloned.

func (*BinOpFloat64) Eval

func (e *BinOpFloat64) Eval(b *batch.RecordBatch, row int) any

func (*BinOpFloat64) EvalFloat64

func (e *BinOpFloat64) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

func (*BinOpFloat64) EvalFloat64Vec

func (e *BinOpFloat64) EvalFloat64Vec(b *batch.RecordBatch, dst []float64, n int) bool

EvalFloat64Vec evaluates left and right operands in bulk, then applies the arithmetic op in a tight loop. Eliminates ~5 function calls per row.

type BinOpInt64

type BinOpInt64 struct {
	Left, Right Int64Expr
	Op          string
	// contains filtered or unexported fields
}

BinOpInt64 is a typed binary op that operates on int64 without boxing. opCode is resolved lazily via opOnce so external construction with only Op populated stays safe; see BinOpFloat64 for the same pattern.

func (*BinOpInt64) Eval

func (e *BinOpInt64) Eval(b *batch.RecordBatch, row int) any

func (*BinOpInt64) EvalFloat64

func (e *BinOpInt64) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

EvalFloat64 allows BinOpInt64 to be used as Float64Expr (int→float promotion).

func (*BinOpInt64) EvalInt64

func (e *BinOpInt64) EvalInt64(b *batch.RecordBatch, row int) (int64, bool)

type BinOpNumeric

type BinOpNumeric struct {
	Left, Right numericOperand
	Op          string
	// contains filtered or unexported fields
}

BinOpNumeric is the mode-resolved arithmetic node.

func (*BinOpNumeric) Eval

func (e *BinOpNumeric) Eval(b *batch.RecordBatch, row int) any

func (*BinOpNumeric) EvalFloat64

func (e *BinOpNumeric) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

EvalFloat64 implements Float64Expr for consumers on the float protocol.

func (*BinOpNumeric) EvalInt64

func (e *BinOpNumeric) EvalInt64(b *batch.RecordBatch, row int) (int64, bool)

EvalInt64 implements Int64Expr. Only meaningful in int mode; float mode reports not-ok so callers fall back to EvalFloat64/Eval.

type BoolExpr

type BoolExpr interface {
	EvalBool(b *batch.RecordBatch, row int) bool
}

BoolExpr evaluates a boolean expression (used for WHERE/HAVING/JOIN conditions). SQL's logic is THREE-valued and EvalBool is the two-valued COLLAPSE a filtering context applies: it answers true only for TRUE — FALSE and UNKNOWN rows are both kept out of a WHERE. The third value is carried by BoolNullExpr, and the two protocols must agree: EvalBool ≡ (val && !null) of EvalBoolNull.

type BoolNullExpr

type BoolNullExpr interface {
	EvalBoolNull(b *batch.RecordBatch, row int) (val, null bool)
}

BoolNullExpr is the three-valued boolean protocol (#370): val is the answer and null reports UNKNOWN, in which case val is meaningless. Every boolean operator implements it — it is what lets NOT distinguish UNKNOWN (stays UNKNOWN, row excluded) from FALSE (becomes TRUE, row kept), and what a projection boxes into SQL NULL.

type Case

type Case struct {
	Operand Expr       // optional: CASE <operand> WHEN ...
	Whens   []CaseWhen // WHEN condition THEN result
	Else    Expr       // optional ELSE clause
}

Case is a CASE WHEN ... THEN ... ELSE ... END expression.

func (*Case) Eval

func (e *Case) Eval(b *batch.RecordBatch, row int) any

type CaseWhen

type CaseWhen struct {
	Cond   Expr // the condition (or value to compare against operand)
	Result Expr
}

CaseWhen is a single WHEN clause in a CASE expression.

type Cast

type Cast struct {
	Operand  Expr
	DestType string // "int", "float", "string", "date", "timestamp"
}

Cast wraps an expression with explicit type conversion.

func (*Cast) Eval

func (e *Cast) Eval(b *batch.RecordBatch, row int) any

type Cmp

type Cmp struct {
	Left, Right Expr
	Op          CmpOp
}

Cmp is a comparison expression.

func (*Cmp) Eval

func (e *Cmp) Eval(b *batch.RecordBatch, row int) any

func (*Cmp) EvalBool

func (e *Cmp) EvalBool(b *batch.RecordBatch, row int) bool

func (*Cmp) EvalBoolNull

func (e *Cmp) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

type CmpFloat64

type CmpFloat64 struct {
	Left, Right Float64Expr
	Op          CmpOp
}

CmpFloat64 is a typed comparison that operates on float64 without boxing.

func (*CmpFloat64) Eval

func (e *CmpFloat64) Eval(b *batch.RecordBatch, row int) any

func (*CmpFloat64) EvalBool

func (e *CmpFloat64) EvalBool(b *batch.RecordBatch, row int) bool

func (*CmpFloat64) EvalBoolNull

func (e *CmpFloat64) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: a not-ok typed operand is a NULL (the operands here are provably float-typed at compile time), so the comparison is UNKNOWN.

type CmpInt64

type CmpInt64 struct {
	Left, Right Int64Expr
	Op          CmpOp
}

CmpInt64 is a typed comparison that operates on int64 without boxing.

func (*CmpInt64) Eval

func (e *CmpInt64) Eval(b *batch.RecordBatch, row int) any

func (*CmpInt64) EvalBool

func (e *CmpInt64) EvalBool(b *batch.RecordBatch, row int) bool

func (*CmpInt64) EvalBoolNull

func (e *CmpInt64) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: a not-ok typed operand is a NULL (the operands here are provably int-typed at compile time), so the comparison is UNKNOWN.

type CmpOp

type CmpOp int

CmpOp represents a comparison operator.

const (
	CmpEq CmpOp = iota
	CmpNe
	CmpLt
	CmpLe
	CmpGt
	CmpGe
)

type CmpTemporalLit

type CmpTemporalLit struct {
	Col  *ColRef
	Lit  string // original literal text (generic-fallback operand)
	Op   CmpOp
	Flip bool // literal was the LEFT operand: evaluate as (lit OP col)
	// contains filtered or unexported fields
}

CmpTemporalLit compares a bare column against a string literal that parses as a date/timestamp, without per-row parsing, cache lookups, or boxing — the generic path spent 3.2% of SF100 worker CPU inside the date-parse memo's sync.Map.Load (interface-key hashing dominated; 2026-07-25 re-rank). The literal is parsed once into BOTH temporal units at compile time; the unit is chosen from the column's resolved type per batch. Every non-fast sub-case (non-temporal column, the epoch-zero literal guard) delegates to the generic compare() with the original operand order, keeping semantics bit-identical with Cmp.

func (*CmpTemporalLit) Eval

func (e *CmpTemporalLit) Eval(b *batch.RecordBatch, row int) any

func (*CmpTemporalLit) EvalBool

func (e *CmpTemporalLit) EvalBool(b *batch.RecordBatch, row int) bool

func (*CmpTemporalLit) EvalBoolNull

func (e *CmpTemporalLit) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

type Coalesce

type Coalesce struct {
	Args []Expr
}

Coalesce returns the first non-null argument.

func (*Coalesce) Eval

func (e *Coalesce) Eval(b *batch.RecordBatch, row int) any

type ColEmptyStr

type ColEmptyStr struct {
	Col      *ColRef
	Not      bool // true for <>
	Fallback *Cmp
}

ColEmptyStr evaluates a column compared for equality or inequality against the empty string literal as a zero-length offsets test. Restricted to TypeString: that is the only type for which the generic Cmp path compares ColRef.Eval's boxed string against "" (a TypeBytes column boxes []byte, which compare() handles through a different branch, and the network/UUID types render their bytes).

NULL handling matches Cmp exactly: a NULL operand makes the comparison UNKNOWN for BOTH = and <> — nil on the boxed path, excluded by EvalBool.

func (*ColEmptyStr) Eval

func (e *ColEmptyStr) Eval(b *batch.RecordBatch, row int) any

func (*ColEmptyStr) EvalBool

func (e *ColEmptyStr) EvalBool(b *batch.RecordBatch, row int) bool

func (*ColEmptyStr) EvalBoolNull

func (e *ColEmptyStr) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

type ColIsNull

type ColIsNull struct {
	Col      *ColRef
	Not      bool
	Fallback *IsNull
}

ColIsNull evaluates `col IS [NOT] NULL` off the null bitmap for a byte-array column, where the generic IsNull node boxes (and for TypeString copies) the value only to test it against nil.

Scoped to TypeString/TypeBytes deliberately: ColRef.Eval returns nil for exactly the null rows of those two types (TypeString via GetString's ok flag, TypeBytes via GetValue's leading null check), so the rewrite is value-identical. Other types keep the generic node.

func (*ColIsNull) Eval

func (e *ColIsNull) Eval(b *batch.RecordBatch, row int) any

func (*ColIsNull) EvalBool

func (e *ColIsNull) EvalBool(b *batch.RecordBatch, row int) bool

func (*ColIsNull) EvalBoolNull

func (e *ColIsNull) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: IS [NOT] NULL never answers UNKNOWN.

type ColRef

type ColRef struct {
	Name string
	// contains filtered or unexported fields
}

ColRef reads a column value from the batch. Caches the column index and type after first resolution for zero-allocation reads on numeric types. The cache is guarded by sync.Once so concurrent callers (parallel pipeline workers sharing this *ColRef via captured expression closures) don't race on the resolution writes.

func (*ColRef) Eval

func (e *ColRef) Eval(b *batch.RecordBatch, row int) any

func (*ColRef) EvalFloat64

func (e *ColRef) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

EvalFloat64 reads the column value as float64 without any boxing. Returns (0, false) if null or column not found. Uses cached column type to dispatch directly to the typed data slice, avoiding the extra function call and redundant type switch in GetNumericFloat64.

func (*ColRef) EvalFloat64Vec

func (e *ColRef) EvalFloat64Vec(b *batch.RecordBatch, dst []float64, n int) bool

EvalFloat64Vec evaluates the column for all rows [0, n) into dst.

func (*ColRef) EvalInt64

func (e *ColRef) EvalInt64(b *batch.RecordBatch, row int) (int64, bool)

EvalInt64 reads the column value as int64 without boxing.

func (*ColRef) EvalString

func (e *ColRef) EvalString(b *batch.RecordBatch, row int) (string, bool)

EvalString reads the column value as string without boxing.

type ColShapeLen

type ColShapeLen struct {
	Col      *ColRef
	Mul      int // 1 for length/octet_length, 8 for bit_length
	Fallback *FuncCall
}

ColShapeLen evaluates length()/octet_length()/bit_length() over a bare column reference by subtracting offsets, never materializing the value. Any column whose stored bytes are not the value ColRef.Eval would box (numeric, temporal, network-rendered, ROW field access) delegates to the generic FuncCall it replaced, so results are unchanged.

func (*ColShapeLen) Eval

func (e *ColShapeLen) Eval(b *batch.RecordBatch, row int) any

func (*ColShapeLen) EvalFloat64

func (e *ColShapeLen) EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)

func (*ColShapeLen) EvalFloat64Vec

func (e *ColShapeLen) EvalFloat64Vec(b *batch.RecordBatch, dst []float64, n int) bool

EvalFloat64Vec fills dst for rows [0, n), reporting whether any row was null (the VecFloat64Expr contract: the caller re-runs EvalFloat64 per row to set the null bits when this returns true).

func (*ColShapeLen) EvalInt64

func (e *ColShapeLen) EvalInt64(b *batch.RecordBatch, row int) (int64, bool)

func (*ColShapeLen) EvalVec

func (e *ColShapeLen) EvalVec(b *batch.RecordBatch, out *batch.Vector, n int)

EvalVec fills out for the whole batch. Mirrors FuncCall.EvalVec's contract: writes Float64Data, marking nulls in out.Nulls.

type Confidence

type Confidence uint8

Confidence says how a resolved type was arrived at: whether the declaration DECIDED it or only GUESSED it. A same-as-argument declaration has to answer even when none of its candidate arguments decided anything, and that answer — its fallback — is a guess. Reporting a guess as fact is what typed

SELECT COALESCE(NULLIF(n_name, 'ALGERIA'), 'fallback') FROM nation

Float64, so every row came back as the integer 0: nullif's argument 0 is a bare column, which decides nothing by design (its type comes from the input schema at runtime), so nullif fell back to its numeric default — and coalesce took that for a decision, stopped, and never consulted the string literal in argument 1 that would have decided it correctly (#331).

The fallback itself is right where there is nothing better: NULLIF(int_col, 1) as a projection is numeric and stays numeric. What Confidence adds is that a caller holding another candidate can tell the two apart.

const (
	// Undecided: nothing here names a type, and the caller keeps its own
	// fallback. RetDynamic answers this way, as does an unregistered name.
	Undecided Confidence = iota
	// Guessed: a polymorphic declaration reached its fallback because no
	// candidate argument decided. Still an answer — it is THE answer at top
	// level — but a caller with a candidate of its own left to ask must
	// prefer that candidate's decision over this.
	Guessed
	// Decided: the declaration names this type outright, or a candidate
	// argument decided it.
	Decided
)

func (Confidence) String

func (c Confidence) String() string

type CorrelatedExistsSubquery

type CorrelatedExistsSubquery struct {
	Runner          SubqueryRunner
	Not             bool
	OuterRefs       []plansql.OuterRef
	OuterTables     map[string]bool
	ParsedInfo      *plansql.SelectInfo
	UnqualOuterCols map[string]string
}

CorrelatedExistsSubquery evaluates a correlated EXISTS subquery per-row.

func (*CorrelatedExistsSubquery) Eval

func (*CorrelatedExistsSubquery) EvalBool

func (e *CorrelatedExistsSubquery) EvalBool(b *batch.RecordBatch, row int) bool

type CorrelatedInSubquery

type CorrelatedInSubquery struct {
	Expr            Expr
	Runner          SubqueryRunner
	Not             bool
	OuterRefs       []plansql.OuterRef
	OuterTables     map[string]bool
	ParsedInfo      *plansql.SelectInfo
	UnqualOuterCols map[string]string
}

CorrelatedInSubquery checks if a value is in the result set of a correlated subquery.

func (*CorrelatedInSubquery) Eval

func (e *CorrelatedInSubquery) Eval(b *batch.RecordBatch, row int) any

func (*CorrelatedInSubquery) EvalBool

func (e *CorrelatedInSubquery) EvalBool(b *batch.RecordBatch, row int) bool

func (*CorrelatedInSubquery) EvalBoolNull

func (e *CorrelatedInSubquery) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull carries SQL's three-valued IN (#370): a NULL probe is UNKNOWN, and a miss against a result set containing a NULL is UNKNOWN — the NOT IN trap, same rule as the uncorrelated InSubquery.

type CorrelatedScalarSubquery

type CorrelatedScalarSubquery struct {
	Runner          SubqueryRunner
	OuterRefs       []plansql.OuterRef // correlated column references
	OuterTables     map[string]bool    // outer table aliases
	ParsedInfo      *plansql.SelectInfo
	UnqualOuterCols map[string]string // unqualified column → table mapping for outer refs
}

CorrelatedScalarSubquery evaluates a correlated scalar subquery per-row. Unlike ScalarSubquery, it cannot cache the result because the inner query depends on values from the outer row.

func (*CorrelatedScalarSubquery) Eval

type ExistsSubquery

type ExistsSubquery struct {
	SQL    string
	Runner SubqueryRunner
	Not    bool
	// contains filtered or unexported fields
}

ExistsSubquery evaluates to true if a subquery returns any rows. Example: WHERE EXISTS (SELECT 1 FROM orders WHERE orders.user_id = users.id) Uncorrelated: executed once and result cached.

func (*ExistsSubquery) Eval

func (e *ExistsSubquery) Eval(b *batch.RecordBatch, row int) any

func (*ExistsSubquery) EvalBool

func (e *ExistsSubquery) EvalBool(_ *batch.RecordBatch, _ int) bool

type Expr

type Expr interface {
	Eval(b *batch.RecordBatch, row int) any
}

Expr evaluates an expression against a record batch row, returning a typed value.

func Compile

func Compile(node plansql.Node) (Expr, error)

Compile converts our AST Node into an Expr tree.

func CompileSelectExpr

func CompileSelectExpr(expr plansql.Node, alias string) (Expr, string, error)

CompileSelectExpr compiles a SELECT column expression from our AST. Returns the compiled expression and the output column name.

func CompileWithFullScope

func CompileWithFullScope(node plansql.Node, runner SubqueryRunner, outerTables map[string]bool, outerCols map[string]string) (Expr, error)

CompileWithFullScope is like CompileWithScope but also accepts a column-to-table mapping for resolving unqualified column references in correlated subqueries.

func CompileWithRunner

func CompileWithRunner(node plansql.Node, runner SubqueryRunner) (Expr, error)

CompileWithRunner converts our AST Node into an Expr tree, with support for subquery expressions (scalar subqueries, IN subquery, EXISTS).

func CompileWithScope

func CompileWithScope(node plansql.Node, runner SubqueryRunner, outerTables map[string]bool) (Expr, error)

CompileWithScope converts our AST Node into an Expr tree with full scope information, enabling correlated subquery detection and per-row execution. outerTables contains the table names and aliases from the outer query.

func CompileWithScopeResolver

func CompileWithScopeResolver(node plansql.Node, runner SubqueryRunner, outerTables map[string]bool, outerCols map[string]string, innerCols plansql.TableColumns) (Expr, error)

CompileWithScopeResolver is CompileWithFullScope plus a resolver for the column namespace of a subquery's own FROM clause. It is what makes an unqualified name inside a subquery bind to the subquery first, so a name that merely also exists in the outer query does not turn an uncorrelated subquery into a per-row correlated one (issue #334). A nil resolver keeps the weaker table-identifier heuristic.

type Float64Expr

type Float64Expr interface {
	EvalFloat64(b *batch.RecordBatch, row int) (float64, bool)
}

Float64Expr evaluates to float64 without boxing.

type FuncCall

type FuncCall struct {
	Name string
	Args []Expr
	// contains filtered or unexported fields
}

FuncCall represents a scalar function call.

Note: this struct holds NO per-call mutable state. A previous version cached an args buffer on the receiver to avoid per-call allocation, but that was unsafe under parallel pipeline execution: aggPreProject closures (and other wrapped-expression paths) capture the same *FuncCall by pointer rather than cloning it per worker, so concurrent goroutines stomped on the shared args buffer and produced non-deterministic Q02 row counts at SF0.01 (and worse at SF100). The fn / vecFn lookup caches are guarded by sync.Once so concurrent first-time lookups don't race either.

func (*FuncCall) Eval

func (e *FuncCall) Eval(b *batch.RecordBatch, row int) any

func (*FuncCall) EvalVec

func (e *FuncCall) EvalVec(b *batch.RecordBatch, out *batch.Vector, n int)

EvalVec evaluates the function for an entire batch, writing results to out. Falls back to per-row Eval if no vectorized implementation exists or if argument types can't be resolved to column vectors.

type FuncRegistry

type FuncRegistry struct {
	// contains filtered or unexported fields
}

FuncRegistry is a concurrent-safe registry of scalar functions.

func NewFuncRegistry

func NewFuncRegistry() *FuncRegistry

NewFuncRegistry creates a new empty function registry.

func (*FuncRegistry) Has

func (r *FuncRegistry) Has(name string) bool

Has returns true if a function with the given name exists.

func (*FuncRegistry) Lookup

func (r *FuncRegistry) Lookup(name string) ScalarFunc

Lookup returns the function with the given name, or nil if not found.

func (*FuncRegistry) LookupVec

func (r *FuncRegistry) LookupVec(name string) VecScalarFunc

LookupVec returns the vectorized function with the given name, or nil if not found.

func (*FuncRegistry) Names

func (r *FuncRegistry) Names() []string

Names returns all registered function names.

func (*FuncRegistry) Register

func (r *FuncRegistry) Register(name string, fn ScalarFunc, ret Ret)

Register adds or replaces a scalar function. ret declares the type the function's results are stored as; the planner types projections from it (see Ret). Registering without a declaration does not compile, and registering the zero value panics here rather than letting a mistyped output vector reach a kernel.

func (*FuncRegistry) RegisterVec

func (r *FuncRegistry) RegisterVec(name string, fn VecScalarFunc)

RegisterVec adds a vectorized implementation for a scalar function. A vec kernel writes a typed slice of the output vector, so the function it accelerates must already be registered with the return type that names that slice — registering a kernel for an undeclared function is the exact setup that panicked the server four times, and panics here instead.

func (*FuncRegistry) RegisterVecReturn

func (r *FuncRegistry) RegisterVecReturn(name string, dimFn func() int)

RegisterVecReturn marks a function as returning a VECTOR. dimFn is evaluated lazily (at plan time) to obtain the output dimensionality — embed(), for example, derives it from the configured embedding provider.

func (*FuncRegistry) ReturnType

func (r *FuncRegistry) ReturnType(name string) Ret

ReturnType returns the declared return type of a function. An unregistered name yields the zero Ret, which reports Declared() == false and resolves to "caller keeps its fallback".

func (*FuncRegistry) Unregister

func (r *FuncRegistry) Unregister(name string) bool

Unregister removes a scalar function. Returns true if it existed.

func (*FuncRegistry) VecReturnDim

func (r *FuncRegistry) VecReturnDim(name string) (dim int, ok bool)

VecReturnDim reports whether the named function returns a VECTOR and, if so, its current output dimension. ok is false for non-vector-returning functions.

type In

type In struct {
	Expr   Expr
	Values []Expr
	Not    bool
}

In checks if a value is in a set.

func (*In) Eval

func (e *In) Eval(b *batch.RecordBatch, row int) any

func (*In) EvalBool

func (e *In) EvalBool(b *batch.RecordBatch, row int) bool

func (*In) EvalBoolNull

func (e *In) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: `x IN (a, b, NULL)` is the chained OR of comparisons, so a match answers TRUE, and a miss with a NULL anywhere in the list is UNKNOWN — never FALSE. NOT IN is its Kleene negation, which is why `1 NOT IN (2, NULL)` must not answer true: PostgreSQL's reading is "I don't know, so no" (#370).

type InSubquery

type InSubquery struct {
	Expr   Expr
	SQL    string
	Runner SubqueryRunner
	Not    bool
	// contains filtered or unexported fields
}

InSubquery checks if a value is in the result set of a subquery. Example: WHERE user_id IN (SELECT user_id FROM active_users) Uncorrelated: executed once and result set cached in a hash set for O(1) lookup.

func (*InSubquery) Eval

func (e *InSubquery) Eval(b *batch.RecordBatch, row int) any

func (*InSubquery) EvalBool

func (e *InSubquery) EvalBool(b *batch.RecordBatch, row int) bool

func (*InSubquery) EvalBoolNull

func (e *InSubquery) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

type Int64Expr

type Int64Expr interface {
	EvalInt64(b *batch.RecordBatch, row int) (int64, bool)
}

Int64Expr evaluates to int64 without boxing.

type IntervalValue

type IntervalValue struct {
	Years   int
	Months  int
	Days    int
	Hours   int
	Minutes int
	Seconds int
}

IntervalValue represents a SQL INTERVAL (e.g., INTERVAL '30' DAY).

type IsBool

type IsBool struct {
	Operand Expr
	Want    bool // TRUE or FALSE spelling
	Not     bool // IS NOT
}

IsBool is `x IS [NOT] TRUE/FALSE`. Distinct from Cmp because it is a NULL-test like IS NULL, not a comparison: NULL IS TRUE answers FALSE and NULL IS NOT TRUE answers TRUE, where a comparison against NULL would be UNKNOWN (#370 — the Cmp spelling was right only while Cmp itself had no UNKNOWN).

func (*IsBool) Eval

func (e *IsBool) Eval(b *batch.RecordBatch, row int) any

func (*IsBool) EvalBool

func (e *IsBool) EvalBool(b *batch.RecordBatch, row int) bool

func (*IsBool) EvalBoolNull

func (e *IsBool) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

type IsDistinctFrom

type IsDistinctFrom struct {
	Left, Right Expr
	Not         bool // true for IS NOT DISTINCT FROM
}

IsDistinctFrom implements PostgreSQL's NULL-safe (in)equality, IS [NOT] DISTINCT FROM (#374). Unlike Cmp, it never answers UNKNOWN: NULL participates as a value here rather than propagating, so two NULLs are NOT DISTINCT (equal) and a NULL against a non-NULL value IS DISTINCT. "NULL IS DISTINCT FROM NULL" is FALSE, never NULL — the one case a COALESCE-based workaround gets wrong for a real sentinel value.

func (*IsDistinctFrom) Eval

func (e *IsDistinctFrom) Eval(b *batch.RecordBatch, row int) any

func (*IsDistinctFrom) EvalBool

func (e *IsDistinctFrom) EvalBool(b *batch.RecordBatch, row int) bool

func (*IsDistinctFrom) EvalBoolNull

func (e *IsDistinctFrom) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull always reports null=false: DISTINCT FROM is total over NULL inputs, which is the entire point of the operator.

type IsNull

type IsNull struct {
	Operand Expr
	Not     bool // IS NOT NULL
}

IsNull checks if an expression is null.

func (*IsNull) Eval

func (e *IsNull) Eval(b *batch.RecordBatch, row int) any

func (*IsNull) EvalBool

func (e *IsNull) EvalBool(b *batch.RecordBatch, row int) bool

func (*IsNull) EvalBoolNull

func (e *IsNull) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: IS [NOT] NULL never answers UNKNOWN — it is the operator SQL provides to ASK about NULL.

type Like

type Like struct {
	Expr    Expr
	Pattern Expr
	Not     bool
}

Like performs SQL LIKE pattern matching.

func (*Like) Eval

func (e *Like) Eval(b *batch.RecordBatch, row int) any

func (*Like) EvalBool

func (e *Like) EvalBool(b *batch.RecordBatch, row int) bool

func (*Like) EvalBoolNull

func (e *Like) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: LIKE with NULL on either side is UNKNOWN, and NOT LIKE stays UNKNOWN with it (#370).

type Lit

type Lit struct {
	Val any
}

Lit returns a constant value.

func (*Lit) Eval

func (e *Lit) Eval(_ *batch.RecordBatch, _ int) any

func (*Lit) EvalFloat64

func (e *Lit) EvalFloat64(_ *batch.RecordBatch, _ int) (float64, bool)

func (*Lit) EvalFloat64Vec

func (e *Lit) EvalFloat64Vec(_ *batch.RecordBatch, dst []float64, n int) bool

EvalFloat64Vec fills dst[0:n] with the literal value.

func (*Lit) EvalInt64

func (e *Lit) EvalInt64(_ *batch.RecordBatch, _ int) (int64, bool)

type MissingOuterColumnError

type MissingOuterColumnError struct {
	Ref       plansql.OuterRef
	Available []string
}

MissingOuterColumnError reports a correlated subquery whose outer column is absent from the batch the outer query hands it — a planning defect (column pruning, projection, or a rename), not a data condition.

func (*MissingOuterColumnError) Error

func (e *MissingOuterColumnError) Error() string

func (*MissingOuterColumnError) FatalEvalError

func (e *MissingOuterColumnError) FatalEvalError() error

FatalEvalError satisfies the marker the pipeline drivers recover on. Expr's Eval/EvalBool have no error return, so a failure that must not be mistaken for a NULL travels as a panic carrying this value and is turned back into a query error at the pipeline boundary (see exec.FatalEvalPanic).

type Not

type Not struct {
	Operand Expr
}

Not is a logical NOT.

func (*Not) Eval

func (e *Not) Eval(b *batch.RecordBatch, row int) any

func (*Not) EvalBool

func (e *Not) EvalBool(b *batch.RecordBatch, row int) bool

EvalBool: NOT must see the third value — collapsing first turned NOT (UNKNOWN) into true and admitted rows SQL excludes, which was the dangerous half of #370 (`1 NOT IN (2, NULL)` answering true).

func (*Not) EvalBoolNull

func (e *Not) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: NOT UNKNOWN stays UNKNOWN.

type Or

type Or struct {
	Left, Right Expr
}

Or is a logical OR.

func (*Or) Eval

func (e *Or) Eval(b *batch.RecordBatch, row int) any

func (*Or) EvalBool

func (e *Or) EvalBool(b *batch.RecordBatch, row int) bool

func (*Or) EvalBoolNull

func (e *Or) EvalBoolNull(b *batch.RecordBatch, row int) (bool, bool)

EvalBoolNull: TRUE OR anything is TRUE; otherwise a NULL operand makes it UNKNOWN. Short-circuits on a TRUE left operand.

type ParamRef

type ParamRef struct {
	Index int
	// contains filtered or unexported fields
}

ParamRef is an expression node that references a UDF parameter by index.

func (*ParamRef) Eval

func (e *ParamRef) Eval(_ *batch.RecordBatch, _ int) any

type Ret

type Ret struct {
	// contains filtered or unexported fields
}

Ret is a scalar function's declared return type: the vector type its results can be stored in. It is declared where the function is registered, and the planner types a projection from the same declaration the kernel writes through.

Before this existed the two halves lived apart — a function was registered in this package while its return type was asserted by a hand-maintained name list in the physical planner (isNumericFunc). A function missing from that list was typed String, so the projection allocated a Bytes output vector and the function's vec kernel wrote Float64Data/BoolData off the end of a zero-length slice, killing the server process for every connection. That happened four times (temporal extractors, vector distances, the length family, and starts_with/contains/ends_with) before the list was replaced by this declaration (#310).

The zero value is *undeclared* and the registry refuses it: Register's signature makes a missing declaration a compile error, and the zero value makes a field-named literal that skips it a panic at init.

func RetSameAsArg

func RetSameAsArg(fallback batch.TypeID, args ...int) Ret

RetSameAsArg declares a polymorphic return: the type of the first listed argument the caller can decide, or fallback when it can decide none of them — and Resolve marks that fallback as a guess, so a CALLER with candidates of its own keeps looking rather than inheriting it (see Confidence). With no indices every argument is a candidate, which is what coalesce, greatest and least want; nullif mirrors argument 0 only.

func RetTypeOf

func RetTypeOf(t batch.TypeID) Ret

RetTypeOf builds a fixed declaration for a type without a named constant above. Kept for callers registering functions over the network-native types.

func (Ret) Declared

func (r Ret) Declared() bool

Declared reports whether this is a real declaration rather than the zero value. Registration rejects an undeclared Ret.

func (Ret) Numeric

func (r Ret) Numeric() bool

Numeric reports whether the function always returns a number. It is the registry-backed replacement for the compiler's own hand-maintained numeric name list: a numeric call can be wrapped so it satisfies Float64Expr/ Int64Expr and binary operators over it take the typed path.

func (Ret) Resolve

func (r Ret) Resolve(nargs int, argType func(i int) (batch.TypeID, Confidence)) (batch.TypeID, Confidence)

Resolve returns the output type for a call with nargs arguments, and how confidently. argType reports the type of argument i and how confidently the caller decided it; it is consulted only by polymorphic declarations and may be nil.

Undecided means the caller should keep its own fallback: the function is RetDynamic, or the name is not registered at all.

A polymorphic declaration takes the first candidate argument that DECIDED a type. A candidate that only guessed does not end the search — it is remembered, in preference order, and answered only if no later candidate decides. A guess stays a guess all the way up, so an argument that guessed at any depth never displaces an argument that knows.

func (Ret) String

func (r Ret) String() string

type ScalarFunc

type ScalarFunc func(args []any) any

ScalarFunc is a scalar function implementation.

type ScalarSubquery

type ScalarSubquery struct {
	SQL    string
	Runner SubqueryRunner
	// contains filtered or unexported fields
}

ScalarSubquery evaluates a subquery that returns a single scalar value. Example: WHERE price > (SELECT AVG(price) FROM products) Uncorrelated: executed once and result cached.

func (*ScalarSubquery) Eval

func (e *ScalarSubquery) Eval(_ *batch.RecordBatch, _ int) any

type SubqueryRunner

type SubqueryRunner func(sql string) ([]map[string]any, error)

SubqueryRunner executes a SQL subquery and returns its result rows. Each row is a map of column name to value.

type UDFCall

type UDFCall struct {
	Name     string
	ArgExprs []Expr // caller-supplied argument expressions
	Body     Expr   // compiled UDF body with ParamRef nodes
	// contains filtered or unexported fields
}

UDFCall evaluates a user-defined function by binding arguments, then evaluating the compiled body expression.

func (*UDFCall) Eval

func (e *UDFCall) Eval(b *batch.RecordBatch, row int) any

type UDFDef

type UDFDef struct {
	Name   string   // function name (lowercase)
	Params []string // parameter names (lowercase)
	Body   string   // SQL expression body (e.g. "param1 * 2 + param2")
	Owner  string   // who created this function (empty = system/unowned)
	Locked bool     // if true, only the owner (or admin) can modify/drop
}

UDFDef defines a user-defined function.

type UDFPersister

type UDFPersister func(udfs []UDFDef) error

UDFPersister is called after UDF register/unregister to persist the current state.

type UDFStore

type UDFStore struct {
	// contains filtered or unexported fields
}

UDFStore holds compiled UDF definitions for use by the expression engine. Thread-safe for concurrent reads and writes.

func NewUDFStore

func NewUDFStore() *UDFStore

NewUDFStore creates a new empty UDF store.

func (*UDFStore) CompileUDFCall

func (s *UDFStore) CompileUDFCall(name string, argExprs []Expr) (Expr, error)

CompileUDFCall creates a UDFCall expression node.

func (*UDFStore) Get

func (s *UDFStore) Get(name string) (UDFDef, bool)

Get returns a UDF definition by name.

func (*UDFStore) List

func (s *UDFStore) List() []UDFDef

List returns all registered UDF definitions.

func (*UDFStore) LoadDefs

func (s *UDFStore) LoadDefs(defs []UDFDef) int

LoadDefs registers pre-existing UDF definitions (e.g., from KV on startup). Skips compilation errors silently so one bad UDF doesn't block startup.

func (*UDFStore) Register

func (s *UDFStore) Register(def UDFDef, isAdmin bool) error

Register compiles and registers a UDF.

func (*UDFStore) SetPersister

func (s *UDFStore) SetPersister(p UDFPersister)

SetPersister sets the function called after UDF mutations to persist state.

func (*UDFStore) Unregister

func (s *UDFStore) Unregister(name, caller string, isAdmin bool) error

Unregister removes a UDF.

type UnaryOp

type UnaryOp struct {
	Operand Expr
	Op      string // -, +
}

UnaryOp is a unary arithmetic expression (negation).

func (*UnaryOp) Eval

func (e *UnaryOp) Eval(b *batch.RecordBatch, row int) any

type UnknownFuncError

type UnknownFuncError struct {
	Name string
	// Aggregate marks a name this engine recognizes as an aggregate from
	// other SQL dialects but does not implement. The distinction matters to
	// the reader: an unimplemented aggregate silently dropped the GROUP BY
	// as well as the value, so the result had the wrong row COUNT.
	Aggregate bool
}

UnknownFuncError names a function the registry cannot resolve.

It is a distinct type because the physical planner's compile sites are forgiving by design: a projection whose AST will not compile falls back to copying an input column of the same name, which is how an aggregate's output column reaches the projection. That fallback is right for every compile failure EXCEPT this one — a name nothing implements has no column to fall back to, so swallowing it converted "unknown function foo" into the far less actionable "column \"foo(x)\" does not exist in the input schema", or, before the check existed, into no message at all. Callers test for this type with errors.As and propagate rather than falling back.

func (*UnknownFuncError) Error

func (e *UnknownFuncError) Error() string

func (*UnknownFuncError) SQLState

func (e *UnknownFuncError) SQLState() string

SQLState returns PostgreSQL's undefined_function code. sqlerr.StateOf picks it up through the Coder interface so the wire reports 42883 rather than the blanket 42000 (#366).

type VecExpr

type VecExpr interface {
	EvalVec(b *batch.RecordBatch, out *batch.Vector, n int)
}

VecExpr evaluates an expression for an entire batch at once, writing results directly to the output vector. This avoids per-row interface dispatch and boxing.

type VecFloat64Expr

type VecFloat64Expr interface {
	EvalFloat64Vec(b *batch.RecordBatch, dst []float64, n int) bool
}

VecFloat64Expr evaluates an expression for all rows [0, n) at once, writing results to dst. Returns true if any output is null. Eliminates per-row function call overhead (~5 calls/row/expression).

type VecScalarFunc

type VecScalarFunc func(args []*batch.Vector, out *batch.Vector, n int)

VecScalarFunc is a vectorized scalar function that operates on entire columns at once, reading from input vectors and writing to an output vector. This avoids per-row interface dispatch and boxing overhead.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL