spaniter

package module
v0.1.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Jun 26, 2026 License: MIT Imports: 7 Imported by: 0

README

spaniter

Go Reference

spaniter adapts Cloud Spanner row streams to Go standard iterators.

The module is deliberately lower-level than github.com/apstndb/spanvalue: it does not format values, choose output formats, or own export policy. Its job is to turn *spanner.RowIterator into iter.Seq2[*spanner.Row, error] while preserving the Spanner iterator lifecycle.

API

  • RowIteratorSeq: adapts *spanner.RowIterator to iter.Seq2[*spanner.Row, error].
  • DrainRowIterator: consumes *spanner.RowIterator without yielding rows and returns metadata/stats.
  • WithResult: captures metadata, stats, and rows read in a RowIteratorResult.
  • WithOnMetadata: invokes a callback once when ResultSetMetadata becomes available.
  • WithOnStats: invokes a callback after completion with query plan, query stats, and DML row count.
  • Stats.ResultSetStats: converts captured stats to *sppb.ResultSetStats for protobuf-oriented downstream code.
  • Stats.HasResultSetStats: reports whether ResultSetStats would encode fields, useful when building an enclosing ResultSet.
  • Stats.ResultSetStatsForDML: converts captured standard DML stats when RowCount must be represented as row_count_exact, including zero.
  • WithDrainOnEarlyStop: optionally drains remaining rows after an early consumer stop so stats can be populated.
  • Rows: adapt already-built rows for tests and virtual result sets.
  • SliceToRowSeq: adapt existing []*spanner.Row fixtures for downstream tests and fakes.

RowIterator lifecycle

Cloud Spanner result metadata becomes available after the first Next call, unless that call returns an error other than iterator.Done. Query plan and query stats become available after Next returns iterator.Done only for an iterator created with QueryWithStats; DML row count becomes available after iterator.Done when the RowIterator represents a DML statement.

RowIteratorSeq keeps those rules explicit while leaving metadata capture optional:

rows := spaniter.RowIteratorSeq(rowIter)

for row, err := range rows {
	if err != nil {
		return err
	}
	_ = row
}

Once the returned sequence is invoked, RowIteratorSeq owns the RowIterator for that run and calls Stop before returning. The sequence is single-use and must not be consumed concurrently. For short call sites, code can pass a freshly created iterator directly instead of binding it only to defer Stop:

stmt := spanner.Statement{SQL: sql}

for row, err := range spaniter.RowIteratorSeq(txn.Query(ctx, stmt)) {
	if err != nil {
		return err
	}
	_ = row
}

The same ownership transfer applies when passing the sequence to a function that immediately consumes iter.Seq2[*spanner.Row, error]. Merely accepting, storing, or forwarding the sequence does not invoke it or stop the underlying RowIterator. Many consumers do not need ResultSetMetadata; they can process rows directly. Use the inline form only when the called function consumes the sequence synchronously before returning; otherwise bind the RowIterator and keep responsibility for Stop.

func consumeRows(rows iter.Seq2[*spanner.Row, error]) error {
	for row, err := range rows {
		if err != nil {
			return err
		}
		if err := processRow(row); err != nil {
			return err
		}
	}
	return nil
}

stmt := spanner.Statement{SQL: sql}

if err := consumeRows(spaniter.RowIteratorSeq(txn.Query(ctx, stmt))); err != nil {
	return err
}

Each yielded pair is either a non-nil row with a nil error, or a nil row with a non-nil terminal error. After yielding a non-nil error, the sequence stops; code should stop processing on the first error.

If the returned lazy sequence is never invoked, callers must still retain responsibility for stopping the original RowIterator.

Use WithResult when code needs result-set metadata, stats, or rows read outside the row loop:

var result spaniter.RowIteratorResult
for row, err := range spaniter.RowIteratorSeq(rowIter, spaniter.WithResult(&result)) {
	if err != nil {
		return err
	}
	_ = row
}
_ = result.Metadata
_ = result.Stats
_ = result.RowsRead

Use Stats.ResultSetStats when downstream code needs Cloud Spanner's *sppb.ResultSetStats protobuf shape. It re-encodes query stats and represents non-zero RowCount as row_count_exact. Because RowIterator exposes row count as a plain int64, an absent row count and an exact zero DML row count cannot be distinguished. Use Stats.ResultSetStatsForDML when the caller knows the stats came from standard DML and needs row_count_exact: 0 to be preserved. Do not use the DML method for ordinary queries: it would synthesize a row_count_exact field even though the Spanner API omits row_count for query stats. Conversely, using ResultSetStats for DML preserves non-zero counts but drops the explicit row_count_exact: 0 case. The usual ReadWriteTransaction.Update and Client.PartitionedUpdate APIs return counts directly rather than through a RowIterator; handle those counts separately. Use stats.HasResultSetStats() before assigning the result of Stats.ResultSetStats to an enclosing ResultSet when an empty stats field should be omitted. When calling Stats.ResultSetStatsForDML, assign the returned message directly, including for a zero count.

Use WithOnMetadata or WithOnStats when code needs hook-style callbacks instead of a captured result value. For fan-in adapters or streaming sinks that must publish row-type metadata before the first downstream row, prefer WithOnMetadata: WithResult is updated at the same lifecycle points, but an empty result set never enters the consumer's loop body, so loop-local polling cannot observe metadata until after the sequence is exhausted.

If application code needs metadata or stats but does not want row values, DrainRowIterator consumes the stream internally and returns only lifecycle results:

result, err := spaniter.DrainRowIterator(rowIter)
if err != nil {
	return err
}
_ = result.Metadata
_ = result.Stats
_ = result.RowsRead

This still reads the Spanner result stream internally because the Go client populates stats only after Next returns iterator.Done. To avoid reading data rows at the query level, execute a statement that returns no data rows. If draining returns an error, the returned result can still contain partial metadata and RowsRead; stats are populated only after a successful full drain.

If the consumer may stop early but the caller still needs stats, enable WithDrainOnEarlyStop. This drains after any early stop, including a range-loop break or an adapter that stops pulling because downstream work failed, so keep it disabled unless that extra read work is acceptable. It has no effect on DrainRowIterator, which already drains. When this option is enabled, RowsRead includes rows read during the post-stop drain, including rows not yielded to the consumer. If the post-stop drain fails, its error cannot be yielded because the consumer has already stopped. In that case, WithOnStats is not called and RowIteratorResult.Stats remains zero.

rows := spaniter.RowIteratorSeq(rowIter,
	spaniter.WithDrainOnEarlyStop(),
	spaniter.WithOnStats(func(got spaniter.Stats) {
		stats = got
	}),
)

Example: spanvalue/writer

Applications that already expose rows as iter.Seq2 can compose spaniter with spanvalue/writer; this is caller-side composition, not a dependency of spaniter. When spanvalue/writer is the only consumer of a raw *spanner.RowIterator, its WriteRowIterator helper is the simpler direct path. The composition below is useful when the surrounding pipeline is already expressed as iter.Seq2.

RunRowSeqDeferredMetadata evaluates its metadata function after the first pull, or after an empty sequence ends, so WithResult has already recorded the metadata. RowIteratorHooksFromWriter then registers the row type and flushes after a successful run; for a delimited writer with headers, this permits header-only output for an empty result set. This call synchronously pulls the sequence before returning, so it can take lifecycle ownership of the inline RowIteratorSeq argument.

stmt := spanner.Statement{SQL: sql}

var spannerResult spaniter.RowIteratorResult
if _, err := writer.RunRowSeqDeferredMetadata(
	func() *sppb.ResultSetMetadata { return spannerResult.Metadata },
	spaniter.RowIteratorSeq(
		txn.Query(ctx, stmt),
		spaniter.WithResult(&spannerResult),
	),
	writer.RowIteratorHooksFromWriter(w),
); err != nil {
	return err
}

When metadata is already available and code needs hook-level control, RunRowSeq is the direct consumer form:

if _, err := writer.RunRowSeq(
	metadata,
	spaniter.Rows(row1, row2),
	writer.RowIteratorHooksFromWriter(w),
); err != nil {
	return err
}

For ordinary writer use, WriteRowSeq wraps the same hook setup:

if _, err := writer.WriteRowSeq(
	metadata,
	spaniter.Rows(row1, row2),
	w,
); err != nil {
	return err
}

This keeps spaniter independent of formatting packages while letting callers reuse the standard iterator stream directly. In this sequence-based composition, the writer.RowIteratorResult returned by RunRowSeq, RunRowSeqDeferredMetadata, or WriteRowSeq has zero Stats; use spaniter.WithResult or WithOnStats for Spanner-specific execution data. The writer result's RowsRead counts successful writes, whereas spaniter.RowIteratorResult.RowsRead counts rows consumed from the Spanner iterator.

Development

make check

Documentation

Overview

Package spaniter adapts Cloud Spanner row streams to Go standard iterators.

The package is intentionally lower-level than formatters and writers: it owns only iterator lifecycle concerns such as RowIterator.Stop, result metadata, and post-drain query stats. Formatting, headers, and export policy stay in callers.

Index

Constants

This section is empty.

Variables

View Source
var ErrNilRow = errors.New("nil row")

ErrNilRow reports that an adapted source produced a nil row with a nil error.

View Source
var ErrNilRowIterator = errors.New("nil row iterator")

ErrNilRowIterator reports that RowIteratorSeq was given a nil iterator.

Because RowIteratorSeq returns an iter.Seq2, the error is yielded when the sequence is consumed rather than returned by the constructor.

Functions

func RowIteratorSeq

func RowIteratorSeq(rowIter *spanner.RowIterator, opts ...Option) iter.Seq2[*spanner.Row, error]

RowIteratorSeq adapts a cloud.google.com/go/spanner.RowIterator to a Go standard iterator.

The returned sequence owns rowIter: once iteration starts it always calls *cloud.google.com/go/spanner.RowIterator.Stop before returning. Metadata and stats are exposed through WithOnMetadata and WithOnStats hooks instead of requiring callers to keep reading fields from the stopped RowIterator. The sequence is single-use and not safe for concurrent consumption; construct a new RowIterator for another pass.

If the returned sequence is never invoked, RowIteratorSeq cannot call Stop. After constructing a sequence, callers must either consume it, pass it to code that will consume or stop it, or retain responsibility for stopping the original RowIterator.

Each yielded pair is either a non-nil row with a nil error, or a nil row with a non-nil terminal error. After yielding a non-nil error, the sequence stops. Consumers should stop processing and return or break on the first non-nil error. On terminal errors, WithResult contains only lifecycle data observed before the error, and WithOnStats is not called.

func Rows

func Rows(rows ...*spanner.Row) iter.Seq2[*spanner.Row, error]

Rows adapts already-built rows to the fallible sequence shape used by RowIteratorSeq. Non-nil rows are yielded with a nil error. A nil row aborts the sequence by yielding ErrNilRow.

Row sources that can fail per row should produce their own iter.Seq2 instead of pre-building a slice for Rows.

func SliceToRowSeq

func SliceToRowSeq(rows []*spanner.Row) iter.Seq2[*spanner.Row, error]

SliceToRowSeq adapts an existing row slice to the fallible sequence shape used by RowIteratorSeq.

It exists for downstream tests, fakes, and virtual result sets that naturally store fixtures as []*spanner.Row. It is equivalent to Rows(rows...), including nil-row handling: nil rows yield ErrNilRow and abort the sequence.

Types

type Option

type Option func(*config)

Option configures RowIteratorSeq and DrainRowIterator.

func WithDrainOnEarlyStop

func WithDrainOnEarlyStop() Option

WithDrainOnEarlyStop configures RowIteratorSeq to consume the remaining rows after the consumer stops early.

Draining is disabled by default to preserve normal iterator early-exit behavior. Use this option when callers need WithOnStats to run after any early stop, including a range-loop break or an adapter that stops pulling because a downstream operation failed. Errors encountered only during this post-stop drain cannot be yielded to the caller and therefore suppress the stats hook. It has no effect on DrainRowIterator, which always drains.

func WithOnMetadata

func WithOnMetadata(f func(*sppb.ResultSetMetadata)) Option

WithOnMetadata registers a hook that runs once when result metadata becomes available.

For a query with rows, the hook runs after the first successful Next call and before that first row is yielded, so metadata captured by the hook is visible inside the first loop body. For an empty result set, the hook runs after Next returns iterator.Done. A nil hook is ignored.

func WithOnStats

func WithOnStats(f func(Stats)) Option

WithOnStats registers a hook that runs after the adapted iterator has reached iterator.Done and has been stopped.

If the consumer stops early, stats are available only when WithDrainOnEarlyStop is also configured. A nil hook is ignored.

func WithResult

func WithResult(result *RowIteratorResult) Option

WithResult stores iterator lifecycle data in result as it becomes available.

The pointed value is reset when iteration starts. Metadata is set before the first row is yielded, RowsRead is updated after each consumed row, and Stats is set only after the iterator reaches iterator.Done. On errors, result contains the partial lifecycle data observed before the error. A nil result is ignored.

type RowIteratorResult

type RowIteratorResult struct {
	Metadata *sppb.ResultSetMetadata
	Stats    Stats
	RowsRead int64
}

RowIteratorResult is the metadata and stats available from a cloud.google.com/go/spanner.RowIterator.

RowsRead counts rows consumed from the iterator. Metadata and Stats values are not deep-copied from the underlying RowIterator; treat returned maps and protos as read-only.

func DrainRowIterator

func DrainRowIterator(rowIter *spanner.RowIterator, opts ...Option) (*RowIteratorResult, error)

DrainRowIterator consumes rowIter to iterator.Done without yielding rows.

The helper owns rowIter and always calls *cloud.google.com/go/spanner.RowIterator.Stop before returning. It is useful when callers need result metadata, query stats, query plan, or DML row count but do not want to expose row values to application code. If iteration fails, the returned result can be non-nil and contain partial metadata and RowsRead observed before the error; stats are only populated after a successful drain to iterator.Done.

Cloud Spanner only populates metadata after the first Next call, and stats after Next returns iterator.Done. DrainRowIterator therefore still consumes the result stream internally; it does not ask Spanner for stats without reading the stream. To avoid reading data rows at the query level, callers must execute a statement that returns no data rows.

type Stats

type Stats struct {
	QueryPlan  *sppb.QueryPlan
	QueryStats map[string]any
	RowCount   int64
}

Stats holds execution information populated on a cloud.google.com/go/spanner.RowIterator after the iterator reaches iterator.Done.

QueryPlan and QueryStats are set when the query used QueryWithStats. RowCount holds the DML row count after iterator.Done. Values are not deep-copied from the underlying RowIterator; treat returned maps and protos as read-only.

Stats mirrors the public fields exposed by cloud.google.com/go/spanner.RowIterator. Use Stats.ResultSetStats when downstream code needs the protobuf ResultSetStats shape. The Go client has already decoded query stats to a map and exposes row count as a single int64, so Stats cannot distinguish an absent row count from row_count_exact:0. The Go client's PartitionedUpdate APIs return counts directly rather than through RowIterator, so partitioned DML counts are outside this type's normal scope.

func (Stats) HasResultSetStats added in v0.1.1

func (s Stats) HasResultSetStats() bool

HasResultSetStats reports whether s has fields that Stats.ResultSetStats would encode.

This is useful when callers build an enclosing ResultSet and want to omit the stats field when the default conversion would be empty. It returns false for DML row_count_exact:0 because Stats cannot distinguish exact zero from an absent row count. Callers that know the Stats came from standard DML and need row_count_exact:0 should call Stats.ResultSetStatsForDML directly.

func (Stats) ResultSetStats added in v0.1.1

func (s Stats) ResultSetStats() (*sppb.ResultSetStats, error)

ResultSetStats returns s in Cloud Spanner protobuf ResultSetStats form.

Most callers should use this method. The row-count caveat below matters only to consumers that distinguish an absent row_count from row_count_exact:0; code that treats an absent row_count as the zero value sees the same count either way.

QueryPlan is reused, QueryStats is re-encoded as a protobuf Struct, and a non-zero RowCount is encoded as row_count_exact. RowIterator exposes row count as a plain int64, so an absent row count and an exact zero row count are indistinguishable; this method omits row count when RowCount is zero. Callers that know RowCount is a standard DML count can use Stats.ResultSetStatsForDML.

ResultSetStats returns an error if QueryStats contains a key or value that cannot be represented by structpb.Struct.

func (Stats) ResultSetStatsForDML added in v0.1.1

func (s Stats) ResultSetStatsForDML() (*sppb.ResultSetStats, error)

ResultSetStatsForDML returns s as ResultSetStats for standard DML and always encodes RowCount as row_count_exact, including zero.

This method is for consumers that need protobuf oneof presence to distinguish DML row_count_exact:0 from an absent row_count. If the consumer treats an absent row_count as the zero value, Stats.ResultSetStats is sufficient.

Use this method only when the caller knows the Stats came from standard DML. Using it for query stats creates a misleading row_count_exact field even though the Spanner API would omit row_count for queries. Using Stats.ResultSetStats for DML is fine for non-zero row counts, but loses the distinction between absent row count and row_count_exact:0. Partitioned DML is outside spaniter's RowIterator path because the Go client exposes it through Client.PartitionedUpdate, not RowIterator.

ResultSetStatsForDML returns an error if QueryStats contains a key or value that cannot be represented by structpb.Struct.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL