confgen

package
v0.17.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 8, 2026 License: MIT Imports: 35 Imported by: 0

README

confgen — Configuration Generation

Converts workbook data (Excel/CSV/XML/YAML) into protobuf messages with concurrent parsing and a hierarchical error collector for multi-level error limiting.

Parsing Hierarchy

Generator
 ├── GenAll / GenWorkbook
 │    └── collector.NewGroup(ctx)                    ← concurrent workbook batch
 │         └── Group.Go(convert)                     ← one goroutine per proto file
 │
 └── convert(fd)                                     ← sequential per-sheet loop within one workbook
      ├── processScatter → ScatterAndExport
      │    ├── parseMessageFromOneImporter(main)      ← main importer: sequential
      │    └── collector.NewGroup(ctx)                ← concurrent scatter batch
      │         └── Group.Go(parseMessageFromOneImporter)
      │
      └── processMerger → MergeAndExport
           └── ParseMessage
                ├── single importer → parseMessageFromOneImporter   ← sequential
                └── multiple importers
                     └── collector.NewGroup(ctx)                    ← concurrent merge batch
                          └── Group.Go(parseMessageFromOneImporter)
parseMessageFromOneImporter (leaf)
parseMessageFromOneImporter(info, messageCollector, impInfo)
 └── sheetCollector = messageCollector.NewChild(maxErrorsPerSheet=5, BookName, SheetName)
 └── sheetParser.Parse(protomsg, sheet)
      ├── [document sheet] → documentParser.Parse
      │    └── parseMessage(node)                    ← recursive tree walk
      │
      └── [table sheet]    → tableParser.Parse
           └── tableParser.parse
                └── RangeDataRows(row callback)
                     └── parseMessage(row)           ← per row

Concurrent Model

flowchart TB
    subgraph Generator
        C["gen.collector (maxErrors=20)"]
    end

    subgraph "Workbook Group (concurrent)"
        direction TB
        G1["goroutine: convert(fd₁)"]
        G2["goroutine: convert(fd₂)"]
        Gn["goroutine: convert(fdₙ)"]
    end

    Generator --> G1 & G2 & Gn

    subgraph "convert(fd) — sequential sheet loop"
        B["bookCollector = gen.collector.NewChild(maxErrorsPerBook=10)"]
        S1["sheet₁: processScatter → ScatterAndExport"]
        S2["sheet₂: processMerger → MergeAndExport"]
        M["messageCollector = bookCollector.NewChild(0)"]
    end

    G1 --> B --> S1 --> S2

    subgraph "parseMessageFromOneImporter"
        SC["sheetCollector = messageCollector.NewChild(maxErrorsPerSheet=5)"]
        TP["tableParser.parse → RangeDataRows"]
        DP["documentParser.Parse"]
    end

    S1 & S2 --> M --> SC --> TP & DP
Level Collector Limit Scope
Generator gen.collector 20 across all concurrent workbooks
Book bookCollector = gen.collector.NewChild(10) 10 across sheets in one workbook
Message messageCollector = bookCollector.NewChild(0) unlimited one worksheet message
Sheet sheetCollector = messageCollector.NewChild(5) 5 one imported sheet

Error Collector

Hierarchy

Errors are counted at field level. The collector tree has limits at book and imported-sheet levels; the message level groups errors without a separate limit. Collect() increments counters on self and all ancestors. Join() recursively assembles the error tree.

A child collector carries fields shared by its subtree. The book collector supplies module and primary workbook context; the message collector supplies the worksheet and protobuf message; each imported-sheet collector supplies the actual workbook and worksheet names. Join() inherits these fields into each error. WrapKV adds fields at the node it wraps; fields on an individual error remain local to that error.

flowchart TB
    subgraph "Generator"
        Root["gen.collector  (limit=20)"]
    end

    subgraph "convert(fd)"
        Book["bookCollector = gen.collector.NewChild(10)"]
    end

    subgraph "convert(fd): worksheet message"
        Message["messageCollector = bookCollector.NewChild(0)"]
    end

    subgraph "parseMessageFromOneImporter"
        Sheet["sheetCollector = messageCollector.NewChild(5)"]
    end

    Root --> Book --> Message --> Sheet
    Sheet -- "Collect(err) → increments ancestors" --> Root
    Sheet -- "IsFull() → fail-fast: skip remaining rows/fields" --> Sheet
    Book -- "Join() → assembles all children errors" --> Book
Fail-fast Behavior
  • Sheet level: sheetCollector.IsFull() is checked before each row; returns early if full.
  • Book level: convert checks the error returned by messageCollector.Collect(); breaks the sheet loop if an ancestor is full.
  • Generator level: collector.NewGroup propagates the first fatal error (book-full) to stop the workbook goroutine.

Sheet Parser Metrics

Use tableauc --profiling to collect these metrics and write confgen-cpu.pprof and confgen-mem.pprof under the configured output directory. Profiling is disabled by default; disabled runs skip pprof collection, shape scans, labels, and metric aggregation.

The CPU profile spans the full generation run and includes all goroutines in the process. Parser samples carry work=sheet_parse plus book, sheet, detail, and canonical sheet_key labels. go tool pprof -tags confgen-cpu.pprof orders label values by sampled processor time; use -tagfocus='sheet_key=<value>' for one parser operation. Labels follow the goroutine across OS-thread scheduling and are inherited by goroutines created inside the labeled parser call. Very short parses may receive no sample, so compare representative full generation runs rather than treating one sample as an exact timer.

The memory file is a heap profile captured after garbage collection at the end of the run, so it primarily reports retained allocations. Only one CPU profile can run in a process at a time, and each run replaces the previous files.

GenAll and GenWorkbook log one row per imported workbook, sheet, and protobuf message, including sheets used by scatter and merger. Repeated parses of the same source are aggregated. The report reads sampled processor time from the labeled CPU profile and sorts by cumulative CPU time, then wall time. This avoids changing scheduler behavior with OS-thread pinning and includes labeled child goroutines. Very short parses may show zero CPU time when they receive no sample.

Wall time covers only sheetParser.Parse. It excludes workbook import, sheet shape collection, validation, merging, and output. Concurrent wall durations can overlap. Failures counts calls whose parser returned an error, so their wall timings may represent partial work.

For table sheets, rows is the total number of imported rows across calls, including headers; maxCols is the widest row. The cells field is the full rectangular grid for each call, summed across calls. Present counts nonempty strings. Absent is the sum of empty (stored empty strings) and missing (implicit cells beyond a short row). EmptyRows have no nonempty cells, and valueBytes sums the UTF-8 byte lengths of nonempty cells. Document sheets report node count, scalar node count, maximum depth, and scalar value bytes instead of table dimensions.

Further tuning could separate import, validation, merge, encoding, and write time; count allocations and allocated bytes per sheet; and count field descriptor cache misses and reference lookups. Those metrics would show whether a slow sheet is limited by parsing work or by adjacent phases.

Documentation

Index

Constants

View Source
const Version = "0.10.0"

Variables

This section is empty.

Functions

func NewExtendedSheetParser added in v0.11.0

func NewExtendedSheetParser(
	ctx context.Context, protoPackage, locationName string,
	bookOpts *tableaupb.WorkbookOptions,
	sheetOpts *tableaupb.WorksheetOptions,
	extInfo *SheetParserExtInfo,
	source SourceLocation,
) *sheetParser

NewExtendedSheetParser creates a new sheet parser with extended info.

func NewSheetExporter

func NewSheetExporter(outputDir string, output *options.ConfOutputOption, validator protovalidate.Validator, collector *xerrors.Collector) *sheetExporter

NewSheetExporter creates a new sheet exporter.

func NewSheetParser

func NewSheetParser(ctx context.Context, protoPackage, locationName string, opts *tableaupb.WorksheetOptions) *sheetParser

NewSheetParser creates a new sheet parser.

func NewValidator added in v0.16.3

func NewValidator(prFiles *protoregistry.Files) (protovalidate.Validator, error)

NewValidator creates a protovalidate validator whose extension type resolver is built from prFiles, so custom predefined rules can be recognized. It is shared by both the confgen (generate) and load (checker) paths to keep validation behavior consistent.

func ParseFileOptions added in v0.9.0

ParseFileOptions parse the options of a protobuf definition file.

func ParseMessage added in v0.10.5

func ParseMessage(info *SheetInfo, messageCollector *xerrors.Collector, impInfos ...importer.ImporterInfo) (proto.Message, error)

ParseMessage parses multiple importer infos into one protomsg.

func ParseMessageOptions added in v0.9.0

ParseMessageOptions parse the options of a protobuf message.

func Validate added in v0.16.3

func Validate(msg proto.Message, validator protovalidate.Validator) error

Validate validates a proto message using the provided validator. Each violation is wrapped as an E2027 error and all violations are joined together.

Types

type Field

type Field struct {
	// contains filtered or unexported fields
}

type Generator

type Generator struct {
	ProtoPackage string // protobuf package name.
	InputDir     string // input dir of workbooks.
	OutputDir    string // output dir of generated files.

	LocationName  string                    // TZ location name.
	InputOpt      *options.ConfInputOption  // Input settings.
	OutputOpt     *options.ConfOutputOption // output settings.
	ErrorLimitOpt *options.ErrorLimitOption // error collection limits.

	// Sheet parser metrics.
	SheetParserMetrics profile.SheetParserMetrics
	// contains filtered or unexported fields
}

Generator converts workbook data for one generation run. Create a new Generator when another run is needed.

func NewGenerator

func NewGenerator(protoPackage, indir, outdir string, setters ...options.Option) *Generator

func NewGeneratorWithOptions added in v0.9.14

func NewGeneratorWithOptions(protoPackage, indir, outdir string, opts *options.Options) *Generator

func (*Generator) GenAll added in v0.10.0

func (gen *Generator) GenAll() error

func (*Generator) GenWorkbook added in v0.10.0

func (gen *Generator) GenWorkbook(bookSpecifiers ...string) error

bookSpecifier can be:

  • only workbook: excel/Item.xlsx
  • with worksheet: excel/Item.xlsx#Item (To be implemented)

func (*Generator) Generate

func (gen *Generator) Generate(bookSpecifiers ...string) error

bookSpecifier can be:

  • only workbook: excel/Item.xlsx
  • specific worksheet: excel/Item.xlsx#Item (To be implemented)

type SheetInfo added in v0.10.7

type SheetInfo struct {
	ProtoPackage    string
	LocationName    string
	PrimaryBookName string
	MD              protoreflect.MessageDescriptor
	BookOpts        *tableaupb.WorkbookOptions
	SheetOpts       *tableaupb.WorksheetOptions

	ExtInfo *SheetParserExtInfo
}

func (*SheetInfo) BookName added in v0.15.0

func (si *SheetInfo) BookName() string

func (*SheetInfo) HasMerger added in v0.10.7

func (si *SheetInfo) HasMerger() bool

func (*SheetInfo) HasScatter added in v0.10.7

func (si *SheetInfo) HasScatter() bool

func (*SheetInfo) SheetName added in v0.15.0

func (si *SheetInfo) SheetName() string

type SheetParserExtInfo added in v0.11.0

type SheetParserExtInfo struct {
	InputDir           string
	SubdirRewrites     map[string]string
	PRFiles            *protoregistry.Files
	BookFormat         format.Format // workbook format
	DryRun             options.DryRun
	ErrorLimit         *options.ErrorLimitOption // error collection limits
	ReferredCache      *fieldprop.ReferredCache
	ImporterCache      *importer.Cache
	SheetParserMetrics *profile.SheetParserMetrics
}

SheetParserExtInfo holds dependencies shared by extended sheet parsers.

type SourceLocation added in v0.17.0

type SourceLocation struct {
	BookName  string
	SheetName string
}

SourceLocation identifies the workbook and worksheet parsed by a sheet parser.

Directories

Path Synopsis

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL