config

package
v0.5.0-rc.2 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 4, 2026 License: MIT Imports: 10 Imported by: 0

Documentation

Index

Constants

View Source
const SelfServiceName = "otelcontext"

SelfServiceName is the OTel service.name attribute the binary attaches to its own self-instrumentation spans. Mirrors the literal in main.initTracerProvider — keep the two in sync.

Variables

This section is empty.

Functions

func ParseAggregateMetricDims added in v0.5.0

func ParseAggregateMetricDims(s string) (map[string][]string, error)

ParseAggregateMetricDims parses the AGGREGATE_METRIC_DIMS config string. Format: "metric_name:key1,key2;metric_name2:key3,key4" Returns a map of metric name -> sorted list of dimension keys. Fails on malformed input (fail-closed): empty metric name, empty key list, duplicate metric, duplicate key within a metric, stray separators. Empty/unset var is valid (returns empty map).

Types

type Config

type Config struct {
	Env               string
	LogLevel          string
	HTTPPort          string
	GRPCPort          string
	DBDriver          string
	DBDSN             string
	DLQPath           string
	DLQReplayInterval string

	// PprofAddr serves net/http/pprof on a dedicated listener (never the
	// public mux). Loopback-only by default; empty disables profiling.
	PprofAddr string

	// Ingestion Filtering
	IngestMinSeverity      string
	IngestAllowedServices  string
	IngestExcludedServices string

	// Storage Filtering. Logs that pass IngestMinSeverity (so they reach the
	// receiver and feed in-memory consumers like GraphRAG / Drain) but fall
	// below StoreMinSeverity are skipped during the DB persist pass — only the
	// row-write is dropped, not the in-memory enrichment. Defaults to "WARN"
	// (all drivers): INFO/DEBUG still inform anomaly detection + clustering but
	// don't grow the DB. Empty falls back to IngestMinSeverity (no second-tier
	// gate); a value <= IngestMinSeverity is a no-op since the receiver already
	// drops below that.
	StoreMinSeverity string

	// DB Connection Pool
	DBMaxOpenConns    int
	DBMaxIdleConns    int
	DBConnMaxLifetime string // e.g. "1h", "30m"

	// Postgres-only opt-in: declarative range partitioning of the logs table by
	// day. When set to "daily", AutoMigrate provisions logs as a partitioned
	// table and the PartitionScheduler creates lookahead partitions and drops
	// expired ones (DROP PARTITION beats DELETE for retention by orders of
	// magnitude). Greenfield only — startup refuses if `logs` already exists
	// as a non-partitioned table. Empty / "none" = legacy unpartitioned schema.
	DBPostgresPartitioning string

	// Number of future daily partitions to maintain ahead of "today" when
	// DBPostgresPartitioning=daily. Defaults to 3. Tune up if your retention
	// policy is short and ingest spikes around a daily boundary.
	DBPartitionLookaheadDays int

	// Retention
	HotRetentionDays int

	// Retention tuning. Defaults (batch=50000, sleep=1ms) work for Postgres at
	// 100k logs/sec sustained. Lower on resource-constrained hosts; raise on
	// dedicated DB machines. 0/negative values use defaults.
	RetentionBatchSize    int
	RetentionBatchSleepMs int

	// RetentionFullVacuum restores the daily full VACUUM during SQLite
	// maintenance. Default false: the daily pass runs
	// PRAGMA incremental_vacuum(10000) instead, because a full VACUUM holds
	// an exclusive lock for 10-60 minutes on multi-GB files and starves
	// ingest into a 429 storm. On-demand full VACUUM remains available via
	// POST /api/admin/vacuum. Ignored on non-SQLite drivers.
	RetentionFullVacuum bool

	// TSDB
	TSDBRingBufferDuration string // e.g. "1h"

	// Smart Observability — Adaptive Sampling
	SamplingRate               float64
	SamplingAlwaysOnErrors     bool
	SamplingLatencyThresholdMs int

	// Smart Observability — Metric Cardinality
	MetricAttributeKeys  string // comma-separated allowlist
	MetricMaxCardinality int

	// Per-tenant cardinality cap. 0 = unlimited (only the global cap
	// applies, preserving legacy single-tenant behavior). Setting this
	// gives every tenant its own series budget so a noisy tenant cannot
	// starve siblings of fresh series in the in-memory TSDB. The global
	// cap (MetricMaxCardinality) remains a backstop and is checked
	// after the per-tenant cap.
	MetricMaxCardinalityPerTenant int

	// DLQ Safety
	DLQMaxFiles   int
	DLQMaxDiskMB  int
	DLQMaxRetries int
	// DLQMaxReplayPerTick caps how many DLQ files the replay worker attempts
	// in a single tick. Without it, an outage that filled the DLQ with 10k
	// files would replay all of them in the first post-restart tick,
	// hammering the (just-restarted) DB and exhausting connections.
	// 0 = unlimited (legacy default).
	DLQMaxReplayPerTick int

	// API Protection
	APIRateLimitRPS int

	// MCP Server
	MCPEnabled bool
	MCPPath    string
	// MCPMaxConcurrent caps the in-flight tools/call invocations server-wide.
	// Beyond this, callers receive a JSON-RPC server-overloaded error. <=0
	// disables the cap. Default 32 — sized for tight agent polling loops
	// without overrunning the GraphRAG in-memory store.
	MCPMaxConcurrent int
	// MCPCallTimeoutMs is the per-invocation deadline for tools/call. A tool
	// that exceeds it gets cancelled and the client receives an RPC timeout
	// error. <=0 disables the deadline. Default 30000 (30s).
	MCPCallTimeoutMs int
	// MCPCacheTTLMs is the lifetime of a memoized tool result for the cheap
	// in-memory GraphRAG tools (get_service_map, impact_analysis, etc.).
	// <=0 disables caching. Default 5000 (5s).
	MCPCacheTTLMs int

	// Compression
	CompressionLevel string // "default", "fast", "best"

	// LogFTSEnabled toggles SQLite FTS5 provisioning + querying. The FTS5
	// inverted index typically consumes 30-40% of SQLite DB disk for
	// log-heavy workloads, while the LIKE fallback (log_repo.go:105) keeps
	// search_logs functional without it. Default false; opt in with
	// LOG_FTS_ENABLED=true. Only meaningful on SQLite; Postgres uses pg_trgm
	// independently of this flag.
	LogFTSEnabled bool

	// GraphRAG worker count (background consumers of the ingestion event channel).
	// Defaults to 4 if unset or <=0. Increase under sustained high ingest.
	GraphRAGWorkerCount int

	// GraphRAG event channel buffer size. Defaults to 10000 if unset or <=0.
	GraphRAGEventQueueSize int

	// GraphRAGTraceTTL bounds how long spans/traces stay in the in-memory
	// TraceStore before the refresh tick prunes them. Duration string, e.g.
	// "1h". Defaults to "1h"; flipped to "30m" on SQLite (the in-memory span
	// window is the largest GraphRAG heap consumer at 120 services). Anomaly
	// and investigation paths look back <=5min, so a 30min window is safe.
	GraphRAGTraceTTL string

	// GraphRAGMaxSpansPerTenant hard-caps the in-memory TraceStore span map
	// per tenant. At the cap, NEW spans are skipped (counted via
	// otelcontext_graphrag_events_dropped_total{signal="span_capacity"});
	// updates to resident spans still apply. The graph is best-effort — the
	// DB remains the source of truth. 0 = default (500000); negative
	// disables the cap.
	GraphRAGMaxSpansPerTenant int

	// GraphRAGTenantIdleTTL evicts a tenant's entire in-memory store slice
	// after this much time without any ingest event or query. Duration
	// string, default "24h". The default tenant is never evicted, and an
	// active tenant is re-created within one refresh tick (60s) from recent
	// DB spans — eviction is self-healing.
	GraphRAGTenantIdleTTL string

	// Async ingest pipeline (Phase 1 robustness work). Decouples OTLP Export
	// from synchronous DB writes. When enabled, Export() returns as soon as
	// the parsed batch is enqueued; persistence runs on a worker pool.
	//
	// Backpressure is hybrid:
	//   <90% queue       — accept all
	//   90%-100% queue   — drop healthy batches (silent), errors/slow always pass
	//   100% queue       — return RESOURCE_EXHAUSTED so OTLP clients back off
	IngestAsyncEnabled      bool // default true; opt out via INGEST_ASYNC_ENABLED=false
	IngestPipelineQueueSize int  // default 50000 batches; per-deployment tunable
	// IngestPipelineMaxBytes caps the approximate bytes held by queued
	// batches. The item-count queue size alone cannot bound memory — one
	// batch may carry arbitrarily large span/log payloads. At the cap the
	// pipeline rejects with RESOURCE_EXHAUSTED / HTTP 429 even for priority
	// (error/slow) batches: a 429 is recoverable, an OOM kill is not.
	// Default 512MB; SQLite default 128MB (see applyDriverDefaults).
	IngestPipelineMaxBytes int
	IngestPipelineWorkers  int // default 8 worker goroutines
	// IngestPipelinePerTenantCap caps in-flight batches per tenant so a noisy
	// tenant cannot starve siblings of fresh queue slots when fullness is
	// below the soft-backpressure threshold. When unset it defaults to ~30% of
	// the resolved queue size (see Load) so multi-tenant deployments are
	// protected out of the box; an explicit INGEST_PIPELINE_PER_TENANT_CAP=0
	// disables the cap for single-tenant deployments. Operators can instead
	// pin it to roughly Capacity/N where N is the expected number of
	// concurrently-active tenants, with headroom for short bursts.
	IngestPipelinePerTenantCap int

	// TLS (HTTP + gRPC). When both paths are set, TLS is enabled on both servers.
	// Empty values (default) keep plaintext behavior.
	TLSCertFile string
	TLSKeyFile  string

	// TLSAutoSelfsigned enables zero-friction self-signed TLS bootstrap for dev /
	// internal deployments. Ignored when TLSCertFile/TLSKeyFile are set (explicit
	// cert-file mode wins). Generated material is cached under TLSCacheDir.
	TLSAutoSelfsigned bool
	TLSCacheDir       string

	// API key authentication. When empty, auth middleware is a pass-through.
	// Loaded from API_KEY env var — never logged.
	APIKey string

	// OTelExporterEndpoint enables self-instrumentation. When set, the platform
	// exports its own spans to the configured OTLP endpoint (e.g. "localhost:4317"
	// for self-ingest, or an external collector).
	OTelExporterEndpoint string

	// DefaultTenant is the tenant ID assigned to rows ingested without an explicit
	// X-Tenant-ID header (HTTP) / x-tenant-id gRPC metadata.
	DefaultTenant string

	// OTLPTrustResourceTenant enables resolving the tenant from the OTLP
	// `tenant.id` resource attribute when no transport-level tenant header
	// was provided. Disabled by default because resource attributes are
	// client-controlled — a compromised SDK could set tenant.id to forge
	// another tenant's data. Only turn this on in closed environments where
	// all OTLP producers are trusted.
	OTLPTrustResourceTenant bool

	// APITenantKeysFile, when non-empty, switches API auth from a single
	// shared API_KEY into per-tenant bearer tokens. JSON or YAML, chosen by
	// extension, mapping bearer key to tenant ID (several keys may map to one
	// tenant). Loaded once at startup; the process holds only SHA-256 digests.
	// A matched key's tenant BINDS the request — client-supplied X-Tenant-ID,
	// gRPC metadata, and OTLP `tenant.id` resource attributes are ignored and
	// counted. Empty = disabled (shared-key mode remains for single-tenant dev).
	APITenantKeysFile string

	// AuthTrustExternal turns on proxy-injected identity: the value of
	// AuthExternalTenantHeader is trusted as an authenticated tenant. It is
	// an authentication BYPASS unless the front proxy authenticates callers,
	// strips inbound copies of that header, and the application ports are
	// unreachable except through it — see CLAUDE.md's Authentication section.
	AuthTrustExternal bool

	// AuthExternalTenantHeader is the dedicated identity header (and gRPC
	// metadata key, lower-cased) honoured only when AuthTrustExternal is set.
	// Deliberately distinct from X-Tenant-ID, which stays client-controlled.
	AuthExternalTenantHeader string

	// WSAllowedOrigins is the WebSocket origin allowlist. Entries may be full
	// origins ("https://app.example.com") or bare hosts. Empty means
	// same-host only. Enforced when authentication is enabled or in
	// production; ignored in an unauthenticated development deployment.
	WSAllowedOrigins []string

	// GRPCReflection controls gRPC server reflection. Defaults to true outside
	// production and false in production — reflection enumerates every service
	// and message type to an unauthenticated peer.
	GRPCReflection bool

	// AllowInsecureGRPC waives the production requirement for TLS and
	// authentication on the OTLP gRPC listener. Explicit acknowledgement that
	// telemetry crosses the network unprotected and unauthenticated.
	AllowInsecureGRPC bool

	// DevMode disables origin checks for WebSocket and enables dev-friendly defaults.
	// Derived from APP_ENV == "development".
	DevMode bool

	// gRPC server tuning — protects against huge OTLP batches and connection abuse.
	GRPCMaxRecvMB            int
	GRPCMaxConcurrentStreams int

	// AllowSqliteProd lets operators explicitly acknowledge that SQLite is
	// being used outside dev/test. Without it, a production Env + SQLite
	// combination refuses to start.
	AllowSqliteProd bool

	// WSMaxClients caps simultaneous WebSocket connections to /ws*
	// endpoints. 0 = unlimited (default). When set, new connections past
	// the cap receive HTTP 503. Sized for the operator's expected dashboard
	// audience — small for ops dashboards, larger for read-heavy public UIs.
	WSMaxClients int
	// Aggregate Engine Configuration (Phase 1, accounting only)
	//
	// AggregateMode controls how the aggregation engine participates:
	// - "legacy": aggregate engine inactive, use existing TSDB/GraphRAG path only
	// - "aggregate-shadow": aggregate accounting runs; counts identical to legacy;
	//   the read path uses legacy TSDB; allows A/B testing before switchover
	// - "aggregate": aggregate engine is the only active path; legacy TSDB
	//   aggregation is retired
	AggregateMode string

	// AggregateMaxSeries is the global budget for materialized active series
	// across all signals. Active = present in ≥1 mutable window.
	AggregateMaxSeries int

	// Per-signal sub-caps (materialized active series). These must sum to ≤
	// AggregateMaxSeries. Sum validation runs at startup.
	AggregateMaxSeriesMetrics int // Default 2400
	AggregateMaxSeriesTraces  int // Default 2400
	AggregateMaxSeriesEdges   int // Default 500
	AggregateMaxSeriesLogs    int // Default 500
	AggregateMaxSeriesSystem  int // Default 200

	// Per-service caps. Enforcement order: tenant → service → signal sub-cap →
	// global. These are isolation ceilings, not reservations; sum of per-service
	// budgets may exceed instance-wide cap.
	AggregateMaxOperationsPerService   int     // Default 20
	AggregateMaxTraceSeriesPerService  int     // Default 50
	AggregateMaxLogTemplatesPerService int     // Default 10
	AggregateMaxMetricSeriesPerService int     // Default 50
	AggregateSeriesPerTenantFraction   float64 // Default 0 (disabled); range [0, 1]

	// Baseline configuration for counter temporality tracking.
	// AggregateMaxProducerBaselinesPerSeries caps the number of baseline
	// entries per series (one per producer). Default 8.
	AggregateMaxProducerBaselinesPerSeries int

	// AggregateMaxBaselines is the global budget for baseline entries.
	// 0 (default) = derive as AggregateMaxSeriesMetrics ×
	// AggregateMaxProducerBaselinesPerSeries. Nonzero overrides.
	// Use ResolvedAggregateMaxBaselines() to get the final value.
	AggregateMaxBaselines int

	// AggregateDBPath is the aggregate database file (ADR 0003: its own
	// file, its own WAL, its own PRAGMA stanza — never the main DB).
	AggregateDBPath string

	// AggregateAllowRebuild permits DESTROYING and recreating the
	// aggregate-owned tables when the on-disk schema is partial or
	// version-mismatched. Off by default: v1 has no automatic migrations and
	// a refused startup is better than silent data loss.
	AggregateAllowRebuild bool

	// AggregateSynchronous is the aggregate DB's SQLite synchronous mode,
	// "NORMAL" or "FULL". NORMAL survives process/container death (the ACK
	// contract of #160); FULL additionally survives host power loss at the
	// cost of one fsync per group commit.
	AggregateSynchronous string

	// Group-commit cadence (#160): the first waiter opens a coalescing
	// window of AggregateCommitCoalesceMs, and the commit fires early once
	// the batch reaches the delta-count or byte target.
	//
	// The default is 25 ms, not #160's provisional 5 ms. Measured on the
	// wave-5 2-vCPU acceptance run at 10k pts/s: 5 ms held the writer at a
	// 37.8% duty cycle for a 286 ms ACK p99, while 25 ms measured 109 ms p99
	// and still absorbed the 2x burst. Wider batches amortise the fixed cost
	// of a WAL commit over more deltas, which is the opposite of what the
	// latency arithmetic suggests until you notice the writer is the queue.
	AggregateCommitCoalesceMs int
	AggregateCommitMaxDeltas  int
	AggregateCommitMaxBytes   int

	// Triple admission bound (#160). A breach returns gRPC
	// RESOURCE_EXHAUSTED / HTTP 429, never a silent drop and never an
	// automatic downgrade to bounded-loss ACK.
	AggregateCommitMaxPendingBytes  int
	AggregateCommitMaxWaiters       int
	AggregateCommitMaxPendingDeltas int

	// AggregateTopologyRestoreHorizon is how much FINALIZED history the
	// engine's topology projection is rebuilt from at startup (#194 finding
	// 15). Without it a restart erases the recent service map until new
	// telemetry re-derives one. Accepts a Go duration ("30m", "1h"); 0
	// disables the restore. Clamped to 24h, and internally to the projection's
	// own retention horizon — reading a window the projection would prune on
	// arrival is startup cost with nothing to show for it.
	AggregateTopologyRestoreHorizon time.Duration

	// AggregateFinalizeIntervalSec is how often the writer looks for windows
	// whose lateness horizon has expired.
	AggregateFinalizeIntervalSec int

	// AggregateGCEnabled runs the daily mark-and-sweep over the aggregate
	// dictionary, series and log-template tables. On by default. Turning it
	// off is a disk-growth decision, not a safety one — a pass that fails
	// leaves memory untouched.
	AggregateGCEnabled bool

	// AggregateMaxValueBytes caps the ENCODED length of a dictionary value
	// for every non-tenant kind: service names, metric names, operations,
	// dimension keys, dimension values and dimension tuples. An over-length
	// value routes to __other__ and is NEVER truncated — a truncated value is
	// a different identity wearing the same name.
	AggregateMaxValueBytes int

	// AggregateMaxTenantBytes is the stricter cap on a tenant name, and
	// AggregateMaxTenants is the instance-wide tenant-identity cap. A tenant
	// that breaches either is REJECTED: the point is refused and counted, and
	// the tenant is never collapsed into a shared identity, because a shared
	// overflow tenant is precisely the cross-tenant merge the cap prevents.
	AggregateMaxTenantBytes int
	AggregateMaxTenants     int

	// Per-tenant and instance-wide dictionary count caps for the namespaces
	// that were uncapped before #200. The per-tenant cap bounds one tenant;
	// the instance cap is the backstop for many tenants each staying just
	// under their own. Overflow routes to __other__, per existing semantics.
	AggregateMaxServicesPerTenant  int
	AggregateMaxServices           int
	AggregateMaxDimKeysPerTenant   int
	AggregateMaxDimKeys            int
	AggregateMaxDimValuesPerTenant int
	AggregateMaxDimValues          int
	AggregateMaxDimTuplesPerTenant int
	AggregateMaxDimTuples          int

	// AggregateMetricDims is the parsed AGGREGATE_METRIC_DIMS config:
	// map of metric name -> sorted list of OTLP attribute keys.
	// Empty map when AGGREGATE_METRIC_DIMS is unset/empty.
	// Populated during Load() by parsing and validating the env var.
	AggregateMetricDims map[string][]string

	// Bounded exemplar retention (#176; budgets frozen in #161).
	//
	// These apply ONLY when AggregateMode == "aggregate", where the adaptive
	// Sampler is retired and the exemplar policy is the sole raw-retention
	// gate. In legacy and aggregate-shadow the Sampler is untouched and none of
	// these values are read.
	//
	// Counts AND bytes bind; first breach wins. SamplingLatencyThresholdMs
	// stays the shared definition of "slow" in every mode.
	ExemplarTracesPerServiceWindow    int     // Default 25, unified budget with priority fill
	ExemplarTracesGlobalWindow        int     // Default 1500
	ExemplarBytesPerServiceWindow     int     // Default 524288 (512 KiB)
	ExemplarBytesGlobalWindow         int     // Default 8388608 (8 MiB)
	ExemplarHealthyRate               float64 // Default 0.005 (0.5% eligibility target)
	ExemplarStratumTopK               int     // Default 5, per (operation × status class)
	ExemplarLogsErrorPerServiceWindow int     // Default 50
	ExemplarLogsWarnEnabled           bool    // Default false — WARN is opt-in
	ExemplarLogsWarnPerServiceWindow  int     // Default 20
	ExemplarMaxSpansPerTrace          int     // Default 500
	ExemplarMaxBytesPerTrace          int     // Default 262144 (256 KiB)

	// Exemplar-tier retention and synthesized-log metering (#201 Q2/Q3).
	//
	// ExemplarRetentionDays is SHORTER than HotRetentionDays on purpose: in
	// aggregate mode the raw rows are exemplars attached to a seven-day
	// aggregate dataset, and 576 five-minute windows (two days) at the 3 MiB
	// global window budget is 1.69 GiB of charged payload — 3.38 GiB at the
	// provisional 2x DB/index/FTS amplification, inside the 4.0 GiB main tier
	// with ~0.62 GiB of margin. Seven days of the same rate does not fit.
	ExemplarRetentionDays    int // Default 2, validated 1..HotRetentionDays
	ExemplarSynthLogsPerSpan int // Default 8
	// ExemplarSynthLogsPerTrace bounds the synthesized logs one retained trace
	// may carry across all its spans.
	ExemplarSynthLogsPerTrace int // Default 64

	// Data-volume budget and disk watchdog (#201 Q1/Q5).
	//
	// DataDiskBudgetMB is the configured ceiling for everything this process
	// writes: main relational tier, aggregate.db, DLQ, WAL/temp, and the
	// mandatory unused headroom. Enforcement uses the LOWER of this and the
	// usable volume capacity — a 4 GiB PVC does not become 8 GiB because the
	// config says so.
	DataDiskBudgetMB int    // Default 8192 (8 GiB)
	DataDiskPath     string // Default ./data — any path on the data volume

	// Aggregate runtime readiness thresholds (#194 finding 18).
	//
	// Startup recovery is not the only way an aggregate deployment stops
	// being able to serve: the store can become unreachable, group commits
	// can fail in a row, admission can saturate, the finalizer can wedge and
	// the delta log can grow without bound. Each threshold below turns one of
	// those into a /ready 503. Degraded-not-dead: none of them touches
	// /live, and none of them stops the process.
	//
	// Every threshold takes 0 as "disable this probe" so an operator who
	// disagrees with a default can switch it off without patching the binary.
	ReadyMaxCommitFailureStreak   int     // READY_MAX_COMMIT_FAILURE_STREAK, default 3
	ReadyMaxFinalizeFailureStreak int     // READY_MAX_FINALIZE_FAILURE_STREAK, default 3
	ReadyMaxDeltaLogAgeS          int     // READY_MAX_DELTA_LOG_AGE_S, default 1800
	ReadyMaxAdmissionRatio        float64 // READY_MAX_ADMISSION_RATIO, default 0.9
	ReadyAggregateDiskBudgetMB    int     // READY_AGGREGATE_DISK_BUDGET_MB, default 2304 (the 2.25 GiB tier)
	ReadyMaxAggregateDiskRatio    float64 // READY_MAX_AGGREGATE_DISK_RATIO, default 1.0 (fail at the tier boundary)
}

func Load

func Load(customPath string) (*Config, error)

func (*Config) AuthEnabled added in v0.5.0

func (c *Config) AuthEnabled() bool

AuthEnabled reports whether any credential source is configured: the shared operator key, a per-tenant key file, or a trusted front proxy.

func (*Config) EnforceWSOrigin added in v0.5.0

func (c *Config) EnforceWSOrigin() bool

EnforceWSOrigin reports whether the WebSocket origin policy applies. It does as soon as authentication is configured, and always in production — an unauthenticated development deployment keeps today's permissive behaviour.

func (*Config) GRPCReflectionEnabled added in v0.5.0

func (c *Config) GRPCReflectionEnabled() bool

GRPCReflectionEnabled reports whether to register gRPC server reflection. Production defaults to off; GRPC_REFLECTION=true re-enables it explicitly.

func (*Config) GuardSelfInstrumentation

func (c *Config) GuardSelfInstrumentation()

GuardSelfInstrumentation prevents an amplification loop when OTEL_EXPORTER_OTLP_ENDPOINT points at the binary's own gRPC port. Without this, every span the OTel SDK emits would re-enter Export, generate more spans (one per Export call), and re-enter again — unbounded fan-out.

Strategy: when the configured endpoint resolves to a loopback address, the own service name is auto-added to IngestExcludedServices so the ingest filter drops self-emitted batches. Operators can still override by setting the variable explicitly — the guard only ADDS, never removes.

No-op when self-instrumentation is disabled (empty endpoint) or the endpoint is non-loopback (a separate collector, the operator's responsibility).

func (*Config) IsProduction added in v0.5.0

func (c *Config) IsProduction() bool

IsProduction reports whether APP_ENV names the production environment.

func (*Config) ResolvedAggregateMaxBaselines added in v0.5.0

func (c *Config) ResolvedAggregateMaxBaselines() int

ResolvedAggregateMaxBaselines returns the effective global baseline entry cap. When AggregateMaxBaselines is 0 (default), derives it as AggregateMaxSeriesMetrics × AggregateMaxProducerBaselinesPerSeries. Nonzero AggregateMaxBaselines overrides.

func (*Config) TLSCertFileMode

func (c *Config) TLSCertFileMode() bool

TLSCertFileMode reports whether explicit cert-file TLS is configured. This path has precedence over self-signed.

func (*Config) TLSEnabled

func (c *Config) TLSEnabled() bool

TLSEnabled reports whether HTTPS + gRPC-TLS should be served using any mode (explicit files or auto self-signed).

func (*Config) TLSSelfsignedMode

func (c *Config) TLSSelfsignedMode() bool

TLSSelfsignedMode reports whether the self-signed bootstrap path should be used. False when explicit cert files are set (cert-file wins).

func (*Config) Validate

func (c *Config) Validate() error

Validate checks that all configuration values are within valid ranges. Call this once after Load() during startup to catch misconfiguration early.

func (*Config) ValidateDBForEnv

func (c *Config) ValidateDBForEnv() error

ValidateDBForEnv refuses the combination of SQLite driver + production environment unless AllowSqliteProd is explicitly set. SQLite's single-writer lock caps sustained throughput to ~5 services; using it in production will silently throttle ingestion.

Call once during startup after Load + Validate.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL