telemetry

package
v0.21.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 5, 2026 License: AGPL-3.0 Imports: 36 Imported by: 0

Documentation

Overview

Package telemetry assembles OpenTelemetry tracing and metrics for a Memoh process.

Tracing answers the question logs cannot: where the time went, and which step of a request failed. A record says "workspace snapshot failed"; a trace says the turn spent eleven seconds waiting for the model and then failed on the third tool call. Metrics answer the aggregate version of the same question — which endpoint is slow, for everyone, over the last hour — which a sampled set of traces cannot.

Nothing here is on by default. With no collector configured the process installs no provider at all, leaving OpenTelemetry's global no-ops in place, and the only thing this package contributes is context propagation — which costs nothing and must be unconditional (see Setup).

Index

Constants

View Source
const ScopeName = "github.com/felinics/memoh"

ScopeName identifies spans this repository creates, as opposed to spans from an instrumentation library.

Variables

This section is empty.

Functions

func ContextWithTrigger

func ContextWithTrigger(ctx context.Context, trigger Trigger) context.Context

ContextWithTrigger returns ctx carrying trigger and no current span.

Both halves matter. The trigger is how a span started later in this context says what caused it; dropping the current span is what stops that span from being adopted into the caller's trace, which is the whole point on a connection whose own span has already ended.

func EchoServer

func EchoServer(next echo.HandlerFunc) echo.HandlerFunc

EchoServer traces inbound HTTP requests and records their duration as http.server.request.duration.

This is written here rather than taken from a library because the contributed Echo instrumentation is deprecated in favour of a replacement that has not been released; depending on either would mean depending on something on its way out or on something with no tagged version. What the middleware has to do is small and stable, and writing it keeps two details under our control that a generic one would get wrong for this codebase: the span name uses the matched route rather than the raw path, and the URL attribute goes through the same sanitizer as the access log.

Install it after middleware.RequestID and httpx.RequestIDContext, so the span can carry the id the client is given.

The span and the metric cover the same requests, with the same exceptions below. The metric carries fewer attributes: every attribute value is a separate series for as long as the process runs, so it keeps only what is bounded. The concrete path, the request id and the client address are per request; server.address is whatever Host the caller sent.

A WebSocket is skipped, and has to be. Echo calls the handler and the handler does not return until the socket closes, which for a chat connection is hours. A span around that measures how long someone left a tab open, stays open for its whole duration, and — because it is the current span while it is open — adopts every turn sent over the connection into one trace. Handlers that upgrade trace the handshake themselves, which is the part with a duration worth having.

func RecordPoolStats

func RecordPoolStats(pool *pgxpool.Pool)

RecordPoolStats reports a pool's state as the pgxpool.* metrics: connections in use against the maximum, how often an acquire found the pool empty and how long it waited. A request that is slow because every connection is taken looks, on its span, like a slow query; these say which one it is.

The metrics come from otelpgx, which is used here for nothing else. They are counts read from pool.Stat(), so the objection that keeps its query tracer out (see PgxTracer) does not apply. Each pool is told apart by db.client.connection.pool.name, host:port/database.

Call it once per pool, after the pool is opened. The meter comes from the global provider, which hands off to the one Setup installs whether the pool was opened before Setup ran or after; with metrics off it is a no-op. A failure to register goes to the OpenTelemetry error handler, because missing pool metrics are not a reason to refuse a database.

func StartDetached

func StartDetached(ctx context.Context, name string, opts ...trace.SpanStartOption) (context.Context, trace.Span)

StartDetached begins a span for work whose cause is recorded in ctx by ContextWithTrigger, and an ordinary child span when it is not.

The two cases are both real and must both work. A turn arriving over a WebSocket or off a queue is detached from whatever asked for it, and links to it. A turn arriving over the internal RPC — Channel calling Server in split mode — is genuinely inside its caller's request, and the trace has to stay continuous across that hop.

func StartLinked

func StartLinked(ctx context.Context, trigger Trigger, name string, opts ...trace.SpanStartOption) (context.Context, trace.Span)

StartLinked begins a new trace for name, linked to the operation that caused it. It is a root even when ctx already carries a span: the caller's span is the thing being deliberately not joined.

func Tracer

func Tracer() trace.Tracer

Tracer returns the tracer for spans created by this repository's own code. Before Setup — or with tracing disabled — this is the global no-op.

Types

type PgxTracer

type PgxTracer struct{}

PgxTracer records a span per query.

It implements pgx's own QueryTracer hook rather than pulling in a third-party bridge: the interface is two methods, and doing it here settles the one decision a generic bridge gets wrong for us — query arguments are never recorded. They are the values of the rows being read and written, so a span carrying them would publish exactly the data `docs/logging.md` forbids putting in a log record. The SQL text is recorded; it is code, not data, and without it a slow query span says only that something was slow.

Install it on the pool config: poolCfg.ConnConfig.Tracer = telemetry.PgxTracer{}. With tracing off the global tracer is a no-op, so this costs a call that returns a shared non-recording span.

func (PgxTracer) TraceQueryEnd

func (PgxTracer) TraceQueryEnd(ctx context.Context, _ *pgx.Conn, data pgx.TraceQueryEndData)

func (PgxTracer) TraceQueryStart

func (PgxTracer) TraceQueryStart(ctx context.Context, conn *pgx.Conn, data pgx.TraceQueryStartData) context.Context

type Service

type Service struct {
	// Name is the service.name every span is attributed to: "memoh-server",
	// "memoh-channel". A backend groups by it, so it has to be stable.
	Name string
	// InstanceID distinguishes replicas of the same service. Optional: when
	// empty, each process gets a random one.
	InstanceID string
}

Service describes the process being traced.

type Shutdown

type Shutdown func(context.Context) error

Shutdown flushes pending spans and metrics and releases the exporters.

func Setup

func Setup(ctx context.Context, cfg config.TelemetryConfig, svc Service, log *slog.Logger) (Shutdown, error)

Setup installs context propagation and, when a collector is configured, a tracer provider and a meter provider exporting to it. The returned Shutdown is always safe to call.

Metrics go to the same collector as traces, over the same transport, unless OTEL_METRICS_EXPORTER=none says otherwise: that is the standard way to tell a process its collector only accepts traces.

Propagation is installed even when export is off, and that is deliberate. A propagator only reads and writes the traceparent header; with no provider there is no span to write, so the cost is a map lookup on a header that is not there. What it buys is that the decision is not baked into the binary: a deployment that turns on tracing in one service gets continuous traces across the others rather than a separate disconnected trace per hop, and nobody has to discover that propagation was the missing piece.

type Trigger

type Trigger struct {
	// contains filtered or unexported fields
}

Trigger names the work that caused a detached operation, without joining it.

A turn does not run inside the request that asked for it. The HTTP handler answers immediately and a worker picks the message up later, or the message arrives on a WebSocket that stays open for hours. Making the turn a child of either produces a trace that is wrong in a specific way: a parent that ends before its children, or one that never ends at all, and in both cases every turn on that connection shares one trace id, so asking "how long did this answer take" returns the length of the conversation.

A link says what caused the work without claiming it contains it. The two traces stay separately queryable and each has an honest duration, and a reader who starts at either end can still reach the other.

The zero Trigger links nothing, which is what a turn started with no traceable cause should do.

func TriggerFrom

func TriggerFrom(ctx context.Context) Trigger

TriggerFrom captures the operation ctx is currently in.

It is a serialized form rather than a trace.SpanContext because the value travels in a queue item next to the message: the producer and the consumer are different goroutines with unrelated lifetimes, and a context is not the thing to hand across that gap.

func TriggerFromContext

func TriggerFromContext(ctx context.Context) Trigger

TriggerFromContext reports the trigger ctx carries, or the zero Trigger.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL