logschema

module
v0.1.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 15, 2026 License: MIT

README

LogSchema

LogSchema provides versioned, storage-neutral schemas for structured log data. It is intended to be shared by log producers, profilers, local query tools, and other consumers without coupling them to a parser or storage engine.

The initial scope is deliberately small:

  • common source, resource, attribute, and trace-correlation types;
  • normalized HTTP request log records;
  • normalized database slow-query log records;
  • JSON Schema Draft 2020-12 definitions;
  • pure Go reference types;
  • valid and invalid conformance fixtures.

LogSchema does not provide parsers, SQL normalization, aggregation algorithms, storage, query execution, networking, or a CLI.

Status

LogSchema is pre-v1. Schemas and Go APIs may change incompatibly while ALP, SLP, and the planned local workspace integration are being validated.

Repository layout

core/v1/       Pure Go common reference types
http/v1/       Pure Go HTTP request reference types
sql/v1/        Pure Go database slow-query reference types
schema/        Normative JSON Schemas and an embedded fs.FS
testdata/      Cross-implementation conformance fixtures

The JSON Schemas are normative. Go packages are reference implementations and must pass the same conformance fixtures.

The Go record types use strict JSON decoding: unknown fields, non-canonical decimal values, missing required values, and semantically invalid records are rejected. Encoding validates records built or modified through direct struct access before producing JSON.

The Go packages also provide small validating constructors for new records. They set schema_version and kind; direct struct construction remains available when a producer needs to populate optional fields before validation.

Resolving schema references offline

Schema $id values are stable identifiers. Validation does not require those HTTPS identifiers to be fetched over the network. Consumers should register all embedded resources with their validator before compiling an HTTP or SQL schema.

Go consumers can discover and read resources without hard-coding identifiers or repository paths:

compiler := jsonschema.NewCompiler()
for _, resource := range schema.Resources() {
    data, err := schema.Read(resource.ID)
    if err != nil {
        return err
    }

    document, err := jsonschema.UnmarshalJSON(bytes.NewReader(data))
    if err != nil {
        return err
    }
    if err := compiler.AddResource(resource.ID, document); err != nil {
        return err
    }
}

requestSchema, err := compiler.Compile(schema.HTTPV1RequestID)

Non-Go consumers can use schema/catalog.json to map the same identifiers to files distributed in this repository. Validators must use the catalog or an equivalent local registry; network retrieval is neither required nor assumed.

Encoding rules

  • Nanosecond timestamps and 64-bit measurements are decimal strings in JSON, avoiding loss of precision in JSON consumers.
  • time_unix_nano is the event time. observed_time_unix_nano is the time at which a producer observed or collected the record. Producers must not place an event timestamp in the observed-time field.
  • Units are part of field semantics and names; durations use nanoseconds and sizes use bytes.
  • For nullable fixed fields, an omitted property and an explicit null both mean that no value is available. Zero remains a real value.
  • Reference encoders emit known nullable fields as null, while decoders accept those fields when omitted.
  • Nil Go attribute maps are valid and encode as {}. Attribute properties may be omitted from input records when they are empty.
  • Attribute values cannot be null. Omit the attribute key when no value is available.
  • Go producers must encode binary attribute values explicitly as strings; []byte values are rejected to avoid implicit base64 conversion.
  • Source attributes describe the input source or how records were obtained from it, such as an adapter version, artifact encoding, compression format, or storage cursor. Record content belongs in the record data attributes, and properties of the producing entity belong in resource attributes.
  • Producers use the other source kind with source attributes for source types that LogSchema has not standardized. They should use stable, collision- resistant attribute keys rather than assigning new meanings to existing keys.
  • Source fingerprints include value, algorithm, and version so consumers do not compare identities produced by incompatible algorithms.
  • trace_state stores the serialized W3C tracestate field value when one is available. LogSchema applies bounded character validation, while producers remain responsible for validating the complete W3C member grammar.
  • HTTP URLs are stored as scheme, authority, path, and query components. The Go URI, URLFull, and URLReference methods derive combined forms so duplicate serialized representations cannot disagree. Producers should redact sensitive query values before populating url_query.
  • SQL query_summary is the normalized, low-cardinality grouping form. query_text is optional because it can contain credentials, personal data, or other literals; producers should retain it only under an explicit policy.
  • SQL fingerprints include value, algorithm, and version as one object so the grouping identity remains self-describing. Fingerprint values are not assumed to be safe: an identity algorithm can retain the complete query text.
  • Unknown fields are rejected within a schema version.

Source metadata can help a workspace or an AI-assisted investigation explain where a record came from. Producers must still apply an explicit disclosure policy: credentials, signed URLs, tenant secrets, and other sensitive adapter configuration must not be copied into source attributes by default.

ALP and SLP integration

LogSchema is designed to be the canonical parsed-record representation in ALP and SLP. Parsers, filters, and aggregators consume the LogSchema records directly; consumer-specific values such as floating-point seconds are derived only at query and rendering boundaries:

  • ALP's URI grouping key uses URLReference(); its response time and body size come from duration_nano and response_body_size_bytes.
  • SLP's default abstract query maps to query_summary. --noabstract input maps to the sensitive query_text field. Existing row and byte metrics map directly, and unavailable metrics remain null rather than becoming measured zeros.

Seconds-to-nanoseconds and floating-point byte conversions are producer responsibilities. ALP and SLP integrations must define and test their rounding, overflow, and invalid-input behavior before replacing their internal types.

See COMPATIBILITY.md for versioning rules.

Directories

Path Synopsis
core
v1
Package corev1 provides storage-neutral common reference types for LogSchema v1.
Package corev1 provides storage-neutral common reference types for LogSchema v1.
http
v1
Package httpv1 provides the normalized HTTP request log record v1.
Package httpv1 provides the normalized HTTP request log record v1.
internal
Package schema exposes the normative LogSchema JSON Schema documents.
Package schema exposes the normative LogSchema JSON Schema documents.
sql
v1
Package sqlv1 provides the normalized database slow-query log record v1.
Package sqlv1 provides the normalized database slow-query log record v1.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL