Documentation
¶
Overview ¶
Package inferred implements the inferred schema catalogue for OpenVaultDB.
When data is written in partial or schemaless mode, OpenVaultDB observes and records inferred schema metadata. "Schemaless means no required pre-declared schema. It does not mean no schema information."
The catalogue is persisted to a JSON file after each Catalogue.Observe call. Persistence is synchronous and uses an atomic temp-file + os.Rename pattern to avoid partial writes. This is deliberately simple for MVP: a future iteration could batch writes or use a write-ahead log to reduce fsync overhead on high-write workloads.
Index ¶
Constants ¶
This section is empty.
Variables ¶
This section is empty.
Functions ¶
This section is empty.
Types ¶
type Catalogue ¶
type Catalogue struct {
// contains filtered or unexported fields
}
Catalogue is a thread-safe inferred schema catalogue. It accumulates field type observations across writes and persists them to a JSON file.
func Load ¶
Load reads the catalogue JSON from filePath. A missing file yields an empty catalogue (not an error). Corrupt or unreadable files are returned as errors.
func (*Catalogue) FieldsForCollection ¶
FieldsForCollection derives schema.Field declarations from observations for the named collection's top-level fields only.
Rules:
- All returned fields have Required=false (inferred fields are never required).
- Type is the single observed non-null type mapped to a schema.FieldType, or schema.TypeAny when conflicting types were observed.
- Returns nil if the collection was never observed.
func (*Catalogue) Observe ¶
Observe records one written record's shape into the catalogue and persists the updated catalogue to disk atomically.
MVP trade-off: persistence is synchronous — every Observe call writes the full JSON file via temp-file + os.Rename. This is safe and simple but adds latency on each write. A future iteration could batch flushes or use a write-ahead log.
type CollectionStats ¶
type CollectionStats struct {
// Fields maps a dotted field path (e.g. "address.city") to its statistics.
Fields map[string]*FieldStats `json:"fields"`
// RecordCount is the total number of records observed for this collection.
RecordCount int `json:"record_count"`
// FirstSeen is the UTC time of the first record observed for this collection.
FirstSeen time.Time `json:"first_seen"`
// LastSeen is the UTC time of the most recent record observed for this
// collection.
LastSeen time.Time `json:"last_seen"`
}
CollectionStats holds aggregate statistics for a single collection.
type FieldStats ¶
type FieldStats struct {
// ObservedTypes maps JSON type name ("string","number","boolean","object",
// "array","null") to the count of times that type was observed for this field.
ObservedTypes map[string]int `json:"observed_types"`
// IsArray is true when the field itself (not an element) is always or
// sometimes an array.
IsArray bool `json:"is_array,omitempty"`
// ElementTypes maps JSON type names to counts for element values observed
// within array values of this field.
ElementTypes map[string]int `json:"element_types,omitempty"`
// FirstSeen is the UTC time of the first observation of this field.
FirstSeen time.Time `json:"first_seen"`
// LastSeen is the UTC time of the most recent observation of this field.
LastSeen time.Time `json:"last_seen"`
// SampleCount is the number of collection records in which this field was
// present.
SampleCount int `json:"sample_count"`
// MissingCount is the number of collection records observed that did NOT
// contain this field.
MissingCount int `json:"missing_count"`
// HasConflict is true when more than one non-null type was observed for this
// field across all records.
HasConflict bool `json:"has_conflict,omitempty"`
}
FieldStats holds inferred type statistics for a single field path within a collection.
type Snapshot ¶
type Snapshot struct {
Collections map[string]*CollectionStats `json:"collections"`
}
Snapshot is a deep-copyable, JSON-serializable view of the entire catalogue, suitable for serving over the HTTP API.