snapshot

package
v0.16.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Oct 9, 2026 License: MIT Imports: 28 Imported by: 0

Documentation

Overview

Package snapshot serialises the durable on-disk representation of a gograph snapshot (CSR + LPG + schema) and reads it back into a fresh process.

A snapshot is a directory containing a manifest.json plus one binary file per kept-on-disk component. Publication is atomic on any POSIX filesystem: the writer assembles the new directory under a sibling .tmp path, fsyncs every file, then renames the .tmp directory to its final name. Concurrent readers continue using the previous directory until they re-open.

Example

Example writes a full (v3) snapshot of a labelled graph to a directory and loads it back into a fresh process view, inspecting the manifest and the parsed CSR readback.

package main

import (
	"fmt"
	"os"
	"path/filepath"

	"github.com/FlavioCFOliveira/GoGraph/graph/adjlist"
	"github.com/FlavioCFOliveira/GoGraph/graph/csr"
	"github.com/FlavioCFOliveira/GoGraph/graph/lpg"
	"github.com/FlavioCFOliveira/GoGraph/store/snapshot"
)

func main() {
	dir, err := os.MkdirTemp("", "snapshot-example")
	if err != nil {
		panic(err)
	}
	defer func() { _ = os.RemoveAll(dir) }()

	// A small labelled, weighted graph plus its frozen CSR snapshot.
	g := lpg.New[string, int64](adjlist.Config{Directed: true})
	if err := g.AddEdge("alice", "bob", 7); err != nil {
		panic(err)
	}
	if err := g.SetNodeLabel("alice", "Person"); err != nil {
		panic(err)
	}
	c := csr.BuildFromAdjList(g.AdjList())

	// WriteSnapshotFull lays out csr.bin + labels.bin + properties.bin +
	// a manifest, and (because the graph is string-keyed) a mapper.bin,
	// stamping the manifest at v3. Publication is atomic.
	snapDir := filepath.Join(dir, "snapshot")
	if err := snapshot.WriteSnapshotFull(snapDir, c, g); err != nil {
		panic(err)
	}

	// Load it back: LoadSnapshotFull verifies every component's CRC and
	// returns the parsed readbacks.
	loaded, err := snapshot.LoadSnapshotFull(snapDir)
	if err != nil {
		panic(err)
	}
	// The readback exposes the parsed edge array and the interned label
	// strings. (CSR.Vertices is the dense row-pointer array sized by the
	// largest interned NodeID, not the live-node count, so it is not
	// asserted here.)
	fmt.Printf("manifest version=%d\n", loaded.Manifest.Version)
	fmt.Printf("csr edges=%d\n", len(loaded.CSR.Edges))
	fmt.Printf("label strings=%d\n", len(loaded.Labels.Strings))

}
Output:
manifest version=4
csr edges=1
label strings=1

Index

Examples

Constants

View Source
const CSRFile = "csr.bin"

CSRFile is the conventional file name carrying the CSR triplet (vertices + edges + optional weights) inside a snapshot directory.

View Source
const ConstraintsFile = "constraints.bin"

ConstraintsFile is the conventional file name carrying the durable schema constraint set inside a snapshot directory. It is a sibling of CSRFile and is referenced by an additional entry in the Manifest.Files slice.

The component is OPTIONAL: the writer emits it only when at least one constraint is declared, so a snapshot of a graph with no constraints is byte-identical to one produced before this component existed. A snapshot without the component loads as an empty constraint set — the backward-compatibility contract.

Forward compatibility is one-directional, matching tombstones.bin: a reader that predates this component ignores the unknown file name and so would lose the constraints (a downgrade hazard); upgrades (older snapshot, newer binary) are always safe.

View Source
const CurrentIndexBuilderEpoch uint64 = 1

CurrentIndexBuilderEpoch identifies the secondary-index BUILDER this build ships, and is what WriteSnapshotFull stamps into Manifest.IndexBuilderEpoch on every full snapshot it writes. A reader hydrates an `indexes/<name>.bin` payload only when the manifest names THIS epoch; see Manifest.IndexBuilderEpoch for why, and store/recovery.indexImageReason for where the refusal is applied.

When to bump it

Whenever a defect is fixed in what a backfill WRITES INTO an index — the values, the node set, or the label gate it resolves — because a payload produced before that fix is durable, is rehydrated verbatim, and no fix to the builder can reach it. Bumping is the only remediation: it makes every store written by the defective builder rebuild its indexes once, from the recovered graph, on first open.

Do NOT bump it for a change that leaves the CONTENT a correct builder would produce unchanged — a serialisation-format change (the payload carries its own magic and version, and an index that refuses a payload already falls back to a rebuild), a performance rewrite, or a change confined to how the payload is framed on disk.

Epoch 1 is the first epoch, introduced for rmp #2797. It is the builder that resolves every backfill read through a snapshot-bound read view (rmp #2778) and seeds a UNIQUE value-set from that same read (rmp #2792). Epoch 0 is not a builder: it is the ABSENCE of the field, which is every snapshot written before this one, and which therefore includes every payload the two defects above could have fabricated.

View Source
const DefaultMaxManifestBytes = 32 << 20 // 32 MiB

DefaultMaxManifestBytes is the upper bound the file-backed manifest readers impose on a manifest.json. A manifest is a small JSON document — a version header plus one FileEntry/IndexFileEntry per snapshot component — so even a graph with thousands of on-disk index files stays in the low single-digit MiB. 32 MiB is far above any legitimate manifest yet stops a hostile or corrupt manifest.json (a giant array or string field) from driving a multi-gigabyte transient decode allocation at recovery before any version or CRC check bounds it.

View Source
const EdgeHandlesFile = "edgehandles.bin"

EdgeHandlesFile is the conventional file name carrying the durable per-handle edge metadata (per-CREATE relationship type and properties keyed by the stable edge handle) inside a snapshot directory. It is a sibling of CSRFile and is referenced by an additional FileEntry in the manifest only when the writer emitted at least one record.

View Source
const EdgeLabelSlotOverflow uint32 = ^uint32(0)

EdgeLabelSlotOverflow is the reserved EdgeLabelEntry.Slot value marking a record that belongs to the pair's OVERFLOW list rather than to one slot's inline label column. A pair's durable type state is exactly those two halves, so the file needs to distinguish them; every real ordinal is a slot index and so cannot collide with the maximum uint32.

View Source
const IndexDefsFile = "indexdefs.bin"

IndexDefsFile is the conventional file name carrying the durable secondary index DEFINITION set inside a snapshot directory. It is a sibling of CSRFile and is referenced by an additional entry in the Manifest.Files slice.

It is DISTINCT from IndexesDir ("indexes"), which holds the serialized per-index PAYLOADS (a best-effort recovery speed-up). The definition set (label, property, kind, name) is the load-bearing component: recovery rebuilds each index by backfilling it from the recovered graph, so the durable thing that must survive a WAL-truncating checkpoint is the definition, not the payload (#1755).

The component is OPTIONAL: the writer emits it only when at least one index is declared, so a snapshot of a graph with no indexes is byte-identical to one produced before this component existed. A snapshot without the component loads as an empty index-definition set — the backward-compatibility contract.

Forward compatibility is one-directional, matching constraints.bin: a reader that predates this component ignores the unknown file name and so would lose the index definitions (a downgrade hazard); upgrades (older snapshot, newer binary) are always safe.

View Source
const IndexesDir = "indexes"

IndexesDir is the conventional sub-directory inside a v2 snapshot that holds one index.Serializer-encoded file per registered secondary index. The file name is <indexName>.bin; the manifest records the size and CRC32C of every entry under Manifest.Indexes.

View Source
const IntegrityCRC32CTrailer = "crc32c-trailer"

IntegrityCRC32CTrailer is the value Manifest.Integrity carries when the writer framed the manifest with a checksum trailer. It is the only integrity scheme this build emits; see Manifest.Integrity for why the field exists at all when the trailer is self-identifying.

View Source
const LabelsFile = "labels.bin"

LabelsFile is the conventional file name carrying the durable LPG label state inside a v2 snapshot directory. It is a sibling of CSRFile and is referenced by an additional entry in the Manifest.Files slice.

View Source
const ManifestVersion = 4

ManifestVersion is the highest on-disk schema version this build understands. The current build writes version 4 manifests via WriteSnapshotFull whenever it emits mapper.bin (CSR + labels + properties + mapper + nodeids.bin, fully self-sufficient on load; the mapper may carry holes, which a build before WAL v2 step 1 cannot load, so such a build refuses the version with ErrManifestUnsupported), version 2 manifests via the same writer when it emits no mapper (requires WAL replay to reconstruct the natural-key mapper), and version 1 manifests via the legacy WriteSnapshotCSR code path (CSR-only snapshots). The loader accepts versions 1 to 4; version 3 is what builds before WAL v2 step 1 wrote with a mapper.

View Source
const MapperFile = "mapper.bin"

MapperFile is the conventional file name carrying the durable (NodeID -> natural key) interning table inside a v3 snapshot directory. It is a sibling of CSRFile, LabelsFile and PropertiesFile and is referenced by an additional entry in Manifest.Files when the writer emitted it.

View Source
const MetricLabelsSkippedEmptyRegistry = "store.snapshot.WriteLabels.skippedEmptyRegistry"

MetricLabelsSkippedEmptyRegistry counts how many times WriteLabels proved, from an empty label registry, that no node and no edge carries a label and therefore skipped the O(V + E) collection walk (rmp #2271).

It exists because the optimisation is otherwise unobservable: skipping a walk whose result was empty changes no output byte and no allocation, so a test cannot distinguish "the short-circuit fired" from "the short-circuit is not there and the walk happened to find nothing". Without this counter the regression test passes against the unfixed code — which is exactly what the first draft of that test did.

View Source
const MetricPropertiesSkippedEmptyRegistry = "store.snapshot.WriteProperties.skippedEmptyRegistry"

MetricPropertiesSkippedEmptyRegistry counts how many times WriteProperties proved, from an empty property-key registry, that no node and no edge carries a property and therefore skipped the O(V + E) collection walk (rmp #2271).

The property counterpart of MetricLabelsSkippedEmptyRegistry; see that constant for why an engagement counter is required rather than optional.

View Source
const NodeIDsFile = "nodeids.bin"

NodeIDsFile is the snapshot component that records the mapper's per-shard high-water marks (WAL v2 step 1, docs/design-wal-v2.md §3.3). Recovery hands them to graph.Mapper.LoadFrom, so a key interned after the restore never receives an id at or below one the image could still name, and the holes the capture left out of mapper.bin stay holes.

Layout, little-endian, 2 056 bytes:

magic   u32  "GNID"
version u16  1
shards  u16  256
next    256 × u64

Integrity is the manifest FileEntry CRC32C, like every other component.

View Source
const PropertiesFile = "properties.bin"

PropertiesFile is the conventional file name carrying the durable LPG typed-property state inside a v2 snapshot directory. It is a sibling of CSRFile and LabelsFile and is referenced by an additional entry in Manifest.Files when the writer emitted any property at all.

View Source
const TombstonesFile = "tombstones.bin"

TombstonesFile is the conventional file name carrying the durable LPG node-tombstone set inside a snapshot directory. It is a sibling of CSRFile and is referenced by an additional entry in the Manifest.Files slice.

The component is OPTIONAL: the writer emits it only when the graph has at least one tombstoned node, so a snapshot of a graph that never deleted a node is byte-identical to one produced before this component existed. A snapshot without the component loads as an empty tombstone set — the backward-compatibility contract.

Forward compatibility is one-directional. A reader that predates this component ignores the unknown file name and so would resurrect deleted nodes; reopening a store written by a current binary with an older binary is therefore a downgrade hazard. Upgrades (older snapshot, newer binary) are always safe.

View Source
const WALFormatSegmented = 2

WALFormatSegmented is the Manifest.WALFormat value of a snapshot paired with a segmented write-ahead log.

Variables

View Source
var ErrCSRCorrupted = errors.New("snapshot: csr.bin corrupted")

ErrCSRCorrupted is returned by ReadCSR when the csr.bin payload is structurally malformed: an implausible vertex/edge count, an out-of-range weight-element size, or a weights-array byte length that overflows. It mirrors the per-component corruption sentinels used by the sibling readers (ErrLabelsCorrupted, ErrPropertiesCorrupted, ErrMapperCorrupted) so callers can classify a corrupt CSR the same way. [readVerifiedCSR] / Open wrap it under ErrCorrupted.

View Source
var ErrConstraintsCorrupted = errors.New("snapshot: constraints.bin corrupted")

ErrConstraintsCorrupted is returned by ReadConstraints when the constraints.bin file is structurally malformed (bad magic, unsupported format version, implausible count or string length, or a truncated record).

View Source
var ErrCorrupted = errors.New("snapshot: directory corrupted")

ErrCorrupted is returned by Open when a component file CRC32C disagrees with the manifest, or when a referenced file is missing or shorter than expected.

View Source
var ErrEdgeHandlesCorrupted = errors.New("snapshot: edgehandles.bin corrupted")

ErrEdgeHandlesCorrupted is returned by ReadEdgeHandles when the edgehandles.bin file is structurally malformed (bad magic, unsupported version, implausible count, label/key index past its table, unknown property kind, or a truncated record).

View Source
var ErrFieldTooLong = errors.New("snapshot: field too long for the reader's cap")

ErrFieldTooLong is the sentinel every snapshot length-prefix refusal wraps. A capture that would have to emit a field the matching reader is required to reject fails with it instead, so the caller sees a typed, testable error rather than a durable file nothing can load (rmp #2743).

It is the snapshot-format counterpart of github.com/FlavioCFOliveira/GoGraph/store/txn.ErrFieldTooLong, deliberately sharing that name: the two durable formats — the WAL and the snapshot — now refuse an over-long field in the same shape, with the same sentinel name and the same message layout, so an operator who has learned one has learned the other. Consistency across the two formats was judged worth more than a locally nicer name.

It reaches the caller from every snapshot component writer, and from there out through the checkpointer, whose phase-2 writeSnapshot failure returns BEFORE the phase-3 WAL prefix truncation. That ordering is the whole point of the guard: the refusal happens at the one choke point where the WAL still holds the data, so an unwritable value costs a failed checkpoint and an un-reclaimed WAL, never a committed write.

View Source
var ErrIndexDefsCorrupted = errors.New("snapshot: indexdefs.bin corrupted")

ErrIndexDefsCorrupted is returned by ReadIndexDefs when the indexdefs.bin file is structurally malformed (bad magic, unsupported format version, implausible count or string length, or a truncated record).

View Source
var ErrLabelsCorrupted = errors.New("snapshot: labels.bin corrupted")

ErrLabelsCorrupted is returned by ReadLabels when the labels.bin file is structurally malformed (bad magic, truncated record, or a label-string index that points beyond the embedded string table).

View Source
var ErrManifestCorrupted = errors.New("snapshot: manifest corrupted")

ErrManifestCorrupted is returned when the manifest does not parse as JSON or its file list disagrees with what is on disk.

View Source
var ErrManifestTooLarge = errors.New("snapshot: manifest exceeds maximum size")

ErrManifestTooLarge is returned by LoadManifest — and so by the file-backed readers ReadManifestFile / ReadManifestFileFS above it — when a manifest.json exceeds DefaultMaxManifestBytes. Inspect with errors.Is. It bounds the transient allocation an attacker-supplied or corrupt snapshot directory can force at store-open time, mirroring the DefaultMaxBytes ceiling every sibling loader (csv/jsonl/graphml, the WAL frame decoder, the csrfile loader) already applies.

The ceiling bounds the BYTES READ, not merely the bytes the JSON decoder consumes. It bounded only consumption before the checksum trailer existed, because the decoder stops at the closing brace and so never read trailing padding — which meant a manifest.json of any length on disk was accepted. Verifying a trailer requires reading to the end of the file, so the ceiling now covers the file, which is both the stricter reading and the one the stated purpose implies.

View Source
var ErrManifestUnsupported = errors.New("snapshot: manifest version unsupported")

ErrManifestUnsupported is returned by LoadManifest when the manifest version is newer than this build understands.

View Source
var ErrMapperApply = errors.New("snapshot: cannot apply mapper")

ErrMapperApply is returned by ApplyMapperToGraph when the supplied readback violates an invariant the writer is responsible for upholding (an intra index at or above the shard's high-water mark, hash/shard mismatch, duplicate key, or a non-empty target mapper). It wraps the underlying [graph.ErrMapper…] sentinels so callers can branch on the typed cause via errors.Is.

View Source
var ErrMapperCorrupted = errors.New("snapshot: mapper.bin corrupted")

ErrMapperCorrupted is returned by ReadMapperString when the mapper.bin file is structurally malformed (bad magic, unsupported format version, truncated record, or an implausible length prefix).

View Source
var ErrNodeIDsCorrupted = errors.New("snapshot: nodeids.bin corrupted")

ErrNodeIDsCorrupted is returned when nodeids.bin does not parse: a wrong magic, an unknown version, a shard count other than 256, or a short file.

View Source
var ErrPropertiesCorrupted = errors.New("snapshot: properties.bin corrupted")

ErrPropertiesCorrupted is returned by ReadProperties when the properties.bin file is structurally malformed (bad magic, truncated record, key index past the embedded key table, unknown kind, or a value length implausibly large).

View Source
var ErrTombstonesCorrupted = errors.New("snapshot: tombstones.bin corrupted")

ErrTombstonesCorrupted is returned by ReadTombstones when the tombstones.bin file is structurally malformed (bad magic, unsupported format version, implausible count, or a truncated record).

View Source
var ErrWeightCodecRequired = errors.New(
	"snapshot: csr.bin carries codec-encoded edge weights but no WeightCodec was supplied to decode them")

ErrWeightCodecRequired is returned when applying a CSR whose weights section is codec-encoded ([weightSizeCodec]) to a graph, without a weight codec to decode it.

Like ErrWeightNotPersistable this fails rather than degrades. Silently applying zero weights would discard values that ARE durably on disk, which is the same class of loss on the read side.

View Source
var ErrWeightNotPersistable = errors.New(
	"snapshot: edge weights cannot be persisted: weight type is not a fixed-width primitive and no WeightCodec was supplied")

ErrWeightNotPersistable is returned by the CSR writer when the graph carries edge weights that cannot be persisted: the weight type W is not one of the fixed-width primitives [csrWeightSize] knows, and no [txn.WeightCodec] was supplied to encode it.

It is deliberately a hard failure and not a degradation. Returning it fails the snapshot write, which fails the checkpoint, which leaves the WAL prefix holding the only surviving copy of those weights UNTRUNCATED. Writing a weightless snapshot instead — what this code did before rmp #2526 — let the checkpoint go on to discard that prefix and made the loss permanent.

The remedy is to supply the store's weight codec to the checkpointer:

checkpoint.New(cfg, g, wlog, &mu,
	checkpoint.WithMapperCodec[N, W](st.Codec()),
	checkpoint.WithWeightCodec[N, W](st.WeightCodec()))

Functions

func ApplyCSRToGraph

func ApplyCSRToGraph[N comparable, W any](g *lpg.Graph[N, W], rb *CSRReadback) error

ApplyCSRToGraph replays the adjacency in rb into g. The pre- condition is that g's underlying mapper has already been populated with every NodeID referenced by rb — typically by an immediately- preceding ApplyMapperToGraph (v3 snapshots) or by a WAL replay (v2 snapshots that pair with a WAL prefix). Records whose endpoints the mapper cannot resolve are skipped and counted via `store.snapshot.ApplyCSR.unresolved`; the function does not return an error for them so a partial mapper degrades cleanly rather than aborting recovery mid-way.

Weights in the DENSE native layout are decoded for the fixed-width primitives (int8/uint8/bool, int16/uint16, int32/uint32/float32, int/uint/int64/uint64/float64/uintptr). Weights in the codec-encoded layout are decoded through the supplied weight codec, which covers every other W.

The metric `store.snapshot.ApplyCSR.weightFallback` counts dense-layout slots that decoded to the zero value for want of bytes. It is NOT a signal that weights were lost to an unsupported type: that loss used to happen at WRITE time, where the file recorded "no weights" and there was nothing here to count. rmp #2526 removed that path — the writer now refuses — so this counter means a malformed file, not an unsupported weight type.

ApplyCSRToGraph is idempotent against a freshly-loaded mapper but not against a graph that already contains edges: re-applying a CSR to a graph with existing edges may duplicate them in multigraph mode or no-op in simple-graph mode. Callers should run this exactly once per recovery, immediately after the mapper restore and before any WAL replay.

rb is passed by pointer to avoid copying the three slices in the readback (vertices, edges, weight bytes) on every call. The function does not mutate rb.

A snapshot whose weights are codec-encoded cannot be applied through this entry point, which has no codec to decode them: it returns ErrWeightCodecRequired. Use ApplyCSRToGraphWithWeightCodec.

func ApplyCSRToGraphWithWeightCodec added in v0.12.0

func ApplyCSRToGraphWithWeightCodec[N comparable, W any](g *lpg.Graph[N, W], rb *CSRReadback, wdec weightDecoder[W]) error

ApplyCSRToGraphWithWeightCodec is the weight-codec-aware variant of ApplyCSRToGraph. wdec decodes the variable-width weights section written by WriteCSRWithWeightCodec; pass the owning store's codec, [txn.Store.WeightCodec].

A nil wdec is accepted and behaves exactly as ApplyCSRToGraph for every snapshot whose weights are in the dense native layout (the fixed-width primitives) or absent. It is refused with ErrWeightCodecRequired for a snapshot whose weights ARE codec-encoded: those weights are durably present on disk, and applying zero weights instead would discard committed data (rmp #2526).

CSR apply walks every src slot, resolves endpoints, decodes weight by W type

func ApplyEdgeHandlesToGraph

func ApplyEdgeHandlesToGraph[N comparable, W any](g *lpg.Graph[N, W], rb EdgeHandlesReadback) error

ApplyEdgeHandlesToGraph replays rb into a live g, re-attaching every per-handle edge label and property keyed by its stable handle and the endpoint NodeID pair. It MUST run AFTER the mapper and CSR (with its handle column) are applied so the handle the record references is already live on the adjacency slot — though the per-handle metadata stores are keyed by (NodeID pair, handle) directly and do not require the adjacency edge to be present, so a record whose edge the CSR did not materialise is still re-attached harmlessly. The handle high-water counter is re-seeded for every record so a post-recovery edge creation never re-mints a live handle (invariant I5).

It returns an error wrapping lpg.ErrTokenTooLong when a record carries a relationship type or property key longer than lpg.MaxTokenLen bytes (rmp #2748): the engine refuses such a token on every write path, so a snapshot carrying one predates that bound, and loading it silently would hide the data it names. The error result is a breaking change: this function used to return nothing.

func ApplyLabelsToGraph

func ApplyLabelsToGraph[N comparable, W any](g *lpg.Graph[N, W], rb LabelsReadback) error

ApplyLabelsToGraph replays rb into a live g. The pre-condition is that g's underlying mapper has already been populated with every NodeID referenced by rb — typically by replaying the WAL prefix covered by the snapshot, or by re-issuing the original AddNode / AddEdge calls. Records whose NodeID cannot be resolved by the mapper are skipped and counted via the `store.snapshot.ApplyLabels.unresolved` metric counter; the function does not return an error for them so a partial mapper degrades cleanly rather than aborting recovery mid-way.

Edge label records whose endpoints are resolvable but whose edge is absent from the adjacency list (e.g., the CSR was not yet applied) are likewise skipped and counted under `store.snapshot.ApplyLabels.edgeMissing`; this matches lpg.Graph.SetEdgeLabel's own no-op-on-missing-edge contract.

Per-slot versus per-pair replay

From labels.bin version 2 ([labelsFormatVersionPerSlot]) each edge record names ONE place on the pair. A record carrying a canonical ordinal is replayed through lpg.Graph.SetEdgeRelTypeAtSlotByID, so a multigraph pair's parallel slots keep the types they actually carried; a record marked EdgeLabelSlotOverflow is replayed through lpg.Graph.AddEdgeRelTypeOverflowByID, restoring the pair-wide half of its type state. A record whose ordinal the live adjacency cannot resolve — the pair now holds fewer slots than the snapshot recorded — is skipped and counted under `store.snapshot.ApplyLabels.slotMissing` rather than being retargeted at some other slot, because guessing would be the very failure this format removes.

A version-1 readback (and any hand-constructed one, which leaves Version at zero) is replayed through lpg.Graph.SetEdgeLabel, which names the PAIR. That is exactly what those records meant when they were written; see [labelsFormatVersionPerSlot] for why an older file is read rather than rejected.

func ApplyMapperToGraph

func ApplyMapperToGraph[N comparable, W any](g *lpg.Graph[N, W], rb MapperReadback) error

ApplyMapperToGraph rebuilds g's underlying graph.Mapper from the snapshot readback. It is only meaningful for string-keyed graphs: any other N type returns nil without touching g, because no v3 mapper.bin is ever produced for non-string graphs. The caller is expected to invoke this function before ApplyCSRToGraph, ApplyLabelsToGraph, or ApplyPropertiesToGraph so subsequent resolution calls see the restored interning table.

Pre-condition: g must hold a fresh (empty) mapper. Calling on a graph that already has interned values returns ErrMapperApply wrapping graph.ErrMapperNotEmpty so the caller can distinguish a programmer error from a corruption error.

Concurrency: ApplyMapperToGraph is not safe to call concurrently with mutations or reads on g. It is intended for the one-shot snapshot-load phase of recovery.

func ApplyMapperToGraphWithCodec

func ApplyMapperToGraphWithCodec[N comparable, W any](g *lpg.Graph[N, W], rb MapperReadback, codec keyDecoder[N]) error

ApplyMapperToGraphWithCodec rebuilds g's underlying graph.Mapper from a version-2 (codec) snapshot readback for ANY comparable key type N. Each MapperRawPair carries the codec-encoded key bytes the snapshot writer produced via WriteMapper; this function decodes them back into N via the supplied codec (the same one the store uses on the WAL) and seeds the interning table through graph.Mapper.LoadFrom.

It is the codec-aware dual of ApplyMapperToGraph: recovery calls this when the loaded readback carries RawPairs (non-string keys) and the string-specialised path when it carries Pairs. An empty readback is a no-op.

Pre-condition and concurrency contract match ApplyMapperToGraph: g must hold a fresh (empty) mapper, and the call must not race with any other access to g. A decode failure surfaces as ErrMapperApply wrapping the codec error; a structural violation surfaces as ErrMapperApply wrapping the relevant [graph.ErrMapper…] sentinel.

func ApplyPropertiesToGraph

func ApplyPropertiesToGraph[N comparable, W any](g *lpg.Graph[N, W], rb PropertiesReadback) error

ApplyPropertiesToGraph replays rb into a live g. The pre-condition is that g's underlying mapper has already been populated with every NodeID referenced by rb — typically by replaying the WAL prefix covered by the snapshot, or by re-issuing the original AddNode / AddEdge calls. Records whose NodeID cannot be resolved by the mapper are skipped and counted via the `store.snapshot.ApplyProperties.unresolved` metric counter; the function does not return an error for them so a partial mapper degrades cleanly rather than aborting recovery mid-way.

Edge property records whose endpoints are resolvable but whose edge is absent from the adjacency list (e.g., the CSR was not yet applied) are likewise skipped and counted under `store.snapshot.ApplyProperties.edgeMissing`; this matches lpg.Graph.SetEdgeProperty's own no-op-on-missing-edge contract.

apply: bounds + mapper resolve + edge resolve + kind decode

func ApplyTombstonesToGraph

func ApplyTombstonesToGraph[N comparable, W any](g *lpg.Graph[N, W], rb TombstonesReadback) error

ApplyTombstonesToGraph replays rb into a live g, re-tombstoning every NodeID the snapshot recorded as removed. It must run AFTER the snapshot nodes are loaded (mapper + CSR) so the ids it restores reference the same stable slots; it re-tombstones by id directly via lpg.Graph.RestoreTombstones and so does not require the natural keys to be resolvable.

A later WAL re-create (OpAddNode) for any of these ids still revives it, preserving the chronology of a delete→recreate cycle that straddles the snapshot boundary.

It returns the error lpg.Graph.RestoreTombstones reports — lpg.ErrIndexedRawWrite when g already has a secondary index registered, which recovery never does because it applies the snapshot to a graph it has just constructed.

func ReadNodeIDs added in v0.16.0

func ReadNodeIDs(r io.Reader) (*[graph.MapperShards]uint64, error)

ReadNodeIDs parses a nodeids.bin from r.

func VerifyMapperDecodable added in v0.14.1

func VerifyMapperDecodable[N comparable](rb MapperReadback, codec keyDecoder[N]) error

VerifyMapperDecodable decodes every codec-encoded key in rb through codec and discards the decoded values, returning nil only when the readback would survive the decode step of ApplyMapperToGraphWithCodec — the call recovery makes on this same readback. It performs NO filesystem I/O: rb is already in memory, so the whole cost is CPU over bytes the caller has already read.

It exists for the checkpointer (rmp #2780). Parsing a published snapshot establishes that the reader accepts it; it does NOT establish that the applier accepts it, and the checkpointer discards the WAL prefix — the only other copy of the data — on the strength of the latter. A key whose encoding the codec refuses, or whose bytes the codec does not consume in full, passes LoadSnapshotFull and fails recovery.

An empty readback (a version-1 string mapper, whose keys carry no codec framing and land in MapperReadback.Pairs, or no mapper.bin at all) is a no-op and returns nil: there is nothing to decode. A nil codec with a non-empty MapperReadback.RawPairs is refused, exactly as ApplyMapperToGraphWithCodec refuses it — the keys cannot be verified and must not be assumed good.

What this does NOT establish

The structural half of the apply — graph.Mapper.LoadFrom's intra-shard gap, hash/shard and duplicate-key invariants — is NOT checked here, because running it requires an interning table to load into, and the caller's purpose is to verify an image rather than to build a graph from it. Nor does it establish media durability: like the parse it accompanies, it reads bytes the process handed to the OS moments earlier.

Errors wrap ErrMapperApply and the underlying codec error, so callers can branch on either via errors.Is.

func WriteCSR

func WriteCSR[W any](w io.Writer, c *csr.CSR[W]) (size int64, crc uint32, err error)

WriteCSR serialises c to w, returning the number of bytes written and the CRC32C of the serialised payload. The on-disk layout is:

uint64 nVertices      (little-endian)
uint64 nEdges
uint8  hasWeights     (1 = weights array present)
uint8  weightSizeBytes (0 when hasWeights = 0; 0xFF = codec-encoded)
[vertices]            (nVertices * 8 bytes)
[edges]               (nEdges * 8 bytes)
[weights]             (nEdges * weightSizeBytes bytes, when present)
uint8  hasHandles      (OPTIONAL trailing block; 1 = handle array present)
[handles]             (nEdges * 8 bytes, when hasHandles = 1)

WriteCSR persists weights only for the fixed-width primitives [csrWeightSize] knows. A graph whose weights are of any other type — a struct, a NAMED integer type such as time.Duration, a string — needs a weight codec, which this entry point does not take: it returns ErrWeightNotPersistable rather than writing a weightless snapshot, because a weightless snapshot would let the caller's checkpoint truncate the WAL prefix holding the only surviving copy (rmp #2526). Use WriteCSRWithWeightCodec to persist those weights.

The trailing handles block (Stage 2 of the stable-edge-handle work) is emitted ONLY when the source CSR carries a per-slot handle column (csr.CSR.HandlesSlice != nil). A graph that never used AddEdgeH produces no trailing block, so its csr.bin is byte-identical to one written before this column existed — the v1 golden and the cross-process byte-equality fixtures are unaffected. [readCSRLimited] detects the block by attempting to read one more byte after the weights array: present → handles follow, EOF → none (the backward-compatible read branch).

func WriteCSRWithWeightCodec added in v0.12.0

func WriteCSRWithWeightCodec[W any](w io.Writer, c *csr.CSR[W], wenc weightEncoder[W]) (size int64, crc uint32, err error)

WriteCSRWithWeightCodec is the weight-codec-aware variant of WriteCSR. It persists edge weights of ANY type W by encoding each one through wenc — pass the owning store's codec, [txn.Store.WeightCodec] — into a variable-width section indexed by an offsets array. See csr_weight_codec.go for the layout and the compatibility argument.

The codec is consulted ONLY for weight types the fixed-width path cannot size. A float64, int64 or int32 weight still takes the original dense native layout and the bytes are unchanged, so a snapshot written with a codec installed is byte-identical to one written without it whenever W is one of those types.

A nil wenc is accepted and behaves exactly as WriteCSR.

func WriteCapture added in v0.11.0

func WriteCapture[W any](dir string, capt *Capture[W], constraints []ConstraintSpec, indexDefs []IndexDefSpec) error

WriteCapture publishes a Capture taken earlier to dir, emitting constraints.bin from constraints and indexdefs.bin from indexDefs (each omitted when nil or empty). It touches no graph: every graph-derived byte was fixed when the Capture was taken, which is what lets a checkpointer publish LOCK-FREE while still writing a single-instant, transaction-boundary image (rmp #2269).

constraints and indexDefs must be captured by the caller in the SAME exclusion window as capt, or the published snapshot mixes two instants again.

func WriteCaptureFS added in v0.11.0

func WriteCaptureFS[W any](fsys fileSystem, dir string, capt *Capture[W], constraints []ConstraintSpec, indexDefs []IndexDefSpec) error

WriteCaptureFS is the filesystem-seam variant of WriteCapture, routing every filesystem operation through fsys. It is what the deterministic-simulation harness (internal/sim) publishes a capture through so it can crash mid-publish. Passing osBackend{} reproduces WriteCapture byte-for-byte.

The fsys parameter type ([fileSystem]) is intentionally unexported: an external package cannot name it but can still supply a value satisfying it (mirroring wal.OpenWith).

func WriteConstraints added in v0.2.0

func WriteConstraints(w io.Writer, specs []ConstraintSpec) (size int64, crc uint32, err error)

WriteConstraints serialises specs into w in the constraints.bin format. It returns the number of bytes written and the CRC32C of the serialised payload — both stored in the manifest's FileEntry for the component so LoadSnapshotFull can verify integrity at load time.

The CRC32C covers the entire on-disk file, including the magic header, matching the tombstones.bin / labels.bin discipline. The records are emitted in deterministic order (kind, label, property, name) so the component is byte-identical across writes of the same logical state.

func WriteEdgeHandles

func WriteEdgeHandles[N comparable, W any](w io.Writer, g *lpg.Graph[N, W], at *lpg.Snapshot) (size int64, crc uint32, emitted bool, err error)

WriteEdgeHandles serialises every per-handle edge label and property attached to g into w in the edgehandles.bin format. It returns the number of bytes written and the CRC32C of the serialised payload — both stored in the manifest's FileEntry so LoadSnapshotFull can verify integrity at load time. It returns (0, 0, nil, false) when the graph carries no per-handle metadata, signalling the caller to omit the component entirely.

Records are emitted in the deterministic order lpg.Graph.WalkEdgeHandles yields (the same source-node order csr.bin / labels.bin use), and within a record the label names and property keys are sorted, so the component is byte-stable across writes of the same logical state — the cross-process byte-equality contract the snapshot relies on.

edgehandles write: collect + two string tables + per-record labels + per-record props, each guarded

func WriteIndexDefs added in v0.6.0

func WriteIndexDefs(w io.Writer, specs []IndexDefSpec) (size int64, crc uint32, err error)

WriteIndexDefs serialises specs into w in the indexdefs.bin format. It returns the number of bytes written and the CRC32C of the serialised payload — both stored in the manifest's FileEntry for the component so LoadSnapshotFull can verify integrity at load time.

The CRC32C covers the entire on-disk file, including the magic header, matching the constraints.bin / tombstones.bin discipline. The records are emitted in deterministic order (kind, name, label, property) so the component is byte-identical across writes of the same logical state.

func WriteLabels

func WriteLabels[N comparable, W any](w io.Writer, g *lpg.Graph[N, W], at *lpg.Snapshot) (size int64, crc uint32, err error)

WriteLabels serialises every node and edge label attached to g into w in the labels.bin format documented at the top of this file. It returns the number of bytes written and the CRC32C of the serialised payload — both stored in the manifest's FileEntry for the labels.bin component so Open / LoadSnapshotFull can verify integrity at load time.

The CRC32C covers the entire on-disk file, including the magic header. This lets the manifest's CRC field validate every byte of labels.bin end-to-end without a separate inner-payload checksum.

The on-disk string table is populated by walking g's lpg.LabelRegistry in interning order; the labelStringIdx written for each (node | edge) record indexes into that table. Because LabelID is itself assigned in interning order, this preserves the registry's identity across save and load: the reader interns each name back in the same order and observes the same LabelID values without an extra remap step.

lpg.LabelRegistry is a lock-free, copy-on-write structure (see its own doc): there is no RLock held across the string-table emission and the later node/edge enumeration, which read the registry and the live graph independently, at different times, via the same lock-free / per-shard-RLock-only primitives the public LPG accessors expose. A name interned strictly between those two reads and immediately attached to a node/edge is therefore visible to the enumeration but absent from the already-captured string table: collectNodeLabelRecords / collectEdgeLabelRecords detect this and return a "not in registry snapshot" error. Rather than abort the whole checkpoint attempt on that race, WriteLabels re-captures and retries as one consistent unit (self-healing, bounded by [maxRegistryCaptureRetries]); because the registry is monotonic and append-only, a re-capture is guaranteed to include the newly-interned name, so the retry converges — in the steady state on the first attempt (#1880). Only sustained, adversarial schema churn that races every attempt exhausts the budget, in which case WriteLabels falls back to the prior fail-stop (return the error with nothing written or truncated — the identical fail-safe posture [Checkpointer.truncatePrefixLocked] uses for a schema DDL racing phase 2, #1774), never a Consistency violation.

func WriteManifest

func WriteManifest(w io.Writer, m Manifest) error

WriteManifest writes m to w in canonical (pretty-printed) JSON followed by a CRC32C trailer over the whole document (see the integrity section on Manifest). It sets Manifest.Integrity on its own copy of m, so callers need not — and must not rely on the value they passed in surviving.

The document is buffered rather than streamed because the checksum must be computed over the finished bytes; it is also what makes the file reach w in a single Write.

func WriteMapper

func WriteMapper[N comparable](w io.Writer, m *graph.Mapper[N], codec keyEncoder[N], include func(graph.NodeID) bool) (size int64, crc uint32, err error)

WriteMapper serialises every (NodeID -> key) pair held by m into w, generalising WriteMapperString to any comparable key type N via the supplied codec. It returns the number of bytes written and the CRC32C of the serialised payload, both recorded in the manifest's FileEntry for the mapper.bin component so LoadSnapshotFull can verify integrity at load time.

Back-compatibility: when N is the canonical string type the function delegates to WriteMapperString, which emits the frozen version-1 layout (raw UTF-8 key bytes, no codec framing). The on-disk image is therefore byte-identical to every string mapper.bin produced before the codec generalisation, regardless of which string codec is supplied. For any other N the function emits a version-2 layout whose per-record key bytes are the output of codec.Encode.

On-disk layout (version 2, all little-endian):

uint32  magic           ('GMAP', 0x50414D47)
uint16  formatVersion   (2)
uint64  pairCount
for each pair:
    uint64  nodeID
    uint32  keyLen       (length of the codec-encoded key bytes)
    [keyLen]byte key     (codec.Encode output for the natural key)

Pairs are emitted in graph.Mapper.Walk order (shard-major, intra-index-major) so the read side reconstructs the mapper deterministically. The CRC32C covers the entire on-disk file, including the magic header.

func WriteMapperString

func WriteMapperString(w io.Writer, m *graph.Mapper[string], include func(graph.NodeID) bool) (size int64, crc uint32, err error)

WriteMapperString serialises every (NodeID -> string key) pair held by m into w in the mapper.bin format documented below. It returns the number of bytes written and the CRC32C of the serialised payload — both stored in the manifest's FileEntry for the mapper.bin component so LoadSnapshotFull can verify integrity at load time.

On-disk layout (all little-endian):

uint32  magic           ('GMAP', 0x50414D47)
uint16  formatVersion   (1)
uint64  pairCount
for each pair:
    uint64  nodeID
    uint32  keyLen
    [keyLen]byte key

Pairs are emitted in graph.Mapper.Walk order (shard-major, intra-index-major) so the read side can reconstruct the mapper deterministically.

The CRC32C covers the entire on-disk file, including the magic header. The reader recomputes the CRC end-to-end at load time.

func WriteNodeIDs added in v0.16.0

func WriteNodeIDs(w io.Writer, next *[graph.MapperShards]uint64) (int64, uint32, error)

WriteNodeIDs serialises next to w and returns the byte count and the CRC32C of what it wrote.

func WriteProperties

func WriteProperties[N comparable, W any](w io.Writer, g *lpg.Graph[N, W], at *lpg.Snapshot) (size int64, crc uint32, err error)

WriteProperties serialises every node and edge property attached to g into w in the properties.bin format documented at the top of this file. It returns the number of bytes written and the CRC32C of the serialised payload — both stored in the manifest's FileEntry for the properties.bin component so LoadSnapshotFull can verify integrity at load time.

The CRC32C covers the entire on-disk file, including the magic header. This lets the manifest's CRC field validate every byte of properties.bin end-to-end without a separate inner-payload checksum.

The on-disk key string table is populated by walking g's lpg.PropertyKeyRegistry in interning order; the keyIdx written for each (node | edge) record indexes into that table.

Concurrency contract: lpg.PropertyKeyRegistry is a lock-free, copy-on-write structure (see its own doc), so the walk relies on the same lock-free / RLock-only primitives the public LPG accessors expose, with no RLock held across the key-table snapshot and the later node/edge enumeration — the identical structure WriteLabels uses for labels, and the identical narrow race applies: a key interned strictly between those two reads and immediately attached to a node/edge is visible to the enumeration but absent from the already-captured key table, which the collectors detect and reject with a "not in registry snapshot" error. As in WriteLabels, WriteProperties re-captures and retries the key-table snapshot plus record collection as one consistent unit (self-healing, bounded by [maxRegistryCaptureRetries]); the monotonic append-only registry guarantees the retry converges (#1880). Budget exhaustion under sustained churn falls back to the prior fail-stop — a clean abort of this checkpoint attempt with nothing written or truncated — never a silently inconsistent record.

func WriteSnapshotCSR

func WriteSnapshotCSR[W any](dir string, c *csr.CSR[W]) error

WriteSnapshotCSR is the legacy high-level helper that lays a snapshot directory containing a v1 manifest plus the CSR. It is retained for backward compatibility: callers that also need LPG label durability must use WriteSnapshotFull which writes a v2 manifest with both csr.bin and labels.bin. Atomic publication is achieved by assembling the snapshot under dir + ".tmp" and renaming it to dir on success.

Example

ExampleWriteSnapshotCSR shows the lighter, CSR-only (v1) path: it writes just the adjacency and reads it straight back with Open.

package main

import (
	"fmt"
	"os"
	"path/filepath"

	"github.com/FlavioCFOliveira/GoGraph/graph/adjlist"
	"github.com/FlavioCFOliveira/GoGraph/graph/csr"
	"github.com/FlavioCFOliveira/GoGraph/store/snapshot"
)

func main() {
	dir, err := os.MkdirTemp("", "snapshot-csr-example")
	if err != nil {
		panic(err)
	}
	defer func() { _ = os.RemoveAll(dir) }()

	a := adjlist.New[string, int64](adjlist.Config{Directed: true})
	if err := a.AddEdge("a", "b", 1); err != nil {
		panic(err)
	}
	if err := a.AddEdge("a", "c", 2); err != nil {
		panic(err)
	}
	c := csr.BuildFromAdjList(a)

	snapDir := filepath.Join(dir, "snapshot")
	if err := snapshot.WriteSnapshotCSR(snapDir, c); err != nil {
		panic(err)
	}

	loaded, err := snapshot.Open(snapDir)
	if err != nil {
		panic(err)
	}
	fmt.Printf("manifest version=%d\n", loaded.Manifest.Version)
	fmt.Printf("csr edges=%d\n", len(loaded.CSR.Edges))

}
Output:
manifest version=1
csr edges=2

func WriteSnapshotCSRCtx

func WriteSnapshotCSRCtx[W any](ctx context.Context, dir string, c *csr.CSR[W]) error

WriteSnapshotCSRCtx is the context-aware variant of WriteSnapshotCSR. ctx.Err() is checked at three stage boundaries: before the CSR write, before the manifest write, and before the atomic rename. On cancellation the temporary staging directory is cleaned up and the wrapped ctx.Err is returned.

func WriteSnapshotFull

func WriteSnapshotFull[N comparable, W any](dir string, c *csr.CSR[W], g *lpg.Graph[N, W]) error

WriteSnapshotFull is the v2/v3 high-level helper: it lays out a snapshot directory containing csr.bin (legacy v1 component), labels.bin (v2 component), properties.bin (v2 component) and a manifest indexing them. When the underlying graph.Mapper is string-keyed (N=string) the writer additionally emits mapper.bin — the durable (NodeID -> natural key) interning table — and the manifest is stamped at ManifestVersion (v4). For any other N the writer falls back to the v2 layout (no mapper.bin) and the manifest records [manifestVersionV2]; recovery from a v2 snapshot continues to rely on WAL replay to re-intern keys.

Atomic publication is achieved by assembling the snapshot under dir + ".tmp" and renaming it to dir on success — the same protocol used by WriteSnapshotCSR.

When g carries a non-nil index.Manager (set via lpg.Graph.SetIndexManager) with at least one registered index that implements index.Serializer, an indexes/ sub-directory is also produced — one file per registered serializable index, each referenced from the manifest's Indexes field. Subscribers that do not implement index.Serializer are skipped (rebuild-on-restart).

Callers that do not need durable LPG labels or properties can keep using WriteSnapshotCSR; it writes a v1-shaped directory that future readers (including this one) accept transparently.

func WriteSnapshotFullCtx

func WriteSnapshotFullCtx[N comparable, W any](
	ctx context.Context,
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
) error

WriteSnapshotFullCtx is the context-aware variant of WriteSnapshotFull. ctx.Err() is checked at five stage boundaries: before the CSR write, before the labels write, before the properties write, before the manifest write, and before the atomic rename. On cancellation the temporary staging directory is cleaned up and the wrapped ctx.Err is returned.

func WriteSnapshotFullWithConstraints added in v0.2.0

func WriteSnapshotFullWithConstraints[N comparable, W any](
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	constraints []ConstraintSpec,
) error

WriteSnapshotFullWithConstraints is WriteSnapshotFull plus a durable constraints.bin component carrying the engine's schema constraint set. It is the snapshot entry point a checkpointer must use when the engine has constraints declared: without the component a checkpoint that truncated the WAL prefix which first declared a constraint would lose that constraint (a durability defect). The mapper.bin component is emitted for string-keyed graphs exactly as WriteSnapshotFull does.

constraints may be nil or empty, in which case no constraints.bin is written and the output is byte-identical to WriteSnapshotFull.

func WriteSnapshotFullWithConstraintsAndIndexDefs added in v0.6.0

func WriteSnapshotFullWithConstraintsAndIndexDefs[N comparable, W any](
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	constraints []ConstraintSpec,
	indexDefs []IndexDefSpec,
) error

WriteSnapshotFullWithConstraintsAndIndexDefs is WriteSnapshotFullWithConstraints plus a durable indexdefs.bin component carrying the engine's secondary-index definition set. It is the snapshot entry point a checkpointer must use when the engine has indexes declared: without the component a checkpoint that truncated the WAL prefix which first declared an index would lose that index definition (a durability defect — #1755), the index analogue of the constraints case.

Both constraints and indexDefs may be nil or empty; when both are empty the output is byte-identical to WriteSnapshotFull.

func WriteSnapshotFullWithMapperCodec

func WriteSnapshotFullWithMapperCodec[N comparable, W any](
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	codec keyEncoder[N],
) error

WriteSnapshotFullWithMapperCodec is the codec-aware variant of WriteSnapshotFull: it threads codec (the same [txn.Codec] the store uses to serialise node identifiers onto the WAL) into the mapper.bin writer so the durable NodeID->key interning table is emitted for ANY comparable key type N, not just string. A snapshot written this way is self-sufficient on load for every key type, which lets the checkpointer truncate the WAL instead of retaining it unboundedly (audit gap F3).

For string-keyed graphs the mapper bytes remain byte-identical to the version-1 layout (see WriteMapper), so this entry point is a safe drop-in for the existing WriteSnapshotFull on string stores too.

codec must not be nil; pass the store's [txn.Store.Codec].

func WriteSnapshotFullWithMapperCodecAndConstraints added in v0.2.0

func WriteSnapshotFullWithMapperCodecAndConstraints[N comparable, W any](
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	codec keyEncoder[N],
	constraints []ConstraintSpec,
) error

WriteSnapshotFullWithMapperCodecAndConstraints is the codec-aware variant of WriteSnapshotFullWithConstraints: it threads codec into the mapper.bin writer (so the snapshot is self-sufficient for any key type) AND persists the constraint set. This is the entry point a checkpointer over a non-string store uses when constraints are declared.

codec must not be nil. constraints may be nil or empty (no constraints.bin).

func WriteSnapshotFullWithMapperCodecAndConstraintsFS added in v0.6.0

func WriteSnapshotFullWithMapperCodecAndConstraintsFS[N comparable, W any](
	fsys fileSystem,
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	codec keyEncoder[N],
	constraints []ConstraintSpec,
) error

WriteSnapshotFullWithMapperCodecAndConstraintsFS is the filesystem-seam variant of WriteSnapshotFullWithMapperCodecAndConstraints: it routes every filesystem operation of the snapshot publish through fsys instead of the default OS backend. It is the entry point the deterministic-simulation harness (internal/sim) uses to back a snapshot with an in-memory disk so it can crash mid-publish and during the WAL prefix-truncate that follows a checkpoint.

The fsys parameter type ([fileSystem]) is intentionally unexported: an external package cannot name it but can still supply a value that satisfies it (mirroring wal.OpenWith). Passing osBackend{} reproduces WriteSnapshotFullWithMapperCodecAndConstraints byte-for-byte.

codec must not be nil. constraints may be nil or empty (no constraints.bin).

func WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefs added in v0.6.0

func WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefs[N comparable, W any](
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	codec keyEncoder[N],
	constraints []ConstraintSpec,
	indexDefs []IndexDefSpec,
) error

WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefs is the codec-aware variant of WriteSnapshotFullWithConstraintsAndIndexDefs: it threads codec into the mapper.bin writer (so the snapshot is self-sufficient for any key type) AND persists BOTH the constraint set and the index-definition set. This is the entry point a checkpointer over a non-string store uses when either constraints or indexes are declared.

codec must not be nil. constraints and indexDefs may each be nil or empty (the corresponding component is then omitted).

func WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefsFS added in v0.6.0

func WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefsFS[N comparable, W any](
	fsys fileSystem,
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	codec keyEncoder[N],
	constraints []ConstraintSpec,
	indexDefs []IndexDefSpec,
) error

WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefsFS is the filesystem-seam variant of WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefs: it routes every filesystem operation of the snapshot publish through fsys instead of the default OS backend, persisting BOTH the constraint set and the index-definition set. It is the entry point the deterministic-simulation harness (internal/sim) uses to back a snapshot with an in-memory disk so the simulated checkpoint also carries durable index definitions (#1755).

The fsys parameter type ([fileSystem]) is intentionally unexported (mirroring wal.OpenWith). Passing osBackend{} reproduces WriteSnapshotFullWithMapperCodecConstraintsAndIndexDefs byte-for-byte.

codec must not be nil. constraints and indexDefs may each be nil or empty.

func WriteSnapshotFullWithMapperCodecCtx

func WriteSnapshotFullWithMapperCodecCtx[N comparable, W any](
	ctx context.Context,
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	codec keyEncoder[N],
) error

WriteSnapshotFullWithMapperCodecCtx is the context-aware variant of WriteSnapshotFullWithMapperCodec. ctx cancellation is honoured at the same stage boundaries as WriteSnapshotFullCtx; the only difference is that the mapper.bin component is emitted for every key type via codec rather than for string alone.

func WriteSnapshotFullWithWeightCodecCtx added in v0.12.0

func WriteSnapshotFullWithWeightCodecCtx[N comparable, W any](
	ctx context.Context,
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	wcodec weightEncoder[W],
) error

WriteSnapshotFullWithWeightCodecCtx is the weight-codec-aware variant of WriteSnapshotFullCtx: it persists edge weights of ANY type W by encoding them through wcodec.

Without a codec this writer can persist only weights whose Go type has a fixed width (the built-in integer, float and bool primitives), and a graph holding any other weight type fails with ErrWeightNotPersistable rather than publishing a weightless snapshot (rmp #2526). A nil wcodec is accepted and behaves exactly as WriteSnapshotFullCtx.

The mapper is persisted for string-keyed graphs only, as on WriteSnapshotFullCtx; pair with a mapper codec via the WriteSnapshotFullWithMapperCodec* family when the key type needs one.

func WriteSnapshotFullWithWeightCodecCtxFS added in v0.16.0

func WriteSnapshotFullWithWeightCodecCtxFS[N comparable, W any](
	ctx context.Context,
	fsys fileSystem,
	dir string,
	c *csr.CSR[W],
	g *lpg.Graph[N, W],
	wcodec weightEncoder[W],
) error

WriteSnapshotFullWithWeightCodecCtxFS is the filesystem-seam variant of WriteSnapshotFullWithWeightCodecCtx: it routes every filesystem operation of the snapshot publish through fsys instead of the default OS backend. It is the entry point store/bulkimport's seamed publish uses, so the deterministic- simulation harness (internal/sim) can inject faults into a bulk-import publish and crash it mid-way (rmp #2518).

The fsys parameter type ([fileSystem]) is intentionally unexported, as on WriteSnapshotFullWithMapperCodecAndConstraintsFS. Passing osBackend{} reproduces WriteSnapshotFullWithWeightCodecCtx byte for byte. A nil wcodec behaves exactly as WriteSnapshotFullCtx over fsys.

func WriteTombstones

func WriteTombstones[N comparable, W any](w io.Writer, g *lpg.Graph[N, W], at *lpg.Snapshot) (size int64, crc uint32, err error)

WriteTombstones serialises g's current tombstone set (the NodeIDs removed via lpg.Graph.RemoveNode) into w in the tombstones.bin format. It returns the number of bytes written and the CRC32C of the serialised payload — both stored in the manifest's FileEntry for the component so LoadSnapshotFull can verify integrity at load time.

The CRC32C covers the entire on-disk file, including the magic header, matching the labels.bin / properties.bin discipline. The id list is emitted in ascending order (lpg.Graph.TombstonedIDs sorts it) so the component is deterministic across writes of the same logical state.

func WriteTombstonesFrom added in v0.11.0

func WriteTombstonesFrom(w io.Writer, ids []graph.NodeID) (size int64, crc uint32, err error)

WriteTombstonesFrom is WriteTombstones over an ALREADY-COMPUTED dead set.

It exists because a concurrent capture must derive every number in its manifest from ONE walk of the mapper (rmp #2310). Recomputing the set here would be a second walk at a second moment, and while writers commit those two moments do not agree: each narrowing of that window reduced the disagreement — 8 nodes, then 10, then 2 — and none could close it, because the window IS the design. The capture now walks once, keeps the interned set and the dead set from that walk, and hands both to the writers that serialise them.

ids must be ascending, which is what the reader expects.

Types

type CSRReadback

type CSRReadback struct {
	Vertices    []uint64
	Edges       []graph.NodeID
	WeightBytes []byte
	// Handles is the optional per-slot stable-edge-handle column, aligned
	// slot-for-slot with Edges (handles[i] is the stable handle of the edge
	// at edges[i]). It is nil when the snapshot predates the column or its
	// source graph carried no handles; when non-nil it has the same length
	// as Edges. See the trailing-block note on [WriteCSR].
	Handles    []uint64
	HasWeights bool
	// WeightSize is the per-weight width in bytes for the dense native layout,
	// or [weightSizeCodec] (0xFF) when the weights are codec-encoded and
	// variable width, in which case WeightOffsets indexes WeightBytes.
	WeightSize uint8
	// WeightOffsets is the (len(Edges)+1) offsets array of a codec-encoded
	// weights section: WeightBytes[WeightOffsets[k]:WeightOffsets[k+1]] is the
	// encoded weight of the edge at Edges[k]. It is nil for the dense native
	// layout and for a weightless snapshot. Validated at parse time to be
	// zero-based, monotonic and within the payload, so those slices are
	// unconditionally in range. See csr_weight_codec.go.
	WeightOffsets []uint64
}

CSRReadback parses a CSR previously serialised by WriteCSR. It returns the parsed vertices, edges, and (optional) raw weight bytes plus the on-disk CRC32C of the payload.

func ReadCSR

func ReadCSR(r io.Reader) (CSRReadback, error)

ReadCSR parses a CSR previously written by WriteCSR from r. The caller is responsible for verifying the surrounding manifest CRC; this function only enforces the structural contract.

Untrusted input: a bare ReadCSR over an io.Reader of unknown length cannot know the true remaining-bytes bound, so the declared vertex, edge, and weight sizes are checked only against the absolute backstop cap [maxCSRCount] plus the overflow-safe weights computation. That cap is deliberately tight (1<<34 records ⇒ a ≈ 128 GiB array): it bounds the worst-case eager reservation on this entry point well below the multi-TiB an unbounded ceiling would permit, while still admitting any CSR this engine could legitimately produce. It is, however, only a backstop, not a precise bound — that requires the file size. Callers loading an untrusted snapshot should prefer the Open / LoadSnapshotFull path, which supplies the manifest-recorded size (FileEntry.Size) so the count is rejected the moment it exceeds what that many bytes could possibly hold; bounding a bare reader more tightly than the backstop otherwise remains the caller's responsibility.

func (*CSRReadback) CodecWeights added in v0.12.0

func (r *CSRReadback) CodecWeights() bool

CodecWeights reports whether this readback carries a codec-encoded, variable-width weights section, which can only be decoded with a weight codec (ApplyCSRToGraphWithWeightCodec).

type Capture added in v0.11.0

type Capture[W any] struct {
	// contains filtered or unexported fields
}

Capture is a point-in-time, fully-serialised image of EVERY snapshot component that derives from the live graph: csr.bin, labels.bin, properties.bin, mapper.bin, tombstones.bin, edgehandles.bin, the indexes/<name>.bin payloads, and the graph's directed/multigraph/weightless shape.

Why this type exists (ACID Atomicity, rmp #2269)

A checkpoint is an OBSERVER of the graph, and the file it publishes is what a crash recovery replays. The snapshot must therefore be an image of ONE transaction-boundary state — never a mixture of two.

Before this type, the checkpointer captured only the CSR adjacency under the commit lock and then handed the LIVE graph to the writer, which walked it for the mapper, labels, properties, tombstones, edge handles and index payloads during the deliberately lock-free snapshot write. Those walks therefore observed a LATER state than the CSR. For a workload of `CREATE (a)-[:R]->(b)` transactions the published snapshot carried the node pairs of transactions that committed DURING the write while carrying none of their edges, so a snapshot-only recovery reconstructed a graph with Order > 2*Size — an artefact no serial schedule could produce, and a partial transaction made durable. The same skew could publish a node's labels or properties from after a mutation whose edge changes were not captured, and (worse) could capture the eagerly-applied writes of a transaction whose WAL commit had not yet succeeded and might still be rolled back.

The invariant this type enforces is therefore: every byte of a snapshot comes from the same instant. The publish step (WriteCapture) touches no graph at all, so publishing stays lock-free because what phase 2 writes is bytes, not a graph.

How the one instant is obtained (rmp #2310)

Originally by EXCLUSION: the caller took the Capture inside its own window, which for the checkpointer meant the store's commit serialisation plus the old lpg.Graph.View, held for the whole serialisation. That made the stall proportional to the graph, and it was the last place in the module where a reader excluded writers.

Now by VERSIONING: the caller opens an MVCC snapshot and passes it as `at`, every component resolves through it, and writers commit throughout. Exclusion is still available for callers that want the present (`at == nil`) and is what the offline one-shot writer uses. See CaptureGraph for the obligations each mode carries.

Concurrency: a Capture is an immutable value once returned; it shares no state with the graph it was taken from and is safe to publish from another goroutine.

Cost: a Capture holds the whole serialised snapshot in memory until it is published. That is the price of an atomic image; see CaptureGraph.

func CaptureGraph added in v0.11.0

func CaptureGraph[N comparable, W any](
	g *lpg.Graph[N, W],
	cs *csr.CSR[W],
	codec keyEncoder[N],
	at *lpg.Snapshot,
) (*Capture[W], error)

CaptureGraph serialises every live-graph-derived snapshot component of g into memory, producing an atomic image that WriteCapture can publish without touching g again.

cs is the CSR adjacency the caller has already built for this same instant (via csr.BuildFromAdjList); it is adopted, not rebuilt, so the caller pays the adjacency cost once.

codec, when non-nil, serialises the NodeID->key interning table for ANY key type N, making the snapshot self-sufficient so a checkpointer may truncate the WAL. When nil the mapper is emitted only for string-keyed graphs (the historical v3 behaviour) and non-string snapshots stay v2.

at is the MVCC instant every component is resolved at. When non-nil, writers may commit freely throughout the capture and the image still describes exactly that instant (rmp #2310). When nil, the capture reads the PRESENT and the caller must hold its own exclusion for the read to be atomic — which is the offline and single-goroutine shape, and the one WriteSnapshotFull uses.

Caller's obligation

CaptureGraph performs NO locking of its own, and which obligation the caller carries depends on whether it supplies an instant.

With at == nil the caller MUST hold whatever exclusion makes the read of g atomic with respect to writers, and must have built cs inside that same window. Capturing the present without such exclusion reintroduces exactly the cross-component skew this type exists to prevent (rmp #2269).

With at != nil the caller does NOT need to exclude writers for the capture — that is the whole point — and it must have built cs at the same instant. Transactions open at the instant are allowed: the image carries only ids ever born as of at, and every other assigned id is a hole (WAL v2 step 1). A snapshot opened with lpg.Graph.BeginCaptureRead also supplies the per-shard high-water marks for nodeids.bin; for any other snapshot they are read after the walk.

The component writers this calls take only their own per-shard read locks and never re-enter the visibility barrier, so calling CaptureGraph from inside a barrier hold is deadlock-free and does not trip the barrier's re-entrancy guard. (This used to name lpg.Graph.View as the enclosing hold; rmp #2344 removed it, and the checkpointer now pins a snapshot with lpg.Graph.BeginRead instead.)

Cost

The returned Capture holds the fully serialised snapshot in memory. This is deliberate: it is what lets the publish step run lock-free while remaining atomic. The peak cost is the on-disk snapshot size, held for the duration of the publish, on top of the CSR the caller already built. A Capture taken through this entry point carries NO weight codec, so publishing it fails with ErrWeightNotPersistable when the graph holds weights of a type the fixed-width layout cannot size. Use CaptureGraphWithWeightCodec for those. See csr_weight_codec.go.

func CaptureGraphWithWeightCodec added in v0.12.0

func CaptureGraphWithWeightCodec[N comparable, W any](
	g *lpg.Graph[N, W],
	cs *csr.CSR[W],
	codec keyEncoder[N],
	wcodec weightEncoder[W],
	at *lpg.Snapshot,
) (*Capture[W], error)

CaptureGraphWithWeightCodec is the weight-codec-aware variant of CaptureGraph. wcodec is adopted by the returned Capture and used by WriteCapture to serialise the CSR's weights column for any weight type W — pass the owning store's codec, [txn.Store.WeightCodec].

The codec is consulted only for weight types the fixed-width layout cannot size, so a float64, int64 or int32 graph publishes byte-identical bytes with or without it. A nil wcodec is accepted and behaves exactly as CaptureGraph.

The codec is captured here rather than passed at publish time so that a Capture stays what it claims to be: a self-contained image that WriteCapture can publish without any further input about how to encode it.

func (*Capture[W]) CommitTS added in v0.11.0

func (c *Capture[W]) CommitTS() uint64

CommitTS reports the MVCC instant this image was captured at, or 0 when the originating graph had no MVCC clock.

It is what Manifest.CommitTS records and what recovery folds into the derived clock floor. Exported so a caller that publishes a capture through its own path can carry the instant, and so a test can assert what an image claims.

func (*Capture[W]) IndexesCommitTS added in v0.12.0

func (c *Capture[W]) IndexesCommitTS() (uint64, bool)

IndexesCommitTS reports the MVCC instant this image's secondary-index payloads are declared to describe, and whether the capture could name one at all.

ok is true ONLY for a capture taken at an explicit instant — the quiesced checkpointer path, whose instant is paired with the WAL watermark, so a recovery can tell from the WAL suffix alone what changed after the payloads were serialised. It is false for a present-time capture (a nil `at`), which has no such pairing; the writer then omits Manifest.IndexesCommitTS and no reader may hydrate from the payloads.

It is at-or-BEFORE the payload bytes, which is the safe direction

[captureIndexes] serialises the live index.Manager rather than reading it as of `at`, so a payload can carry entries for a transaction that committed during this serialisation. Those transactions are all ABOVE the watermark, so they are in the WAL suffix a recovery replays, and the recovery's own per-index staleness filter refuses the payload whenever the suffix touched that index's (label, property). The reported instant is therefore a floor on what the payloads contain, never a ceiling — which is the direction that makes the filter conservative rather than wrong.

func (*Capture[W]) Order added in v0.11.0

func (c *Capture[W]) Order() uint64

Order reports how many NODES this image carries.

It is the mapper's count, not the CSR's vertex-array length (rmp #2310)

The CSR's array is sized from the present id space, because node ids are packed as (intra << shardBits) | shard and an id-indexed array must span that space whatever instant is being read. A concurrent capture therefore has vertex SLOTS for ids interned after its instant — empty slots naming no node, because the mapper the image carries is filtered at the instant and does not hold them.

Reporting the array length here made the manifest disagree with the image it describes: measured as manifest Order=2178 against a reconstructed 2176, two slots belonging to one transaction that was still in flight. What a consumer means by "Order" is how many nodes it will get back, so that is what this returns.

func (*Capture[W]) SetWALPosition added in v0.16.0

func (c *Capture[W]) SetWALPosition(storeID uint64, redoPos int64)

SetWALPosition records the store id of the write-ahead log this image pairs with and the log position it covers — every frame below redoPos is folded into the image. A capture that publishes a mapper (manifest version 4) then writes them as the manifest's store_id, wal_redo_pos and wal_format = 2, which recovery checks against the log (rmp #3014). A storeID of 0 records nothing. Call it before the capture is published; a Capture is not safe for concurrent mutation.

func (*Capture[W]) Size added in v0.11.0

func (c *Capture[W]) Size() uint64

Size reports the edge count of the captured CSR adjacency.

type ConstraintSpec added in v0.2.0

type ConstraintSpec struct {
	// Label is the constrained node label.
	Label string
	// Property is the constrained property key.
	Property string
	// Name is the user-defined constraint name.
	Name string
	// Kind is the constraint-kind tag: 0 = UNIQUE, 1 = NOT NULL.
	Kind uint8
}

ConstraintSpec is one durable constraint definition: its kind tag (0 = UNIQUE, 1 = NOT NULL, matching store/txn.ConstraintKind), the constrained node label, the constrained property key, and the user-defined name. The snapshot layer keeps its own constraint type so it does not import the txn or recovery packages.

type ConstraintsReadback added in v0.2.0

type ConstraintsReadback struct {
	Specs []ConstraintSpec
}

ConstraintsReadback is the structural parse of a constraints.bin file: the recovered constraint definitions in the order they were written (the writer sorts them deterministically). The caller maps them back into its engine constraint registry.

func ReadConstraints added in v0.2.0

func ReadConstraints(r io.Reader) (ConstraintsReadback, error)

ReadConstraints parses a constraints.bin payload produced by WriteConstraints. It performs strict structural validation: a missing or wrong magic, a future format-version word, an implausible record count, an over-long string field, or a truncated record all surface as ErrConstraintsCorrupted.

The caller is responsible for verifying the surrounding manifest CRC matches the file bytes (LoadSnapshotFull does this); this function only enforces the structural contract.

type EdgeHandleRecord

type EdgeHandleRecord struct {
	Properties map[string]lpg.PropertyValue
	Labels     []string
	Src        uint64
	Dst        uint64
	Handle     uint64
}

EdgeHandleRecord is one persisted per-handle edge metadata record: the endpoint NodeIDs, the stable handle, the per-CREATE label names, and the per-CREATE properties. NodeIDs are stored verbatim (the snapshot mapper is restored before this component is applied), matching csr.bin's NodeID references.

type EdgeHandlesReadback

type EdgeHandlesReadback struct {
	Records []EdgeHandleRecord
}

EdgeHandlesReadback is the structural parse of an edgehandles.bin file. The caller materialises it back into a live lpg.Graph via ApplyEdgeHandlesToGraph once the underlying mapper is populated.

func ReadEdgeHandles

func ReadEdgeHandles(r io.Reader) (EdgeHandlesReadback, error)

ReadEdgeHandles parses an edgehandles.bin payload produced by WriteEdgeHandles. It performs strict structural validation: a missing or wrong magic, an unsupported version, an implausible count, a label/key index past its table, an unknown property kind, or a truncated record all surface as ErrEdgeHandlesCorrupted.

edgehandles read: header + two string tables + per-record labels + per-record props, each bounds-checked

type EdgeLabelEntry

type EdgeLabelEntry struct {
	Src       uint64
	Dst       uint64
	Slot      uint32
	StringIdx uint32
}

EdgeLabelEntry names ONE relationship type of the directed pair (Src, Dst): Slot says WHERE on that pair the type lives, and StringIdx indexes the file's string table.

Slot is what makes a multigraph pair representable. A relationship type belongs to the relationship INSTANCE, so two parallel edges between the same endpoints may carry different types — or one may carry a type and the other none. Version 1 had no such field and keyed a record by (Src, Dst) alone, so those slots folded together on disk and a checkpoint could silently lose a committed type or invent one that was never attached (rmp #2262).

Slot is either:

  • a canonical slot ordinal — the slot's position among the pair's slots after a STABLE sort by stable-edge handle ascending, the order BOTH recovery paths converge on. The record carries that slot's INLINE adjacency label column entry. lpg.Graph.ForEachPairSlotRelTypeByID produces it on the write side and lpg.Graph.SetEdgeRelTypeAtSlotByID resolves it on the apply side; or
  • EdgeLabelSlotOverflow, meaning the type is in the pair's overflow list — a type that could not be placed in any slot's column and that every column-typed slot of the pair therefore carries.

Slot is meaningless in a version-1 readback and is left zero there.

A slot's inline entry is recorded whether or not a by-handle type record also covers it. The by-handle store has its own durable component (edgehandles.bin) and stays authoritative for what such a slot IS, but the column is what lpg.Graph.EdgeLabels and lpg.Graph.RelationshipTypesInUse read, so omitting it here would leave every Cypher-created relationship's type absent from those answers after a restart.

type EdgePropertyEntry

type EdgePropertyEntry struct {
	ValueBytes []byte
	Src        uint64
	Dst        uint64
	KeyIdx     uint32
	Kind       lpg.PropertyKind
}

EdgePropertyEntry pairs an (src, dst) NodeID couple with the key string-table index, the kind tag, and the encoded value bytes for one property attached to that edge. An edge carrying P properties yields P entries; as with labels, parallel edges between the same endpoints fold into the same edgeKey on disk just as they do in lpg.Graph's in-memory shards.

type File added in v0.6.0

type File interface {
	io.Writer
	Sync() error
	Close() error
	Stat() (fs.FileInfo, error)
}

File is the minimal handle the snapshot writers need of an open component file. *os.File satisfies it directly; the in-memory DST backend supplies an equivalent handle. Stat is required by the per-index writer, which records the file's on-disk size in the manifest.

It is exported so an external filesystem backend (the deterministic- simulation harness) can name it as the return type of its Create method and thereby satisfy the unexported [fileSystem] interface; production code never references it directly.

Concurrency: a File is used by a single writer goroutine while one snapshot component is written and is not safe for concurrent use; any implementation's concurrency guarantees are its own.

type FileEntry

type FileEntry struct {
	Name   string `json:"name"`
	Size   int64  `json:"size"`
	CRC32C uint32 `json:"crc32c"`
}

FileEntry records one component file inside a snapshot directory.

type GraphConfig added in v0.2.0

type GraphConfig struct {
	// Directed records whether AddEdge was a directed insertion in the
	// originating graph.
	Directed bool `json:"directed"`
	// Multigraph records whether the originating graph allowed parallel
	// edges between the same ordered endpoint pair.
	Multigraph bool `json:"multigraph"`
	// Weightless records whether the originating graph stored no per-edge
	// weight column (adjlist.Config.Weightless, #1650). It is omitempty and
	// backward-compatible: a snapshot written before this field, or by a
	// weighted graph, omits it, so it decodes to false (weighted) — the prior
	// behaviour. A recovered weightless graph stays weightless, preserving the
	// per-edge memory saving across a restart rather than re-allocating a
	// zero-filled weight column.
	Weightless bool `json:"weightless,omitempty"`
}

GraphConfig is the JSON-persisted shape of the originating graph's adjacency-list configuration. It mirrors the directed/multigraph flags of [adjlist.Config] without importing that package, so the snapshot manifest stays decoupled from the graph backend. The snapshot writer fills it from the live graph; recovery reads it to reconstruct the same variant.

Only the shape-defining flags are persisted. [adjlist.Config.MaxShardCapacity] is deliberately omitted: it is a runtime growth bound, not a property of the stored graph, and re-imposing it at recovery time could make recovery itself fail with [adjlist.ErrShardFull] while replaying data that legitimately exceeds the cap. A recovered graph is therefore always reconstructed unbounded.

type IndexDefSpec added in v0.6.0

type IndexDefSpec struct {
	// Name is the user-defined index name.
	Name string
	// Label is the indexed node label.
	Label string
	// Property is the indexed property key.
	Property string
	// Kind is the index-kind tag: 0 = hash, 1 = btree.
	Kind uint8
}

IndexDefSpec is one durable index definition: its kind tag (0 = hash, 1 = btree, matching store/txn.IndexKind), the user-defined name, the indexed node label, and the indexed property key. The snapshot layer keeps its own index type so it does not import the txn or recovery packages.

type IndexDefsReadback added in v0.6.0

type IndexDefsReadback struct {
	Specs []IndexDefSpec
}

IndexDefsReadback is the structural parse of an indexdefs.bin file: the recovered index definitions in the order they were written (the writer sorts them deterministically). The caller maps them back into its engine index registry.

func ReadIndexDefs added in v0.6.0

func ReadIndexDefs(r io.Reader) (IndexDefsReadback, error)

ReadIndexDefs parses an indexdefs.bin payload produced by WriteIndexDefs. It performs strict structural validation: a missing or wrong magic, a future format-version word, an implausible record count, an over-long string field, or a truncated record all surface as ErrIndexDefsCorrupted.

The caller is responsible for verifying the surrounding manifest CRC matches the file bytes (LoadSnapshotFull does this); this function only enforces the structural contract.

type IndexFileEntry

type IndexFileEntry struct {
	Name   string `json:"name"`
	Size   int64  `json:"size"`
	CRC32C uint32 `json:"crc32c"`
}

IndexFileEntry pairs an index file's logical name (the name it was registered under with index.Manager.CreateIndex) with its on-disk size and CRC32C. It is the secondary-index analogue of FileEntry and travels in Manifest.Indexes.

func WriteIndexes

func WriteIndexes(dir string, m *index.Manager) ([]IndexFileEntry, error)

WriteIndexes serialises every registered index in m to one file per index under dir/IndexesDir. Returns one IndexFileEntry per successfully serialised index, which the caller threads into the manifest. Subscribers that do not implement index.Serializer are silently skipped (rebuild-on-restart contract).

It has no production caller in this module

The snapshot writer publishes index payloads through [writeCapturedIndexes] instead, from bytes serialised earlier at capture time, so the image is transaction-boundary consistent. This function reads the LIVE manager and so cannot offer that guarantee; it is retained as exported API for an embedder that publishes a snapshot through its own path, and is exercised by this package's tests. Prefer the capture-based writers.

On any I/O error the partial directory under dir/IndexesDir is removed (best effort) so the caller does not need to clean up.

After all index files are written, WriteIndexes calls [dirFsync] on idxDir. On POSIX, fsyncing an individual file only makes that file's data and inode durable; the directory entry (dirent) that links the file name to the inode is not guaranteed to be durable until the parent directory is fsynced. Without this call, a crash between the individual file fsyncs and the kernel's directory-writeback window could leave the indexes/ directory entry for one or more index files absent on the next mount, making those indexes appear missing to recovery even though their data is intact on disk.

type IndexReadback

type IndexReadback struct {
	Name  string
	Bytes []byte
}

IndexReadback is the raw byte payload of one secondary index file returned by LoadSnapshotFull, or a nil Bytes when the file was missing or its CRC32C did not match the manifest. The snapshot loader does not interpret the bytes.

It is NOT deserialised by store/recovery. Recovery classifies each readback into a per-payload reason code and reports it (store/recovery.IndexPayload); the Cypher engine, which owns the index bindings and knows which concrete implementation each name belongs to, is what calls index.Serializer.Deserialize — and only for a payload recovery certified usable. See the "Recovery semantics" section of docs/persistence.md.

func LoadIndexes

func LoadIndexes(dir string, entries []IndexFileEntry) ([]IndexReadback, error)

LoadIndexes reads every entry in entries from dir/IndexesDir and returns the raw bytes for each. Files whose on-disk CRC32C does not match the manifest record surface a metric warning via `store.snapshot.indexes.corrupted` and are reported with IndexReadback.Bytes == nil; the caller treats nil bytes as "rebuild from LPG" rather than as a fatal error.

A missing indexes/ directory is not an error: it simply means the snapshot does not carry persisted indexes (forward compat with snapshots produced before this format extension).

type LabelsReadback

type LabelsReadback struct {
	Strings    []string
	NodeLabels []NodeLabelEntry
	EdgeLabels []EdgeLabelEntry
	// Version is the on-disk labels.bin format version this readback was parsed
	// from. It selects how [ApplyLabelsToGraph] replays EdgeLabels: at or above
	// [labelsFormatVersionPerSlot] each record names one SLOT, below it each
	// record names the PAIR. A hand-constructed readback leaves it zero and is
	// therefore replayed with the conservative per-pair semantics.
	Version uint32
}

LabelsReadback is the structural parse of a labels.bin file. The caller materialises it back into a live lpg.Graph via ApplyLabelsToGraph once the underlying mapper is populated.

func ReadLabels

func ReadLabels(r io.Reader) (LabelsReadback, error)

ReadLabels parses a labels.bin payload produced by WriteLabels. It performs strict structural validation: a missing or wrong magic, a future format-version byte, a truncated record, or an out-of-range string-table index all surface as ErrLabelsCorrupted.

The caller is responsible for verifying the surrounding manifest CRC matches the file bytes (the Open / LoadSnapshotFull helpers do this); this function only enforces the structural contract.

labels read: header + string table + node records + edge records, each bounds-checked

type LoadedCSR

type LoadedCSR struct {
	CSR      CSRReadback
	Manifest Manifest
}

LoadedCSR is the result of [LoadCSR] / Open: the parsed CSR arrays plus the manifest entry that produced them.

func Open

func Open(dir string) (LoadedCSR, error)

Open verifies and loads the snapshot rooted at dir. It reads the manifest, then reads csr.bin and verifies its CRC32C matches the manifest entry. Future versions may load additional components (labels.bin, properties.bin, schema.bin) by extending Manifest.Files.

type LoadedSnapshot

type LoadedSnapshot struct {
	Labels     LabelsReadback
	Properties PropertiesReadback
	Mapper     MapperReadback
	Indexes    []IndexReadback
	// Tombstones is the node-removal set restored from tombstones.bin. It
	// is empty for snapshots that carry no tombstones.bin entry (older
	// snapshots, or any snapshot of a graph that never removed a node) —
	// the backward-compatibility contract.
	Tombstones TombstonesReadback
	// EdgeHandles is the per-handle edge metadata restored from
	// edgehandles.bin (each parallel edge's per-CREATE relationship type and
	// properties keyed by its stable handle). It is empty for snapshots that
	// carry no edgehandles.bin entry (older snapshots, or any snapshot of a
	// graph that never used the handle-keyed metadata stores) — the
	// backward-compatibility contract.
	EdgeHandles EdgeHandlesReadback
	// Constraints is the durable schema constraint set restored from
	// constraints.bin. It is empty for snapshots that carry no constraints.bin
	// entry (older snapshots, or any snapshot taken with no constraints
	// declared) — the backward-compatibility contract.
	Constraints ConstraintsReadback
	// IndexDefs is the durable secondary-index definition set restored from
	// indexdefs.bin (each index's label/property/kind/name). It is empty for
	// snapshots that carry no indexdefs.bin entry (older snapshots, or any
	// snapshot taken with no indexes declared) — the backward-compatibility
	// contract. It is DISTINCT from [LoadedSnapshot.Indexes], which holds the
	// optional per-index byte payloads; the definition set is what recovery
	// rebuilds each index from (#1755).
	IndexDefs IndexDefsReadback
	CSR       CSRReadback
	Manifest  Manifest
}

LoadedSnapshot is the result of LoadSnapshotFull: the parsed CSR arrays, the parsed labels readback (empty for v1 snapshots), the parsed properties readback (empty when properties.bin is absent), the parsed mapper readback (empty when mapper.bin is absent, e.g. a v1 CSR-only snapshot or a v2 snapshot written without a codec for a non-string key type), the optional per-index byte payloads (one entry per indexes/<name>.bin file referenced by the manifest), and the manifest that produced them.

When mapper.bin is present, exactly one of MapperReadback.Pairs (version-1 string layout) and MapperReadback.RawPairs (version-2 codec layout) is populated; see MapperReadback.

Each IndexReadback.Bytes may be nil even when the manifest references the index — that signals the file was missing or its CRC32C did not validate. Callers must treat nil bytes as "rebuild from LPG" rather than as a fatal error; the corruption was already metered by LoadIndexes under `store.snapshot.indexes.corrupted`.

func LoadSnapshotFull

func LoadSnapshotFull(dir string) (LoadedSnapshot, error)

LoadSnapshotFull verifies and loads the snapshot rooted at dir, returning the CSR, the labels readback, and the properties readback. v1 snapshots are accepted transparently: their manifest has no labels.bin or properties.bin entry, and the returned LoadedSnapshot.Labels / LoadedSnapshot.Properties are zero values (empty tables, no records). v2 snapshots may carry any combination of labels.bin and properties.bin; each component is CRC-validated only when its manifest entry is present.

CSR CRC verification mirrors Open; labels and properties CRC verification use the same TeeReader pattern so a corrupted component surfaces as ErrCorrupted.

func LoadSnapshotFullFS added in v0.6.0

func LoadSnapshotFullFS(fsys fileSystem, dir string) (LoadedSnapshot, error)

LoadSnapshotFullFS is the filesystem-seam variant of LoadSnapshotFull: it routes every read of the snapshot through fsys instead of the default OS backend. It is the entry point the deterministic-simulation harness (internal/sim) uses to load a snapshot backed by an in-memory disk. Passing osBackend{} reproduces LoadSnapshotFull exactly.

type Manifest

type Manifest struct {
	CreatedAt   time.Time        `json:"created_at"`
	GraphConfig *GraphConfig     `json:"graph_config,omitempty"`
	Files       []FileEntry      `json:"files"`
	Indexes     []IndexFileEntry `json:"indexes,omitempty"`
	Version     int              `json:"version"`
	Order       uint64           `json:"order"`
	Size        uint64           `json:"size"`
	// CommitTS is the MVCC instant the image was captured at, or 0 / absent when
	// the originating graph had no MVCC clock or the writer had no graph in hand
	// (the legacy CSR-only writer, which omits GraphConfig for the same reason).
	//
	// Recovery folds it into the derived clock floor, so a reopened graph never
	// re-mints an instant the image already contains (rmp #2309, MVCC C3d). It is
	// the quantity Memgraph reads back as info.start_timestamp.
	//
	// NO MANIFEST VERSION BUMP. The manifest is JSON, so an older reader ignores an
	// unknown field and a newer reader on an older manifest decodes the zero value —
	// which is exactly the "absent means no timestamp" policy the OpCommit body
	// uses. `omitempty` keeps a timestamp-less manifest byte-identical to what
	// previous builds wrote, so no fixture or golden file moves.
	//
	// Corruption can no longer zero this field silently: it lies inside the
	// region the trailer checksums, so a flip in the `commit_ts` KEY — which
	// leaves valid JSON whose renamed key encoding/json would drop, decoding the
	// timestamp as 0 and skipping RestoreMVCCClock — now fails the manifest
	// checksum and fail-stops recovery. See the integrity section above.
	CommitTS uint64 `json:"commit_ts,omitempty"`

	// IndexesCommitTS is the MVCC instant the `indexes/<name>.bin` payloads in
	// this manifest describe, or absent when the writer could not name one.
	//
	// # Absent means NEVER HYDRATE, and that is the back-compat guarantee
	//
	// An index payload is only loadable if a reader can decide whether the WAL it
	// is about to replay on top of the image invalidates it. That decision needs
	// the instant the payload was taken at. Only the quiesced checkpointer path
	// has one: it opens an MVCC snapshot, pairs it with the WAL watermark, and
	// captures at that instant, so everything committed afterwards is in the WAL
	// suffix a recovery replays. The present-time writers
	// ([WriteSnapshotFull] and friends) call the capture with a nil instant —
	// "read the present under the caller's own exclusion" — and have no such
	// pairing, so they write NO watermark and every payload they publish is
	// unhydratable by construction.
	//
	// Every snapshot that existed before this field was added therefore keeps its
	// exact previous meaning: a recovery rebuilds its indexes from the graph, as
	// it always did. There is no migration and no format break.
	//
	// NO MANIFEST VERSION BUMP, for the reason [Manifest.CommitTS] gives one
	// field up: the manifest is JSON, an older reader ignores the unknown key, a
	// newer reader on an older manifest decodes the zero value, and `omitempty`
	// keeps a watermark-less manifest byte-identical to what previous builds
	// wrote — so no fixture and no golden file moves. It sits inside the region
	// the trailer checksums like every other field, so a flip in the
	// `indexes_commit_ts` KEY (which would silently zero it and merely lose the
	// optimisation) still fails the manifest checksum.
	IndexesCommitTS uint64 `json:"indexes_commit_ts,omitempty"`

	// IndexBuilderEpoch identifies the secondary-index BUILDER that produced the
	// `indexes/<name>.bin` payloads in this manifest, or is absent when they were
	// produced by a build that predates the field.
	//
	// # Absent means NEVER HYDRATE, and it is the only remediation there is
	//
	// [Manifest.IndexesCommitTS] one field up answers "does the WAL replayed on
	// top of this image invalidate the payloads?". This field answers a question
	// that instant cannot reach: "was the builder that produced them one whose
	// output can be trusted at all?".
	//
	// The two are independent because a payload is DURABLE. A defect in what a
	// backfill writes into an index — rmp #2778 wrote entries for values an open
	// transaction had eagerly written and rolled back, rmp #2792 did the same on
	// the UNIQUE constraint path — survives a checkpoint, and a reader that
	// hydrates it reinstates the fabricated entry verbatim. Fixing the builder
	// cannot heal such a store: the entry is not rebuilt, it is loaded. Measured
	// on a two-build experiment before this field existed, a store written by the
	// pre-#2778 build and reopened on the fixed build reported hydrated=2
	// rebuilt=0 and answered a seek for the rolled-back value with the row of a
	// node that never held it, while a seek for that node's real value returned
	// nothing.
	//
	// So the epoch is what makes the fix reach the disk. A reader hydrates only
	// when the manifest names [CurrentIndexBuilderEpoch]; anything else — absent,
	// older, or newer than this build knows — falls into the per-index rebuild
	// path that already exists for a payload with no watermark, and the index is
	// reconstructed from the recovered graph by the builder this build ships.
	// Refusing a NEWER epoch too is deliberate: this build cannot know what a
	// future builder writes, and a rebuild is always correct, so the comparison
	// is equality and not "at least".
	//
	// The cost is bounded and one-time: every upgraded store rebuilds its indexes
	// once, O(N) in the mapper length, on its first open, after which it carries
	// this epoch and hydrates again. It was measured rather than assumed, at
	// 50 000 nodes with four registered indexes, on an Apple M4 without -race:
	//
	//	whole first open   pre-epoch 73.79m ±1%  vs current 56.99m ±2%
	//	                   +16.80 ms one-time  (+29.48%, p=0.000, n=10)
	//	engine construction only, rebuild vs hydrate
	//	                   26.59m ±3% vs 10.54m ±4%  (+152%, p=0.000, n=10)
	//
	// against a same-code noise floor of about 3% on the same host. The two
	// benchmarks are BenchmarkIndexBuilderEpochFirstOpen (which drives this field
	// directly) and BenchmarkRecoveredIndexPopulation (which isolates the
	// rebuild-versus-hydrate step), both in package cypher.
	//
	// NO MANIFEST VERSION BUMP and no migration, for the reason
	// [Manifest.CommitTS] gives two fields up: the manifest is JSON, an older
	// reader ignores the unknown key, a newer reader on an older manifest decodes
	// the zero value — which is exactly the "absent means never hydrate" policy —
	// and `omitempty` keeps an epoch-less manifest byte-identical to what
	// previous builds wrote, so no fixture and no golden file moves. It sits
	// inside the region the trailer checksums like every other field, so a flip
	// in the `index_builder_epoch` KEY zeroes it into a REBUILD rather than into
	// a hydration, and fails the manifest checksum on top of that.
	IndexBuilderEpoch uint64 `json:"index_builder_epoch,omitempty"`

	// StoreID, WALRedoPos and WALFormat pair the snapshot with a segmented
	// write-ahead log (docs/design-wal-v2.md §3.3): the store identity as 16
	// hex digits, the WAL position this snapshot covers (every frame below it
	// is folded into the image), and [WALFormatSegmented]. Manifest version 4
	// defines them; the checkpointer writes them whenever the capture carries a
	// mapper and was given a position ([Capture.SetWALPosition]). Absent, the
	// snapshot records no position, and recovery accepts it beside a segmented
	// log only when that log still begins at position 0.
	StoreID    string `json:"store_id,omitempty"`
	WALRedoPos uint64 `json:"wal_redo_pos,omitempty"`
	WALFormat  int    `json:"wal_format,omitempty"`

	// Integrity names the framing scheme the writer used, or is empty for a
	// manifest written before the trailer existed. The current writer always sets
	// [IntegrityCRC32CTrailer]; `omitempty` keeps a legacy manifest — and the
	// frozen v1 fixture — byte-identical to what previous builds wrote.
	//
	// The trailer identifies itself, so this field is not how a reader FINDS it.
	// It exists for the one case the trailer cannot speak for: a manifest whose
	// trailer has been lost entirely (a zeroed tail block, a truncating copy).
	// Without it that file is indistinguishable from a legacy manifest and would
	// be accepted unverified; with it, a manifest that says it was framed and
	// arrives unframed is refused. Losing the protection therefore requires TWO
	// independent damages — the trailer AND this marker — rather than one.
	//
	// Whenever a trailer is present this field sits inside the checksummed region
	// like any other, so reading it is an assertion, never a trust decision.
	Integrity string `json:"integrity,omitempty"`

	// IntegrityVerified reports that [LoadManifest] verified a checksum trailer
	// over these bytes. It is a decode-time result, never serialised and never
	// read from the file (`json:"-"`), so a corrupt manifest cannot assert its
	// own soundness. False means the manifest predates the trailer and was
	// accepted on its JSON syntax alone.
	IntegrityVerified bool `json:"-"`
}

Manifest is the JSON-encoded index of a snapshot directory.

Integrity, and why it does not fight forward compatibility

The manifest is read FIRST, before any other component can be trusted: it carries the CRC32C of every component file, so a manifest that is itself wrong makes every downstream check meaningless. Integrity is therefore enforced at the FILE FRAMING layer and never at the JSON SCHEMA layer, and keeping those two layers apart is what lets both properties hold at once:

  • The FRAMING layer is closed. WriteManifest appends a fixed 16-byte trailer holding a CRC32C over the entire JSON document — every key name, every value, every byte of indentation. LoadManifest verifies it before it looks at a single field. Nothing in the JSON is trusted to decide whether the JSON is intact.
  • The SCHEMA layer stays open. Unknown fields are still ignored, and an absent field still decodes to its zero value — the "absent means no value" policy Manifest.CommitTS documents is unchanged.

A required-field check would have pulled against that policy, because it makes absence itself an error and so forbids the very evolution the policy grants. The trailer does not, because it never asks what the fields MEAN. It only establishes that the bytes are the bytes the writer produced. Once that holds, "absent" can only mean the writer genuinely omitted the field — never that corruption renamed its key — so the two properties stop being in tension and become complementary: the framing makes the schema's permissiveness SAFE.

The trailer costs no compatibility in either direction:

  • Older snapshot, newer reader. A manifest written before the trailer existed has no trailer region, is accepted unverified, and reports Manifest.IntegrityVerified false. No version step, no migration; the frozen v1 fixture in testdata still loads byte-for-byte unchanged.
  • Newer snapshot, older reader. json.Decoder stops at the end of the first complete value, so a build that predates the trailer never reads past the closing brace and is unaffected by what follows it. ManifestVersion is therefore deliberately NOT bumped: the trailer is a framing change, not a schema change, and bumping the schema version would make older readers reject the file with ErrManifestUnsupported for no reason.

Prior art

The layering above is RocksDB's MANIFEST, generalised. RocksDB is the one engine surveyed that gets BOTH properties, and it gets them exactly this way: a VersionEdit is a tag-length-value document whose unknown tags are skipped (kTagSafeIgnoreMask, db/version_edit.h — its own comment calls these "forward compatible (aka ignorable) records"), while the CRC32c that protects it lives in the log-record framing OUTSIDE the payload (db/log_format.h: "Header is checksum (4 bytes), length (2 bytes), type (1 byte)"), so a reader that skips a field it does not understand still validates every byte of it.

The trailer shape is Lucene's. CodecUtil.writeFooter appends a fixed 16-byte footer — magic, algorithm identifier, checksum — to every file including segments_N, and CodecUtil.footerLength() is a compile-time constant. That constancy is what lets checkFooter detect a file that was EXTENDED as well as one that was truncated, which is why manifestTrailerSize is a constant here too. Lucene's CRC additionally covers its own magic and algorithm identifier; this trailer instead adjudicates those two fields by shape, which detects the same flips (proven byte-by-byte in manifest_integrity_test.go) and reports them more precisely.

Two negative results shaped the decision as much as the positive ones. No engine surveyed puts a checksum in a NAMED FIELD inside a variable-shaped, self-describing serialization and obtains full byte integrity from it: etcd comes closest and its walpb.Record.Crc covers only the opaque Data blob, not the record type and not the protobuf tags. And RocksDB's own backup metadata file — a text manifest with genuine ignore-unknown-field machinery — reserves a "// FOOTER" section documented for "a checksum of the meta file" that is emitted only under test options; the slot was designed and never filled, so in production that manifest has no integrity over itself. Both are the shape this manifest was in before the trailer.

PostgreSQL's pg_control is the closest fixed-shape analogue (crc is the last struct field, covering offsetof(ControlFileData, crc)) and it deliberately checks the VERSION before the CRC, because the number of bytes to checksum depends on the struct layout. The order is reversed here on purpose: this checksum covers the file minus a constant, so it does not need the schema to be interpreted first, and the stronger ordering — bytes before fields — is available.

Scope of the guarantee

CRC32C detects accidental corruption — a flipped bit, a torn or partially rewritten block, a bad cable. It is NOT a message authentication code: an attacker who can rewrite manifest.json can also recompute the trailer. The defence against a hostile store directory remains the surrounding controls (O_NOFOLLOW component opens, DefaultMaxManifestBytes, the per-component allocation bounds), not this checksum.

Indexes is the secondary-index sub-manifest: it carries one IndexFileEntry per file written under indexes/<name>.bin. The field is omitted from the JSON form when empty so v2 manifests produced before this extension are byte-identical to the ones produced by current builds when no indexes are registered.

GraphConfig records the originating graph's directed/multigraph shape. It is a pointer with omitempty so it is dropped from the JSON form entirely when nil — every snapshot written before this field existed (and the CSR-only legacy writer, which has no live graph to read) is therefore byte-identical to what it would have been. A reader that finds the field absent must default the configuration to the historical recovery behaviour ([adjlist.Config]{Directed: true, Multigraph: true}); see store/recovery.Open. Only NEW snapshots produced by the full writer carry the real config.

func LoadManifest

func LoadManifest(r io.Reader) (Manifest, error)

LoadManifest parses a manifest from r.

The order of the checks is load-bearing: the bytes are bounded, then their framing is verified, and only then is any FIELD consulted. A manifest is the first thing a store open reads, so no decision may rest on a field whose bytes have not yet been established — the version check included.

Returns ErrManifestTooLarge above DefaultMaxManifestBytes, ErrManifestCorrupted when the JSON does not parse or the trailer does not verify, and ErrManifestUnsupported when the version is newer than this build.

func ReadManifestFile

func ReadManifestFile(path string) (Manifest, error)

ReadManifestFile is a convenience wrapper around an O_NOFOLLOW open plus LoadManifest. The file is opened via [openSnapshotComponent] so a manifest.json that is a symlink in an untrusted snapshot directory is rejected rather than dereferenced.

func ReadManifestFileFS added in v0.6.0

func ReadManifestFileFS(fsys fileSystem, path string) (Manifest, error)

ReadManifestFileFS is the filesystem-seam variant of ReadManifestFile: it opens the manifest through fsys instead of the default OS backend. It is the entry point the deterministic-simulation harness (internal/sim) uses to read a manifest backed by an in-memory disk. Passing osBackend{} reproduces ReadManifestFile exactly.

type MapperPair

type MapperPair struct {
	Key string
	ID  graph.NodeID
}

MapperPair is one (NodeID, natural key) record as parsed from the on-disk mapper.bin payload. The slice exposed by MapperReadback.Pairs is enumerated in shard-major / intra-index- major order — the same order graph.Mapper.Walk produces, which is the order the writer serialised.

type MapperRawPair

type MapperRawPair struct {
	Key []byte
	ID  graph.NodeID
}

MapperRawPair is one (NodeID, raw key bytes) record as parsed from a codec-encoded (version 2) mapper.bin payload. The bytes are the opaque output of [txn.Codec.Encode] for the natural key; recovery decodes them back into the concrete key type N via the matching codec. The slice exposed by MapperReadback.RawPairs is enumerated in the same Walk order the writer serialised.

type MapperReadback

type MapperReadback struct {
	Pairs    []MapperPair
	RawPairs []MapperRawPair
	// Next is the per-shard high-water marks read from the snapshot's
	// nodeids.bin, or nil for a snapshot without one (written before WAL v2
	// step 1). The apply functions pass it to [graph.Mapper.LoadFrom].
	Next *[graph.MapperShards]uint64
}

MapperReadback is the structural parse of a mapper.bin file. The caller materialises it back into a live graph.Mapper via graph.Mapper.LoadFrom once a fresh mapper has been constructed.

Exactly one of Pairs / RawPairs is populated, selected by the on-disk format version:

  • Pairs holds string keys, parsed from a version-1 (string) file. ApplyMapperToGraph consumes this directly for string-keyed graphs.
  • RawPairs holds codec-encoded key bytes, parsed from a version-2 file produced for a non-string key type. ApplyMapperToGraphWithCodec decodes these via the supplied codec.

Both are empty when mapper.bin was absent (v1/v2 snapshots written by the no-codec writer for non-string keys, or any v1 CSR-only snapshot).

func ReadMapperBytes

func ReadMapperBytes(r io.Reader) (MapperReadback, error)

ReadMapperBytes parses a mapper.bin payload produced by WriteMapper for a non-string key type (version 2): the per-record key bytes are the opaque codec output and are returned verbatim in MapperReadback.RawPairs for the caller to decode with the matching codec. A version-1 (string) payload is accepted too — its UTF-8 key bytes are returned as RawPairs so a single reader path can serve both layouts when a codec is in hand.

Structural validation matches ReadMapperString: bad magic, an unsupported version, a truncated record, or an implausible length prefix all surface as ErrMapperCorrupted. The caller verifies the surrounding manifest CRC ([readVerifiedMapperBytes] does this); this function enforces only the structural contract.

func ReadMapperString

func ReadMapperString(r io.Reader) (MapperReadback, error)

ReadMapperString parses a mapper.bin payload produced by WriteMapperString. It performs strict structural validation: a missing or wrong magic, an unsupported format version, a truncated record, or an implausible key length all surface as ErrMapperCorrupted.

The caller is responsible for verifying the surrounding manifest CRC matches the file bytes (LoadSnapshotFull does this); this function only enforces the structural contract.

type NodeLabelEntry

type NodeLabelEntry struct {
	NodeID    uint64
	StringIdx uint32
}

NodeLabelEntry pairs a NodeID with the string-table index of one label name attached to that node. A node carrying N labels yields N entries.

type NodePropertyEntry

type NodePropertyEntry struct {
	ValueBytes []byte
	NodeID     uint64
	KeyIdx     uint32
	Kind       lpg.PropertyKind
}

NodePropertyEntry pairs a NodeID with the key string-table index, the kind tag, and the encoded value bytes for one property attached to that node. A node carrying P properties yields P entries.

type PropertiesReadback

type PropertiesReadback struct {
	Keys           []string
	NodeProperties []NodePropertyEntry
	EdgeProperties []EdgePropertyEntry
}

PropertiesReadback is the structural parse of a properties.bin file. The caller materialises it back into a live lpg.Graph via ApplyPropertiesToGraph once the underlying mapper is populated.

func ReadProperties

func ReadProperties(r io.Reader) (PropertiesReadback, error)

ReadProperties parses a properties.bin payload produced by WriteProperties. It performs strict structural validation: a missing or wrong magic, a future format-version byte, a truncated record, an unknown kind tag, or a key-table index that points beyond the embedded string table all surface as ErrPropertiesCorrupted.

The caller is responsible for verifying the surrounding manifest CRC matches the file bytes (the LoadSnapshotFull helper does this); this function only enforces the structural contract.

properties read: header + key table + node records + edge records, each bounds-checked

type ReadFile added in v0.6.0

type ReadFile interface {
	io.Reader
	Close() error
	// Stat returns file info for the OPEN descriptor (an fstat, not a path
	// re-stat), so a size derived from it reflects exactly the bytes this
	// handle reads — used by the CSR reader's allocation bound to avoid any
	// stat/open TOCTOU (see safeCSRAllocBound). Both backends already provide
	// it: osBackend's *os.File natively, and the simulator's *SimFileHandle.
	Stat() (fs.FileInfo, error)
}

ReadFile is the minimal handle the snapshot readers need of an open component file. *os.File satisfies it directly; the in-memory backend supplies an equivalent read handle. It is exported for the same reason as File.

Concurrency: a ReadFile is consumed by a single reader goroutine while one snapshot component is read and is not safe for concurrent use; any implementation's concurrency guarantees are its own.

type TombstonesReadback

type TombstonesReadback struct {
	IDs []graph.NodeID
}

TombstonesReadback is the structural parse of a tombstones.bin file: the sorted set of removed NodeIDs. The caller materialises it back into a live lpg.Graph via ApplyTombstonesToGraph.

func ReadTombstones

func ReadTombstones(r io.Reader) (TombstonesReadback, error)

ReadTombstones parses a tombstones.bin payload produced by WriteTombstones. It performs strict structural validation: a missing or wrong magic, a future format-version word, an implausible count, or a truncated record all surface as ErrTombstonesCorrupted.

The caller is responsible for verifying the surrounding manifest CRC matches the file bytes (the LoadSnapshotFull helper does this); this function only enforces the structural contract.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL