archive

package
v0.5.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 24, 2026 License: MIT Imports: 24 Imported by: 0

Documentation

Overview

Package archive reads and writes Seamless instance archives: one gzipped tar carrying the markdown corpus, a consistent snapshot of seam.db, and a manifest describing what is inside.

The archive is the whole instance minus its identity. Config and the MCP API key are never included -- a restored instance gets a new key and its own config, so an archive can be copied around without carrying a credential.

Layout, in tar order:

manifest.json                     always FIRST, so a reader can refuse an
                                  archive before extracting a single byte
seam.db                           absent when exported with NoDB
memory/<project|_global>/<name>.md
notes/<project|_global>/<slug>.md

Every entry is a regular file with mode 0600 and uid/gid zeroed: an archive restored as another user must not carry the exporting machine's ownership, and the corpus is owner-only on disk for the same reason the data dir is.

The package imports core, files, store, and validate. It must NOT import internal/config: an archive is described by its manifest and its data dir argument, never by the config of whatever process happens to be holding it, which is what lets an export run against a data dir the caller names and an import target a directory that has no config yet.

Index

Constants

View Source
const (
	// ManifestName is the first entry of every archive.
	ManifestName = "manifest.json"
	// DBName is the snapshot of seam.db, absent from a NoDB export.
	DBName = "seam.db"
	// MemoryTree and NotesTree are the two markdown trees carried verbatim.
	MemoryTree = "memory"
	NotesTree  = "notes"
)

The fixed entry names inside an archive. They are constants rather than literals at each site because the writer and the reader have to agree exactly, and a transcribed name drifts in silence (AGENTS.md).

View Source
const FormatVersion = 1

FormatVersion is the archive container version. It describes the tar layout and manifest shape, NOT the database schema: schema_version travels separately in the manifest because a v1 archive is written by every seamlessd regardless of how far its migrations have run.

Variables

View Source
var (
	// ErrSchemaTooNew reports an archive whose schema_version exceeds this
	// binary's store.LatestSchemaVersion. Migrating forward cannot help: the
	// migrations that produced the snapshot are not in this build, so the only
	// remedy is a newer seamlessd.
	ErrSchemaTooNew = errors.New("archive was made by a newer seamlessd")

	// ErrUnsafeEntry reports a tar entry that must never be written to disk --
	// a symlink or hardlink, a non-regular file, a path escaping the
	// destination, or a name outside the archive layout. It also covers the
	// export side of the same rule: a memory/ or notes/ tree root that is not a
	// real directory is refused rather than silently exported as empty.
	ErrUnsafeEntry = errors.New("unsafe archive entry")

	// ErrNotArchive reports a file that is not a Seamless archive: not a
	// gzipped tar, or a tar whose first entry is not manifest.json.
	ErrNotArchive = errors.New("not a seamless archive")

	// ErrDaemonRunning reports that a daemon is answering on the target
	// instance's address. A fresh restore replaces seam.db underneath it, so it
	// refuses rather than corrupting a live instance.
	ErrDaemonRunning = errors.New("seamlessd is running against this data dir")
)

Sentinels. Each one names a distinct refusal, so callers can tell "this file is not one of ours" from "this file is ours and I am too old to read it" without matching on message text.

Functions

This section is empty.

Types

type Counts

type Counts struct {
	// MemoryFiles and NoteFiles count the .md entries in each tree.
	MemoryFiles int `json:"memory_files"`
	NoteFiles   int `json:"note_files"`

	// Tables maps table name to row count in the seam.db snapshot, read from
	// the snapshot after it verified rather than from the live database, so the
	// numbers describe the bytes actually in the archive. Empty for a NoDB
	// export. Enumerated from sqlite_master rather than a hand-written list, so
	// a table added by a later migration appears without anyone remembering to
	// add it here.
	Tables map[string]int `json:"tables,omitempty"`
}

Counts is the manifest's inventory: what a reader should expect to find, so a truncated or hand-edited archive is detectable before it is trusted.

type ExportOptions

type ExportOptions struct {
	// DataDir is the instance root: the directory holding seam.db, memory/ and
	// notes/. It is named explicitly rather than read from config so an export
	// can target a data dir this process does not serve.
	DataDir string

	// Out receives the gzipped tar. It is an io.Writer, not a path, because the
	// CLI writes to a temp file it renames into place and to stdout for `-o -`;
	// choosing the destination is the caller's job.
	Out io.Writer

	// NoDB exports the markdown trees only. The manifest then carries no
	// schema version, no table counts, and no embedding models: those describe
	// a database this archive does not contain, and stamping them from the live
	// file would describe bytes that are not here. seam.db is not even opened,
	// so a knowledge-only export works on a data dir with no database at all.
	NoDB bool

	// SeamlessdVersion is recorded verbatim in the manifest. It is passed in
	// because the build version is linked into package main and this package
	// sits below cmd/.
	SeamlessdVersion string

	// Host overrides the recorded source host. Empty means os.Hostname.
	Host string
}

ExportOptions configures one export. Only DataDir and Out are required.

type ImportOptions

type ImportOptions struct {
	// DataDir is the destination instance root. It may not exist yet.
	DataDir string

	// Src is the archive stream. It is an io.Reader, not a path, so `import
	// --from -` can read the archive from stdin; the reader is consumed once,
	// front to back, and never seeked.
	Src io.Reader

	// DryRun reports what the import would do and imports nothing. A fresh
	// destination is left untouched entirely (nothing is even extracted); a
	// merge still stages the archive in a temp dir and opens the destination
	// database, because the counts it promises are the real ones and the only
	// way to compute them is against both schemas.
	DryRun bool

	// Embedder vectorizes memories and notes as they are written during a
	// merge. Nil is supported and warned about: the merge imports no embeddings
	// row (vectors belong to whichever model this machine runs), so without an
	// embedder the imported items are lexically searchable and not semantically
	// until a re-embed pass runs.
	Embedder llm.Embedder

	// Logger receives the counts. Nil uses slog.Default.
	Logger *slog.Logger
}

ImportOptions configures one import run.

type ImportReport

type ImportReport struct {
	Mode    Mode
	DryRun  bool
	DataDir string

	// Manifest is the archive's own description of itself, carried through so
	// the caller can print what it was handed without re-reading the file.
	Manifest Manifest

	// Memories and Notes count corpus files written: restored verbatim for a
	// fresh import, written through the files layer for a merge. Skipped counts
	// items whose id was already present (the idempotent re-merge).
	Memories int
	Notes    int
	Skipped  int

	// Projects counts projects-table rows registered for a slug the archive's
	// corpus referenced but its projects table did not carry.
	Projects int

	// Rows is the per-table inserted count of the merge row pass, keyed by
	// table name. Empty for a fresh restore, where the whole database arrives
	// as a snapshot and the manifest's counts describe it.
	Rows map[string]int

	PathCollisions []PathCollision
	NameCollisions []NameCollision
	Warnings       []string
}

ImportReport is what one import did, or (DryRun) would do.

func Import

func Import(ctx context.Context, opts ImportOptions) (*ImportReport, error)

Import restores or merges the archive in opts.Src into opts.DataDir.

Which of the two it is comes from DetectMode, is recorded in the report, and is never overridden: restoring over a populated instance would replace its database, and merging into an empty one would rewrite every markdown file through the YAML marshaller instead of restoring it byte-for-byte.

The report is returned even on failure, so a caller can print what happened before the error that stopped it.

func (ImportReport) String

func (r ImportReport) String() string

String renders the report for a CLI, in the shape importer.Report uses.

type Manifest

type Manifest struct {
	// FormatVersion is the container version (FormatVersion).
	FormatVersion int `json:"format_version"`

	// SchemaVersion is the migration version of the snapshot, 0 for a NoDB
	// export. A value above the reader's store.LatestSchemaVersion is
	// ErrSchemaTooNew; anything at or below it migrates forward on import.
	SchemaVersion int `json:"schema_version"`

	// SeamlessdVersion is the build that wrote the archive, supplied by the
	// caller: the version is linked into the binary, and this package sits
	// below cmd/.
	SeamlessdVersion string `json:"seamlessd_version"`

	CreatedAt  time.Time `json:"created_at"`
	SourceHost string    `json:"source_host"`

	// IncludesDB is false for a knowledge-only (NoDB) export, where the tar
	// carries the two markdown trees and nothing else.
	IncludesDB bool `json:"includes_db"`

	// EmbeddingModels is DISTINCT model FROM embeddings in the snapshot. The
	// importer compares it with the destination's configured embedder: vectors
	// from another model are not comparable, so a mismatch is the signal to
	// re-embed rather than to trust the imported rows.
	EmbeddingModels []string `json:"embedding_models,omitempty"`

	Counts Counts `json:"counts"`
}

Manifest is manifest.json, the first entry of every archive.

func Export

func Export(ctx context.Context, opts ExportOptions) (Manifest, error)

Export writes a complete archive of the instance at opts.DataDir to opts.Out and returns the manifest it recorded.

Ordering is deliberate. The database snapshot is taken FIRST and the markdown trees are walked second, so the file set is a superset of what the snapshot's indexes describe: a memory written during the export is present as a file with no index row, which the import's Reconcile heals. The other order loses the file.

The snapshot is `VACUUM INTO` on a store.OpenExisting handle. OpenExisting (not Open) because a newer binary must never migrate a live database as a side effect of backing it up, and VACUUM INTO (not a file copy) because WAL means the bytes on disk are not a consistent database on their own: the snapshot is taken inside a read transaction, so a write in flight from the running daemon is simply not in it.

type Mode

type Mode string

Mode is what an import into a given destination turns out to be. It is a property of the DESTINATION, not a flag: there is no --mode override, because the only thing an override could do is turn a merge into a restore that overwrites a populated instance.

const (
	// ModeFresh restores the archive bit-for-bit into an empty destination.
	ModeFresh Mode = "fresh"
	// ModeMerge folds the archive into an instance that already has data,
	// first-writer-wins by ULID.
	ModeMerge Mode = "merge"
)

func DetectMode

func DetectMode(dataDir string) (Mode, error)

DetectMode reports whether an import into dataDir would be a fresh restore or a merge.

Fresh iff the data dir is absent, or it holds neither seam.db nor seam.db-wal AND neither markdown tree holds a regular non-dot file. The -wal file is part of the test because a database can legitimately be sitting in WAL with its main file freshly checkpointed; treating that as "no database" would restore on top of a live instance. Dot-prefixed names do not count, so a .DS_Store or a stray .seamless-tmp-* from an interrupted write does not make a pristine destination look populated.

type NameCollision

type NameCollision struct {
	Table      string
	Column     string
	Value      string
	IncomingID string
	ExistingID string
}

NameCollision is a row the merge could not insert because a UNIQUE column (sessions.name, projects.slug) is already taken by a different row. Reported, never resolved: minting "cc/ab12cd34-2" would invent a session name that nothing else in either instance refers to.

type PathCollision

type PathCollision struct {
	// Path is the data-dir-relative path both items want.
	Path string
	// IncomingID is the archive's item; ExistingID is the one already there,
	// empty when the file exists on disk with no readable id.
	IncomingID string
	ExistingID string
}

PathCollision is a corpus file the merge refused to write because the path it belongs at is already held by a different item. Both ids are reported because the resolution is a human decision -- keep B's, rename A's, or delete the tombstone holding the name -- and it cannot be made without knowing which two items are involved. The importer never renames around a collision.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL