hive

package
v0.0.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 22, 2026 License: MIT Imports: 38 Imported by: 0

Documentation

Index

Constants

View Source
const SchemaVersion = "v0.5.1"

SchemaVersion is the version stamped into the embedded schema. Kept as a constant so callers can compare against archives without executing SQL.

Variables

View Source
var (
	// ErrNotFound is returned when a lookup fails because no row matches.
	// HTTP handlers map this to 404.
	ErrNotFound = errors.New("core: not found")

	// ErrValidation is returned for user-visible validation failures.
	// HTTP handlers map this to 422 with per-field details when available.
	ErrValidation = errors.New("core: validation failed")

	// ErrConflict is returned for optimistic-concurrency mismatches
	// (If-Match / col__modified) and unique-constraint violations.
	// HTTP handlers map this to 409.
	ErrConflict = errors.New("core: conflict")

	// ErrReadOnly is returned when a write operation is attempted against
	// an archive opened with the ReadOnly option (or one whose underlying
	// file is read-only on disk). HTTP handlers map this to 403.
	ErrReadOnly = errors.New("core: archive is read-only")

	// ErrExists is returned by Create when the target path already exists.
	ErrExists = errors.New("core: archive already exists")
)
View Source
var (
	Version = "v0.0.1"
	Build   = "n/a"
)

Version and Build identify the binary. Release builds set them at link time with -ldflags "-X github.com/sfborg/hive/pkg.Version=<tag> -X github.com/sfborg/hive/pkg.Build=<commit>"; other builds use the defaults below. Update Version for each release.

AllRoles is the canonical order in which the WUI/TUI render agent sections. Kept in one place so a new role (if sfga ever adds one) is a one-file change here + a UI addition; nothing loops over the tables directly.

View Source
var ErrBHLnamesNoMatch = bhlnames.ErrNoMatch

ErrBHLnamesNoMatch is re-exported so callers relying on the sentinel keep working.

View Source
var ErrBibTeXParse = bibtex.ErrParse

ErrBibTeXParse is re-exported so callers can errors.Is-check the existing core sentinel without importing pkg/bibtex.

View Source
var ErrOpenAlexNotFound = openalex.ErrNotFound

ErrOpenAlexNotFound is re-exported so existing callers keep compiling against a stable core-package symbol.

View Source
var RuleFixtures = buildRuleFixtures()

RuleFixtures is the canonical registry — every rule hive ships with should have an entry here. Order is loose; grouped by scoping (role-table check-digit family, then name-scoped rules, then taxon-scoped, then reference-scoped).

View Source
var Schema string

Schema is the sfga SQLite schema DDL, embedded at build time. Source: github.com/sfborg/sfga (schema.sql). Update by re-copying from upstream when sfga cuts a new version. The SchemaVersion constant must be kept in sync with the version stored in the sfga `version` table.

Functions

func ActorFromContext

func ActorFromContext(ctx context.Context) string

ActorFromContext returns the actor stored in ctx, or empty string if none was set. An empty actor is legal — some flows (imports, tests) run without one — but callers that require attribution should check.

func BHLnames

func BHLnames() *bhlnames.Client

BHLnames returns the shared process client, pointed at the production BHLnames endpoint. Cached across calls; construct a fresh client in tests via bhlnames.New(httptestURL).

func GetVersion

func GetVersion() gnvers.Version

GetVersion returns hive's version and build information.

func ItalicForRank

func ItalicForRank(rankID string) bool

ItalicForRank is the exported alias that front-ends use to apply the same italicization rule as BuildLabel does for its HTML output. The TUI calls it to wrap canonical name spans in lipgloss italic before rendering, giving the terminal the same visual hierarchy as the WUI's HTML.

func NameRawStatus

func NameRawStatus(nameID string) string

NameRawStatus returns the raw col__status_id string for a name we previously read via GetName, or "" if we haven't seen it. Callers that need to display or PATCH a NOMEN URI use this instead of n.Status.ID().

func NomenCodePrefix

func NomenCodePrefix(sfgaCodeID string) string

NomenCodePrefix maps a sfga nom_code enum ID (ZOOLOGICAL, BOTANICAL, BACTERIAL, VIRUS, CULTIVARS, PHYTOSOCIOLOGICAL) to the code abbreviation NOMEN labels use (ICZN, ICN, ICNP, ICVCN, ICNCP, ICPN). The status picker filters NOMEN terms by this prefix so a name under zoological code only sees ICZN options. Unknown / empty input → "".

func NomenDowncast

func NomenDowncast(uri string) string

NomenDowncast returns the ChecklistBank-generalized nom_status bucket for a NOMEN URI, walking the OWL subClassOf hierarchy up to a mapped root. Used at CoLDP export time (and at write time once tw__name_status lands — see PROJECTS.md § tw__name_status future direction). Empty return means "no CoLDP-generalized target for this URI"; the caller writes NULL to col__status_id.

func NomenLabelFor

func NomenLabelFor(uri string) string

NomenLabelFor returns the human label for a NOMEN URI, or "" if the URI isn't in the ontology. Read paths use this so a name's stored col__status_id renders as "ICZN objective synonym" rather than the raw URI. Legacy col__status_id values (CoLDP-generalized enums like "ESTABLISHED") won't be in the map — callers should fall back to their own display derivation.

func OpenAlex

func OpenAlex() *openalex.Client

OpenAlex returns a shared openalex.Client built from the current process identity's OpenAlexEmail (see config.CurrentIdentity, set by main.go after resolving flag > env > config). Cached — repeated calls return the same client. Change requires a process restart.

func ParseBibTeX

func ParseBibTeX(input string) (*coldp.Reference, error)

ParseBibTeX is a thin wrapper preserving the existing caller API (core.ParseBibTeX) while the parser itself lives in pkg/bibtex.

func ParseNomStatus

func ParseNomStatus(s string) coldp.NomStatus

ParseNomStatus mirrors ParseRank for the nomenclatural-status enum. Empty stays empty; a name whose col__status_id is NULL stays NULL on round-trip.

func ParseRank

func ParseRank(s string) coldp.Rank

ParseRank returns coldp.NewRank(s) when s is non-empty, and the zero-value Rank (ID="") otherwise. Use for every place hive materializes a Rank from a database cell, HTTP field, or picker id.

func PrimaryReferenceID

func PrimaryReferenceID(raw string) string

PrimaryReferenceID returns the first entry from name.col__reference_id. sfga stores this column as a comma-separated list (a name can be cited in multiple references), but hive v1 treats it as a single primary value — the picker and PATCH surface both operate on one id at a time. Multi-reference support is deferred; existing multi-entry values on read are preserved through Update (only the primary is exposed to the picker; the raw column round-trips untouched otherwise).

func ReferenceLabel

func ReferenceLabel(r *coldp.Reference) string

ReferenceLabel composes a compact "Author (Year) Title" display from a coldp.Reference — the same format the /api/reference/search hits use and the shape the TUI/WUI pickers show curators.

Falls back to Citation, then ID, when the structured fields are empty so the picker never shows a blank line for a stored reference.

func SeedNomenIntoNomStatus

func SeedNomenIntoNomStatus(ctx context.Context, db *sql.DB) error

SeedNomenIntoNomStatus inserts NOMEN classification URIs into the sfga nom_status vocab table, and NOMEN relationship URIs into nom_rel_type — the two FK targets for col__status_id (on name / synonym) and col__type_id (on name_relation) respectively. Idempotent via INSERT OR IGNORE; safe to call on every Archive.Open.

Only URIs with a TW-known Kind ("classification" or "relationship") get seeded — ranks, name-parts, and other NOMEN concepts don't belong in either FK-checked vocabulary. Existing CoLDP-generalized rows in both tables are untouched; NOMEN URIs coexist alongside them.

func WithActor

func WithActor(ctx context.Context, actor string) context.Context

WithActor returns a context carrying the given actor string. Mutations run through *Tx will write this actor to col__modified_by (and other actor columns like col__scrutinizer_id where applicable) on every affected row.

In v0 the actor is typically the local username. Post-ORCID-SSO the actor will be the ORCID iD, injected by auth middleware. The core package does not care where the actor comes from.

Types

type Agent

type Agent struct {
	ID           int
	Role         Role
	Orcid        string
	Given        string
	Family       string
	RorID        string
	Department   string
	Organisation string
	City         string
	State        string
	Country      string
	Email        string
	URL          string
	Note         string
}

Agent mirrors sflib's coldp.Actor plus the sfga row id and the role that identifies which table the row lives in. All 5 sfga role tables share this exact column shape; carrying Role in the value keeps CRUD polymorphic without a table string sneaking into the HTTP wire format.

Note: sfga names the column col__note (singular). Curators use it for a role/contribution note ("primary curator", "database contact"); the WUI/TUI label reflects that even though the sfga column name stays as-is.

func FromActor

func FromActor(role Role, a coldp.Actor) Agent

FromActor copies a coldp.Actor into an Agent under the given role. Convenience for callers that already have an Actor (e.g. from an import flow) and want to write it via the hive edit API. ID is left zero so CreateAgent generates one.

func (Agent) ToActor

func (a Agent) ToActor() coldp.Actor

ToActor materializes a coldp.Actor from this Agent — useful for round-tripping into sflib's export path without leaking hive's Role wrapper.

type Archive

type Archive struct {
	// contains filtered or unexported fields
}

Archive is an sfga SQLite archive opened for reading and, unless the ReadOnly option was used, writing.

Reads are methods on *Archive. Writes are methods on *Tx, obtained via WithTx. This split makes it impossible to accidentally mutate the archive outside a transaction.

A read-write Archive wraps a sflib sfga.Archive handle so we can share version/compat checks, schema migration, and import/enrichment flows with the rest of the SFBorg ecosystem. Read-only archives skip sflib because sflib's Connect() sets journal_mode=WAL, which requires write access.

func Create

func Create(path string) (*Archive, error)

Create creates a new sfga archive at path by applying the embedded schema DDL. It refuses to overwrite an existing file (returns ErrExists).

This bypasses sflib's Create(dir) workflow, which is geared toward archive-conversion pipelines and writes to a fixed schema.sqlite name in a working directory. Hive's Create-at-path is simpler: fresh file, apply the embedded schema.sql, then hand off to the standard Open path so the resulting Archive is indistinguishable from any other read-write open.

func Open

func Open(path string, opts ...OpenOption) (*Archive, error)

Open opens an existing sfga archive at path.

func (*Archive) Ancestors

func (a *Archive) Ancestors(ctx context.Context, id string) ([]string, error)

Ancestors returns the parent chain of a taxon in root-down order, excluding the taxon itself. Empty return means id is at root level (or unknown — callers checking existence should GetTaxon first).

Used by the tree panes to expand the path from the root down to a freshly-moved taxon so it appears in the correct place without a full tree reload. Runs in O(depth) via a recursive CTE.

Prefer Classification for callers that also need display data — Ancestors is the id-only shim kept for the legacy /ancestors endpoint and any caller that truly just needs the id chain.

func (*Archive) Classification

func (a *Archive) Classification(ctx context.Context, id string) ([]TaxonHit, error)

Classification returns the full ancestor chain of a taxon in root-down order, INCLUDING the taxon itself as the final element. Each hit carries the display name, rank, authorship, and status — enough for callers to render breadcrumbs, expand a tree path, or feed a picker without a second round trip per ancestor.

Empty return means id is unknown. A one-element return means id has no parent (root-level taxon).

One recursive-CTE query with an outer JOIN against name replaces what the old reveal path took N sequential Ancestors + ListChildren calls to compute. See DEFERRED.md § Bulk classification endpoint (now resolved by this method).

func (*Archive) ClearRuleConfig

func (a *Archive) ClearRuleConfig(ctx context.Context, ruleID string) error

ClearRuleConfig removes a rule's override entirely. Silent no-op when the row doesn't exist.

func (*Archive) ClearRulesetConfig

func (a *Archive) ClearRulesetConfig(ctx context.Context, name string) error

ClearRulesetConfig removes a ruleset override. Silent no-op when the row doesn't exist.

func (*Archive) Close

func (a *Archive) Close() error

Close releases the underlying database handle.

func (*Archive) CodeForParent

func (a *Archive) CodeForParent(ctx context.Context, parentID string) (string, error)

CodeForParent returns the nomenclatural code (col__code_id on the parent taxon's associated name row) so the new-taxon form can seed its code picker from the parent by default. Empty parentID → "". Missing parent, no name, or no code all return "" — the caller is expected to treat an empty result as "no default known, ask the user."

func (*Archive) CreateNamePrefix

func (a *Archive) CreateNamePrefix(ctx context.Context, parentID string) (string, error)

CreateNamePrefix returns the string a new child of parentID should start with — the parent's canonical name plus a trailing space, so the curator only types the new epithet. Empty string means "no prefix, curator types the full name" (for parents above the genus-group where the child is a fresh uninomial).

Examples:

parent Felis (GENUS)         → "Felis "
parent Felis catus (SPECIES) → "Felis catus "     (child is subspecies)
parent Felidae (FAMILY)      → ""                 (child is a genus)

Uses gn__canonical_full (from gnparser) — no authorship, subgenus parens preserved — so a curator adding a subspecies of "Felis catus Linnaeus, 1758" gets "Felis catus " to type after, not "Felis catus Linnaeus, 1758 " (each child usually has its own authorship). Falls back to canonical_simple if _full is unset, then to col__scientific_name only if both parser caches are empty (very legacy archive).

Sfga's rank vocab flags (col__genus_group, col__infraspecific) plus a special case for SPECIES / SPECIES_AGGREGATE decide when to prefix at all.

func (*Archive) DB

func (a *Archive) DB() *sql.DB

DB returns the underlying *sql.DB for read operations that don't fit any higher-level API. Callers must not begin transactions directly through this handle — use WithTx.

func (*Archive) GetAgent

func (a *Archive) GetAgent(ctx context.Context, role Role, id int) (*Agent, error)

GetAgent returns a single agent by role + id. ErrNotFound when the row is missing so the HTTP layer maps cleanly to 404.

func (*Archive) GetDistribution

func (a *Archive) GetDistribution(ctx context.Context, rowid int64) (*DistributionHit, error)

GetDistribution returns the row with the given rowid, or ErrNotFound when nothing matches. Frontends usually have the row from a list call already; this method exists for direct- address paths (deep links, PATCH round-trips that want the pre-image, tests).

func (*Archive) GetMetadata

func (a *Archive) GetMetadata(ctx context.Context) (*Metadata, error)

GetMetadata returns the archive's single metadata row (id=1 by convention). Returns ErrNotFound when no row exists — callers on legacy archives should treat that as "prompt curator to seed one." Every hive-created archive is seeded on Create so fresh archives always have a row.

func (*Archive) GetName

func (a *Archive) GetName(ctx context.Context, id string) (*coldp.Name, error)

GetName returns the name with the given col__id, including all gn__* GlobalNames fields.

COALESCE folds NULL columns back to ” for the plain-string coldp.Name fields (see pkg/taxon.go for the "" ↔ NULL convention).

func (*Archive) GetReference

func (a *Archive) GetReference(ctx context.Context, id string) (*coldp.Reference, error)

GetReference returns the full coldp.Reference for id.

func (*Archive) GetRuleConfig

func (a *Archive) GetRuleConfig(ctx context.Context, ruleID string) (*RuleConfig, error)

GetRuleConfig returns the override row for a rule, if one exists. Returns nil (no error) when the rule has no override — caller treats that as "use bundle defaults."

func (*Archive) GetRulesetConfig

func (a *Archive) GetRulesetConfig(ctx context.Context, name string) (*RulesetConfig, error)

GetRulesetConfig returns the toggle state for a ruleset. Nil (no error) when no override row exists — caller treats that as "use default (enabled)."

func (*Archive) GetSpeciesInteraction

func (a *Archive) GetSpeciesInteraction(ctx context.Context, rowid int64) (*SpeciesInteractionHit, error)

GetSpeciesInteraction returns the row with the given rowid, or ErrNotFound when nothing matches. Frontends usually have the row from a list call already; this method exists for direct- address paths (deep links, PATCH round-trips that want the pre-image, tests).

func (*Archive) GetSynonym

func (a *Archive) GetSynonym(ctx context.Context, taxonID, nameID string) (*coldp.Synonym, error)

GetSynonym returns the single synonym link between taxonID and nameID. A single call fetches one row even if the synonym is pro-parte; see SynonymPartners to enumerate the taxa a pro-parte synonym links to.

func (*Archive) GetTaxon

func (a *Archive) GetTaxon(ctx context.Context, id string) (*coldp.Taxon, error)

GetTaxon returns the taxon with the given col__id.

Denormalized classification columns (col__genus, col__family, col__kingdom, …) and their sf__* ID siblings are NOT populated here — they are a query-time cache managed by MoveTaxon, not part of the taxon's identity. Callers wanting a full classification should walk the parent chain via TaxonPath (once implemented) or query the columns directly for legacy archives where they were pre-populated by an importer.

func (*Archive) GetTypeMaterial

func (a *Archive) GetTypeMaterial(ctx context.Context, rowid int64) (*TypeMaterialHit, error)

GetTypeMaterial returns the row with the given rowid or ErrNotFound when nothing matches. Frontends typically already have the row from a list call; this method exists for direct- address paths (deep links, PATCH round-trips, tests).

func (*Archive) GetVernacular

func (a *Archive) GetVernacular(ctx context.Context, rowid int64) (*VernacularHit, error)

GetVernacular returns the row with the given rowid, or ErrNotFound when nothing matches. Frontends generally have the row from a list call already; this method exists for direct-address paths (deep links, PATCH round-trips that want the pre-image, tests).

func (*Archive) IsReadOnly

func (a *Archive) IsReadOnly() bool

IsReadOnly reports whether the archive was opened read-only.

func (*Archive) IssueSummary

func (a *Archive) IssueSummary(ctx context.Context) ([]IssueSummaryRow, error)

IssueSummary returns every (table, rule, severity, count) triple present in the archive. Ordered by count desc so the top of the list is the biggest bar in any chart.

func (*Archive) ListAgents

func (a *Archive) ListAgents(ctx context.Context, role Role) ([]Agent, error)

ListAgents returns every row from the role's table, ordered by family then given so the WUI/TUI card list is stable across reloads. sfga role tables are small (dataset agents, not taxonomic records) — pagination isn't needed at v0.

func (*Archive) ListChildren

func (a *Archive) ListChildren(ctx context.Context, parentID string) ([]TaxonHit, error)

ListChildren returns every direct child of parentID. Convenience wrapper over ListChildrenPage with limit == 0. Suits the TUI, which wants the full sibling list to render an expandable tree node.

func (*Archive) ListChildrenPage

func (a *Archive) ListChildrenPage(ctx context.Context, parentID string, limit, offset int) ([]TaxonHit, int, error)

ListChildrenPage returns direct children with SQL-level LIMIT/OFFSET pagination. limit == 0 disables the LIMIT clause and returns every row (matching ListChildren). The full sibling count (independent of the page window) is also returned so pagers can render "N of M".

Ordered by col__ordinal (NULLs last), then by scientific name — the same stable order in every page so callers can advance a cursor safely.

See GetTaxon for the "" ↔ NULL parent_id COALESCE convention.

func (*Archive) ListDistributions

func (a *Archive) ListDistributions(ctx context.Context, taxonID string) ([]DistributionHit, error)

ListDistributions returns every distribution row attached to the given taxon, ordered by gazetteer then area for a stable display where curators expect to see the geographic groups (ISO / TDWG / TEOW / TEXT) laid out together.

func (*Archive) ListIssues

func (a *Archive) ListIssues(ctx context.Context, f IssueFilter, limit, offset int) ([]Issue, int, error)

ListIssues returns a page of stored issues matching the filter, with record labels + navigation hints pre-resolved so the frontend renders in one round-trip. Ordered by severity (error first) then rule then most-recent-first, so the top of the list is the highest- priority open issue.

func (*Archive) ListNameRelations

func (a *Archive) ListNameRelations(ctx context.Context, id string) ([]NameRelationHit, error)

ListNameRelations returns every name_relation row where the given name is either the subject or the object. Used by the detail-view pane to surface "basionym: X" and "has-basionym-of: Y" links.

The projection carries the counterpart name id (whichever side isn't `id`) plus a `direction` marker so callers can render the relation naturally: "→" when `id` is the subject, "←" when it's the object. RowID is the addressable handle for update / delete calls — sfga's name_relation has no col__id column, so hive uses SQLite's implicit rowid (same pattern as vernacular / distribution).

func (*Archive) ListReferencesPage

func (a *Archive) ListReferencesPage(ctx context.Context, limit, offset int) ([]ReferenceHit, int, error)

ListReferencesPage returns references with SQL-level LIMIT/OFFSET pagination. limit == 0 disables the LIMIT clause. Ordered by author then year for a stable, curator-scannable list.

func (*Archive) ListRuleConfigs

func (a *Archive) ListRuleConfigs(ctx context.Context) ([]RuleConfig, error)

ListRuleConfigs returns every rule override the archive holds, ordered by rule id for deterministic display.

func (*Archive) ListRulesetConfigs

func (a *Archive) ListRulesetConfigs(ctx context.Context) ([]RulesetConfig, error)

ListRulesetConfigs returns every ruleset toggle the archive holds.

func (*Archive) ListSpeciesInteractionTypes

func (a *Archive) ListSpeciesInteractionTypes(ctx context.Context) ([]VocabTermDetail, error)

ListSpeciesInteractionTypes returns every row in the species_interaction_type vocab with full detail — the editor's list source. Ordered by name for scannability (43-ish entries; a two-page scroll at most).

func (*Archive) ListSpeciesInteractions

func (a *Archive) ListSpeciesInteractions(ctx context.Context, taxonID string) ([]SpeciesInteractionHit, error)

ListSpeciesInteractions returns every interaction row where the given taxon is the subject (col__taxon_id). Ordered by type then resolved related-name so the WUI list groups by interaction kind and reads alphabetically within a kind.

Left-joins to the related taxon's name row when RelatedTaxonID is populated, so the returned Label carries a rendered display string without a follow-up per-row lookup. Free-text interactions (RelatedTaxonID empty, RelatedTaxonScientificName populated) leave Label empty; the frontend renders the free-text field.

func (*Archive) ListSynonymHits

func (a *Archive) ListSynonymHits(ctx context.Context, taxonID string) ([]SynonymHit, error)

ListSynonymHits returns synonyms for a taxon with each row's name rendered as a Label (canonical + authorship, italicized per rank). Same ordering as ListSynonyms — alphabetical by canonical name — so the two list views stay consistent across the TUI and WUI.

func (*Archive) ListSynonyms

func (a *Archive) ListSynonyms(ctx context.Context, taxonID string) ([]coldp.Synonym, error)

ListSynonyms returns all synonym rows pointing at the given taxon. Ordered by the associated name's canonical form for stable UI display.

func (*Archive) ListTypeMaterials

func (a *Archive) ListTypeMaterials(ctx context.Context, nameID string) ([]TypeMaterialHit, error)

ListTypeMaterials returns every type_material row attached to the given name, ordered by status then rowid so holotype / lectotype records sort near the top of the display.

func (*Archive) ListVernaculars

func (a *Archive) ListVernaculars(ctx context.Context, taxonID string) ([]VernacularHit, error)

ListVernaculars returns every vernacular row attached to the given taxon, ordered preferred-first, then by language then name so the WUI list renders the "one preferred per language" pattern curators expect near the top.

func (*Archive) MetadataWarnings

func (a *Archive) MetadataWarnings(ctx context.Context, metadataID int) []ValidationWarning

func (*Archive) NameDependencies

func (a *Archive) NameDependencies(ctx context.Context, nameID string) (NameDependencies, error)

NameDependencies returns the count-projection for what references the given name. Cheap — three indexed COUNT queries, ~microseconds on a warm cache. Used by the WUI synonym-delete flow to decide whether to offer a "cascade to name" option; would refuse the cascade DELETE if a subsequent write shows fresh dependencies (the pkg-level DeleteName does its own check inside the tx).

func (*Archive) NameRef

func (a *Archive) NameRef(ctx context.Context, id string) (Ref, error)

NameRef returns id + rendered label for a name — the analogue of TaxonRef but keyed on the name row directly (no taxon lookup). Used by name_relation display paths where the counterpart is a Name that may or may not have its own Taxon row (basionyms are usually synonyms; there's no accepted taxon to route through).

Missing name → Ref containing just the id as the label text.

func (*Archive) NameWarnings

func (a *Archive) NameWarnings(ctx context.Context, nameID string) []ValidationWarning

NameWarnings / TaxonWarnings / MetadataWarnings return the persisted non-blocking issue set for a single record of the corresponding table. Backed by __gsvalidator_results; write paths keep the cache fresh via post-commit syncXIssues. Read path filters to warn/info severities — hard errors would surface via a different channel (RFC 7807 problem) and aren't attached to a successful GET response; debug is diagnostic-only and hidden unless the curator explicitly opts in from the Issues screen.

Empty result set may mean "clean" or "not yet synced" (legacy row). The store makes no distinction; a hive validate reindex backfills.

func (*Archive) NomenclaturalHistory

func (a *Archive) NomenclaturalHistory(ctx context.Context, taxonID string) (*NomenclaturalHistory, error)

NomenclaturalHistory returns the basionym-anchored cluster projection for a taxon. Reads only — safe to call from any request path. Returns an empty (non-nil) NomenclaturalHistory with zero clusters when the taxon has no accepted name row and no synonyms; ErrNotFound when the taxon id itself doesn't exist.

Cost is bounded by (synonym count + basionym cluster size) name rows to fetch. Two batched-IN queries (one for outgoing BASIONYM, one for incoming) plus one to hydrate every name's display fields — cheap on typical taxa (a handful of names) and linear in cluster size on the rare monotypic-genus case (dozens of recombinations).

func (*Archive) ParseNamePreview

func (a *Archive) ParseNamePreview(codeID, verbatim string) (*coldp.Name, string)

ParseNamePreview runs the archive's gnparser on a verbatim scientific name and returns a coldp.Name populated with everything the parser can derive — gn__* cache + atomized col__* fields + a rank guess (via RankGuess) scoped to the given nom_code — plus the unparsed Tail gnparser rejected (empty for clean parses). The returned Name is a *preview* only: no row is written, ID is left empty, Modified/ModifiedBy are not stamped. Tail is a diagnostic aid for the create/edit form; it is not persisted on the name row (gn__* is a cache — tail is re-derivable from the verbatim string on demand).

Backs the /api/name/parse endpoint that drives the two-step name-add form. The same Archive gnparser instance is used (single-flight via parserMu) so the preview and the eventual CreateName agree on the atomization — same version, same options, same result.

func (*Archive) Path

func (a *Archive) Path() string

Path returns the on-disk path of the archive.

func (*Archive) RefreshTimeBasedIssues

func (a *Archive) RefreshTimeBasedIssues(ctx context.Context) error

RefreshTimeBasedIssues evaluates every enabled time-based rule whose cooldown has elapsed and re-syncs its target records. Called from Archive.Open (covers restart), periodic tickers in the server and TUI (covers long-running sessions), and implicitly by ReindexValidation (which walks everything unconditionally).

State lives in __gsvalidator_rule_state.last_run_at, keyed by rule id. A rule missing from that table (never run) is treated as due. Successful evaluation updates the timestamp; failure leaves the previous value in place so a transient error doesn't push the next check by another RecheckDays.

Currently limited to rules whose TableName is one of the tables hive has issue infrastructure for (name / taxon / metadata) — a rule targeting an unknown table is silently skipped.

func (*Archive) ReindexValidation

func (a *Archive) ReindexValidation(ctx context.Context, progress func(ReindexProgress)) error

ReindexValidation walks every row in every hive-validated table and rewrites its __gsvalidator_results rows to match the current rule set. Backfills archives edited before persistence landed and repairs caches when a rule is added, tuned, or removed. Progress fires once per row so a CLI or SSE stream can render a live counter; pass nil to skip reporting.

Iteration order: name → taxon → metadata. Each table walks one row at a time in its own tx via syncIssues — the write lock is held briefly, curators editing the archive in parallel see only per-row contention. Cancellation via ctx stops between rows; rows already synced stay synced.

func (*Archive) SchemaVersion

func (a *Archive) SchemaVersion(ctx context.Context) (string, error)

SchemaVersion returns the version stamped in the archive's `version` table. This is the version the archive was created with; it may differ from the core package's SchemaVersion constant if the archive is older.

func (*Archive) SearchNames

func (a *Archive) SearchNames(ctx context.Context, q string, limit int) ([]NameHit, error)

SearchNames returns up to `limit` names matching q as a substring against gn__canonical_simple, and falls back to col__scientific_name for legacy rows whose gn__* fields are empty. Case-insensitive. Ordered by canonical simple form.

This is deliberately simple for v0: no rank filters, no authorship, no FTS operators. Advanced search is a follow-up; the CLAUDE.md § Deliberately deferred list calls this out.

func (*Archive) SearchReferences

func (a *Archive) SearchReferences(ctx context.Context, q string, limit int) ([]ReferenceHit, error)

SearchReferences returns references whose author, title, or citation matches q as a case-insensitive substring. Ordered by author + year for stable browsing.

func (*Archive) SearchTaxa

func (a *Archive) SearchTaxa(ctx context.Context, q string, opts SearchOpts) ([]TaxonHit, error)

SearchTaxa returns up to opts.Limit taxa whose associated name canonical or scientific-name string matches q under opts.Mode. Returns thin TaxonHit projections — same shape as ListChildren so the WUI's tree components can render either result set uniformly.

When opts.IncludeSynonyms is true, the result set also includes accepted taxa reached via a matching synonym: for each synonym whose name matches q, the accepted taxon it points at appears in the results with IsSynonym=true and MatchedName set to the synonym's text. Pro-parte synonyms — one synonym row family pointing at multiple accepted taxa — produce one hit per resolved accepted taxon so curators see every destination. If the same accepted taxon matches both by its own name and via a synonym, the accepted row wins; the synonym row is dropped.

Ordering is alphabetical by matched_name so synonym and accepted hits interleave in the order curators would look for them.

Every hit carries parent context (ParentName / ParentLabel / ParentRank) so front-ends can disambiguate homonyms and give synonym rows a "you land under X" cue. Root-level accepted taxa leave the parent fields empty.

See DEFERRED.md § Substring name search (FTS follow-up) for the substring-semantics extension that composes with partial mode.

func (*Archive) SetRuleConfig

func (a *Archive) SetRuleConfig(ctx context.Context, cfg RuleConfig) error

SetRuleConfig upserts a rule's override. Both Enabled and SeverityOverride are optional — nil / empty means "clear that dimension of the override, use bundle default." When both are nil/empty the accessor deletes the row entirely (no-op override takes no space).

func (*Archive) SetRulesetConfig

func (a *Archive) SetRulesetConfig(ctx context.Context, name string, enabled bool) error

SetRulesetConfig upserts a ruleset toggle. name must be one of the known bundle names ("hive" | "clb" | "tw") — enforced here to avoid orphan rows for typos of nonexistent bundles.

func (*Archive) Sflib

func (a *Archive) Sflib() sfga.Archive

Sflib returns the underlying sflib sfga.Archive for read-write archives. Returns nil for read-only archives. Callers use this to drive sflib flows (import via Reader/Writer, migration via Migrator, enrichment via Enricher) without hive owning a parallel implementation.

func (*Archive) SynonymPartners

func (a *Archive) SynonymPartners(ctx context.Context, synonymID string) ([]string, error)

SynonymPartners returns every taxon ID that a pro-parte synonym points at. For a non-pro-parte synonym the result is a single-element slice. Returns an empty slice if synonymID is unknown or is "" (the schema allows anonymous synonyms, but hive always assigns an ID so a hive-managed archive will not contain them).

func (*Archive) TaxonDeletePreview

func (a *Archive) TaxonDeletePreview(ctx context.Context, id string) (TaxonDeletePreview, error)

TaxonDeletePreview computes the counts a cascade delete would touch. Read-only; runs in its own connection since it's a read path.

func (*Archive) TaxonRef

func (a *Archive) TaxonRef(ctx context.Context, id string) (Ref, error)

TaxonRef returns id + rendered label for a taxon in one query. Missing taxa return a Ref containing just the id (as the label text), so callers using this as a display resolver get graceful fallback rather than a hard failure. Empty ID short-circuits with a zero-value Ref.

func (*Archive) TaxonWarnings

func (a *Archive) TaxonWarnings(ctx context.Context, taxonID string) []ValidationWarning

func (*Archive) ValidChildRanks

func (a *Archive) ValidChildRanks(ctx context.Context, parentID string) ([]ui.ChildRank, error)

ValidChildRanks returns the rank IDs (plus typical_use flags) a new child of parentID should be allowed to pick from — filtered per the parent's own rank and nomenclatural code via the TaxonWorks-derived rank hierarchy (pkg/ui.ValidChildRanks).

An empty parentID or an archive whose parent lacks a rank / code returns nil, signaling "no filter" so the frontend shows every rank in the vocab. Curator overrides the guess on the form as needed; filtering just removes the obviously wrong picks.

func (*Archive) ValidateName

func (a *Archive) ValidateName(ctx context.Context, nameID string) ([]*domain.Result, error)

ValidateName runs every rule that applies to the name table against the row with the given id. Result includes both passes and failures — callers separate hard errors (Result.IsError) from soft warnings (Result.IsWarning). The first slice of the save-with-acknowledgment UX uses this to attach warnings to successful create/update responses; a follow-up hooks it into Tx.CreateName / Tx.UpdateName for pre-commit hard-fail rollback.

func (*Archive) Vocabulary

func (a *Archive) Vocabulary(ctx context.Context) (*Vocabulary, error)

Vocabulary returns the cached vocabulary, loading it on the first call. Subsequent callers share the same *Vocabulary — safe because the tables don't change during a session unless the vocab editor mutates them. Editor writes call invalidateVocab() so the next Vocabulary() call reloads from disk.

Errors from the load are cached too; callers seeing an error should treat it as terminal for the process (only the initial-load path can error; subsequent loads only happen after successful writes).

func (*Archive) WithTx

func (a *Archive) WithTx(ctx context.Context, fn func(*Tx) error) error

WithTx runs fn inside a transaction. If fn returns nil the transaction is committed; otherwise it is rolled back and fn's error is returned unchanged. The *Tx passed to fn carries the actor from ctx (see WithActor).

WithTx returns ErrReadOnly on a read-only archive without opening a transaction. It always uses the provided context for both begin and commit.

After a successful commit, WithTx runs each name-side validation sync requested via Tx.markNameDirty. Sync errors are swallowed: the write already committed, and a stale issue set is preferable to failing the caller's request. A future reindex flow will backfill any misses.

type BHLnameHit

type BHLnameHit = bhlnames.Hit

Type aliases so existing callers of core.BHLnameHit / core.BHLnameLookupOpts keep compiling without an import change.

type BHLnameLookupOpts

type BHLnameLookupOpts = bhlnames.LookupOpts

Type aliases so existing callers of core.BHLnameHit / core.BHLnameLookupOpts keep compiling without an import change.

type Country

type Country struct {
	ID        string `json:"id"`
	Name      string `json:"name"`
	Alpha3    string `json:"alpha3,omitempty"`
	Continent string `json:"continent,omitempty"`
}

Country is one ISO 3166-1 entry surfaced by /api/vocab/countries. `ID` carries the alpha-2 code (matching sfga's vernacular.col__country storage convention). Alpha3 and Continent are extra context the frontend can render in dropdown rows.

func Countries

func Countries() ([]Country, error)

Countries returns the ISO 3166-1 alpha-2 catalog, alphabetized by name. Parsed from the embedded JSON on first call and cached for the process lifetime — the data is immutable.

type DistributionHit

type DistributionHit struct {
	RowID       int64
	TaxonID     string
	SourceID    string
	Area        string
	AreaID      string
	Gazetteer   coldp.GazetteerEnt
	Status      coldp.DistrStatus
	ReferenceID string
	Remarks     string
	Modified    string
	ModifiedBy  string
	// IssueCount is the number of open (not-yet-acknowledged)
	// validation issues currently filed against this distribution
	// row. Populated per-hit by ListDistributions in a single
	// batched grouping query — cheap even with a large per-taxon
	// distribution set. Zero when the row is clean; front-ends
	// render a warn icon on the row when > 0.
	IssueCount int
	// MaxSeverity is the highest severity ("error" > "warn" > "info"
	// > "debug") among the row's open issues. Empty string when
	// IssueCount is 0. Drives the badge color on the WUI row.
	MaxSeverity string
}

DistributionHit is the projection returned by ListDistributions and the read side of the distribution CRUD. RowID carries SQLite's implicit primary key so update/delete calls can address a single row unambiguously even when multiple rows share (taxon, area).

type FixtureCase

type FixtureCase struct {
	Note  string
	Setup func(context.Context, *Archive) error
}

FixtureCase is one data scenario for a rule. Setup inserts the row(s) that define the scenario into a fresh archive. Bad cases are expected to trigger their parent rule; Good cases are expected to NOT trigger it (controls that prove the rule discriminates).

type Issue

type Issue struct {
	ID             string
	TableName      string // "name" (taxon/reference/... later)
	RecordID       string // col__id of the flagged row
	RuleID         string
	RuleName       string
	FieldName      string
	Severity       string
	Enforcement    string
	Message        string
	ActualValue    string
	ExpectedValue  string
	CreatedAt      string
	AcknowledgedBy string
	AcknowledgedAt string

	// RecordLabel is a human-readable label for the flagged record —
	// scientific name for names, taxon path for taxa, citation for
	// references. Server-resolved so the Issues list renders in one
	// round-trip instead of N follow-ups.
	RecordLabel string

	// LinkTaxonID is the taxon id a click should navigate to, when
	// resolvable. For a name-scoped issue, this is the taxon using
	// that name (nil when the name has no taxon or has more than one).
	LinkTaxonID string
}

Issue is a stored validation result read back from __gsvalidator_results. Frontends consume this directly on both the per-record detail banner (via NameWarnings) and the Issues screen (via ListIssues). Includes the record's owning-taxon id when the rule fired on a name row and that name is used by exactly one taxon — the common case, and enough for the Issues screen to route row clicks into the tree.

type IssueFilter

type IssueFilter struct {
	TableName string
	// RecordID scopes the query to a single row (paired with TableName
	// most usefully — a bare RecordID lookup is fine but would match
	// any table sharing that id, which sfga's UUID space makes
	// vanishingly unlikely). Empty = no filter on record.
	RecordID         string
	RuleID           string
	Severities       []string
	HideAcknowledged bool
}

IssueFilter narrows a ListIssues call. Empty fields mean "no filter on that axis." Severities is OR-matched (any of); TableName and RuleID are exact matches. HideAcknowledged drops rows with a non-empty acknowledged_at, for the default "still-open" view.

type IssueSummary

type IssueSummary struct {
	// Count is the number of open (not-yet-acknowledged) issues
	// currently filed against the record.
	Count int
	// MaxSeverity is the highest severity present ("error" > "warn"
	// > "info" > "debug"). Empty string when Count is 0. Drives
	// the badge color on the WUI side via the shared
	// validationSeverityBadge helper.
	MaxSeverity string
}

IssueSummary is the per-record aggregate returned by batchIssueSummaries — the two facts the WUI needs to render a severity-colored validation badge on a picker or list row.

type IssueSummaryRow

type IssueSummaryRow struct {
	TableName string
	RuleID    string
	RuleName  string
	Severity  string
	Count     int
}

IssueSummaryRow is one bar in the issue-count breakdown: how many issues of a given rule + severity exist on a given table. Frontends group or filter this list however suits their layout.

type Label

type Label struct {
	Text string `json:"text,omitempty"`
	HTML string `json:"html,omitempty"`
}

Label carries both plain-text and HTML-styled forms of a display string. HTML is server-rendered with ICZN-appropriate italics + extinct dagger, so front-ends render consistently without duplicating the styling logic. Plain text has the dagger but no styling — the TUI uses this directly, or can re-apply italics via lipgloss using the same italicForRank rules.

func BuildLabel

func BuildLabel(canonical, authorship, rankID string, extinct bool) Label

BuildLabel produces the (text, html) pair for a name row. Called by TaxonRef, NameRef, and the synonym hit builder to ensure identical formatting across every endpoint.

Formatting:

  • Extinct taxa get a "† " prefix (both forms).
  • The canonical part is wrapped in <i>…</i> when italicForRank(rank) is true — genus group and below, plus infraspecific ranks. Everything above the family group (family, order, class, …) stays roman per ICZN convention.
  • Authorship is never italicized.
  • HTML output escapes canonical and authorship so weird data (unlikely, but possible in legacy archives) can't inject markup.

The rank ID is the raw sfga enum value ("SPECIES", "GENUS", "FAMILY", …). Unknown ranks default to non-italic — safe fallback that keeps the label rendering even for values italicForRank doesn't recognize.

type Language

type Language struct {
	ID   string `json:"id"`
	Name string `json:"name"`
}

Language is one ISO 639-3 entry surfaced by /api/vocab/languages. `ID` is the three-letter code (matching sfga's vernacular.col__language storage convention).

func Languages

func Languages() ([]Language, error)

Languages returns the ISO 639-3 catalog, alphabetized by name. Parsed from the embedded JSON on first call.

type Metadata

type Metadata struct {
	ID              int
	DOI             string
	Title           string
	Alias           string
	Description     string
	Issued          string
	Version         string
	Keywords        string
	GeographicScope string
	TaxonomicScope  string
	TemporalScope   string
	Confidence      *int
	Completeness    *int
	License         string
	URL             string
	Label           string
	Citation        string
	Private         *bool
}

Metadata is the dataset-level metadata row (title, description, license, logo, private flag, …). Hive-authored because sflib doesn't expose a coldp.Metadata type — the sfga `metadata` table is a hive-facing concept for now.

Sfga models it as an INTEGER-primary-key table that other tables reference via col__metadata_id (defaulting to 1). Convention across the ecosystem is one row per archive; hive follows that and treats the metadata row as a singleton — GetMetadata returns the id=1 row (or the sole row present, if any). Seeded on Create so the default FK references from taxon / name / etc. resolve cleanly.

Confidence / Completeness map SQL NULL into an *int on the Go side so unset (never scored) stays distinguishable from 0 (worst score). Private is a tri-state (nil / true / false) for the same reason.

type NameDependencies

type NameDependencies struct {
	// TaxonCount is the number of taxon rows whose col__name_id is
	// this name — >0 means this name is (or was) an accepted name.
	TaxonCount int
	// SynonymCount is the number of synonym rows whose col__name_id is
	// this name — pro-parte synonyms count once per taxon link.
	SynonymCount int
	// NameRelationCount is the number of name_relation rows referencing
	// this name via either col__name_id or col__related_name_id (basionym,
	// homonym, etc.).
	NameRelationCount int
}

NameDependencies is the count-projection returned by Archive.NameDependencies. Each field records how many rows in the respective table reference the given name via its col__name_id (or, for name_relation, either side of the relation). A name with all zero counts is a "bare name" — it exists in the archive but no taxon claims it and no synonym relation points at it. See CLAUDE.md § pkg/ package conventions for the "delete refuses on dependencies" contract that DeleteName enforces.

func (NameDependencies) Total

func (d NameDependencies) Total() int

Total returns the sum of all counts — > 0 when the name is referenced by anything at all. Callers use this to short-circuit the "is this name bare?" check without touching individual fields.

type NameHit

type NameHit struct {
	ID         string
	Scientific string // gn__canonical_simple falling back to col__scientific_name
	Full       string // col__scientific_name (raw)
	Authorship string
	Rank       string // rank ID
	Code       string // nom_code ID
	Status     string // nom_status ID
}

NameHit is the search-result projection returned by SearchNames. Callers fetch the full coldp.Name only when opening a detail view.

type NameRelationHit

type NameRelationHit struct {
	RowID         int64
	CounterpartID string
	Type          string // nom_rel_type ID (BASIONYM, SPELLING_CORRECTION, …)
	Direction     string // "outgoing" (id is subject) or "incoming" (id is object)
	ReferenceID   string
	Page          string
	Remarks       string
}

NameRelationHit is the projection returned by ListNameRelations. It carries the counterpart name id (never `id` itself) plus a direction marker so the caller can render "this name → basionym X" vs. "this name ← recombined-into Y" without another query. RowID is the opaque handle for update / delete calls (see UnlinkNameRelation).

type NomenCluster

type NomenCluster struct {
	Role  string // "accepted" or "synonym"
	Names []NomenName
}

NomenCluster is one basionym-anchored family of names surfaced on a taxon's nomenclatural-history section.

type NomenName

type NomenName struct {
	NameID      string
	Label       Label
	Authorship  string
	Rank        string
	Year        string
	IsBasionym  bool
	Involvement string
	SynonymID   string
	ReferenceID string
	// Atomized authorship — surfaced alongside the pre-formatted
	// Authorship string so the WUI's "Standardized authorship" render
	// can compose the hybrid "(basionym_author, basionym_year)
	// combination_author, combination_year" form neither ICZN nor ICN
	// produces on its own. Empty fields fall back to whichever pieces
	// exist. Same source as apiName's equivalent fields — comes from
	// gnparser's atomization on write (or curator override).
	BasionymAuthorship        string
	BasionymAuthorshipYear    string
	CombinationAuthorship     string
	CombinationAuthorshipYear string
	// IssueCount is the number of __gsvalidator_results rows currently
	// open against this name row (table_name='name', any severity).
	// Populated per-name by a single batched query in
	// Archive.NomenclaturalHistory so the frontend can render a warn
	// icon on synonym / basionym rows that need attention without a
	// per-row round trip.
	IssueCount int
	// MaxSeverity is the highest severity among the open issues on
	// this name ("error" > "warn" > "info" > "debug"). Empty when
	// IssueCount is 0. Drives the WUI badge color via the shared
	// validationSeverityBadge helper.
	MaxSeverity string
}

NomenName is one name entry in a nomenclatural-history cluster. Involvement records how this row relates to the taxon we're viewing: "accepted" (this IS the taxon's accepted name), "synonym" (linked as a synonym; SynonymID is populated), or "unlinked" (a sibling recombination discovered via name_relation but not currently attached to this taxon — useful context, not something the curator needs to act on).

type NomenTerm

type NomenTerm struct {
	ID         string `json:"id"`
	Local      string `json:"local"`
	Label      string `json:"label"`
	Code       string `json:"code,omitempty"`
	Kind       string `json:"kind,omitempty"`
	Class      string `json:"class,omitempty"`
	GbifStatus string `json:"gbif_status,omitempty"`
	ColStatus  string `json:"col_status,omitempty"`
	Parent     string `json:"parent,omitempty"`
	Popular    bool   `json:"popular,omitempty"`
	Comment    string `json:"comment,omitempty"`
}

NomenTerm is one class from the NOMEN ontology, projected to the fields hive's status picker needs.

Code — nomenclatural-code prefix detected from the label (ICZN, ICN,

ICNP, ICVCN, ICNCP, ICPN) so the frontend can filter by the current
name's code. Empty when the term is code-general.

Kind — TW-authoritative partition (see below). Parent — parent URI from OWL's rdfs:subClassOf / rdfs:subPropertyOf;

drives the CoL-style downcast walk.

GbifStatus — TW's Latin annotation (`invalidum`, `conservandum`, …).

Useful for display; not the CoLDP source of truth.

ColStatus — the CoLDP-generalized bucket resolved by walking the

Parent chain up to a mapped root. Populated at parse time.
Matches ChecklistBank's nom_status vocabulary (ESTABLISHED,
NOT_ESTABLISHED, ACCEPTABLE, UNACCEPTABLE, CONSERVED, REJECTED,
DOUBTFUL). Empty when no mapped ancestor exists.

Popular — true for the ~24 CoL-mapped root URIs (ICZN Available,

ICZN Unavailable, ICN validly published, etc.). The picker floats
these to the top of the empty-query results so the common
90%-of-cases choices land in front of the curator without
scrolling — matches TW's "popular statuses on the first tab" UX.

Kind values:

"classification" — attaches to a single name (col__status_id target)
"relationship"   — links two names (col__type_id on name_relation)
""               — not modeled by TW (ranks, name-parts, ...); NOT
                   surfaced in the status picker

func Nomen

func Nomen() ([]NomenTerm, error)

Nomen returns the parsed NOMEN vocabulary. Parses the embedded OWL file on first call, merges in TW's kind/gbif_status partition from nomen_tw.json, resolves the CoLDP-generalized downcast for every term via CoL's mapping table + subClassOf walk, and caches the result — vocab is process-scoped and immutable, so a mutex isn't needed on subsequent reads.

type NomenclaturalHistory

type NomenclaturalHistory struct {
	Clusters []NomenCluster
}

NomenclaturalHistory is the multi-cluster projection returned by Archive.NomenclaturalHistory. A taxon's history is grouped by **basionym anchor**: every name in a cluster shares an original combination (either explicitly via name_relation type=BASIONYM, or implicitly by BEING the basionym itself). This mirrors how taxonomists actually read a synonymy — original name first, then subsequent combinations of the same root.

Clusters carry a role:

  • "accepted" — the cluster contains the taxon's currently accepted name. There is exactly one such cluster per response.
  • "synonym" — the cluster contains only synonyms of the taxon (and any recombinations we discovered via name_relation that aren't themselves linked to this taxon).

Names within a cluster are ordered basionym-first, then by authorship year ascending, then alphabetical. That's the read order for a synonymy paragraph: original combination, then each recombination chronologically.

type OpenOption

type OpenOption func(*openOpts)

OpenOption configures Open. Options are functional; pass them variadically.

func ReadOnly

func ReadOnly() OpenOption

ReadOnly opens the archive in read-only mode. Any WithTx call on a read-only archive returns ErrReadOnly without touching the database.

type Ref

type Ref struct {
	ID    string `json:"id"`
	Label Label  `json:"label,omitzero"`
}

Ref is a lightweight pointer to another entity — id + rendered label. Used wherever a response would otherwise ship a raw UUID (parent taxa, name references, according-to links). Front-ends render the label directly; the id is present for follow-up navigation.

type ReferenceHit

type ReferenceHit struct {
	ID       string
	Author   string
	Year     string
	Title    string
	Citation string
	Type     string // reference_type ID (raw enum ID; front-end can pretty-print)
	// IssueCount is the number of open (not-yet-acknowledged)
	// validation issues currently filed against this reference row.
	// Zero → elided from the wire projection (apiReferenceHit uses
	// omitempty). Populated by scanReferenceHits via a batched
	// __gsvalidator_results lookup.
	IssueCount int
	// MaxSeverity is the highest severity among the record's open
	// issues ("error" > "warn" > "info" > "debug"). Empty when
	// IssueCount is 0. Drives the WUI badge color via the shared
	// validationSeverityBadge helper — same signal appears on every
	// picker/list rendering this record.
	MaxSeverity string
}

ReferenceHit is a thin projection for list / search endpoints. Full coldp.Reference is fetched only when opening a detail view. Mirrors TaxonHit / NameHit so front-end list components can render either aggregate uniformly.

Display: hive picks a short form (author + year + title-short) when the full citation is verbose. The `Citation` field falls back to a composed "Author (Year) Title" when col__citation is empty — many imports leave citation blank and rely on the structured fields.

type ReindexProgress

type ReindexProgress struct {
	Table   string
	Done    int
	Total   int
	Current string // record ID that just finished; empty on the summary tick
}

ReindexProgress reports the state of a long-running reindex to the caller. Sent once per row plus a final call with Done == Total.

type Role

type Role string

Role identifies which of the sfga role tables (creator, contact, editor, contributor, publisher) a hive.Agent lives in. sfga models each role as its own table with an identical column layout, so hive treats them uniformly and parameterizes CRUD on the role value. Persistence stays local to a single table per role — no join table, no shared "person" abstraction, matching the sfga data model exactly.

const (
	RoleCreator     Role = "creator"
	RoleContact     Role = "contact"
	RoleEditor      Role = "editor"
	RoleContributor Role = "contributor"
	RolePublisher   Role = "publisher"
)

func ParseRole

func ParseRole(s string) (Role, error)

ParseRole validates a role string. Returns an empty Role and ErrValidation when the input isn't one of the known role names. Used by the HTTP layer where the role arrives as a URL path parameter.

type RuleConfig

type RuleConfig struct {
	RuleID           string
	Enabled          *bool
	SeverityOverride string // "error" | "warn" | "info" | "debug" | "" for none
	UpdatedAt        string
	UpdatedBy        string
}

RuleConfig captures a curator's overrides for one rule. A nil Enabled means "use ruleset default"; likewise for SeverityOverride. Zero value = "no override at all" (rule uses bundle defaults).

type RuleFixture

type RuleFixture struct {
	RuleID      string
	Description string
	Bad         []FixtureCase
	Good        []FixtureCase
}

RuleFixture couples a rule id with the fixture data that exercises it — canonical Bad case(s) that must fire the rule and optional Good case(s) that must not. Both slices support extension over time as edge cases are discovered.

Two consumers:

  • pkg tests iterate every fixture, applying each case to its own fresh archive and asserting the fire / no-fire invariant.
  • tools/mkdemo applies every fixture (bad + good) to demo.db so developers can inspect a canonical archive where every rule has representative data.

type RulesetConfig

type RulesetConfig struct {
	RulesetName string // "hive" | "clb" | "tw"
	Enabled     bool
	UpdatedAt   string
	UpdatedBy   string
}

RulesetConfig is a coarser toggle: whole bundle on or off. Only records rows for rulesets a curator has explicitly changed; missing row means "use default (enabled)".

type SearchMode

type SearchMode string

SearchMode selects the matching algorithm SearchTaxa uses to find candidate names. Prefix — the default — is fast and unambiguous but only matches names that start with the query. Partial uses an FTS5 mirror to match on any word-prefix in the canonical or scientific name string, so typing an epithet like "rusci" finds "Ceroplastes rusci". Fuzzy (added in a follow-up step) will use a trigram FTS5 mirror to tolerate typos.

const (
	SearchModePrefix  SearchMode = "prefix"
	SearchModePartial SearchMode = "partial"
	SearchModeFuzzy   SearchMode = "fuzzy" // TODO(step-3)
)

type SearchOpts

type SearchOpts struct {
	// Mode selects the matching algorithm. Empty is treated as
	// SearchModePrefix so calls that don't care about mode stay
	// compatible.
	Mode SearchMode
	// IncludeSynonyms extends results with accepted taxa reached
	// via a matching synonym (see SearchTaxa for the resolution
	// semantics).
	IncludeSynonyms bool
	// Limit caps the number of returned hits. Zero → 50.
	Limit int
}

SearchOpts bundles the knobs SearchTaxa accepts. Zero value (empty Mode, IncludeSynonyms=false, Limit=0) is a valid call that behaves as prefix-only, accepted-only, limit=50 — the default combobox contract.

type SexTerm

type SexTerm struct {
	ID          string `json:"id"`
	Name        string `json:"name"`
	Symbol      string `json:"symbol,omitempty"`
	Description string `json:"description,omitempty"`
}

SexTerm is one entry in the rich sex vocabulary — same three values sfga's `sex` table carries (FEMALE / MALE / HERMAPHRODITE) plus the ChecklistBank-sourced human-facing name, glyph symbol, and definition. The plain /api/vocab bundle still ships the flat sfga vocab; this endpoint serves the enriched form pickers use.

func SexTerms

func SexTerms() ([]SexTerm, error)

SexTerms returns the enriched sex vocabulary (symbol + description per term) — same set as sfga's flat `sex` table but suitable for a picker that wants to show the ♀ / ♂ / ⚥ glyph and hover-tooltip definition.

type SpeciesInteractionHit

type SpeciesInteractionHit struct {
	RowID                      int64
	TaxonID                    string
	RelatedTaxonID             string
	RelatedTaxonScientificName string
	RelatedTaxonLabel          Label
	Type                       coldp.SpInteractionType
	// TypeRaw is the col__type_id verbatim from the DB — preserved
	// so freeform values (e.g. "eats" from datasets that don't
	// follow the SCREAMING_SNAKE_CASE convention) survive the round
	// trip. sflib's NewSpInteractionType maps unknown strings to
	// UnknownSpIntT whose .ID() is "", which would otherwise both
	// hide the value from the WUI display and silently zero-out the
	// column on any PATCH that doesn't touch the type. Callers
	// prefer Type.ID() (normalized enum) if non-empty, else fall
	// back to TypeRaw so freeform values render as-authored.
	TypeRaw     string
	SourceID    string
	ReferenceID string
	Remarks     string
	Modified    string
	ModifiedBy  string
	// IssueCount is the number of open (not-yet-acknowledged)
	// validation issues currently filed against this row. Populated
	// per-hit in a single batched grouping query — mirrors
	// vernacular / distribution.
	IssueCount int
}

SpeciesInteractionHit is the projection returned by ListSpeciesInteractions and the read side of the CRUD. RelatedTaxonLabel is server-resolved from the related taxon + name row so the front-end row can render "Panthera leo" without a per-row fetch. Empty when RelatedTaxonID is empty or the referenced taxon can't be resolved (free-text-only interactions rely on RelatedTaxonScientificName instead).

type SynonymHit

type SynonymHit struct {
	ID         string
	TaxonID    string
	NameID     string
	Label      Label
	NamePhrase string
	Status     string // taxonomic_status ID (raw enum ID; front-end can pretty-print)
	Link       string
	Remarks    string
	Modified   string
	ModifiedBy string
	// ReferenceID is the synonym-relation's own reference (col__reference_id
	// on the synonym row) — the citation that establishes the synonymy,
	// often distinct from the reference that originally published the
	// synonym's name. The nomenclatural-history / footnote accumulator on
	// the taxon detail page uses this to number every citation appearing
	// on the page. May contain a comma-separated list per the sfga schema
	// note ("ids about this synonym") — hive treats it as opaque for now
	// and lets the frontend split when it renders.
	ReferenceID string
}

SynonymHit is a synonym row plus a server-rendered Label for its name — the projection front-ends want so they don't render raw name UUIDs. Populated by ListSynonymHits in one query (no N+1 lookup per row). The Label carries text + HTML forms; see core.BuildLabel for the rules.

type TaxonDeletePreview

type TaxonDeletePreview struct {
	DirectChildCount          int
	DescendantCount           int // includes self
	SynonymCount              int
	VernacularCount           int
	DistributionCount         int
	MediaCount                int
	TreatmentCount            int
	SpeciesEstimateCount      int
	TaxonPropertyCount        int
	SpeciesInteractionCount   int
	TaxonConceptRelationCount int
	// ParentID of the taxon being previewed. Empty when the taxon is a
	// root — reparent-flow callers use this to decide the destination
	// (children become new roots when their parent was already a root).
	ParentID string
}

TaxonDeletePreview summarizes what a cascade delete of the given taxon would remove: the descendant set (including the taxon itself) and the per-taxon association rows attached to any member of that set. Returned to the WUI's delete-confirm modal so curators see concrete numbers before they type DELETE.

DirectChildCount is broken out separately so the modal can pick the right flow (leaf → simple confirm; parent → three-option UI). DescendantCount includes the taxon itself.

type TaxonHit

type TaxonHit struct {
	ID          string
	ParentID    string
	NameID      string
	Name        string // best-effort display: gn__canonical_simple falling back to col__scientific_name
	Authorship  string
	Rank        string // rank ID (references rank.col__id)
	Status      string // taxonomic_status ID
	Extinct     sql.NullBool
	HasChildren bool
	Label       Label

	// IsSynonym is true when this hit was resolved via a synonym
	// pointing at the accepted taxon (ID). Set only by SearchTaxa
	// when its includeSynonyms parameter is true; false elsewhere.
	IsSynonym bool
	// MatchedName is the name string that actually satisfied the
	// query — the synonym's canonical when IsSynonym, else the
	// accepted taxon's own canonical (same as Name). Callers
	// display it alongside Name so curators can see why a result
	// appeared when it came in via a synonym.
	MatchedName string

	// ParentName / ParentAuthorship / ParentRank / ParentLabel carry
	// the immediate parent taxon's display context so search
	// front-ends can disambiguate homonyms (two accepted taxa sharing
	// a canonical name) and give synonym rows a "this is where you
	// land" cue. Populated only by SearchTaxa; list endpoints
	// (ListChildrenPage, roots, Classification) leave them zero and
	// the wire converter elides the parent object when ParentName is
	// empty. Root-level taxa (no parent) also leave them zero.
	ParentName       string
	ParentAuthorship string
	ParentRank       string
	ParentLabel      Label
}

TaxonHit is a thin projection returned by list/search endpoints. Callers fetch the full coldp.Taxon only when opening a detail view. This split matters for the WUI over HTTP as much as for TUI tree rendering.

Label is server-rendered from Name+Authorship+Rank+Extinct via BuildLabel; the individual fields are still exposed for callers that want to compose their own display (e.g., the TUI's tree, which applies italics via lipgloss instead of HTML).

type Tx

type Tx struct {
	// contains filtered or unexported fields
}

Tx is a transactional edit context. All mutations in pkg/ take *Tx as a receiver, ensuring writes cannot happen outside a transaction. Tx carries a back-pointer to its Archive so aggregate methods can reach shared services like the gnparser instance and cached enums.

func (*Tx) Actor

func (t *Tx) Actor() string

Actor returns the actor string carried by this transaction, taken from the context passed to WithTx at begin time.

func (*Tx) AddDistribution

func (t *Tx) AddDistribution(d coldp.Distribution) (int64, error)

AddDistribution inserts a new distribution row and returns the new rowid as the opaque handle for later PATCH/DELETE.

  • d.TaxonID is required; empty → ErrValidation.
  • Every other field is optional. sfga's schema has NOT-NULL defaults of ” on col__area / col__area_id, and the FK columns (source_id, gazetteer_id, status_id, reference_id) accept ” → NULL so unset values round-trip cleanly.
  • col__modified / col__modified_by stamped from tx context.

func (*Tx) AddSpeciesInteraction

func (t *Tx) AddSpeciesInteraction(s coldp.SpeciesInteraction, typeRaw string) (int64, error)

AddSpeciesInteraction inserts a new interaction row and returns the new rowid as the opaque handle for later PATCH/DELETE.

  • s.TaxonID and s.RelatedTaxonID are both required. sfga's schema declares col__related_taxon_id NOT NULL with a FK to taxon(col__id), so a purely free-text interaction (with only RelatedTaxonScientificName populated) would fail the FK check. The scientific-name field remains available as an annotation alongside the FK, matching sfga's shape.
  • Nullable FKs ("" → NULL): source_id, type_id, reference_id.
  • typeRaw takes precedence over s.Type when non-empty — lets callers preserve freeform vocab values (e.g. "eats" from datasets that don't follow the controlled-vocab SCREAMING_ SNAKE_CASE convention) instead of coercing them to UnknownSpIntT.ID() = "". Pass "" to fall back to s.Type.ID().
  • col__modified / col__modified_by stamped from tx context.

func (*Tx) AddSpeciesInteractionType

func (t *Tx) AddSpeciesInteractionType(term VocabTermDetail) error

AddSpeciesInteractionType writes a new vocab term. ID and Name are required; the rest are optional. Uniqueness is enforced by sfga's PRIMARY KEY on col__id — a duplicate returns ErrConflict. Invalidates the cached vocab bundle so /api/vocab reflects the new term on next read.

func (*Tx) AddSynonym

func (t *Tx) AddSynonym(s coldp.Synonym) (string, error)

AddSynonym inserts a new synonym link.

  • s.TaxonID and s.NameID are required; empty either → ErrValidation.
  • s.TaxonID == s.ID is refused by the sfga CHECK constraint; hive forbids explicit self-links via ErrValidation up-front.
  • If s.ID is empty a UUID is generated so the row can participate in pro-parte relationships later.
  • Nullable FKs ("" → NULL): source_id, according_to_id, status_id.
  • col__modified / col__modified_by stamped from the transaction context.

Returns the synonym's ID (generated or preserved) so the caller can use it with AddSynonymTaxon for pro-parte.

func (*Tx) AddSynonymTaxon

func (t *Tx) AddSynonymTaxon(synonymID, additionalTaxonID string) error

AddSynonymTaxon extends an existing synonym into a pro-parte relationship by adding another accepted-taxon link with the same synonym ID.

  • synonymID must already exist in the archive; unknown → ErrNotFound.
  • additionalTaxonID must not be the synonym's own ID (schema CHECK).
  • The new link inherits source_id, name_id, name_phrase, according_to_id, status_id, reference_id, and link from the FIRST existing row with the given synonym ID. Curators wanting a different name or status per taxon should use AddSynonym with a fresh ID rather than extending pro-parte.
  • col__modified / col__modified_by on the NEW row are stamped from context.

func (*Tx) AddTypeMaterial

func (t *Tx) AddTypeMaterial(m coldp.TypeMaterial) (int64, error)

AddTypeMaterial inserts a new type_material row and returns the new rowid as the opaque handle for later PATCH/DELETE.

  • t.NameID is required; empty → ErrValidation.
  • Nullable FKs ("" → NULL): source_id, status_id, reference_id, sex_id.
  • Nullable numerics: Latitude, Longitude, Altitude — Valid=false stores NULL, Valid=true stores the given value (0 included).
  • col__id (SpecimenID) passes through verbatim including empty.
  • col__modified / col__modified_by stamped from tx context.

func (*Tx) AddVernacular

func (t *Tx) AddVernacular(v coldp.Vernacular) (int64, error)

AddVernacular inserts a new vernacular row and returns the new rowid as the opaque handle for later PATCH/DELETE.

  • v.TaxonID and v.Name are required; empty either → ErrValidation.
  • Nullable FKs ("" → NULL): source_id, sex_id, reference_id. Empty non-FK strings pass through as empty strings.
  • Preferred is stored verbatim from sql.NullBool — invalid stays NULL, valid stores 0/1.
  • col__modified / col__modified_by stamped from tx context.

func (*Tx) BackfillAuthorshipYears

func (t *Tx) BackfillAuthorshipYears(nameID string) error
  synonyms first — hive never orphans a taxon into a nameless state.
- Cascades auxiliary rows keyed by the name: type_material, name_match,
  and name_relation (both directions — col__name_id and
  col__related_name_id). The sfga schema does not declare ON DELETE
  CASCADE, so hive does the cleanup explicitly.
- Missing name → ErrNotFound.

BackfillAuthorshipYears fills the two atomized year fields on the given name from the citation graph — its own reference and, if linked, its basionym's reference. Runs after write paths so curators who don't expand the atomized-fields section still get years populated when the graph makes them derivable.

Two shapes:

  • Recombination (name has an outgoing BASIONYM relation): combination_year ← this name's reference.issued basionym_year ← the linked basionym name's reference.issued Plausibility gate: basionym_year MUST be strictly less than combination_year. Equal (same paper cited both — likely wrong) or reversed (impossible ordering) → skip entirely; the validation rule surfaces the mismatch for curator review.

  • Original combination (no BASIONYM relation): basionym_year ← this name's own reference.issued There's no separate combination act, so combination_year is left alone. The Standardized-authorship render falls back to basionym_year for originals anyway.

Never overwrites curator-set values. Idempotent — running twice with the same graph is a no-op.

func (*Tx) CascadeDeleteTaxon

func (t *Tx) CascadeDeleteTaxon(id string) error

CascadeDeleteTaxon removes the taxon, every descendant, and every per-taxon association attached to any member of that set. Runs in a single transaction with deferred FK checks so we can delete rows in any order — the self-referential col__parent_id constraint waits until COMMIT, by which point every descendant is gone.

The taxa's associated `name` rows are NOT deleted — names are shared across taxa and often referenced elsewhere. Curators who also want to remove names go through DeleteName after the cascade.

func (*Tx) Context

func (t *Tx) Context() context.Context

Context returns the context this transaction was begun with.

func (*Tx) CopyAgentToRole

func (t *Tx) CopyAgentToRole(
	fromRole Role, id int, toRole Role, blankNote bool,
) (int, error)

CopyAgentToRole creates a new agent in toRole from the source (fromRole, id). Returns the new id.

blankNote controls whether the source agent's col__note carries over. The note field is role-specific in curator practice (e.g. "primary curator" for a creator doesn't survive a copy to publisher), so the WUI/TUI default is to blank it on copy and let the curator author a fresh role-appropriate note.

func (*Tx) CreateAgent

func (t *Tx) CreateAgent(ag Agent) (int, error)

CreateAgent inserts a new row into the role's table. Returns the new col__id assigned by SQLite (AUTOINCREMENT).

  • ag.Role must be a valid role.
  • Per-role NOT NULL fields (see validateAgent) must be non-empty.
  • col__metadata_id defaults to 1 via the schema; hive assumes the singleton metadata convention and does not override it.

func (*Tx) CreateName

func (t *Tx) CreateName(n coldp.Name) (string, error)

CreateName inserts a new name row.

Behavior per CLAUDE.md § gn__* column policy:

  • gnparser is always invoked on the incoming scientific-name string.
  • All gn__* fields are populated from the parse result. Unparseable names still save; gn__parse_quality = 0 is the signal, not an error.
  • Structural col__* fields (uninomial, genus, subgenus, species, infraspecies, cultivar_epithet, authorship, combination_authorship, combination_authorship_year, basionym_authorship, basionym_authorship_year, combination_ex_authorship, basionym_ex_authorship) are auto-filled from the parse result IF the caller didn't supply a value — same "fill gaps, don't overwrite" rule that CoLDP import uses for gn__*. Authorship-in-title now works for any name the parser can atomize (which is nearly all of them); curators who need to override a bad parse can send explicit col__ values on Create.

If n.ID is empty a UUID v4 is generated and returned. Imported string IDs are preserved verbatim.

If both n.ScientificNameString and n.ScientificName are empty, the write is rejected — sfga's schema requires both NOT NULL. gnparser needs a verbatim string to work with; hive uses ScientificNameString as the source of truth and falls back to ScientificName if the caller only supplied the canonical form.

func (*Tx) CreateReference

func (t *Tx) CreateReference(r coldp.Reference) (string, error)

CreateReference inserts a new reference row.

  • If r.ID is empty a UUID v4 is generated and returned; imported string IDs are preserved verbatim (same convention as CreateTaxon / CreateName).
  • col__source_id and col__type_id have FK constraints against content / enum tables; empty values are translated to SQL NULL via nullIfEmpty so PRAGMA foreign_keys stays happy.
  • col__modified is stamped with the current time in RFC3339 UTC. col__modified_by comes from the tx's actor context (WithActor).

func (*Tx) CreateTaxon

func (t *Tx) CreateTaxon(taxon coldp.Taxon) (string, error)

CreateTaxon inserts a new taxon row.

If t.ID is empty a UUID v4 is generated and returned. If t.ID is set (e.g., from an import where the source ID must be preserved verbatim) it is used as-is; sfga treats col__id as an opaque string, so numeric-looking or slug-like values round-trip untouched.

col__modified is stamped with the current time in RFC3339 UTC. col__modified_by is stamped with the actor from the transaction's context (see WithActor).

Denormalized classification columns (col__genus, col__family, sf__genus_id, …) are set to empty regardless of what t contains. They are managed by MoveTaxon-driven reclassification, never by direct writes.

func (*Tx) DeleteAgent

func (t *Tx) DeleteAgent(role Role, id int) error

DeleteAgent removes a role row. Missing row → ErrNotFound.

Role tables have no dependents that FK back to them, so unlike DeleteReference this doesn't need to guard against downstream citations — the row can always be removed cleanly.

func (*Tx) DeleteDistribution

func (t *Tx) DeleteDistribution(rowid int64) error

DeleteDistribution removes the row at rowid. Unknown rowid → ErrNotFound so the handler surfaces a 404.

func (*Tx) DeleteName

func (t *Tx) DeleteName(id string) error

func (*Tx) DeleteReference

func (t *Tx) DeleteReference(id string) error

DeleteReference removes a reference. Refuses if any downstream row still points at it via one of the FK columns in referenceCitationTables — safer than silently NULLing those refs. Curator must reassign the dependents (or delete them) first.

  • Missing reference → ErrNotFound.
  • Any dependent row → ErrConflict with the offending table name.

func (*Tx) DeleteSpeciesInteraction

func (t *Tx) DeleteSpeciesInteraction(rowid int64) error

DeleteSpeciesInteraction removes the row at rowid. Unknown rowid → ErrNotFound so the handler surfaces a 404.

func (*Tx) DeleteSpeciesInteractionType

func (t *Tx) DeleteSpeciesInteractionType(id string) error

DeleteSpeciesInteractionType removes a term. sfga's FK from species_interaction.col__type_id to species_interaction_type.col__id rejects the delete if any species_interaction row still cites the term — returns ErrConflict in that case so the caller can prompt the curator to reassign the affected rows first.

func (*Tx) DeleteTaxon

func (t *Tx) DeleteTaxon(id string) error

DeleteTaxon removes a taxon and its per-taxon associations (synonyms, vernaculars, distributions, media, treatments, species estimates, taxon properties, and both directions of species-interaction and taxon-concept-relation rows).

The taxon's associated `name` row is NOT deleted — names are shared entities across the archive and often referenced by other taxa or synonyms. Curators wanting to delete a name go through DeleteName.

  • If the taxon has children, DeleteTaxon returns ErrConflict without modifying anything. Curators must reparent or delete children first. A bulk DeleteSubtree(id) is a future addition.
  • Missing taxon returns ErrNotFound.

This method exists because the sfga schema does not declare ON DELETE CASCADE on FK columns pointing at taxon; without hive doing the cascade explicitly, the DELETE would fail with a foreign-key error.

func (*Tx) DeleteTypeMaterial

func (t *Tx) DeleteTypeMaterial(rowid int64) error

DeleteTypeMaterial removes the row at rowid. Unknown rowid → ErrNotFound so the handler surfaces a 404 rather than silently no-op'ing.

func (*Tx) DeleteVernacular

func (t *Tx) DeleteVernacular(rowid int64) error

DeleteVernacular removes the row at rowid. Unknown rowid → ErrNotFound so the handler surfaces a 404 rather than silently no-op'ing (the frontend's confirm modal already assumes the row existed).

func (*Tx) LinkNameRelation

func (t *Tx) LinkNameRelation(n coldp.NameRelation) error

LinkNameRelation writes a name_relation row. Fields on n:

  • NameID (required): the subject name.
  • RelatedNameID (required): the object name.
  • Type (required): a NomRelType enum value. Empty ID is rejected — an untyped relation has no meaning.
  • ReferenceID (optional): the reference that describes this nomenclatural act.
  • Page, Remarks, SourceID: optional.

col__modified is stamped with the current time; col__modified_by comes from the tx's actor context (WithActor).

Callers wanting a whole basionym-add flow (create the basionym name + link it as a synonym of the current accepted taxon + write the relation) should do all three inside a single WithTx so partial failure rolls back cleanly.

func (*Tx) MoveAgentToRole

func (t *Tx) MoveAgentToRole(
	fromRole Role, id int, toRole Role, blankNote bool,
) (int, error)

MoveAgentToRole reassigns an agent from one role to another, atomically: creates the row in toRole and deletes it from fromRole in the same transaction. The source id is gone after the call; the returned id is the newly-created target row.

Same blankNote semantics as CopyAgentToRole — the caller decides whether to keep the source's col__note or clear it. Move is typically triggered when a curator realises a person was filed under the wrong role from the start; clearing the note is a sensible default for that flow.

func (*Tx) MoveSynonym

func (t *Tx) MoveSynonym(id, newTaxonID string) error

MoveSynonym reassigns a synonym row to a different accepted taxon in place — synonym.col__id stays stable so any external reference (audit log, undo history) survives the move. Stamps col__modified / col__modified_by from the tx actor.

Refuses (ErrValidation) when newTaxonID is empty; refuses (ErrNotFound) when the synonym row doesn't exist. Doesn't validate that newTaxonID exists in the taxon table — leaves that as an FK concern (sfga does declare the FK, so a bad id surfaces as a DB error). Pro-parte synonyms: MoveSynonym addresses a single row by col__id; sibling rows sharing the col__id aren't touched.

func (*Tx) MoveTaxon

func (t *Tx) MoveTaxon(id, newParentID string) error

MoveTaxon changes the parent of the taxon identified by id.

  • newParentID == "" moves the taxon to root level (stored as SQL NULL).
  • id == newParentID returns ErrValidation.
  • If newParentID is already a descendant of id, the move would create a cycle; the operation returns ErrValidation.
  • Missing target taxon returns ErrNotFound.

The taxon's col__modified and col__modified_by are stamped.

**Denormalized classification is NOT refreshed in v0.** Columns like col__genus/family/…/sf__*_id on the moved taxon and its descendants may become stale relative to the new parent chain. This is a deliberate v0 scoping — a future Reclassify(id) operation will rebuild the cache; callers who care about those columns should invoke it after a move (once it lands) or rely on downstream re-import. The parent_id link is the source of truth; the classification columns are a query-time convenience. See CLAUDE.md § pkg/ package conventions.

func (*Tx) RemoveSynonym

func (t *Tx) RemoveSynonym(taxonID, nameID string) error

RemoveSynonym deletes the synonym link between taxonID and nameID.

For a pro-parte synonym, this removes only the specified accepted-taxon link; other links with the same synonym ID are preserved. To delete the entire pro-parte structure, iterate SynonymPartners and call RemoveSynonym for each pair, or use RemoveSynonymByID for single-row deletes.

func (*Tx) RemoveSynonymByID

func (t *Tx) RemoveSynonymByID(id string) (taxonID, nameID string, err error)

RemoveSynonymByID deletes a single synonym row addressed by its col__id (opaque string per CLAUDE.md's ID rule). Returns the deleted row's taxon_id + name_id so callers doing conditional cascade work (e.g. "if the name is now bare, also DeleteName") don't need a second query to reconstruct the pair.

Pro-parte synonyms are represented as multiple rows sharing the same col__id; RemoveSynonymByID removes only the specific row (by PK on rowid isn't available at this layer — see the query below). The WUI's per-row trash affordance addresses one synonym row at a time; bulk pro-parte teardown uses SynonymPartners + RemoveSynonym.

func (*Tx) ReparentAndDeleteTaxon

func (t *Tx) ReparentAndDeleteTaxon(id string) error

ReparentAndDeleteTaxon moves the taxon's direct children up one level (to its parent — empty/NULL when the taxon is a root, so children become new roots) and then deletes the taxon as a leaf. Flat one-level reparent — grandchildren stay under their parents, which stay under the newly-promoted children.

Curators who want the deeper "delete this subtree entirely" flow use CascadeDeleteTaxon.

func (*Tx) SQL

func (t *Tx) SQL() *sql.Tx

SQL returns the underlying *sql.Tx for aggregate-specific mutation code in pkg/taxon.go, pkg/name.go, etc. Not intended for callers outside pkg/.

func (*Tx) SetNameStatus

func (t *Tx) SetNameStatus(nameID, status string) error

DeleteName removes a name and its auxiliary rows.

  • Refuses with ErrConflict if any taxon or synonym still references the name via col__name_id. Curators must delete or reassign the taxa /

SetNameStatus writes col__status_id directly as a raw string, bypassing sflib's coldp.NomStatus enum. Necessary because the enum silently maps NOMEN URIs to UnknownNomStatus (which ID()'s back to "") — any UpdateName that went through the enum would blank a URI-shaped status on every edit.

Empty status writes NULL. Non-empty must exist in nom_status (SeedNomenIntoNomStatus takes care of the NOMEN URIs on Open). col__modified / col__modified_by are stamped so audit trails work for status-only changes too.

func (*Tx) UnlinkNameRelation

func (t *Tx) UnlinkNameRelation(rowid int64) error

UnlinkNameRelation removes the row at rowid. Unknown rowid → ErrNotFound so the handler surfaces a 404 rather than silently no-op'ing. Mirrors the vernacular / distribution delete shape; callers wanting to "change the type of an existing relation" do unlink + Link (the row's identity IS its type + endpoints, so updating any of those is semantically a new row).

func (*Tx) UpdateAgent

func (t *Tx) UpdateAgent(ag Agent) error

UpdateAgent overwrites the row at (ag.Role, ag.ID) with ag's field values. Missing row → ErrNotFound.

func (*Tx) UpdateDistribution

func (t *Tx) UpdateDistribution(rowid int64, d coldp.Distribution) error

UpdateDistribution rewrites every editable column of the row at rowid. TaxonID is NOT editable — reparenting a distribution is semantically a delete+add on a different taxon. Zero-row affected (unknown rowid) surfaces as ErrNotFound so the handler maps to 404.

func (*Tx) UpdateMetadata

func (t *Tx) UpdateMetadata(m Metadata) error

UpdateMetadata writes a metadata row. If m.ID is 0 the row is either inserted (when no rows exist) or the existing sole row's fields are replaced (when one exists) — callers using the singleton convention don't need to know the id.

Title is required by the sfga schema (NOT NULL); an empty title returns ErrValidation.

func (*Tx) UpdateName

func (t *Tx) UpdateName(n coldp.Name) error

UpdateName writes a coldp.Name back to an existing row keyed by n.ID.

Semantics:

  • n.ID is required; empty → ErrValidation.
  • Missing row → ErrNotFound.
  • Optimistic concurrency: if n.Modified is non-empty, it must match the row's current col__modified or ErrConflict is returned.
  • gnparser always re-runs on the incoming verbatim string. All gn__* fields are refreshed from the parse result, regardless of what the caller supplied. Parsing is microseconds; keeping the cache invariant tight is more valuable than saving those microseconds.
  • Symmetric fallback between ScientificName and ScientificNameString (same rule as CreateName).
  • col__modified / col__modified_by stamped from context.

func (*Tx) UpdateReference

func (t *Tx) UpdateReference(r coldp.Reference) error

UpdateReference writes r back to the existing row keyed by r.ID.

  • r.ID required (ErrValidation on empty).
  • Missing row → ErrNotFound.
  • Optimistic concurrency: non-empty r.Modified is compared against the current col__modified; mismatch → ErrConflict. Empty Modified skips the check (blind write).
  • col__modified / col__modified_by are re-stamped from the tx.

func (*Tx) UpdateSpeciesInteraction

func (t *Tx) UpdateSpeciesInteraction(rowid int64, s coldp.SpeciesInteraction, typeRaw string) error

UpdateSpeciesInteraction rewrites every editable column of the row at rowid. TaxonID is NOT editable — reparenting is semantically a delete+add on a different taxon. Zero-row affected (unknown rowid) surfaces as ErrNotFound so the handler maps to 404. typeRaw takes precedence over s.Type when non-empty (see AddSpeciesInteraction for the rationale).

func (*Tx) UpdateSpeciesInteractionType

func (t *Tx) UpdateSpeciesInteractionType(term VocabTermDetail) error

UpdateSpeciesInteractionType rewrites every editable column of the term at term.ID. ID itself is not editable — deleting + re-adding is the way to rename an id (also rare — the id is what other rows reference by FK).

func (*Tx) UpdateSynonym

func (t *Tx) UpdateSynonym(s coldp.Synonym) error

UpdateSynonym updates a synonym link identified by (s.TaxonID, s.NameID).

Optimistic concurrency: if s.Modified is non-empty it is treated as an If-Match token vs. col__modified; mismatch → ErrConflict.

Editable fields: NamePhrase, AccordingToID, Status, ReferenceID, SourceID, Link, Remarks. The natural-key columns (TaxonID, NameID) and structural ID are NOT changed here — a curator switching the target taxon should remove this synonym and add a fresh one, keeping the change auditable.

func (*Tx) UpdateTaxon

func (t *Tx) UpdateTaxon(taxon coldp.Taxon) error

UpdateTaxon writes a coldp.Taxon back to an existing row, keyed by taxon.ID.

Semantics:

  • taxon.ID is required; empty returns ErrValidation.
  • Missing row returns ErrNotFound.
  • Optimistic concurrency: if taxon.Modified is non-empty, it is treated as an If-Match token and compared against the current col__modified. Mismatch returns ErrConflict. Empty Modified skips the check (blind write — the HTTP layer enforces the header requirement per request).
  • col__parent_id is NOT updated here. Reparenting goes through MoveTaxon (an explicit verb with cycle checks). Any ParentID in the input is silently ignored — the field is present on coldp.Taxon for round-trip purposes but hive owns the tree structure.
  • Denormalized classification columns (col__genus/family/…/sf__*_id) and read-only cache fields (col__branch_length) are also silently ignored; they are managed by the tree-structure code path.

col__modified is stamped with the current time in RFC3339 UTC. col__modified_by with the actor from the transaction's context.

func (*Tx) UpdateTypeMaterial

func (t *Tx) UpdateTypeMaterial(rowid int64, m coldp.TypeMaterial) error

UpdateTypeMaterial rewrites every editable column of the row at rowid. NameID is NOT editable — reparenting a type_material row is semantically a delete+add on a different name. Zero rows affected (unknown rowid) surfaces as ErrNotFound so the handler maps to 404.

func (*Tx) UpdateVernacular

func (t *Tx) UpdateVernacular(rowid int64, v coldp.Vernacular) error

UpdateVernacular rewrites every editable column of the row at rowid. TaxonID is NOT editable — reparenting a vernacular is semantically a delete+add on a different taxon. Zero-row affected (unknown rowid) surfaces as ErrNotFound so the handler maps to 404.

The full-row rewrite is intentional: sfga vernaculars are small and the whole-row PATCH matches how the frontend edit-form submits (all fields at once). Callers wanting field-level PATCH semantics load-then-modify-then-update.

type TypeMaterialHit

type TypeMaterialHit struct {
	RowID               int64
	SpecimenID          string
	NameID              string
	SourceID            string
	Citation            string
	Status              coldp.TypeStatus
	InstitutionCode     string
	CatalogNumber       string
	ReferenceID         string
	Locality            string
	Country             string
	Latitude            sql.NullFloat64
	Longitude           sql.NullFloat64
	Altitude            sql.NullInt64
	Host                string
	Sex                 coldp.Sex
	Date                string
	Collector           string
	AssociatedSequences string
	Link                string
	Remarks             string
	Modified            string
	ModifiedBy          string
	// IssueCount is the number of open validation issues currently
	// filed against this type_material row. Populated per-hit by
	// ListTypeMaterials in a batched grouping query. Zero when
	// clean; front-ends render a warn icon when > 0.
	IssueCount int
	// MaxSeverity is the highest severity ("error" > "warn" > "info"
	// > "debug") among the row's open issues. Empty string when
	// IssueCount is 0. Drives the badge color the WUI paints.
	MaxSeverity string
}

TypeMaterialHit is the read projection returned by ListTypeMaterials and GetTypeMaterial. RowID carries the SQLite rowid handle; SpecimenID is the optional col__id value the curator may fill (e.g. the collection / catalogue identifier).

type ValidationWarning

type ValidationWarning struct {
	RuleID    string
	RuleName  string
	FieldName string
	Severity  string // "warn" | "info" | "debug"
	Message   string
}

ValidationWarning is the frontend-facing shape of a non-blocking validation result. Both TUI and WUI consume this so they can render warnings the same way without touching gsvalidator's domain package.

type VernacularHit

type VernacularHit struct {
	RowID           int64
	TaxonID         string
	SourceID        string
	Name            string
	Transliteration string
	Language        string
	Preferred       sql.NullBool
	Country         string
	Area            string
	Sex             coldp.Sex
	ReferenceID     string
	Remarks         string
	Modified        string
	ModifiedBy      string
	// IssueCount is the number of open (not-yet-acknowledged)
	// validation issues currently filed against this vernacular row.
	// Populated per-hit by ListVernaculars in a single batched
	// grouping query — cheap even with a large per-taxon vernacular
	// set. Zero when the row is clean; front-ends render a warn icon
	// on the row when > 0.
	IssueCount int
	// MaxSeverity is the highest severity ("error" > "warn" > "info"
	// > "debug") among the row's open issues. Empty string when
	// IssueCount is 0. Drives the badge color on the WUI row.
	MaxSeverity string
}

VernacularHit is the projection returned by ListVernaculars and the read side of the vernacular CRUD. RowID carries SQLite's implicit primary key so update/delete calls can address a single row unambiguously even when multiple rows share (taxon, name).

type VocabTerm

type VocabTerm struct {
	ID   string `json:"id"`
	Name string `json:"name,omitempty"`
}

VocabTerm is one entry in a controlled vocabulary. `Name` is the human-facing label. For flat vocabularies (nom_code, gender, sex, …) the schema stores only the ID; hive derives a display Name by lower-casing and replacing underscores with spaces (matching the convention coldp.ToStr already uses).

The empty-string term (id == "") is always the first entry in every vocabulary — the schema seeds it as the "unspecified" value — and frontends can present it as "(unset)".

type VocabTermDetail

type VocabTermDetail struct {
	ID          string
	Name        string
	Description string
	// Obo is an ontology URI (e.g. Relations Ontology PURL
	// "http://purl.obolibrary.org/obo/RO_0002442"). Named for sfga's
	// col__obo column even though the value is any URI, not
	// necessarily an OBO one — sfga's column naming is legacy.
	Obo         string
	Inverse     string
	Symmetrical bool
	// SuperTypes is a comma-separated list of parent term ids from
	// sfga's col__superTypes. Kept as a raw string on the wire; the
	// UI splits / rejoins.
	SuperTypes string
}

VocabTermDetail is the full-fidelity term shape for the vocab editor — the picker bundle at /api/vocab ships only id + name so the round trip stays small; the editor's list + form need the remaining columns (description, ontology URI via obo, inverse, symmetrical, superTypes) to display and edit rich vocabs like species_interaction_type.

Not every rich vocab populates every field. Callers should treat empty strings and false as "unset" — the display path elides them.

type Vocabulary

type Vocabulary struct {
	NomCode                []VocabTerm `json:"nom_code"`
	NomStatus              []VocabTerm `json:"nom_status"`
	TaxonomicStatus        []VocabTerm `json:"taxonomic_status"`
	Rank                   []VocabTerm `json:"rank"`
	Gender                 []VocabTerm `json:"gender"`
	Sex                    []VocabTerm `json:"sex"`
	DistributionStatus     []VocabTerm `json:"distribution_status"`
	NomRelType             []VocabTerm `json:"nom_rel_type"`
	EstimateType           []VocabTerm `json:"estimate_type"`
	NamePart               []VocabTerm `json:"name_part"`
	MatchType              []VocabTerm `json:"match_type"`
	ReferenceType          []VocabTerm `json:"reference_type"`
	TypeStatus             []VocabTerm `json:"type_status"`
	Gazetteer              []VocabTerm `json:"gazetteer"`
	GeoTime                []VocabTerm `json:"geo_time"`
	SpeciesInteractionType []VocabTerm `json:"species_interaction_type"`
	TaxonConceptRelType    []VocabTerm `json:"taxon_concept_rel_type"`
	Licenses               []VocabTerm `json:"licenses"`
}

Vocabulary is the JSON-serializable bundle GET /api/vocab returns. One field per sfga enum table; more can be added without breaking the wire contract since JSON is additive.

Licenses is a hive-curated suggestion list — sfga's col__license is a free TEXT column with no FK, so this vocab is display-only. The WUI/TUI use it to populate a suggestions dropdown; curators can still type a value that isn't in the list.

Directories

Path Synopsis
Package bhlnames is an HTTP client for the BHLnames service (bhlnames.globalnames.org/api/v1).
Package bhlnames is an HTTP client for the BHLnames service (bhlnames.globalnames.org/api/v1).
Package bibtex converts a single BibTeX entry into a coldp.Reference.
Package bibtex converts a single BibTeX entry into a coldp.Reference.
Package config owns hive's user-config file: a small YAML document at $XDG_CONFIG_HOME/sfborg/hive/config.yml (or the platform equivalent), plus the merged flag > env > config Identity that downstream packages read via CurrentIdentity.
Package config owns hive's user-config file: a small YAML document at $XDG_CONFIG_HOME/sfborg/hive/config.yml (or the platform equivalent), plus the merged flag > env > config Identity that downstream packages read via CurrentIdentity.
Package openalex is a thin HTTP client for api.openalex.org that resolves DOIs and runs cross-field search, mapping the returned Work records into coldp.Reference values.
Package openalex is a thin HTTP client for api.openalex.org that resolves DOIs and runs cross-field search, mapping the returned Work records into coldp.Reference values.
Package orcid is a Go client for the ORCID public API (https://pub.orcid.org/v3.0).
Package orcid is a Go client for the ORCID public API (https://pub.orcid.org/v3.0).
config
Package config holds the Config and its functional Option setters used to construct an orcid Client.
Package config holds the Config and its functional Option setters used to construct an orcid Client.
Package sfgarules holds the SFGAMapper — hive's implementation of gsvalidator's SchemaMapper and joins.PrimaryKeyProvider interfaces for the sfga schema (col__ / gn__ / sf__ prefixes, col__id as the primary key across every table).
Package sfgarules holds the SFGAMapper — hive's implementation of gsvalidator's SchemaMapper and joins.PrimaryKeyProvider interfaces for the sfga schema (col__ / gn__ / sf__ prefixes, col__id as the primary key across every table).
Package ui holds hive's shared UI descriptions — data that both the TUI and the WUI render from, so a single edit propagates to both frontends.
Package ui holds hive's shared UI descriptions — data that both the TUI and the WUI render from, so a single edit propagates to both frontends.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL