chats

package
v0.7.3 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 21, 2026 License: MIT Imports: 29 Imported by: 0

README

Chat orchestration

internal/pkg/core/chats owns durable turns: model/fixed-response selection, prompt construction, MCP tool execution, SSE streaming, and persistence. HTTP handlers validate transport input and delegate here.

Contents

Turn contract

For a real turn, the order is intentional:

  1. Resolve the model and acquire a per-chat lock.
  2. Build the system message, bounded history, and current user message, the exact list sent to the upstream.
  3. Save the user message as an incomplete durable row before opening SSE.
  4. Checkpoint useful assistant text/reasoning at most once per interval under the same turn ID. Checkpoints contain no partial tool call/result data.
  5. Stream ephemeral chat_status progress frames only after the matching provider/tool transition, then stream reasoning, text, tool-use, and tool-result blocks in arrival order.
  6. On success, atomically remove the checkpoint, save assistant/tool rows as one completed turn, and record usage. Provider-native reasoning is stored as opaque JSON and restored byte-for-byte for providers that require signed thinking blocks to be replayed.

Visible history includes completed user, assistant, and tool rows plus an incomplete user row and a useful incomplete assistant checkpoint. The browser labels that checkpoint as an interrupted response. Incomplete tool rows stay hidden. Incomplete rows are excluded from loadHistory, so neither a submitted prompt nor a partial assistant answer can be replayed into a later model turn. Do not reorder these writes.

The frontend store at web/src/lib/stores/conversation.svelte.ts adopts a new chat id from message_start without clobbering a live stream with stale history. It treats the closed chat_status enum as UI-only metadata: connecting, waiting_first_token, streaming, running_tool, and retrying are never persisted as assistant content and are cleared when the turn ends, fails, or is stopped. The browser measures elapsed time locally while a status is active. On a terminal provider failure it can retry the retained prompt and model as a new durable turn; cancellation deliberately offers no retry.

History and settings

The chat service supplies the system message, validated stored history, and current user prompt to elelem. The engine applies the outbound shape:

sticky system message → newest fitting history suffix → current user message

The system message includes the base instruction and generated GenUI guide. It is invisible in history but always counted; the current user message is pinned too. The default cap is 100,000 tokens. Elelem's default o200k_base counter counts portable message/tool content and framing. It drops the oldest complete unit until the transcript fits. An assistant tool call and its contiguous results are one unit, so the upstream never gets an orphaned result. A live unresolved exchange stays pinned, while a completed older exchange is droppable.

The composer controls temperature, topP, reasoningEffort, maxOutputTokens, and maxHistoryTokens. Updating settings fully replaces the object; omitted generation settings use provider defaults. maxHistoryTokens is a prompt budget, not an upstream generation parameter. See api/api.yml and web/src/lib/components/ChatSettings.svelte.

Before a user sends, POST /chats/{chatId}/context-preview runs this same loading, tokenization, and complete-unit selection path against the unsent draft. The composer renders that response rather than estimating locally, so its system/history/draft counts and omitted-turn indicator describe the prompt the next stream will actually receive. The endpoint is ownership-checked and does not persist the draft.

MCP, demos, and showcase

MCP tools are namespaced as <server>__<tool>. A chat can disable a globally enabled server in its composer, hiding that server's tools from the model. This is a preference, not a security boundary. See ../../mcp/README.md for transport and credential rules.

Trigger Live model/MCP Durable Purpose
Exact showcase prompt No for the match Yes Paced dashboard turn with deterministic thinking + synthetic tool cards.

CHATZ_SHOWCASE_MODE=true (make run-showcase) only intercepts exact catalog prompts. It never hides models, changes MCP setup, or captures near matches.

Debugging and verification

history_logging.go receives each round's exact post-limiting transcript from elelem immediately before the driver call. At LOG_LEVEL=debug it records model, budget, token count, ordered messages, reasoning, and tool arguments. RedactText masks secret-shaped values, but normal chat content can remain; keep debug local.

Path Responsibility
chats.go / stream.go Turn entry points, SSE production, persistence.
history.go / history_logging.go Stored-history validation and per-round debug trace.
turn_locks.go / mcp_tools.go Serialization and tool filtering.
prompts/ / fixedresponses/ GenUI guide and demo/showcase fixtures.

Run make test, make test-web, and make test-api from the repository root. The showcase render + streamed thinking/tool-cards/GenUI is covered by the Go browser drivers in the api tier tests/api/ (showcase_test.go, smoke_test.go).

Documentation

Overview

Package chats is the chat business-logic layer: it owns the chat + message repositories and the model registry, and drives a turn end-to-end (demo interception, model resolution, the agentic turn loop, SSE streaming, and persistence). HTTP handlers stay dumb — they parse the request, call a method here, and map the result (or a sentinel error) to the API response.

Index

Constants

View Source
const (
	MaxChatSearchRunes = 256
)

Variables

View Source
var (
	// ErrUnknownModel is returned when a real (non-demo) turn names a model no
	// configured upstream serves. Callers map it to a 400.
	ErrUnknownModel = errors.New("chats: unknown model")

	// ErrEmptyTitle is returned when a rename is given a blank title. Callers
	// map it to a 400.
	ErrEmptyTitle = errors.New("chats: title must not be empty")
)

Functions

func ChatMCPServerToAPI

func ChatMCPServerToAPI(sv ChatMCPServer) api.ChatMCPServer

ChatMCPServerToAPI maps a chat's view of an MCP server to the wire shape.

func ChatSearch

func ChatSearch(search string) string

ChatSearch normalizes a title search before passing it into the generated parameterized query. A blank return means no search filter.

func ChatSettingsToAPI

func ChatSettingsToAPI(c *models.Chat) *api.ChatSettings

ChatSettingsToAPI projects a chat's stored generation settings to the API shape. Unset numeric settings stay nil; an empty ReasoningEffort maps to a nil pointer (the wire omits it) rather than an invalid empty enum value.

func ChatSummaryToAPI

func ChatSummaryToAPI(c *models.Chat) api.ChatSummary

ChatSummaryToAPI projects a stored chat row to the wire summary shape.

func MessageToAPI

func MessageToAPI(m *models.Message) (api.Message, error)

MessageToAPI projects a stored message row to the wire shape. Thinking, toolCallId, and isError are omitted (nil) when the row carries none of them, so a plain user/text-only row's JSON is unchanged from before these fields existed. ToolCalls is unmarshalled from the row's stored JSON; a decode failure returns an error rather than silently dropping the tool calls a reload is specifically trying to reconstruct.

func PromptContextToAPI

func PromptContextToAPI(context PromptContext) api.PromptContextPreview

PromptContextToAPI projects the backend-authoritative prompt selection to the compact composer meter shape.

Types

type ChatMCPServer

type ChatMCPServer struct {
	ID      uuid.UUID
	Name    string
	Status  string
	Enabled bool
}

ChatMCPServer is one MCP server as a chat sees it: identity, live global connection status, and whether its tools are enabled for this chat. Per-chat enablement is independent of the server's global enabled flag.

type ListOptions

type ListOptions struct {
	Search string
}

ListOptions narrows one caller's chat history without changing ownership. A blank Search matches every chat.

type PromptContext

type PromptContext struct {
	BudgetTokens         int
	SystemTokens         int
	HistoryTokens        int
	CurrentMessageTokens int
	TotalTokens          int
	OmittedMessages      int
	OmittedTurns         int
	RetainedMessages     int
	RetainedTurns        int

	History []elelem.Message
}

PromptContext describes the exact history selection Chatz made before an outbound model request. Component counts use the same tiktoken codec as the total; TotalTokens is counted over the fully assembled prompt representation.

type Service

type Service struct {
	// contains filtered or unexported fields
}

Service coordinates the chat repositories, the model upstreams, and the MCP tool manager to fulfill the chat API operations.

func New

func New(
	query *repositories.Query,
	models *upstreams.Registry,
	mcpMgr *mcp.Manager,
	showcaseMode bool,
) *Service

New builds the chat service over the given repositories, model registry, and MCP manager. Each outbound prompt has a 100000-token default cap; a chat may override it with max_history_tokens (see UpdateSettings).

func (*Service) Continue

func (s *Service) Continue(
	ctx context.Context,
	chatID, userID uuid.UUID,
	message string,
	requestedModel *string,
) (io.Reader, error)

Continue continues an existing chat and returns a reader that streams the next assistant turn (SSE). requestedModel, when non-empty, switches the chat to that model (persisted so a reopened chat comes back to its last-used model); demo turns bypass the LLM and never change the chat's model. A chat with no title yet (its first real message — the composer creates the chat via GetOrCreateEmpty before the user types anything, so this is the only place a title gets set) is titled from this message, mirroring what the standalone Create() does for its one-shot create+message flow. Returns commerr.ErrNotFound if the chat is missing, or ErrUnknownModel for an unservable real-turn model.

func (*Service) Create

func (s *Service) Create(
	ctx context.Context,
	userID uuid.UUID,
	message, model string,
) (io.Reader, error)

Create creates a chat for the user and returns a reader that streams the first assistant turn (SSE). A demo command works with any model id; a real turn returns ErrUnknownModel when no configured upstream serves model.

func (*Service) Delete

func (s *Service) Delete(
	ctx context.Context,
	chatID, userID uuid.UUID,
) error

Delete removes a caller-owned chat from normal reads through the existing soft-delete model. The durable message rows remain outside normal history.

func (*Service) Get

func (s *Service) Get(
	ctx context.Context,
	chatID, userID uuid.UUID,
) (*models.Chat, error)

Get returns a chat's metadata owned by userID (no messages — fetch those via ListMessages). Returns commerr.ErrNotFound when the chat does not exist or belongs to another user.

func (*Service) GetOrCreateEmpty

func (s *Service) GetOrCreateEmpty(
	ctx context.Context,
	userID uuid.UUID,
) (*models.Chat, error)

GetOrCreateEmpty returns the caller's reusable "empty chat" — the chat a fresh "New chat" click lands on — creating one if none exists. At most one empty chat is ever kept per user: repeated "New chat" clicks before typing anything all resolve to the SAME chat instead of piling up unused rows.

A benign race exists: two concurrent calls (a double-click, two tabs) can both see zero existing empty chats and each create one. This is accepted rather than transaction-guarded — the worst case is a second, harmless empty row that gets reused (or ages out unused) rather than any data loss.

func (*Service) List

func (s *Service) List(
	ctx context.Context,
	userID uuid.UUID,
	limit, offset int,
) ([]*models.Chat, int, error)

List returns a page of the user's non-empty chats (at least one user message), newest activity first, plus the scoped total count. A brand-new chat with no user message yet — see GetOrCreateEmpty — is hidden from history until the user actually sends something in it.

func (*Service) ListMCPServers

func (s *Service) ListMCPServers(
	ctx context.Context,
	chatID, userID uuid.UUID,
) ([]ChatMCPServer, error)

ListMCPServers returns every MCP server with its live global status and whether it is enabled for this chat (ownership-checked). Returns commerr.ErrNotFound when the chat is missing or owned by another user.

func (*Service) ListMessages

func (s *Service) ListMessages(
	ctx context.Context,
	chatID, userID uuid.UUID,
	limit, offset int,
) ([]*models.Message, int, error)

ListMessages returns a page of a chat's visible messages, oldest-first, plus the scoped total count (indexed by chat_id). A row is visible if it carries anything to render — content, a reasoning trace, or tool calls; a row with none of those is purely internal plumbing (e.g. a completed round whose only output was consumed elsewhere) and is omitted from both the page and the total. Rows are also restricted to the three durable roles (user/assistant/tool) — a role="system" tool-steering injection should never have been persisted (persistTurn skips it going forward), but this filters out any that slipped in before that fix rather than rendering as a stray fake user bubble. An incomplete user row is visible so a refresh after stopping a stream retains the submitted message; an incomplete assistant checkpoint is also visible, while incomplete tool rows remain hidden. Returns commerr.ErrNotFound when the chat does not exist or belongs to another user.

func (*Service) ListWithOptions

func (s *Service) ListWithOptions(
	ctx context.Context,
	userID uuid.UUID,
	options ListOptions,
	limit, offset int,
) ([]*models.Chat, int, error)

ListWithOptions returns the requested page of the user's non-empty chats. A blank search has no filter.

func (*Service) PreviewContext

func (s *Service) PreviewContext(
	ctx context.Context,
	chatID, userID uuid.UUID,
	message string,
) (PromptContext, error)

PreviewContext returns the exact prompt selection a turn would use for the caller's unsent message. It shares the production history loader, sticky system prompt, tiktoken counting, and complete-tool-turn selection logic.

func (*Service) Rename

func (s *Service) Rename(
	ctx context.Context,
	chatID, userID uuid.UUID,
	title string,
) (*models.Chat, error)

Rename sets a chat's title (trimmed + capped to maxTitleRunes) and returns the updated chat. Returns ErrEmptyTitle for a blank title and commerr.ErrNotFound when the chat is missing or owned by another user.

func (*Service) SetMCPServerEnabled

func (s *Service) SetMCPServerEnabled(
	ctx context.Context,
	chatID, userID, serverID uuid.UUID,
	enabled bool,
) (*ChatMCPServer, error)

SetMCPServerEnabled enables or disables one MCP server's tools for a chat (ownership-checked). Returns commerr.ErrNotFound when the chat or the server does not exist.

func (*Service) UpdateSettings

func (s *Service) UpdateSettings(
	ctx context.Context,
	chatID, userID uuid.UUID,
	settings Settings,
) (*models.Chat, error)

UpdateSettings replaces the chat's generation settings and returns the updated chat. Full replacement: a field left unset clears that setting. Returns commerr.ErrNotFound when the chat is missing or owned by another user.

type Settings

type Settings struct {
	Temperature      *float64
	TopP             *float64
	ReasoningEffort  string
	MaxOutputTokens  *int
	MaxHistoryTokens *int
}

Settings is the mutable per-chat generation config. A nil pointer / empty ReasoningEffort leaves that setting unset (provider default / no cap).

Directories

Path Synopsis
Package fixedresponses maps exact, recording-ready chat prompts to canned assistant turns embedded at build time.
Package fixedresponses maps exact, recording-ready chat prompts to canned assistant turns embedded at build time.
Package prompts holds generated, embedded LLM system-prompt fragments for the chat turn loop.
Package prompts holds generated, embedded LLM system-prompt fragments for the chat turn loop.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL