harnesses

package
v0.3.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 18, 2026 License: GPL-3.0 Imports: 16 Imported by: 0

Documentation

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type ClaudeCode

type ClaudeCode struct{}

func (*ClaudeCode) Run

func (c *ClaudeCode) Run(ctx context.Context, conversation *agent.Conversation, prompt string) (*RunResult, error)

func (*ClaudeCode) Validate

func (c *ClaudeCode) Validate(ctx context.Context, model string, config json.RawMessage) error

type CodexCLI

type CodexCLI struct{}

func (*CodexCLI) Run

func (c *CodexCLI) Run(ctx context.Context, conversation *agent.Conversation, prompt string) (*RunResult, error)

func (*CodexCLI) Validate

func (c *CodexCLI) Validate(ctx context.Context, model string, config json.RawMessage) error

type Harness

type Harness interface {
	Validate(ctx context.Context, model string, config json.RawMessage) error
	Run(ctx context.Context, conversation *agent.Conversation, prompt string) (*RunResult, error)
}

func New

func New(kind agent.Harness) (Harness, error)

type HarnessInfo

type HarnessInfo struct {
	ID        agent.Harness `json:"id"`
	Binary    string        `json:"binary"`
	Available bool          `json:"available"`
	Models    []string      `json:"models"`
}

HarnessInfo describes one harness: whether its binary is installed and which models it can run.

func ListHarnessInfo

func ListHarnessInfo(ctx context.Context) []HarnessInfo

ListHarnessInfo reports every known harness with availability and models. Results are cached briefly — pi's catalog only changes on `pi update`.

type Permissions

type Permissions string

Permissions is the harness-agnostic access tier declared in the DSL under config.harness.permissions. It is the single knob that controls what an agent may do to the working tree and the host; both the Claude Code and Codex harnesses translate the same tier into their own backend flags via the resolvers below, so the mapping for both lives in one reviewable place.

const (
	// PermissionsReadOnly is the default. The agent may read and inspect the
	// workspace but may not modify it.
	//
	// Enforcement differs by backend and is intentionally documented here:
	//   - Claude Code has no filesystem sandbox, so we gate by tool: the write
	//     and shell tools are hard-disallowed, leaving only read tools. That
	//     means a read_only Claude agent cannot run shell commands at all.
	//   - Codex read-only sandbox blocks filesystem writes and network but still
	//     lets commands run, so a read_only Codex agent CAN run read-only probes
	//     (grep, and builds/tests that don't need to write). Both honor the core
	//     guarantee "cannot modify the workspace or reach the network".
	PermissionsReadOnly Permissions = "read_only"

	// PermissionsExec lets the agent modify the workspace and run shell commands
	// scoped to it (tests, builds, formatters).
	//
	// Codex enforces this as a real sandbox: workspace writes and shell, no
	// network, no writes outside the workspace. Claude Code has no network
	// sandbox, so once shell is allowed it can also reach the network — on
	// Claude, exec and dangerously-exec are therefore nearly identical in
	// practice. The tier means "as locked-down as this backend can express".
	PermissionsExec Permissions = "exec"

	// PermissionsDangerouslyExec removes all guardrails: network access, writes
	// anywhere, and no approvals. Use only for trusted, disposable environments.
	PermissionsDangerouslyExec Permissions = "dangerously-exec"
)

func ParsePermissions

func ParsePermissions(raw string) (Permissions, error)

ParsePermissions normalizes and validates a raw permissions string. An empty value resolves to the safe default (read_only).

type PiCLI

type PiCLI struct{}

PiCLI runs conversations through the pi coding agent (`pi -p --mode json`), which emits a JSONL event stream.

func (*PiCLI) Run

func (c *PiCLI) Run(ctx context.Context, conversation *agent.Conversation, prompt string) (*RunResult, error)

func (*PiCLI) Validate

func (c *PiCLI) Validate(ctx context.Context, model string, config json.RawMessage) error

type RunResult

type RunResult struct {
	LastAssistantMessage string
	SessionRef           string
	State                json.RawMessage
	RawOutput            string
	ExitCode             int
	HarnessError         string
	InputTokens          int64
	OutputTokens         int64
	CachedTokens         int64
}

type TraceEvent

type TraceEvent struct {
	Kind    string `json:"kind"` // reasoning | message | command | tool | error
	Content string `json:"content"`
	Detail  string `json:"detail,omitempty"`
}

TraceEvent is one step of a harness run, parsed from the raw output stream: the model's reasoning summaries, executed commands, emitted messages, and errors.

func ParseTrace

func ParseTrace(harness agent.Harness, rawOutput string) []TraceEvent

ParseTrace extracts trace events from a conversation's raw harness output. Codex emits a JSONL event stream (`codex exec --json`). Claude Code currently runs with a single JSON result — no per-step events are recorded, so its trace is empty.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL