prompt-diff πβ‘
A fast, git-native CLI & local Web UI to version, diff tokens/costs, and benchmark LLM system prompts across local and cloud models.

π‘ Why prompt-diff?
Prompt engineering has shifted from basic trial-and-error to core software engineering. However, standard software tooling treats system prompts as plain text strings, ignoring critical factors like token consumption, cost per 1k runs, variable binding, and behavioral model drift.
prompt-diff treats system prompts as first-class source code. It bridges git version control with real-time token/cost diffing and automated multi-model test harness evaluations.
β¨ Key Features
- βοΈ Git-Native Prompt Tracking: Automatically parses
.prompt templates, identifies template variable injections, and tracks structural evolution across commits.
- π Semantic & Token Diffing: Visualizes added/deleted tokens in your terminal and calculates immediate cost impact projections per 1,000 invocations.
- π§ͺ Batch Evaluation Harness: Runs test matrices across local (Ollama), OpenAI, Anthropic, and Gemini models, with
contains/list/$in/$gte/$lte/$regex/$schema/$llm_judge assertion operators.
- π₯οΈ Zero-Config Local Workspace: Spin up an embedded local web dashboard (
prompt-diff ui) to tweak variables and compare side-by-side model outputs in real time. Running the binary with no arguments (or double-clicking it on Windows) launches the dashboard directly.
- π¦ Single-Binary Zero Setup: Compiles into a single self-contained binary with zero external runtime dependencies.
π οΈ Architecture Overview
| Component |
Stack |
Purpose |
| CLI Engine |
Go / Cobra |
Lightning-fast execution, sub-millisecond startup, native git tree parsing |
| Dashboard UI |
Svelte + Tailwind (Vite) |
Compiled to static assets and embedded into the binary via embed.FS for zero-install local visualization |
| Data Layer |
SQLite (modernc.org/sqlite) |
Pure-Go zero-CGO embedded store for local evaluations and runs |
| Tokenizers |
Embedded cl100k_base + o200k_base BPE (tiktoken vocabs), auto-selected per model |
Pure-Go byte-level BPE token counting without network calls |
π Quick Start
Installation
Using Homebrew (macOS & Linux)
Prebuilt binaries via the 320exh/homebrew-prompt-diff tap:
brew install 320exh/prompt-diff/prompt-diff
Pre-built binary (Linux, without Homebrew)
Download prompt-diff-linux-amd64 or prompt-diff-linux-arm64 from the latest release, then:
chmod +x prompt-diff-linux-* && sudo mv prompt-diff-linux-* /usr/local/bin/prompt-diff
Pre-built binary (Windows)
Download prompt-diff-windows-amd64.exe or prompt-diff-windows-arm64.exe from the latest release, then add it to your PATH (or run it in place). Double-clicking the .exe also works β it boots the dashboard directly instead of printing CLI usage.
go install github.com/320exh/prompt-diff@latest
Build from source
git clone https://github.com/320exh/prompt-diff.git
cd prompt-diff
# build the Go binary (embeds the compiled dashboard UI)
make build
# equivalent to:
# npm --prefix web install && npm --prefix web run build
# go build -o prompt-diff .
π» CLI Usage & Commands
0. Scaffold a New Project
Starting from scratch? init drops a starter .prompt file, eval suite, and config so the rest of the commands below have something to run against immediately:
prompt-diff init
# created example.prompt
# created example.eval.json
# created .prompt-diff.yml
prompt-diff lint example.prompt
Existing files are left untouched unless you pass --force.
1. Inspect & Diff System Prompts
Compare your current prompt draft against HEAD or a specific branch/commit:
# Compare local working prompt against HEAD
prompt-diff diff system_prompt_example.prompt
# Compare against a specific commit or tag
prompt-diff diff prompts/agent_v2.prompt --v1=v1.2.0 --v2=HEAD
# Machine-readable output, e.g. for a CI job posting a PR comment
prompt-diff diff prompts/agent_v2.prompt --json
# Semantic similarity via Voyage AI embeddings (needs VOYAGE_API_KEY)
prompt-diff diff prompts/agent_v2.prompt --semantic
Terminal Output Preview:
Prompt: prompts/agent_v2.prompt
Target Models: gpt-4o, claude-3-5-sonnet
Semantic Similarity: 0.947 (cosine, Voyage AI embeddings)
Token Delta: +142 tokens (+18.4%)
Cost Projection (100k invocations):
- gpt-4o: $0.35 -> $0.41 (+$0.06)
with prompt caching: $0.18 -> $0.21 (+$0.03)
- claude-3-5-sonnet: $0.42 -> $0.50 (+$0.08)
with prompt caching: $0.04 -> $0.05 (+$0.01)
(caching assumes the whole prompt is cached: 1st call pays the cache-write price, the other 99,999 pay the cache-read price; models without a modeled discount are omitted)
Structural Diffs:
+ Added section: [Output Constraints]
~ Modified variable: {{ user_context }} -> {{ augmented_user_profile }}
2. Run Test-Suite Benchmarks
Evaluate prompt changes across a test-case matrix before pushing to production:
# Run local test matrix using Ollama and OpenAI
prompt-diff eval --prompt prompts/agent_v2.prompt --tests tests/eval_suite.json \
--models llama3.1:8b,gpt-4o-mini --output report.html
# Reviews this run's results later
prompt-diff runs
# Compare two stored runs, or export a markdown/PDF benchmark report
prompt-diff runs compare 12 13
prompt-diff runs export 12 13 --out report.md
prompt-diff runs export 12 13 --format pdf --out report.pdf
The eval harness records every run in a local SQLite store (runs.db). See prompt-diff runs to list past runs.
In-flight provider calls are capped at 5 by default (--concurrency N to change) to avoid tripping rate limits on large suites.
prompt-diff runs export writes a report table for the given run(s); pass exactly two ids to also get a delta section. --format is md (default, stdout or --out file) or pdf (--out required).
3. Auto-Compress a Prompt
Ask an LLM to rewrite a prompt more tersely, guarded against quality regressions:
# Report the rewrite and its token savings, print the compressed body (no write)
prompt-diff compress --prompt prompts/agent_v2.prompt --model claude-3-5-sonnet
# Guard the rewrite against a pass-rate drop on your existing suite, then apply it
prompt-diff compress --prompt prompts/agent_v2.prompt --model claude-3-5-sonnet \
--tests tests/eval_suite.json --max-loss 2 --apply
Rejects the rewrite (nonzero exit, nothing written) if it drops a {{ variable }} placeholder, or β with --tests β if its pass rate falls more than --max-loss percentage points below the original's. --apply writes the result back to --prompt (or --out); frontmatter is re-serialized rather than preserved byte-for-byte.
4. Launch Local Web Workspace
Open an instant local playground in your browser:
prompt-diff ui --port 8080
Opens http://localhost:8080 with interactive variable tuning, live token counter, side-by-side stream comparison, and an eval-run history table (click two rows to see a pass-rate/passed/failed/total delta).
5. Convert LangChain / LlamaIndex Prompts
Diff, eval, or compress a prompt authored in another framework β or export a .prompt file back to one:
prompt-diff convert --in chain_prompt.json --from langchain --to prompt --out agent.prompt
prompt-diff convert --in agent.prompt --from prompt --to llamaindex --out agent.json
Covers the common single-template shape (LangChain PromptTemplate.save()/load_prompt(), LlamaIndex PromptTemplate) with f-string {var} placeholders converted to/from {{ var }}. Multi-message ChatPromptTemplate and other schema variants aren't supported.
6. Configuration
Drop a .prompt-diff.yml in the repo root (or ~/.prompt-diff.yml for a global default) to override built-in per-model pricing β useful for negotiated rates or non-USD projections:
price_overrides:
gpt-4o: 2.10 # USD per 1M input tokens
claude-sonnet-4-5: 2.75
prompt-diff supports standard frontmatter metadata for declaring variables, model targets, and system rules:
---
name: Customer Support Classifier
version: 2.1.0
models:
- gpt-4o-mini
- llama3.1:8b
variables:
- user_query
- historical_context
---
You are an expert customer support agent.
Analyze the user query: {{ user_query }}
Context history:
{{ historical_context }}
Return JSON with classification and confidence score.
π€ Contributing
Contributions are warmly welcome! Please see CONTRIBUTING.md for guidelines on setting up a development environment and submitting pull requests.
- Fork the Project
- Create your Feature Branch (
git checkout -b feature/AmazingFeature)
- Commit your Changes (
git commit -m 'Add some AmazingFeature')
- Push to the Branch (
git push origin feature/AmazingFeature)
- Open a Pull Request
π Security
Provider API keys are read from environment variables only, never persisted. See SECURITY.md for details and how to report vulnerabilities.
π License
Distributed under the MIT License. See LICENSE for details.