kibble

command module
v0.17.0 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Aug 31, 2026 License: MIT Imports: 31 Imported by: 0

README

kibble

kibble

Dogfood your docs.

Release Go version License

Eating your own dog food means using what you ship the way a stranger would. Nobody does it for documentation, because your machine already has everything installed and the instructions pass by inspection. kibble is the bowl: it runs your documented steps in a clean container from zero, as a reader with nothing would, so a broken install fails in CI instead of in their terminal.

Your README tells people to run go install ..., then some setup, then a quickstart. Every one of those rots the moment the code moves, and you are the last to know.

What it does

The report is colored when a person is watching and plain the moment it is piped, so a build log never fills with escape sequences. NO_COLOR turns it off everywhere.

kibble reads a repository's README, finds the install commands, and runs each one in a fresh container with nothing preinstalled. It smoke-tests the installed binary and reports which steps a brand-new user could actually complete.

  • Extracts install commands from fenced, inline, and indented code, so a step written inline in prose is not missed.
  • Runs each go install in a clean golang container from zero.
  • Runs documented package installs the same way, verbatim as written: cargo install, npm install -g, yarn global add, pipx install, and uv tool install.
  • Runs each git clone recipe too: the clone and the build lines that follow it in the same code block, with GitHub SSH remotes rewritten to HTTPS for the keyless container.
  • Picks the image from the toolchain the step assumes, so a Rust project builds with cargo and a Node project with npm, and finds the installed binary wherever that toolchain puts it, even when it is not named after the package.
  • Never blames a document for a tool kibble lacks: a missing toolchain is a skip with a reason, so a red result means the documented steps are broken, not that kibble was short a compiler.
  • Verifies each documented brew formula exists in its tap, without installing it.
  • Smoke-tests the binary (--version, then --help) to confirm it runs, not just builds.
  • Checks that every flag and subcommand the README cites still exists in the binary's help output, and reports what has drifted.
  • Replays the README's quickstart and usage blocks in one clean session after install, so a documented example that no longer works fails in CI, not in a user's terminal.
  • Prints a table or JSON, and exits non-zero when a documented install fails.

Install

go install github.com/dcadolph/kibble@latest

Requires Docker, or a compatible runtime, on the host.

Usage

Point it at one or more repository directories, or none at all:

kibble
kibble ./myrepo
kibble ./repo-a ./repo-b

With no path it checks the directory you are standing in, so cd into a project and run it. There are no prompts. kibble's home is a CI job, and a tool that stops to ask a question there either hangs on a closed pipe or needs a flag to defeat it.

Example output:

REPO    KIND        STATUS  TIME  DETAIL
myrepo  brew        PASS    1s    formula exists (install not attempted)
myrepo  example     PASS    22s   15 lines ran, 9 skipped
myrepo  flag-check  PASS    0s    9 cited flags ok, 4 subcommands cited
myrepo  git-clone   PASS    41s   myrepo version 1.4.0
myrepo  go-install  PASS    28s   myrepo version 1.4.0

5 pass, 0 fail, 0 other of 5 checks
Flag Default What  
-image golang:1.26 Fallback image when no toolchain is detected.
-timeout 240s Per-step build timeout.
-workers 3 Max concurrent installs.
-json false Emit results as JSON to stdout.
-version false Print the version and exit.
-strict false Also fail on timeouts, smoke failures, drift, and gaps.
-examples true Replay each document's example blocks in the container.
-plan false Print the example plans as JSON and exit.
-suggest false Propose a .kibble.yml using a model and exit.
-mcp false Serve the Model Context Protocol over stdio.

What it checks today

kibble verifies go install steps end to end: the module resolves, it builds from zero, and the binary runs. A git clone step runs as the documented recipe, meaning the clone line plus the lines that follow it in the same code block, such as cd and make install, and whatever the recipe produces is smoke-tested. A brew step is verified against its tap, so a renamed or missing formula is caught, but nothing is installed. A build that exceeds the timeout is reported as TIMEOUT, never as a failure, so a slow network does not fail a build that would otherwise pass.

A package install runs the documented line verbatim, so what kibble verifies is the command a reader would actually type, flags and all. The bin directory is compared before and after, so the binary is found even when the package does not name it: a cargo install of ripgrep is checked by running rg, and an npm install -g of typescript by running tsc. A local build such as cargo install --path . is left to the clone recipe, which already covers it.

The image follows the step. Commands such as cargo, npm, and pip name the toolchain a project builds with, and the repository's own manifests settle what a bare make install leaves open, so a Rust or Python project is built with the tools its docs were written for instead of failing in a Go image. When nothing identifies a toolchain the run falls back to -image, and if the recipe then reaches for a command that is not there, the step is a SKIP naming the missing tool rather than a FAIL. A verdict about your docs is never a verdict about kibble's own gaps. When kibble itself cannot run a step, because the daemon is unreachable or an image will not pull, that is ERROR, kept separate from FAIL for the same reason. A repository too large to stream whole is an ERROR too: a file kibble failed to deliver is indistinguishable from one the docs never create, and guessing between them is the one mistake this tool must not make. Generated directories such as node_modules, vendor, build, and target are left out of that stream, so the budget goes to source a document might actually name.

A SKIP says kibble could not judge the line. A GAP says it did judge it, and the document is incomplete: the line names a file, directory, or setting that no documented step creates. cp skill/SKILL.md ~/.claude/skills/tool/SKILL.md fails for every reader when nothing creates that directory first, and a command that exits reporting a setting the document never mentions is the same kind of hole. Both used to disappear into the same silent skip as a placeholder the reader is meant to fill in, which is the difference a gap draws: a placeholder is your reader's job, a gap is yours. Because a document can also expect a reader to bring their own file, a gap reports and counts but does not fail a run unless -strict is set.

Instructions outlive the README they started in. A docs tree is where install and usage steps go once the front page fills up, and it rots faster because nobody reads it on the way past, so every document that walks a reader through commands is replayed in its own session and reported under its own name. Kibble picks the README, the docs/ tree, and top-level guides named for the job they do, such as GETTING_STARTED.md or UPGRADING.md. It leaves alone the documents that describe a project rather than instruct a reader: contributing guides, changelogs, security policies, decision records, and anything written for an agent. A document with no runnable recipe is skipped before a container is spent on it. Settle the rest in .kibble.yml with docs to add one and skipDocs to drop one.

Some documents name a file the reader is meant to supply, such as a calendar export or a profile they write themselves. That reads as a gap because kibble cannot tell it from a step somebody forgot. Settle it in .kibble.yml: give the session a fixture when a small valid file makes the line run for real, or a step rule with a skip reason when only the reader can supply the thing.

Flags and subcommands are read from the whole documentation set, since a reference page cites more of a tool than its front page does and drifts faster. Two things keep that from inventing failures. A flag is judged against the screen of the subcommand it was cited on, so one cited on a subcommand whose help never arrived is reported as unverified rather than missing. And a subcommand counts as rejected only when a parser names it as the unknown one: a subcommand that takes arguments rather than flags exits nonzero on --help while plainly existing, and says so about the argument, not itself.

Drift runs in both directions. The check below asks whether every flag and subcommand the docs cite exists in the binary. The reverse question, whether every command the binary advertises appears in the docs, is doc-coverage, and it catches the more common rot: a feature ships and nobody writes it down. The public surface is whatever --help prints, so a command a parser hides is deliberately private and never counted, and no allowlist is needed to say so. A command missing from every document is a GAP. A command documented somewhere but not in the README is reported as a count beside the total, never as a failure, because which commands earn README space is an editorial call and one finding per command would bury the one worth reading.

After a successful install, kibble compares the documentation against the binary itself. Every flag cited on a line that invokes the binary, every flag documented in a markdown flag table, and every subcommand those lines call, is checked against the collected --help output. A flag the binary no longer has, or a subcommand it rejects, is reported as DRIFT.

A rejection has to be a parser naming that subcommand as the one it does not have. The exit code is not enough on its own: a subcommand that takes arguments rather than flags answers --help by complaining about the argument and exiting nonzero, while plainly existing. Reading the code alone condemns every such subcommand, and matching the words loosely condemns them too, because the complaint says "unknown subcommand" about the argument. So the name in the message is what decides, and a probe that settles nothing leaves the subcommand alone.

The comparison needs a name to attribute cited flags to, so it covers the installs that name their binary up front: every go install, whose binary follows from the module path, and every package install whose package names the binary it provides. The two cases it does not cover are honest gaps rather than silent ones. A git clone recipe names no binary in advance, and a package that installs a differently named binary, as ripgrep provides rg, gives the docs nothing to key on; both are still installed and smoke-tested, they are just not compared against the README. The check is conservative in the other direction too: command lines only count when they invoke the binary by name, so flags shown for other tools do not count, and DRIFT fails the run only under -strict. kibble runs this check on its own README, so the flag table above rots loudly, not silently.

A repository kibble cannot read is a failure, not a quiet skip. A path with no README, a path that does not exist, and a malformed .kibble.yml each report ERROR and exit non-zero, so a typo in a workflow cannot pass as a green check that verified nothing.

Examples

An install that builds is only half the promise. The other half is the quickstart: the lines a new user actually types next. kibble replays them. After running the documented install, whichever ecosystem it belongs to, it copies the repository into the container and runs the README's example blocks in one session, in document order, so files and environment carry between blocks the way they do in a real terminal. The session verifies the documented tool actually landed on PATH before replaying anything, and a line that calls a documented tool the install does not provide, such as a conda alternative next to a cargo install, is skipped rather than failed. A block that no longer works, a flag that changed, a command that prints an error where the docs promised output, all fail as example in CI.

The judgment of which lines to run is deterministic and conservative, because a check that cries wolf is worse than no check. A line is skipped, never failed, when it needs something a clean container cannot honestly provide: a placeholder the reader must fill in (<api-key>, age1bob...), an interactive sign-in, a terminal, an API key, a local server, a file the docs reference but never create, a variable such as $HISTFILE that only an interactive shell sets, or a subcommand that opens a shell or serves forever. Skips are reported with their reason, so the coverage is honest about what it did and did not run. A command the docs say exits nonzero, such as a linter that fails when it finds something, is recognized and passes on that exit.

When the heuristics cannot settle a call, a .kibble.yml at the repository root does. It writes fixtures with real contents, exports environment, substitutes placeholder text, installs extra packages, chooses which documents are replayed, and forces a specific line to run, skip, or run in the background with a readiness probe. Every choice lives in the file, so the run stays reproducible and the engine stays the thing that decides pass or fail. Set disable: true to turn example checks off for a repository entirely.

A tool that watches or serves is read from how its documentation introduces it. A README opening with "nodemon is a tool that helps develop Node.js based applications by automatically restarting the node application when file changes are detected" has said that every invocation runs until something stops it, so kibble skips them rather than waiting out the timeout on each one and learning nothing. Only prose counts: a code block containing tool serve says the tool has a serve subcommand, which is a different fact. Questions about the tool itself, --version and --help, still run. To have a watcher actually verified rather than skipped, give it background: true and a readyLog, which is the one thing kibble cannot infer.

The most common use is a scanner or linter whose whole job is to exit nonzero when it finds something. Point it at real code and it does exactly that, which is the tool working, not the docs breaking. One nonzeroOk line settles it.

version: 1
examples:
  packages: [age]
  docs: [walkthrough.md]        # replay a guide the naming convention misses
  skipDocs: [docs/wip.md]       # leave one the convention picks up
  env:
    TOOL_PROFILE: ci            # exported for every document's session
  substitutions:
    "<api-key>": test-key-1234
  fixtures:
    - path: config.yaml
      contents: |
        setting: value
  steps:
    - match: mytool serve
      background: true
      readyLog: listening on
    - match: mytool scan          # a scanner exits nonzero on findings by design
      nonzeroOk: true
    - match: mytool *.go          # a documented form the clean session cannot run
      skip: file-glob loading is unreliable here
    - match: mytool demo          # force a line the planner would skip
      run: true

Drive it from an agent

kibble -mcp serves the Model Context Protocol over stdio, so an agent checks a repository's documentation with the same engine and the same verdicts the command line gives. Point a client at the binary:

{
  "mcpServers": {
    "kibble": { "command": "kibble", "args": ["-mcp"] }
  }
}

Two tools, and which one to reach for is the whole design. plan_docs answers what kibble would run and why it would leave the rest alone. It takes milliseconds, needs no Docker, and is the right first call: an agent asking "why is this line skipped" gets the reason without spending a container. check_docs runs the documented steps for real, which pulls images and builds the project, so it costs minutes and belongs in the same place a CI run does.

There is deliberately no tool for writing a .kibble.yml. plan_docs returns every skip with its reason, and the caller is already a model, so it can write a better config than kibble's own prompt would. Wrapping a model call inside a tool call for a model is a layer that only adds cost.

Let a model write the config

Writing that file by hand means reading every skip reason and deciding which ones you disagree with. -suggest does the reading. It sends the lines the engine could not settle, and only those, to a model you configure, then prints a .kibble.yml with one entry per disagreement for you to review and commit.

export ANTHROPIC_API_KEY=...          # or OPENAI_API_KEY, or KIBBLE_ADVISOR=ollama
kibble -suggest ./myrepo > .kibble.yml

The model never decides whether anything passed. It classifies documented lines, once, into a file you read and commit; after that the run is deterministic again and no model is consulted. kibble with no key configured verifies exactly as much as kibble with one, which is why the flag is the only place a model appears at all. Claude, ChatGPT, and a local Ollama are supported, chosen in that order from the environment, and KIBBLE_ADVISOR picks one explicitly. A local Ollama keeps the whole thing on your machine.

Preview the plan without running anything with -plan, which prints, per repository, the exact lines kibble would run, the ones it would skip and why, and the fixtures and packages the session needs. The preview is the floor, not the ceiling: at run time the session may downgrade more lines to skips for reasons only execution can see, such as a command that turns out to need a terminal. Turn the whole layer off with -examples=false.

Use it in CI

Add a workflow that fails a pull request when a documented install breaks:

name: docs
on: pull_request
jobs:
  kibble:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: dcadolph/kibble@v0.16.0
        with:
          repo: .
          # args: -strict   # fail on timeouts, smoke failures, drift, and gaps too

The runner already has Docker, so kibble spins its clean-room containers there.

The action downloads the released binary for the version you pinned and verifies its checksum against the release's checksums.txt before running it, so a CI log can say which kibble produced its verdict. With no version input, the version is the ref the action was pinned at: @v0.16.0 runs kibble v0.16.0. Pinning @v1 follows the latest release, which is convenient and is also the tradeoff it sounds like.

In a workflow, kibble speaks GitHub natively. A failed install or example is annotated on the exact README line that broke, so it shows up inline in the pull request the way a failing test does. Doc drift becomes a warning annotation, and the job summary gets the full results table, readable without opening a log.

Security model

kibble executes commands it read out of a repository's documentation. That is the product, and it means a README is untrusted input: anyone who can change the docs can change what runs. Treat a kibble run the way you treat a build script.

Every command runs inside a fresh container that is removed afterward, with bounded memory and process counts, no-new-privileges, and every Linux capability dropped except the handful apt needs to install the packages the docs depend on. Nothing from the host is mounted in; the repository is streamed in as a tar of its working tree.

The boundary that stays open is the network, because verifying an install is fetching it: go install, cargo install, and npm install -g are network operations. A malicious documented command can therefore reach out from inside the container. The container is disposable and unprivileged, but Docker isolation is a wall, not a guarantee. So:

  • Running kibble on your own repository in CI is the designed case: anyone who can edit your README can already edit your workflows.
  • Running kibble against a repository you do not trust is running that repository's chosen commands on your Docker daemon. Do it the way you would run their Makefile: in a sandbox you are prepared to lose.
  • Default kibble never contacts a model provider. Only -suggest does, it sends only the documented lines the engine could not settle, and the Ollama route keeps even that on your own machine.

The corpus

corpus/repos.tsv pins real repositories, ripgrep and nodemon and friends, each to a commit and to the verdict counts a person verified by hand. A scheduled workflow runs kibble against all of them and fails when any count moves, because a moved count means kibble's judgment changed, not the repository. This is the regression suite for the part unit tests cannot reach: the rules can all do exactly what they say and still be wrong about a repository nobody anticipated. Adding an entry means reading that repository's documentation and verifying the expected counts first.

Roadmap

  • Install brew formulas for real instead of only verifying they exist.
  • JUnit XML output for CI systems that are not GitHub.

Why "kibble"

Dogfooding means using your own product before you ship it. kibble is the bowl: it feeds your docs back to a fresh machine and tells you whether they still go down.

More tools

  • preen, split a messy working tree into clean, atomic git commits
  • slop-chop, strip the AI tells out of your writing
  • vamoose, route time off through approval, then tell the team
  • whodar, find who to talk to about X across your work tools

License

MIT. See LICENSE.

Documentation

Overview

Command kibble verifies that a project's documented install steps actually work for a fresh user, by running each in a clean container from zero. Eating your own dog food means using what you ship the way a stranger would, and kibble is the bowl: it reads your README so your users do not have to find out it is stale.

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL