brindle / Guides

Done means the tests pass: milestone checks for autonomous coding agents

An autonomous coding agent needs a finish line it can verify. Break a larger request into milestones, give each one a check that exercises the promised behavior, and review the final work against the original goal. brindle runs milestone checks as work progresses and adds a completion audit after every check passes.

Screen recording of brindle tracking a goal while workers run and checks report pass or fail for each milestone

Write a goal that can be checked

A request such as “improve the settings page” leaves too much room for an agent to decide what improvement means. Name the outcome instead: users can update their display name, invalid email addresses show an inline message, and changes remain after reload. These statements describe observable behavior. They also give you a way to decide whether the work is complete.

Keep the goal at the level of the result. Then separate independent parts into milestones. A settings API and its form could be separate milestones if each can be built and verified on its own. If the form depends on the API, state that dependency so the work proceeds in the right order. Avoid splitting one behavior across several milestones when no part can be meaningfully checked alone.

For each milestone, choose a command that exits successfully only when the behavior is present. A focused test is often a stronger check than a broad command that happens to return zero. For example, a test for the save path can submit valid data, inspect the response and confirm the saved value. A UI check can enter an invalid email and verify that the user sees an error. The command should exercise the result, not merely confirm that a file or label exists.

# Settings page

Users can change their name and email.

## Settings API
check: uv run pytest tests/test_settings_api.py -q

## Settings UI
check: npm test -- settings
The form saves and shows errors inline.

This is the shape brindle uses for a goal loaded from .brindle/goals.md: a goal heading, milestone sections, and a check: command for each milestone. The details below a milestone can clarify its expected behavior. In an autopilot session, the supervisor can also set a goal and milestones through its tools. In either case, put the checks next to the outcomes they prove so each milestone has a clear finish line.

Make every check prove its milestone

A command that exits zero is only useful when it tests the intended behavior. Read what the test actually asserts. A test that imports a module but never calls the new behavior can pass while the feature is missing. A check that searches source text can pass even when the code path is broken. A skipped test can make a suite green without checking anything. Keep checks close to the behavior and assert results that matter to a user or a caller.

Choose the narrowest reliable command for the milestone. It should finish in a reasonable time and fail when the behavior regresses. When a behavior crosses layers, consider a focused integration check that follows the real path rather than several unrelated checks. Keep repository-level checks configured for branch merges as a separate safeguard: those checks protect integration, while milestone checks measure progress toward the goal.

brindle treats checks as evidence, not a status message from the model. It runs a milestone's check and marks it verified only when the command exits with code zero. After later work is merged, the supervisor calls check_milestone and brindle reruns the earlier passing milestone checks with it, so a change for one part of the goal can reveal a regression in another. If a check fails, the milestone remains unverified and the output gives the supervisor information to continue working.

This approach makes progress visible without treating a worker's report as proof. The supervisor can assign independent tasks to workers, and each worker gets its own branch and git worktree. Under autopilot, a reviewer approves a worker's exact commit and brindle runs the configured checks before merging it. Once branches are integrated, milestone checks cover the goal's behavior on the combined result.

Why passing milestones may not finish the goal

Even well-chosen checks have limits. A test might miss a requirement, assert too little, or run a path that does not represent the user's request. Passing every milestone is good evidence, but it does not guarantee that the delivered work covers the original goal. A completion review should ask whether every part of the request exists in the code and whether the checks genuinely support that conclusion.

brindle's autopilot ends with this final completion audit. Once all milestone checks pass, brindle starts a read-only reviewer to compare the finished work with the original goal and its milestones. The reviewer examines the code and the checks, looking for missing behavior, stubs, incomplete pieces and weak or skipped tests. It considers the work since the goal began, rather than approving just one worker's branch.

The audit is an additional review gate, not another milestone command. If the reviewer approves the audited commit, the goal is reached. If it finds gaps, the supervisor receives the findings and continues the work. After fixing them, the supervisor commits the changes and calls check_milestone again; brindle reruns the milestone checks and starts a fresh audit for the new commit. Uncommitted changes cannot be audited because there is no fixed commit to review.

The audit can be disabled with "goal_audit": false in .brindle/config.json. With the audit enabled, a passing check suite still matters: it provides the reviewer with concrete evidence and keeps regressions visible during the work. The reviewer checks the evidence against the request instead of assuming that a green status means every requirement was covered.

Keep the process moving while you step away

Autopilot is intended for a goal with multiple parts that benefit from milestones and parallel workers. It can continue assigning work, reviewing and merging branches, and rerunning checks without waiting for you after every step. It pauses when it needs a decision only you can make. If you quit a session, brindle continue resumes it later.

There is a practical balance to set. More milestones make progress easier to inspect, but each should represent a meaningful outcome rather than a tiny implementation step. A single large milestone hides which requirement is incomplete; dozens of small milestones add bookkeeping without adding useful evidence. Group the goal around behaviors that can be verified independently and keep each check focused.

Before handing off a goal, read it as if you were the reviewer. Can every sentence be mapped to a milestone? Does each milestone have a command that fails when its promised behavior is absent? Are the expected results stated clearly enough to distinguish a complete implementation from a placeholder? A few minutes spent tightening these statements makes it easier for both workers and the completion auditor to find the right gaps.

For the details of autopilot's milestones, checks and pauses, read Autopilot. For guidance on when to use brindle and how to shape a task, see Recommended use; the configuration reference explains repository settings.

Try brindle. Free to use, source-available (Brindle License 1.0), on macOS or Linux.
curl -fsSL pawdelta.com/brindle/install | sh