brindle / Guides

A coding agent team on a laptop with Ollama, no API keys

A local model can handle small, clearly scoped coding tasks and review modest diffs. With brindle, workers can share one Ollama server while each works in a separate git worktree.

Screen recording of brindle starting coding agents that each work in a separate git worktree, the same setup local Ollama workers use

What runs locally

Ollama serves open-weight models on your machine. brindle’s native provider talks to the model’s chat endpoint and runs its tool calls, so the worker can read and edit files, run commands, and report its result without a third-party agent CLI. The worker does not need an API key when the endpoint is local. A local model can also review another agent’s branch. The supervisor chat still runs on Claude Code, Codex or Antigravity.

The built-in local developer and reviewer profiles use qwen3-coder:30b. The README lists that model at about 19 GB and says it runs on a 32 GB machine. That is a substantial download and memory requirement, so check your laptop’s available memory and disk space before pulling it. A smaller or different tool-capable model may be a better fit for your hardware; brindle also documents gpt-oss:20b and glm-4.7-flash as models that support tools in Ollama.

Install and check the local model

Install Ollama, download a model, and let brindle check that it is available:

brew install ollama
ollama pull qwen3-coder:30b
brindle doctor

The built-in profiles point at Ollama using the OpenAI-compatible local endpoint. Because a 30B model holds about 20 GB of memory, brindle doesn’t start Ollama unless you ask it to: run ollama serve yourself, or set "local_models": true in .brindle/config.json. With that set, if Ollama on this machine is not answering when brindle starts, brindle starts it in the background with the context length the profiles need and loads their models. It stops a server it started when the last brindle session that uses it ends. A server you started yourself is left alone.

In a repository, start a brindle session as usual. The supervisor can assign work to the built-in local developer, and a repository can use the local reviewer profile for branch reviews:

cd ~/code/myapp
brindle init
brindle
# In the supervisor chat, assign a bounded task to developer-local

To make the local model the configured reviewer, the README gives this setting for .brindle/config.json:

{"review_profile": "reviewer-local"}

These local profiles use brindle’s native agent loop. Native agents run headless; brindle agent peek shows their turns, tool calls, and answers, and brindle send sends them a message.

Give each worker a small task

Local models tend to be most useful when the task has one clear outcome and a check you can name up front. For example, ask one worker to fix a specific validation bug in one module and run its named test file. Ask another to review a modest change for a concrete class of problems. brindle gives workers separate worktrees and branches, so parallel changes do not overwrite each other’s working files.

Avoid handing a local model a vague project-wide goal and expecting it to plan a long sequence of dependent changes. Split that goal into milestones or separate assignments with explicit files, behavior, and checks. The supervisor can coordinate those pieces, then inspect a worker’s result and merge its branch after review and checks pass. A model saying it is finished does not make the work verified: brindle marks a milestone complete when its check command exits successfully.

Tool use varies by model. If a model repeatedly forms malformed tool calls, try a different Ollama model with tool support. Keep task instructions short and specific, and name the check command rather than asking the model to decide what counts as done.

Plan for context and memory

The model’s context window holds the task, repository material, tool results, and its reply. A larger context can make a long conversation possible, but it also takes memory. The local profile specifies context_tokens; Ollama needs a context length at least that large, plus room for the reply, or it may silently truncate the conversation. brindle sets the context length it needs when it starts Ollama itself (with local_models on). The built-in profiles use 40960 tokens for the context and reply together.

Several parallel agents still share the laptop’s CPU, memory, and disk bandwidth. Start with one local worker and one reviewer, then add parallel work only if the machine remains responsive. Each worktree may also need its project dependencies installed. Keep your checks quick and targeted for worker tasks, and use the repository’s normal full checks before accepting a set of changes.

When local is a good fit

A local team is useful when you want workers and reviewers to run without sending code to a hosted model provider, have a capable machine available, and can divide work into focused assignments. The supervisor chat itself still runs on Claude Code, Codex or Antigravity; only air-gap mode (brindle Enterprise) runs it on a local profile too. The model and Ollama run on your computer; brindle handles worktrees, messages, reviews, and checks. Local inference is not automatically as capable or as fast as a hosted model, and a laptop can become memory-bound when several large model requests run together. Begin with a small task and see how your chosen model performs on your repository.

For paid features, brindle Pro adds optional per-worktree databases; it is not required to use Ollama or the native local profiles.

See Providers and models for profile details and Recommended use for choosing and checking agent work.

Try brindle. Free to use, source-available (Brindle License 1.0), on macOS or Linux.
curl -fsSL pawdelta.com/brindle/install | sh