Developer tool
Witan
A local daemon and desktop app that reviews my pull requests with a council of AI agents, and argues with its own findings before posting any of them.


- Type
- Developer tool
- Platform
- macOS app, or any Node 22 host
- Status
- In active development
- Built with
- TypeScript / Claude Agent SDK / Codex SDK / SQLite / Electron
- Six specialist reviewers, led by a chair agent
- A skeptic tries to refute every finding before it's posted
- Learns each repo's standards and keeps them in editable markdown
The problem
Coding agents open pull requests faster than I can read them. Most of my projects now get several PRs a day, and a careful review of each one is the step I skip when I'm busy. Off-the-shelf AI reviewers didn't help much: they post a lot of comments, many of them wrong, and after a week you stop reading them.
So I wanted a reviewer with the opposite habits: few comments, each one worth reading, running on my own machine with my own credentials.
A council, not a single reviewer
Witan is the Old English name for the council that advised the king. Here the maintainer is the king and the council only advises: Witan posts one COMMENT review per push and can never approve, request changes or block a merge.
A daemon polls the repos I've onboarded. When a PR targeting a watched branch gets a new commit, Witan checks the head out into a fresh git worktree and builds a review packet: the diff, the PR description, the repo's standards and anything it has already said on that PR. Then the council meets:
- The chair reads the packet, decides which members the change needs (a docs-only PR doesn't need the performance reviewer), briefs them and collects their findings.
- Six members each look for one kind of problem: correctness, security, performance, maintainability, tests, and the repo's own standards. They can read the whole checkout, not just the diff, so they can follow a call to its definition or check whether a test already exists.
- A skeptic then cross-examines every finding against the actual code and drops the ones it can refute. This is the main defence against false positives.

The review page shows all of this live. Each agent's card has the chair's brief, the file it's reading or the pattern it's searching for, and its token count. Underneath, a timeline draws one lane per agent with a tick for every tool call. Because the progress is stored in SQLite, a review started by the background service shows up live in the app too.
Every finding has to survive
What gets posted goes through three steps. Duplicates from different members are merged; the skeptic tries to refute what's left; then a filter applies a minimum confidence, a cap on findings per review, and any preferences I've expressed. Findings that were already posted on an earlier commit are matched by a fingerprint of the file, the normalised code and the category, so they aren't repeated when lines move.

Nothing is thrown away silently. The portal keeps every finding with its outcome (posted, refuted by the skeptic, or filtered out) and the reason, which is how I tune the thresholds.
Learning what each repo cares about
When a repo is onboarded, a standards scout reads its CONTRIBUTING.md, CLAUDE.md, linter configs, a sample of recent code and past review comments, and writes down the rules it finds, each with evidence and a confidence. They live in a markdown file under ~/.witan/memory/ that I can edit, and my edits always win.
Preferences come from how I react to Witan's comments. A ๐, a thread resolved without a code change, or a reply like "we don't care about this here" all count as signals. Once they agree a few times, they become a rule that suppresses or downgrades that kind of finding. Explicit commands work immediately: @witan-review ignore naming for me or @witan-review remember: <rule>.
Safety
PR content is untrusted input: a title, a comment or the code itself can contain a prompt injection. So the agents get read-only tools (Read, Grep and Glob, confined to the checkout), no shell, no network and no GitHub token. Every output is validated against a schema, so free text never turns into an action. Only the publisher, which is plain code, can talk to GitHub. New repos start in dry-run mode, where reviews are written to disk instead of posted.
Bring your own credentials
Witan can run as a GitHub App with its own bot account, or as me, using my gh login, in which case it marks its comments with hidden markers so it can tell them apart from mine. On the AI side, each role can use a different provider and model: a Claude subscription, an Anthropic API key, Bedrock, Vertex, an Anthropic-compatible gateway, or my ChatGPT plan through Codex. Codex gets the same rules as Claude Code: a read-only sandbox and only Witan's own tools, served over a local MCP server.
The app
Everything is configured from a local portal: providers, the model and effort for each role, repos and their standards, reviews and logs. It runs as an Electron app that keeps watching from the Dock, or as witan portal in a browser tab. It listens on 127.0.0.1 only, behind a session cookie, with host and origin checks. The UI is plain ES modules with Preact and no build step, and it comes in light and dark themes, including popular editor palettes.
The best test so far has been pointing it at itself: since the first working version, Witan has reviewed its own pull requests before I merged them. The screenshots on this page come from a demo setup with made-up reviews of my other projects.