Skip to content

Repository files navigation

gatehouse

Local AI code review using 8 concurrent LLM agents, called through OpenRouter. Inspired by diffray's multi-agent architecture and anti-noise prompting.

Install

uv tool install gatehouse-crunchtools

Usage

# Review current branch vs main
gatehouse

# Review staged changes only
gatehouse --staged

# Review against a specific base
gatehouse --base develop

# Run specific agents only
gatehouse --agents bugs,security

# Use a different model
gatehouse --model google/gemini-3.8-flash

# Advisory mode (never exit non-zero)
gatehouse --advisory

# Write per-agent model, token and cost usage as JSON
gatehouse --usage-json usage.json

Each agent logs one line to stderr with the model that served it, its token counts and cost, and whether a fallback model answered; a total line follows. --usage-json writes the same data as {"agents": [...], "total": {...}}: each agent record has agent, model, prompt, completion, reasoning, cost (USD) and fallback; the total has calls, the summed tokens and cost, and the fallback count.

Agents

Agent Focus Blocking
Bug Hunter Null safety, logic errors, edge cases, async bugs, resource leaks Yes (high/critical)
Security Scan Injection, auth bypass, hardcoded secrets, data exposure Yes (always)
Performance Check O(n^2), N+1 queries, memory leaks, blocking I/O Yes (high/critical)
Test Coverage Missing unit/integration tests, untested edge cases and APIs Advisory only
Documentation Missing/stale docstrings, undocumented public APIs Yes (critical/high)
Constitution Violations of project constitution/spec rules Yes (critical/high)
Consistency Check Naming patterns, API consistency, error handling patterns Advisory only
General Review Over-abstraction, unclear naming, hidden dependencies Advisory only

All 8 agents run concurrently. Findings below 80% confidence are filtered out, and so are findings whose quoted evidence is not in the diff. At most 5 LOW findings from advisory agents (all but Bug Hunter and Security Scan) are reported, the most confident first.

How Agents See Changes

Since v0.3.0, agents do not receive raw unified diffs. Each hunk is rendered as a structured view: a BEFORE block (the old code, removed lines marked [-]) and an AFTER block (the new code with real file line numbers, added lines marked [+]), plus instructions to judge the direction of a change. Protections added by a change are treated as fixes, not findings; protections removed by a change are flagged. Diffs that cannot be parsed fall back to the raw unified format.

Exit Codes

Code Meaning
0 No issues or advisory-only findings
1 Blocking findings detected (critical/high)
2 Usage error (missing API key, bad arguments), or an agent could not finish (reported, but exit 0, under --advisory)

GitHub Actions

Drop in examples/gatehouse.yml to review every PR (forks included) via the reusable review.yml workflow — the diff is piped as data, never checked out or executed.

The example also runs Gatehouse triage, a deterministic check that every required inline finding has a reply (fixed in <sha> or not a bug: <reason>). Findings without a file and line appear only in the review summary and are not counted. CRITICAL, HIGH and MEDIUM findings are required from every agent; LOW findings only from Bug Hunter and Security Scan, and the required_low_agents input changes that list (pass the same value to the review and triage jobs, so the review never caps a LOW that triage requires). Other LOWs still post, as advisory.

On a push to an open PR, only the commits since the last complete Gatehouse review are reviewed (--incremental), fetched through the compare API. A PR's first review, a force-push or rebase, a review that had an unfinished agent, or a failed compare all fall back to reviewing the whole PR. The review summary says which it was, and says when a fallback model served any agent; the job's step summary shows per-agent tokens and cost.

Before posting, gatehouse drops four kinds of finding, and the review summary counts each: a finding whose quoted evidence appears nowhere in the diff (or the styleguide and constitution the agents were given), a finding on a line outside the PR diff (GitHub would reject the whole review), a MEDIUM or LOW that repeats an answered thread by the same agent within 5 lines, and advisory LOWs beyond the 5 most confident. Every finding comment carries a hidden <!-- gatehouse agent=… confidence=… --> marker for tuning.

Add examples/gatehouse-retriage.yml as well if Gatehouse triage is a required check. A reply to a finding starts a run whose checks branch rules ignore; the retriage workflow re-runs triage inside the pull_request_target run, where the result counts.

The review is advisory by default: findings post as PR comments and the check always passes (an agent that could not finish is named in the review instead of failing it), so a non-deterministic LLM finding can never block a merge. Do not mark it a required status check. To let critical/high findings fail the check (still not recommended as a required gate), opt in:

uses: crunchtools/gatehouse/.github/workflows/review.yml@v0.15.2
with:
  blocking: true

Configuration

Set OPENROUTER_API_KEY, or OPENROUTER_API_KEY_FILE pointing at a file that holds the key (the file wins when both are set). Either can live in ~/.config/mcp-env/gatehouse.env. If .gemini/styleguide.md exists in the reviewed project, it is injected as context.

Ignoring files

List paths that should never be reviewed in .gatehouse-ignore at the repo root, in gitignore syntax (*, **, trailing /, ! negation). Their diff sections and file-listing entries are dropped before any agent sees them, so data files, fixtures, lockfiles and generated output stop costing tokens. A change is skipped only when every path it touches is ignored, so a rename out of an ignored directory is still reviewed, and so is any change to .gatehouse-ignore itself.

uv.lock
tests/fixtures/
*.fp

In CI the file is read from the base branch, like the styleguide and constitution: a pull request cannot widen it to hide its own changes. The posted review says how many files were skipped; when every changed file is ignored, gatehouse prints No changes to review., exits 0 and posts nothing.

Model

The default model is openai/gpt-6-luna, with google/gemini-3.1-flash-lite as an automatic fallback when it is rate-limited or down. Every request requires zero data retention and forbids training on prompts, so only providers that keep nothing may serve it. --model takes any OpenRouter slug; the fallback still applies.

Luna was chosen on a 33-diff replay of real crunchtools changes (see crunchtools RT #1505): it caught as many reintroduced bugs as gemini-2.5-flash, with far less noise on clean PRs, at about a sixth of the cost.

Upgrading from 0.8.x: gatehouse no longer reads GEMINI_API_KEY. Replace it with OPENROUTER_API_KEY in your env file and GitHub secrets.

Constitution Discovery

The Constitution agent auto-discovers a project constitution in priority order:

  1. --constitution <path> (explicit override)
  2. .specify/memory/constitution.md (spec-kit)
  3. AGENTS.md (cross-runtime standard)
  4. CLAUDE.md (Anthropic project instructions)

If no constitution file is found, the agent is silently skipped.

License

AGPL-3.0-or-later

About

Local AI code review CLI using Gemini agents

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages