Skip to content

Repository files navigation

Multiagent

Multiagent is a Rust control plane for coordinating existing coding agents. It does not implement another coding agent or model loop. It runs Codex, Claude Code, and Qwen Code in explicit roles, records durable workflow state, and accepts work only when reviewer evidence matches the exact final Git diff.

Stateful control server

The container image runs a same-origin web UI and authenticated WebSocket gateway as PID 1. Each task owns an isolated tmux orchestrator session. Paused, completed, and archived tasks retain their workflow state and terminal transcript without retaining an active tmux process; resuming reconstructs the session from that state.

The production container runs its authenticated control server as trusted UID 10000. A root-owned setuid launcher accepts privileged bootstrap and session launch only from that UID. The orchestrator, writer, readers, authority supervisor, and operations agent run as fixed UIDs 10001 through 10005. Only the setuid-gated role-agent-exec path may launch a registered role process; Landlock and Unix ownership enforce its filesystem boundary. The tmux environment is allowlisted so KMS, prod-mcp, GitHub, and AWS workload credentials remain available only to the authority supervisor and trusted control process.

Production operations are driven by authoritative Markdown runbooks rather than compiled into multiagent. The isolated operations agent materializes a generic JSON prod-mcp execution envelope from the selected .md runbook, and an independent read-only reviewer must bind an accepted verdict to the exact request, original goal, and runbook before the supervisor will sign it with KMS and forward it with the bearer token. The operations agent has logical authority to request any operation allowed by prod-mcp, but it never receives KMS, AWS, bearer-token, Grafana, or Kubernetes credentials. Prod-mcp remains the final operation and target policy boundary, and a separate post-execution reviewer inspects the persisted request and receipt.

Users are configured in a mounted JSON file. Passwords must be scrypt hashes, never plaintext:

node bin/hash-password.mjs operator

The mounted file has this shape:

{
  "sessionSecret": "at-least-32-random-characters",
  "users": [
    {"username": "operator", "passwordHash": "scrypt$16384$8$1$..."}
  ]
}

Task repositories must be provisioned as Git worktrees below MULTIAGENT_REPOSITORY_ROOT. Repository provisioning is owned by the deployment rather than the control server, so restarting the UI never fetches or mutates source checkouts.

Important container variables:

  • MULTIAGENT_USERS_FILE: mounted login configuration, default /run/secrets/multiagent/users.json.
  • MULTIAGENT_REPOSITORY_ROOT: deployment-provisioned Git worktrees available for new tasks.
  • MULTIAGENT_IDLE_TIMEOUT_SECONDS: inactivity period after which a running task is checkpointed and paused.
  • MULTIAGENT_PUBLIC_URL: canonical HTTPS origin accepted for browser and WebSocket requests.
  • PROD_MCP_URL: internal MCP endpoint used only by the authority supervisor.
  • MULTIAGENT_KMS_KEY_ID: AWS KMS P-256 key alias or ARN used only by the authority supervisor.

The PVC mounted at /var/lib/multiagent is the primary store for repositories, CLI conversation history, checkpoints, and session metadata. Final reports, a bounded terminal tail, and the transcript index live under each task's existing logs trace root. They reference immutable agent event traces instead of duplicating full transcripts. Deployment infrastructure may export this state to durable storage without coupling the control server to a storage provider.

Requirements

From a source checkout you need Rust 1.75+, Cargo, Bash, Git, and tmux. Install and authenticate at least one supported coding-agent CLI. Python 3.8+ is used only by evaluation and evidence-analysis tools, not the production control plane.

Quick Start

Run:

./launch.sh --session multiagent --root /absolute/path/to/target-repo

launch.sh is only a compatibility bootstrap. It locates or builds the Rust binary and immediately executes:

multiagent launch --session multiagent --root /absolute/path/to/target-repo

Launches are clean by default. Resume durable state after an interrupted run with:

./launch.sh --resume --session multiagent --root /absolute/path/to/target-repo

The default role backends are Codex for orchestration and verification and Claude Code for workers. To use one backend for every role:

ORCHESTRATOR_CLI=codex \
WORKER_CLI=codex \
SUBAGENT_CLI=codex \
VERIFIER_CLI=codex \
./launch.sh --root /absolute/path/to/target-repo

Supported backend names are codex, claude, and qwen.

What Runs

flowchart LR
    U["Task"] --> O["Read-only orchestrator"]
    O --> D["Decision + contract"]
    D --> W["Path-scoped writer"]
    W --> S["Canonical Git snapshot"]
    S --> V["Read-only reviewers"]
    V --> G{"Supervisor gates pass?"}
    G -- no --> D
    G -- yes --> C["Atomic completion"]
Loading

The Rust binary owns decisions, workflow phases, assignments, snapshots, findings, todos, reviewer evidence, process lifecycle, status, and recovery. Tmux owns PTYs and interactive terminal lifecycle. Python is restricted to evaluation and provenance; it does not implement a second production workflow or acceptance gate.

On production Linux, separate Unix identities isolate the orchestrator, the single active writer, read-only agents, and the authority supervisor. The orchestrator can read worker and reviewer state but cannot write the target repository or protected lifecycle state. Completion is a request to the supervisor, which checks every gate under the lifecycle lock before changing the phase to complete.

Common Commands

multiagent status
multiagent watch
multiagent decision list
multiagent workflow status "$MULTIAGENT_WORKFLOW_ID"
multiagent subagent list
multiagent subagent gate-check
multiagent orchestrator complete

Normally the orchestrator issues lifecycle and subagent commands. Operators use the status, watch, recovery, and inspection commands to supervise a run.

Documentation

  • Decisions — why the control plane and backend boundary have this shape.
  • Architecture — components, authority boundaries, lifecycle, state, and evaluation boundary.
  • Getting started and operations — configuration, normal operation, decisions, agents, recovery, traces, and troubleshooting.

Test

cargo test
bash tests/run.sh

Linux authority-boundary coverage is exercised by:

bash tests/malicious-orchestrator.sh

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages