Ops agents, installed with guardrails. We've seen every way agents fail, so ours don't.
AgentPostmortem installs support and ops agents for teams: Resolvd triages tickets and executes guarded actions against live systems, Webhands runs portal ops with screenshot proof. One workflow live in 14 days: book the pilot.
Underneath is a public registry of 65+ documented AI-agent failures at agentpostmortem.com, plus the verification, eval, and security tooling built from reading those cases: scoped tool access, audit trails, human approvals, and a unit-tested policy engine.
| Repo | What it is |
|---|---|
| agentpostmortem | The public case registry itself, www.agentpostmortem.com |
| Casebook-MCP | Remote MCP server exposing the registry as tools any agent can query, plus an investigator agent that drafts postmortems from real precedents |
| Casebook-Chat | Streaming chat UI that searches the live registry over MCP and answers with cited case IDs |
| Repo | What it is |
|---|---|
| MCP-audit | Security scanner and linter for MCP servers, 18 rules, SARIF output |
| Skill-audit | Scans agent skills for prompt injection, dangerous shell, secret access, and exfiltration before you install them, 31 rules |
| Injection-arena | Self-hostable prompt-injection challenge game with a leaderboard |
| Answerproof | Tamper-evident receipts for RAG answers, Merkle inclusion proofs and Ed25519 signatures |
| VaultRAG | Permission-aware RAG with access control enforced inside the retrieval query, with a gold-set eval that fails CI on any leak |
| Repo | What it is |
|---|---|
| Evalgate | Prompt and agent regression CI, a GitHub Action that fails the build when a prompt gets dumber |
| Tracecase | Record agent runs and replay them against prompt and model changes to catch regressions and unsafe tool calls |
| Voiceeval | Evaluation for voice agents: mis-hearing, missing confirmation, latency, barge-in |
| Agentrace | Observability for Claude Code subagents, reads session transcripts and flags results you should not trust |
| Repo | What it is |
|---|---|
| Ctxlens | Context-window profiler for AI agents, shows what is eating your tokens |
| Ctxtrim | Finds the files ballooning your coding-agent context and writes ignore files to cut it |
| tokencut | Measures and cuts the token cost of LLM and agent message payloads, no model calls |
| Repo | What it is |
|---|---|
| Bridgekit | Scoped MCP server exposing company tools with per-client permission boundaries and an append-only audit log |
| Webhands | Computer-use agent for tools with no usable API, refuses write actions without explicit confirmation |
| Greenlite | Mobile approval cockpit for AI agents, one-tap approve or deny routed back to the agent |
| Resolvd | Support-inbox operator that triages tickets and executes guarded actions against live systems (Shopify orders and refunds), escalating the rest |
| RelayG | Support ticket triage agent as a LangGraph state machine with a human-in-the-loop interrupt and SQLite checkpointing |
| Tenantq | Multi-tenant hybrid-search reference on Qdrant, dense plus sparse RRF fusion with Recall@K and p95 benchmarks |
Browse every repo at github.com/orgs/AgentPostmortem/repositories and filter a repo's issues by the good first issue label.
Two contributions are always welcome. First, a new documented case in agentpostmortem: a real, sourced agent failure written up in the registry format. Second, a new detection rule for MCP-audit or Skill-audit, ideally with a fixture that fails before the rule and passes after.
Read the CONTRIBUTING.md in the repo you are changing (for example MCP-audit/CONTRIBUTING.md) before opening a pull request.