Skip to content

fix(mcp): truthful Extensions MCP rows; reconnect queues and reports - #6724

Merged
Hmbown merged 1 commit into
mainfrom
fix/mcp-panel-reconnect-truth
Sep 29, 2026
Merged

Hmbown merged 1 commit into
mainfrom
fix/mcp-panel-reconnect-truth

Conversation

@Hmbown

@Hmbown Hmbown commented Sep 29, 2026

Copy link
Copy Markdown
Owner

No-Issue: founder live-test report on the TUI Extensions > MCP tab (v0.10.1 lane mcp-panel)

What the founder saw

Extensions > MCP showed "Needs attention (8)". Six servers read "reconnect · configured" with 0 tools, linear (disabled in mcp.json) offered "reconnect" instead of "enable", and pressing Enter on it did nothing visible and wrote nothing to ~/.codewhale/logs. The CLI at the same moment connected github, playwright and chrome-devtools fine.

Causes

  • "configured" forever: lazy boot (MCP: connect servers lazily at point of use instead of eagerly at boot #6033) never starts a server nobody selected, and the panel called that state "configured". McpServerSnapshot::recovery_kind always passed inspected = true, so a server that had never started got "reconnect", and an OAuth-capable one got "re-auth".
  • disabled server offered reconnect: the row read enabled from the live snapshot before the config file it had just loaded, so a server switched off after the last pool event kept its old live state.
  • Enter did nothing: /mcp retry awaited the engine inline on the UI loop. While a turn was running it gave up with "run /mcp retry X again after the turn finishes". A success added no receipt at all, and neither outcome was logged.

Fix

  • Truthful rows: McpServerSnapshot::started() covers connected, auth-required, a recorded error, or observed capabilities. An unstarted row reads not started with the connect action (/mcp retry X). Other states: connected, ◆ auth required (login group, /mcp login X), error (the reason is in the detail line), disconnected, disabled. "Needs attention" now holds only rows whose tone is Attention or Failure. Idle rows (not started or disabled) keep their action but sort with the other servers.
  • Config wins on/off: if the snapshot and the file disagree about whether a server is on, the snapshot is stale and is not shown. A server off in the file offers enable. A server on in the file but still off in the pool offers a reload, because the single-server retry deliberately does not re-read config.
  • Reconnect is observable: the retry op is sent from a background task (PendingMcpRetry + poll_mcp_retries, the same pattern as /mcp login). A running turn just queues it in the engine mailbox: the row reads queued and a toast says it will reconnect when the turn finishes. When it lands, a toast gives the outcome: X connected: N tool(s)., X needs a login: run /mcp login X., or X did not connect: <reason>. Engine::retry_mcp_server logs the outcome at INFO or WARN, so every surface leaves the same entry in the log.
  • Enable: set_server_enabled already writes enabled and disabled together. A test now covers a file where the two disagree.
  • aws -32602 on initialize: reproduced locally against uvx mcp-proxy-for-aws@1.6.4. The proxy returns the same -32602 Invalid request parameters for our current capabilities and for a spec-minimal {}. Its stderr shows LoginRefreshRequired … reauthenticate using 'aws login', so our handshake is not what it rejects. On our side, the error now reads MCP server 'aws' rejected initialize (command \uvx`): {…}; server stderr: `. Stderr is not retained for reviewed plugins, and plugin errors stay suppressed as before.

Evidence (local, this branch)

  • cargo test -p codewhale-tui --lib -- views::extensions::tests mcp_retry_ mcp_enable_persists mcp_mutations_while_turn_running mcp_show_while_turn_running initialize_rejection_names set_server_enabled_makes never_started_server recovery mcp_diagnose_reports test_mcp_config_crud → test result: FAILED. 127 passed; 2 failed. The two failures were core::engine::tests::sse_turn_recovery::*, which timed out under machine load (the broad recovery filter pulled them in). Re-run alone: test result: ok. 2 passed; 0 failed.
  • cargo test -p codewhale-localization → test result: ok. 51 passed; 0 failed. The Extensions key-count guard went from 99 to 100 for the new ExtensionsStateDisconnected key.
  • cargo clippy -p codewhale-tui -p codewhale-localization --all-targets --all-features --locked -- -D warnings … → clean.
  • cargo fmt --all -- --check → clean.
  • New strings are translated in all 15 complete packs. The English text of McpRetryDeferredWhileTurnRuns changed, and its translations were updated to match.

Not done

  • Client capabilities in our initialize still lists tools/resources/prompts, which are server-capability names. They did not cause the aws failure, so they were left unchanged.
  • A successful /mcp login still tells the person to run /mcp reload. It does not auto-retry the server.
  • No manual TUI dogfood run: the checks here are the unit and integration tests above.

🤖 Generated with Claude Code

https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks

Founder run, Extensions > MCP: eight servers under "Needs attention",
six reading "reconnect · configured" with 0 tools, a disabled linear
offering "reconnect", and Enter doing nothing visible or logged.

- Row state comes from the same engine pool the turn uses. Lazy boot
  (#6033) leaves unselected servers unstarted; they now read "not
  started" with "connect" (McpServerSnapshot::started) instead of
  "configured"/"reconnect", and sort with idle rows. "Needs attention"
  holds only failing or waiting rows (tone Attention/Failure).
- Config is the authority for on/off: a snapshot that disagrees with
  the file is stale and not shown, so a disabled server offers
  "enable"; one enabled in the file but not the pool offers a reload.
- /mcp retry no longer awaits on the UI loop or tells the person to
  retry later: the op goes to the engine from a background task, so a
  running turn just queues it (row "queued", toast says so). The outcome
  is toasted (connected: N tools / needs login: /mcp login X / failed:
  reason) and logged at INFO/WARN in Engine::retry_mcp_server.
- initialize JSON-RPC errors now read "MCP server 'X' rejected
  initialize (command `uvx`): {...}; server stderr: <last line>".
  Reproduced the aws -32602 locally: identical with our capabilities and
  with `{}`; the proxy's stderr says its AWS login expired, so the
  handshake shape is not the cause.
- set_server_enabled already writes enabled/disabled together; covered.

Tests (targeted, local): cargo test -p codewhale-tui --lib with
extensions/mcp_retry/mcp enable/initialize filters: 127 passed, 2 failed
(sse_turn_recovery timeouts under load; both pass alone: 2 passed).
cargo test -p codewhale-localization: 51 passed. cargo fmt --check ok.

No-Issue: founder live-test report (v0.10.1 lane mcp-panel)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Copilot AI balanced review requested due to automatic review settings September 29, 2026 06:26
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@Hmbown
Hmbown merged commit f3f2e2d into main Sep 29, 2026
35 checks passed
@Hmbown
Hmbown deleted the fix/mcp-panel-reconnect-truth branch September 29, 2026 09:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants