Skip to content

Fix macro classification for Copilot-only MCP tools - #3097

Merged
George Ng (GeorgeNgMsft) merged 4 commits into
mainfrom
georgengmsft-macro-replay-classification
Sep 30, 2026
Merged

George Ng (GeorgeNgMsft) merged 4 commits into
mainfrom
georgengmsft-macro-replay-classification

Conversation

@GeorgeNgMsft

@GeorgeNgMsft George Ng (GeorgeNgMsft) commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Captured Copilot tools such as github-mcp-server.web_search can have an MCP server name without being accessible to TypeAgent's deterministic replay runtime. This change checks replay-host availability when creating a draft, so those workflows can be reviewed, approved, and handed to the TypeAgent Macro Runner instead of failing replay-tool inspection during approval.

  • Make trace induction asynchronous and inspect captured MCP tools in the recorded workspace context. Only tools found by the replay host are classified as replayable; native tools, unavailable tools, and tools without a replay host are agentRequired.
  • Preserve captured tool identities and hand off mixed workflows as a whole, without executing a replayable prefix.
  • Let the macro runner access installed Copilot MCP tools under live permissions. Its instructions restrict use to the approved procedure, inspection/discovery, and successful candidate submission; unavailable or denied tools stop the run.
  • Distinguish genuinely unavailable capabilities from connection, authentication, tool-list, and relevant configuration-discovery failures. Do not invoke recorded tools during classification.
  • Refresh advertised tool catalogs during inspection so approval and preflight do not reuse stale induction schemas. Keep existing versions immutable and retain fail-closed replay checks.
  • Document recreating an older misclassified draft from its saved trace and explicitly reviewing/approving it.

Adversarial review: two independent principal-engineer passes reviewed base 61c4415235fb72797832845353ea31f38f886428 to initial head bc3e0804f9b41058e68cc715884f88b8a66a5584. Confirmed and fixed runner tool access, discarded discovery diagnostics, and stale catalogs in b46ad19a93dd240a6ea169ffa315aa5759e9bf39. The runtime did not expose reviewer model identities, so model diversity was not verified.

Validation: 39 macro tests, 14 replay-host tests, and 17 focused plugin tests passed, including production-host removal/schema-drift coverage. Dependency-aware server/plugin builds, bundle asset/import validation, formatting, and lint/complexity/circular-dependency/test-debt gates passed.

At the user's request, deployed the reviewed inbox server build and updated the runner profile in the installed and live marketplace plugins, preserving existing routing configuration and newer hook/MCP/extension bundles. Backed up and hash-verified 32 changed files. A live synthetic RPC check passed capture, Copilot-only MCP classification, validation, approval, and auto agent handoff on the restarted installed daemon. Test macro/trace/handoff artifacts were cleaned up. No live weather/search tool or Macro Runner execution was performed; a fresh Copilot session is needed to load the updated runner profile.

Route captured Copilot-only MCP tools through the macro runner without weakening approval or deterministic replay checks.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Expose live tools to the constrained runner, propagate discovery failures, and refresh tool catalogs for approval and preflight.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This follow-up to #3097 fixes two coupled problems in agent-guided macro
execution: the runner could mistake MCP backend provenance for a
callable namespace, and induced result guards described Copilot's
event/UI wrapper rather than the response the runner actually receives.
The change preserves the original tool identity, permissions, raw
evidence, and immutable approved versions.

- Resolve the exact captured callable before deferred discovery.
Document the verified `web_search` / `github-mcp-server/web_search`
bridge without allowing arbitrary aliases or provider substitution.
- Capture separately redacted model-facing results from the SDK's
`result.content`, preserving valid JSON primitives and ordinary text as
well as the raw event result.
- Use the model-facing representation for normal guards and prior-step
result bindings throughout agent-required procedures, including mixed
procedures. Deterministic-only induction continues using raw results.
- Keep legacy guards intact, warn when model-facing evidence is absent,
and require recapture and explicit approval when an older macro needs
unobservable fields.
- Add regression coverage for callable/provenance preservation,
redaction, absent versus null/false/zero/empty results, mixed-step
references, deterministic compatibility, and runner safeguards.

## Local validation before final adversarial review

Actual standalone Copilot CLI worker **1.0.89** executed a benign IANA
search, captured it, and ran its normally induced/validated/approved
version 2 through the real `typeagent:typeagent-macro-runner`. The
runner inspected the immutable version and executed exactly one live
`web_search` with the original arguments. Execution events report
`success: true`; an independent harness checked **all 11 normal inferred
type/path guards**, including answer and citation paths, against the
actual returned model-facing content. No guard pruning, synthetic
result, retry, provider substitution, or candidate submission was used
in this final test.

Harness boundary: real headless CLI events were streamed through
production `SessionCapture`; unrelated conversation-history sinks were
isolated, and the headless CLI final-result marker was adapted to SDK
`session.idle`. This was not an unmodified interactive
extension-recording test. The fixture and its persisted
macro/trace/handoff were removed after retaining evidence and verifying
the approved version was unchanged.

Evidence identifiers: run `36b64f4a-bc0d-4002-8b1e-950aeec44773`, runner
`b1f17ff8-34a8-4554-8820-052e2d04b457`, live search
`call_4pXqFIfIxtxnas16NpdfhPf9`.

- Dependency-aware plugin, macro, and server builds passed.
- 46 macro tests, 178 plugin unit tests, 14 replay-host tests, and the
recording RPC test passed.
- Formatting and all four PR gates passed against the exact parent base,
including tests in lint/complexity checks.
- Two independent final adversarial reviewers examined the expanded
change **after** full guarded CLI validation; neither reported a
substantive issue. Reviewer model identities were not reliably exposed,
so model diversity is unverified.
- Reviewed base: `b46ad19a93dd240a6ea169ffa315aa5759e9bf39`; reviewed
content committed unchanged as
`3aab21a035767af6d3573243ce719a2f6093c137`.

## Compatibility and recovery

The original access refusal was observed in an actual user run, but
baseline resolution was nondeterministic: another baseline fixture found
the tool and then failed the separate wrapper guards. This PR does not
claim a proven permission or subagent-inheritance defect.

The user's existing **Consider Stay Home Day** version 2 was neither
changed nor replayed. Its old wrapper-field guards may still be
unverifiable. After updating the server/plugin and starting a fresh
Copilot session, **record a new interaction, inspect the new draft's
inputs and guards, and explicitly approve it**. Do not waive old guards
or rewrite the approved version. Parameterization/reasoning features are
out of scope.

## Local deployment and rollback

Only five artifacts were updated: the installed server bundle, two
runner profiles, and two extension bundles. Each extension received only
the verified capture-module change; every other installed module was
preserved byte-for-byte. The rebuilt server bundle differed only in the
intended induction logic. Backups and hashes were recorded; the
installed daemon was restarted and verified responsive. No routing
settings, conversation bindings, Azure identities, or authentication
helpers were changed.

Local evidence/rollback directory:

`C:\Users\georgeng\.copilot\session-state\af340a10-e91b-4f95-8997-d99fb88a000d\files`

Key files: `guarded-verification.json`, `real-capture-definition.json`,
`cli-tool-metadata.json`, `fixture-cleanup.json`,
`result-deployment.json`, `extension-module-provenance.json`, and
`runtime-before-results`. Original runner-only backups are
`runner-before-0.agent.md` and `runner-before-1.agent.md`. Restore only
receipt-listed files after checking for subsequent updates; stop/start
the daemon when restoring its bundle.

This PR targets the inherited parent branch, leaves #3097 unchanged, and
is for human review; it has not been merged or self-approved.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@GeorgeNgMsft
George Ng (GeorgeNgMsft) added this pull request to the merge queue Sep 30, 2026
Merged via the queue into main with commit e4d2d2e Sep 30, 2026
28 of 33 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants