Skip to content

Add GHCP end-to-end evaluation harness - #3072

Merged
George Ng (GeorgeNgMsft) merged 15 commits into
mainfrom
georgengmsft-ghcp-eval-implementation
Sep 30, 2026
Merged

George Ng (GeorgeNgMsft) merged 15 commits into
mainfrom
georgengmsft-ghcp-eval-implementation

Conversation

@GeorgeNgMsft

@GeorgeNgMsft George Ng (GeorgeNgMsft) commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Adds a reproducible Copilot end-to-end evaluation in the dedicated copilot-plugin-eval package, comparing TypeAgent routing strategies with native Copilot. This is a simple initial evaluation for ballpark estimates, not a statistically powered benchmark. The latest corpus replaces list-dependent tasks with ordinary file tasks that all candidates can perform.

  • Keeps seven candidates and twenty examples in four five-case cohorts. Protocol 5 (common-files-v1) replaces nine list-dependent cases with file inventory, append, backup/overwrite, comparison, conditional edits and clarification tasks. All candidates have twenty applicable cases: 140 executions per repetition, no native N/A slots.
  • Uses existing registered PowerShell file actions for TypeAgent and ordinary native edit/create tools for Copilot. Per-case fixture permissions, pre-clarification gates, exact backup prerequisites and independent file-state checks protect unrelated state. Arbitrary shell writes are not autoapproved; GitHub and network remain read-only.
  • Creates fresh conversations, bindings and restored fixtures; contract reuse is earned through separate same-binding discovery and preparation is accounted separately. Preserves provenance-checked SDK overflow artifacts and one scripted final-text clarification within the original deadline.
  • Pins future outer sessions, contract preparation and nested Copilot reasoning to Luna 5.6 (gpt-5.6-luna). Rejects mismatched ledgers and stale preflights; freezes fresh corpus/model/order specifications rather than relabeling historical runs.
  • Records E2E timing, model/tool activity, actual internal failed-translation fallback and interaction evidence. Unknown timing stays unknown. Retains bounded credit admission, uncertain reservations and cumulative accounting; the ceiling is 50,000 cumulative credits, not a fresh allowance. External-provider accounting remains separate and incomplete.
  • Moves eval scripts/tests and maintained methodology into ts/packages/copilot-plugin-eval; runtime guards remain at dispatcher enforcement boundaries. Explicitly documents residual host/plugin/environment-isolation risks: a separate directory is not a sandbox.
  • Preserves strict no-replay behavior in this base layer. The dependent Refine prospective GHCP evaluation recovery, applicability and ambiguity case #3077 layer integrates positive-evidence safe-read recovery while retaining denial/cancellation/uncertainty precedence.

Historical evidence, not new file-corpus measurements

The completed 140-trial historical pass used protocol 2 at e847c7a00907a1f4e2c6c4426c935b964fd9997c with gpt-5.6-sol and the original list workload. All combinations were attempted; not all succeeded. Original strict successes were 6, 6, 11, 11, 6, 8, 3 /20 for C1-C7. The later protocol-3 native terminal-gate correction was tested offline, not live-measured.

Removing only the blanket continuation penalty produced retrospective content scores of 6, 6, 11, 11, 7, 8, 8 /20. This is not a new execution or certification of recovery safety: some native failure evidence is incomplete, and network final-presentation failures remain. Excluding nine list-dependent native cases gave a historical native 8/11, not a matched-workload ranking. Original success-conditioned timing populations and uncertainty remain separate; these observations establish neither a general winner nor an internal-fallback benefit.

No Luna or common-files-v1 live evaluation has run. New corpus outcomes must not be pooled with old results. Fresh readiness, model-specific accounting bounds and authorization remain prerequisites. The maintained methodology records the approved corpus amendment; historical reports, sanitized exports, original grades, frozen specifications and accounting artifacts remain external and unchanged. Private network evidence, credentials, billing traces and machine-specific paths are not committed.

Latest offline validation

  • Dependency-inclusive evaluation-package build passed; final dispatcher build passed.
  • 23 eval tests passed, covering all20/all7 scheduling, stale readiness, file/backup oracles, clarification, scoped confirmations, link rejection, Luna admission and relocated entry points.
  • 174 relevant dispatcher/structured/permission tests passed for the corpus commit; the final directory-read clarification guard reran all 17 affected policy/artifact tests successfully.
  • Final-head formatting and all four ratchets passed against origin/main: no added lint/complexity violations, cycles unchanged at 220, no focused/newly skipped tests.
  • Earlier offline artifact validation covered all 140 historical conversations/reviews, passing fixture/clarification/route checks, native post-failure exclusions, network-export sanitization, historical preservation and cumulative accounting. This is historical validation, not a run of the new corpus.
  • Published head: f2f045c2066605f69c990f3b99bfd38cdb257bfe. Hosted CI on this new head is not claimed complete. No shared registrations, main checkout, Azure identities or stack metadata were changed.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve original trials and enforce the user-amended cumulative 50000-credit ceiling.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve measured protocol-two evidence; bump future specifications to protocol three and stop after failed native domain tools even without SDK error details.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@GeorgeNgMsft George Ng (GeorgeNgMsft) changed the title Add credit-capped GHCP end-to-end evaluation harness Add GHCP end-to-end evaluation harness Sep 25, 2026
@GeorgeNgMsft
George Ng (GeorgeNgMsft) added this pull request to stack #3074 September 25, 2026 01:40
Accept the configured temporary-directory alias and its canonical spelling without admitting other aliases, nested files, or modified artifacts. Use host-native paths in confirmation fixtures.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Promote the updated methodology to the eval README, document initial ballpark scope and isolation limitations, and pin future outer/nested Copilot evaluation to Luna 5.6 with fail-fast ledger identity checks.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
George Ng (GeorgeNgMsft) added a commit that referenced this pull request Sep 25, 2026
Merge the published #3072 migration, keep Luna ledger admission and isolated evaluation package boundaries, and reconcile current protocol-four methodology and A4 provenance.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Version the common file corpus, scope native and typed fixture effects, require fresh readiness, and preserve historical measurements.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent read/search permissions on the fixture directory from bypassing the unresolved-file gate.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
jebrans pushed a commit to jebrans/TypeAgent that referenced this pull request Sep 25, 2026
This standalone product fix addresses intermediate-result handoff
failures surfaced by the GHCP evaluation. Successful actions no longer
fail simply because translation assigned a result label without an
entity, deferred requests remain in the execution queue, and concrete
result references are resolved before their consumers run. The PR
targets `main` directly, without eval-harness changes or a dependency on
microsoft#3072.

- Separate optional entity-name references, explicit `resultValue`
references, and deferred translation context. Missing consumed data
still fails; display text and structured display `rawData` are not
silently substituted as typed values.
- Preserve request-local completed-action snapshots across deferred
continuations, even without conversation history/memory extraction.
Translate only the remaining request without replaying producers or
automatically switching to reasoning.
- Restore concrete-value lookup and strictly validate substituted values
against the consumer schema. Preserve empty values and treat
reference-looking returned strings as data. Reject undeclared/duplicate
references.
- Stop remaining legacy actions while a user choice is pending,
including independent and returned additional actions. Keep the choice
available, explicitly report that remaining steps will not resume
automatically, and disable reasoning fallback. Preserve standalone
choices and structured execution's awaited choice handling.
- Bound deferred context to a 64 KiB serialized UTF-8 envelope
containing the remaining request and full history. Project result data
without execution metadata, display alternates, or identical
representations; retain distinct display data even when history text is
only a summary. Oversized context stops before continuation translation
without truncation, summarization, or replay. This is not a total
model-token budget; schema prompts are separate. Concrete `$result`
substitution remains lossless and outside this limit.
- Preserve the caller's active-schema/schema-family scope during
deferred translation. Recheck execution eligibility before enqueueing
translated continuations; unavailable scope, unknown actions, and
execution-disabled actions cannot silently continue or trigger producer
replay.
- Document the contracts and add 29 handoff cases, 17 bounded-context
cases, 7 deferred-scope cases, and 11 schema-validation cases.

### Evidence and scope

The recorded evaluation contains nine explicit missing-result-entity
errors across `readFile`, `getList`, `clearList`, `prFiles`, and
`prChecks`, each after handler success. One `clearList` mutation
persisted before the failure. Full translated plans were not retained,
so the traces do not establish which labels were unused versus consumed.
This PR fixes the eager invariant and independently reproduced
queue/resolver defects without inventing handler entities or claiming
all nine workflows now complete. No live eval was rerun and no
success-rate improvement is claimed.

Automatic queue resumption after a legacy choice and large-output
retrieval/chunking are intentionally not implemented.

### Validation and review

- The original product patch was isolated onto `main` without changing
its contents; eval commits were removed. The original isolated head
passed the full dispatcher suite (135 suites / 2,152 tests, 12 existing
skips) and action-schema suite (4 suites / 332 tests, 62 existing
skips).
- Two independent adversarial reviews examined base `5d5e23fa6` through
bounded-fix head `725a624a6`. Two newly exposed deferred-path issues
were independently validated: missing execution-eligibility checks and
dropped caller schema restrictions. Both were then corrected in
`2029d4fc1`; six regression failures were reproduced before the
corrections. Review comments were not posted or resolved.
- The latest dependency-aware build and 136 targeted tests across five
suites pass, covering handoffs, size boundaries, scope restrictions,
structured execution, and chained choices. Pinned formatting and all
four ratchets (lint, complexity, circular dependencies, test debt) pass
against `origin/main`.
- Size tests cover exactly 64 KiB, one byte over, UTF-8 encoding,
aggregate/inherited context, empty values, and no producer replay. Scope
tests stop at an offline translator boundary; no model calls are needed.
- Hosted verification for `2029d4fc1029cf4e9e34872a076f5943c88f5983`:
[build-ts run
36187276971](https://github.com/microsoft/TypeAgent/actions/runs/36187276971)
passed all six Windows/Linux/macOS × Node 22/24 jobs. All applicable PR
checks also pass, including Linux/Windows Shell & CLI smoke tests, shell
packaging on all three operating systems, .NET, CodeQL, formatting,
repository policy, and documentation generation. The fork-only
formatting check is skipped as expected.

The original Linux/macOS CI failures were inherited eval portability
issues, fixed separately in microsoft#3072 (`b47f8bacd`), whose six-job OS/Node
matrix is green. microsoft#3072 and microsoft#3077 remain in their own eval stack.

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Resolve request-handler integration while retaining upstream reasoning learning and evaluation tracing. Reconcile the lockfile with merged workspace manifests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

@datduyng Dominic Nguyen (datduyng) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Image

The grader has too much regex more than what. Iwould like to see but this is a good start. please work on cleaning up

@GeorgeNgMsft
George Ng (GeorgeNgMsft) added this pull request to the merge queue Sep 30, 2026
Merged via the queue into main with commit 29408f5 Sep 30, 2026
28 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants