Add GHCP end-to-end evaluation harness - #3072
Merged
George Ng (GeorgeNgMsft) merged 15 commits intoSep 30, 2026
Merged
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve original trials and enforce the user-amended cumulative 50000-credit ceiling. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve measured protocol-two evidence; bump future specifications to protocol three and stop after failed native domain tools even without SDK error details. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
George Ng (GeorgeNgMsft)
added this pull request to stack #3074
September 25, 2026 01:40
Accept the configured temporary-directory alias and its canonical spelling without admitting other aliases, nested files, or modified artifacts. Use host-native paths in confirmation fixtures. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
George Ng (GeorgeNgMsft)
removed this pull request from stack #3074
September 25, 2026 19:26
George Ng (GeorgeNgMsft)
added this pull request to stack #3080
September 25, 2026 19:27
Promote the updated methodology to the eval README, document initial ballpark scope and isolation limitations, and pin future outer/nested Copilot evaluation to Luna 5.6 with fail-fast ledger identity checks. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
George Ng (GeorgeNgMsft)
added a commit
that referenced
this pull request
Sep 25, 2026
Merge the published #3072 migration, keep Luna ledger admission and isolated evaluation package boundaries, and reconcile current protocol-four methodology and A4 provenance. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Version the common file corpus, scope native and typed fixture effects, require fresh readiness, and preserve historical measurements. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent read/search permissions on the fixture directory from bypassing the unresolved-file gate. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
jebrans
pushed a commit
to jebrans/TypeAgent
that referenced
this pull request
Sep 25, 2026
This standalone product fix addresses intermediate-result handoff failures surfaced by the GHCP evaluation. Successful actions no longer fail simply because translation assigned a result label without an entity, deferred requests remain in the execution queue, and concrete result references are resolved before their consumers run. The PR targets `main` directly, without eval-harness changes or a dependency on microsoft#3072. - Separate optional entity-name references, explicit `resultValue` references, and deferred translation context. Missing consumed data still fails; display text and structured display `rawData` are not silently substituted as typed values. - Preserve request-local completed-action snapshots across deferred continuations, even without conversation history/memory extraction. Translate only the remaining request without replaying producers or automatically switching to reasoning. - Restore concrete-value lookup and strictly validate substituted values against the consumer schema. Preserve empty values and treat reference-looking returned strings as data. Reject undeclared/duplicate references. - Stop remaining legacy actions while a user choice is pending, including independent and returned additional actions. Keep the choice available, explicitly report that remaining steps will not resume automatically, and disable reasoning fallback. Preserve standalone choices and structured execution's awaited choice handling. - Bound deferred context to a 64 KiB serialized UTF-8 envelope containing the remaining request and full history. Project result data without execution metadata, display alternates, or identical representations; retain distinct display data even when history text is only a summary. Oversized context stops before continuation translation without truncation, summarization, or replay. This is not a total model-token budget; schema prompts are separate. Concrete `$result` substitution remains lossless and outside this limit. - Preserve the caller's active-schema/schema-family scope during deferred translation. Recheck execution eligibility before enqueueing translated continuations; unavailable scope, unknown actions, and execution-disabled actions cannot silently continue or trigger producer replay. - Document the contracts and add 29 handoff cases, 17 bounded-context cases, 7 deferred-scope cases, and 11 schema-validation cases. ### Evidence and scope The recorded evaluation contains nine explicit missing-result-entity errors across `readFile`, `getList`, `clearList`, `prFiles`, and `prChecks`, each after handler success. One `clearList` mutation persisted before the failure. Full translated plans were not retained, so the traces do not establish which labels were unused versus consumed. This PR fixes the eager invariant and independently reproduced queue/resolver defects without inventing handler entities or claiming all nine workflows now complete. No live eval was rerun and no success-rate improvement is claimed. Automatic queue resumption after a legacy choice and large-output retrieval/chunking are intentionally not implemented. ### Validation and review - The original product patch was isolated onto `main` without changing its contents; eval commits were removed. The original isolated head passed the full dispatcher suite (135 suites / 2,152 tests, 12 existing skips) and action-schema suite (4 suites / 332 tests, 62 existing skips). - Two independent adversarial reviews examined base `5d5e23fa6` through bounded-fix head `725a624a6`. Two newly exposed deferred-path issues were independently validated: missing execution-eligibility checks and dropped caller schema restrictions. Both were then corrected in `2029d4fc1`; six regression failures were reproduced before the corrections. Review comments were not posted or resolved. - The latest dependency-aware build and 136 targeted tests across five suites pass, covering handoffs, size boundaries, scope restrictions, structured execution, and chained choices. Pinned formatting and all four ratchets (lint, complexity, circular dependencies, test debt) pass against `origin/main`. - Size tests cover exactly 64 KiB, one byte over, UTF-8 encoding, aggregate/inherited context, empty values, and no producer replay. Scope tests stop at an offline translator boundary; no model calls are needed. - Hosted verification for `2029d4fc1029cf4e9e34872a076f5943c88f5983`: [build-ts run 36187276971](https://github.com/microsoft/TypeAgent/actions/runs/36187276971) passed all six Windows/Linux/macOS × Node 22/24 jobs. All applicable PR checks also pass, including Linux/Windows Shell & CLI smoke tests, shell packaging on all three operating systems, .NET, CodeQL, formatting, repository policy, and documentation generation. The fork-only formatting check is skipped as expected. The original Linux/macOS CI failures were inherited eval portability issues, fixed separately in microsoft#3072 (`b47f8bacd`), whose six-job OS/Node matrix is green. microsoft#3072 and microsoft#3077 remain in their own eval stack. --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Resolve request-handler integration while retaining upstream reasoning learning and evaluation tracing. Reconcile the lockfile with merged workspace manifests. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Dominic Nguyen (datduyng)
approved these changes
Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Adds a reproducible Copilot end-to-end evaluation in the dedicated
copilot-plugin-evalpackage, comparing TypeAgent routing strategies with native Copilot. This is a simple initial evaluation for ballpark estimates, not a statistically powered benchmark. The latest corpus replaces list-dependent tasks with ordinary file tasks that all candidates can perform.common-files-v1) replaces nine list-dependent cases with file inventory, append, backup/overwrite, comparison, conditional edits and clarification tasks. All candidates have twenty applicable cases: 140 executions per repetition, no native N/A slots.edit/createtools for Copilot. Per-case fixture permissions, pre-clarification gates, exact backup prerequisites and independent file-state checks protect unrelated state. Arbitrary shell writes are not autoapproved; GitHub and network remain read-only.gpt-5.6-luna). Rejects mismatched ledgers and stale preflights; freezes fresh corpus/model/order specifications rather than relabeling historical runs.ts/packages/copilot-plugin-eval; runtime guards remain at dispatcher enforcement boundaries. Explicitly documents residual host/plugin/environment-isolation risks: a separate directory is not a sandbox.Historical evidence, not new file-corpus measurements
The completed 140-trial historical pass used protocol 2 at
e847c7a00907a1f4e2c6c4426c935b964fd9997cwithgpt-5.6-soland the original list workload. All combinations were attempted; not all succeeded. Original strict successes were 6, 6, 11, 11, 6, 8, 3 /20 for C1-C7. The later protocol-3 native terminal-gate correction was tested offline, not live-measured.Removing only the blanket continuation penalty produced retrospective content scores of 6, 6, 11, 11, 7, 8, 8 /20. This is not a new execution or certification of recovery safety: some native failure evidence is incomplete, and network final-presentation failures remain. Excluding nine list-dependent native cases gave a historical native 8/11, not a matched-workload ranking. Original success-conditioned timing populations and uncertainty remain separate; these observations establish neither a general winner nor an internal-fallback benefit.
No Luna or common-files-v1 live evaluation has run. New corpus outcomes must not be pooled with old results. Fresh readiness, model-specific accounting bounds and authorization remain prerequisites. The maintained methodology records the approved corpus amendment; historical reports, sanitized exports, original grades, frozen specifications and accounting artifacts remain external and unchanged. Private network evidence, credentials, billing traces and machine-specific paths are not committed.
Latest offline validation
origin/main: no added lint/complexity violations, cycles unchanged at 220, no focused/newly skipped tests.f2f045c2066605f69c990f3b99bfd38cdb257bfe. Hosted CI on this new head is not claimed complete. No shared registrations, main checkout, Azure identities or stack metadata were changed.