Skip to content

Refine prospective GHCP evaluation recovery, applicability and ambiguity case - #3077

Open
George Ng (GeorgeNgMsft) wants to merge 28 commits into
mainfrom
georgengmsft-ghcp-eval-followup-fixes
Open

George Ng (GeorgeNgMsft) wants to merge 28 commits into
mainfrom
georgengmsft-ghcp-eval-followup-fixes

Conversation

@GeorgeNgMsft

@GeorgeNgMsft George Ng (GeorgeNgMsft) commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

This follow-up above #3072 adds bounded safe-read recovery, frozen-run checks and scoped file-handler consent, integrated with the protocol-5 common-file corpus and Luna-pinned evaluation package. It preserves seven candidate semantics and four five-case cohorts without rewriting historical measurements. The independent result-handoff fixes in #3073 are not a dependency.

  • Integrate parent f2f045c20 through a normal merge, preserving bot history. common-files-v1 replaces nine list-dependent cases with common-capability file tasks: all 20 cases apply to every candidate, yielding 140 executions per full repetition. This supersedes protocol 4's 131 executions plus nine native N/A slots, not historical scores.
  • Derive counts from the frozen schedule and reject changed protocols, corpus versions or orders. Require fresh corpus-marked preflight evidence; preserve file snapshots, byte-exact backups, unrelated-state checks and failed-trial evidence.
  • Retain narrow per-case permissions, clarification gates and M3's original-backup prerequisite. No arbitrary shell-write approvals or measured list mutations. Preserve historical A4's original prompt, failure and replacement rationale without claiming general product ambiguity is solved.
  • Recover only from affirmative read-I/O evidence plus a complete read-only backend trace for TypeAgent. Write/copy failures, denials, cancellations, uncertainty, missing details and conflicting error fields remain terminal. No retry loop or expanded artifact access.
  • Handle the file handler's second Run/Cancel consent after dispatcher confirmation. Require the exact prompt/choices, unambiguous current-tool admitted action and independent case oracle. Bind structured consent to scope/operation/interaction, consume once, invalidate on unrelated/terminal work, and compare response fields independent of JSON key order. Consent never supplies an ambiguous referent.
  • Preserve the sibling evaluation package, short plugin README link, initial-ballpark/environment-isolation caveats, Luna pin, early ledger admission and ledger-derived limits.

Validation: dependency-inclusive build and 137 dispatcher tests passed during integration. Final consent corrections pass 34 evaluation tests, package syntax build, pinned formatting and all four committed-head ratchets against 79fad5b4f: lint 0→0, cyclomatic/cognitive over-budget counts 0→0, circular dependencies 220→220 with 11 existing exceptions, test debt 0→0. Fresh preflight completed actual inventory/read/copy/append and read-only GitHub/network operations. Structured S4 pilot verified both consent stages and exact state preservation. Other pilot denials remain failures; pilot evidence is separate from measured outcomes. Hosted CI for the latest head is not yet claimed complete.

Fresh live round — paused after 21/140 trial records: the user separately authorized 40,000 additional Copilot AI credits. Execution is frozen at 96e904546acc389cf64284db9b6fbcbe57191344. Three complete seven-candidate batches ran for M5, M1 and M4; 119 trials remain unstarted. The runner paused before batch four because a timed-out M4/C3 request lacks billing settlement. Its full 250-credit reservation remains held; no automatic retry or trial replay occurred. This is a reconciliation pause, not budget exhaustion. Partial evidence and accounting snapshots are preserved privately. Independent faithfulness/accuracy review is outstanding; no new accuracy or latency conclusion is claimed.

Historical ledgers remain closed and the protocol-2 run at e847c7a009 remains unchanged. Private network payloads, artifact paths and billing traces are not published. Stack #3080 remains #3072 followed by this PR; no stack metadata changes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve original trials and enforce the user-amended cumulative 50000-credit ceiling.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve measured protocol-two evidence; bump future specifications to protocol three and stop after failed native domain tools even without SDK error details.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@GeorgeNgMsft
George Ng (GeorgeNgMsft) added this pull request to stack #3074 September 25, 2026 09:34
Accept the configured temporary-directory alias and its canonical spelling without admitting other aliases, nested files, or modified artifacts. Use host-native paths in confirmation fixtures.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep safe read recovery fail-closed, skip native list slots explicitly, and replace A4 with an unresolved-item clarification case without revising frozen measurements.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Derive runnable trial counts from the frozen paired schedule and reject old protocol resumes. Keep explicit cancellation and conflicting denial codes terminal even with ordinary I/O text.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@GeorgeNgMsft
George Ng (GeorgeNgMsft) force-pushed the georgengmsft-ghcp-eval-followup-fixes branch from 3dfdee5 to 5411499 Compare September 25, 2026 19:24
@GeorgeNgMsft
George Ng (GeorgeNgMsft) removed this pull request from stack #3074 September 25, 2026 19:26
@GeorgeNgMsft
George Ng (GeorgeNgMsft) changed the base branch from georgengmsft-result-entity-handoff-fixes to georgengmsft-ghcp-eval-implementation September 25, 2026 19:26
@GeorgeNgMsft
George Ng (GeorgeNgMsft) added this pull request to stack #3080 September 25, 2026 19:27
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Promote the updated methodology to the eval README, document initial ballpark scope and isolation limitations, and pin future outer/nested Copilot evaluation to Luna 5.6 with fail-fast ledger identity checks.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Merge the published #3072 migration, keep Luna ledger admission and isolated evaluation package boundaries, and reconcile current protocol-four methodology and A4 provenance.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep generated license and repository metadata and describe the evaluation package accurately after repository normalization.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Version the common file corpus, scope native and typed fixture effects, require fresh readiness, and preserve historical measurements.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent read/search permissions on the fixture directory from bypassing the unresolved-file gate.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Replace protocol-four native applicability and list oracles with protocol-five common-files-v1. Keep frozen scheduling, Luna admission and fail-closed recovery across the new file actions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
jebrans pushed a commit to jebrans/TypeAgent that referenced this pull request Sep 25, 2026
This standalone product fix addresses intermediate-result handoff
failures surfaced by the GHCP evaluation. Successful actions no longer
fail simply because translation assigned a result label without an
entity, deferred requests remain in the execution queue, and concrete
result references are resolved before their consumers run. The PR
targets `main` directly, without eval-harness changes or a dependency on
microsoft#3072.

- Separate optional entity-name references, explicit `resultValue`
references, and deferred translation context. Missing consumed data
still fails; display text and structured display `rawData` are not
silently substituted as typed values.
- Preserve request-local completed-action snapshots across deferred
continuations, even without conversation history/memory extraction.
Translate only the remaining request without replaying producers or
automatically switching to reasoning.
- Restore concrete-value lookup and strictly validate substituted values
against the consumer schema. Preserve empty values and treat
reference-looking returned strings as data. Reject undeclared/duplicate
references.
- Stop remaining legacy actions while a user choice is pending,
including independent and returned additional actions. Keep the choice
available, explicitly report that remaining steps will not resume
automatically, and disable reasoning fallback. Preserve standalone
choices and structured execution's awaited choice handling.
- Bound deferred context to a 64 KiB serialized UTF-8 envelope
containing the remaining request and full history. Project result data
without execution metadata, display alternates, or identical
representations; retain distinct display data even when history text is
only a summary. Oversized context stops before continuation translation
without truncation, summarization, or replay. This is not a total
model-token budget; schema prompts are separate. Concrete `$result`
substitution remains lossless and outside this limit.
- Preserve the caller's active-schema/schema-family scope during
deferred translation. Recheck execution eligibility before enqueueing
translated continuations; unavailable scope, unknown actions, and
execution-disabled actions cannot silently continue or trigger producer
replay.
- Document the contracts and add 29 handoff cases, 17 bounded-context
cases, 7 deferred-scope cases, and 11 schema-validation cases.

### Evidence and scope

The recorded evaluation contains nine explicit missing-result-entity
errors across `readFile`, `getList`, `clearList`, `prFiles`, and
`prChecks`, each after handler success. One `clearList` mutation
persisted before the failure. Full translated plans were not retained,
so the traces do not establish which labels were unused versus consumed.
This PR fixes the eager invariant and independently reproduced
queue/resolver defects without inventing handler entities or claiming
all nine workflows now complete. No live eval was rerun and no
success-rate improvement is claimed.

Automatic queue resumption after a legacy choice and large-output
retrieval/chunking are intentionally not implemented.

### Validation and review

- The original product patch was isolated onto `main` without changing
its contents; eval commits were removed. The original isolated head
passed the full dispatcher suite (135 suites / 2,152 tests, 12 existing
skips) and action-schema suite (4 suites / 332 tests, 62 existing
skips).
- Two independent adversarial reviews examined base `5d5e23fa6` through
bounded-fix head `725a624a6`. Two newly exposed deferred-path issues
were independently validated: missing execution-eligibility checks and
dropped caller schema restrictions. Both were then corrected in
`2029d4fc1`; six regression failures were reproduced before the
corrections. Review comments were not posted or resolved.
- The latest dependency-aware build and 136 targeted tests across five
suites pass, covering handoffs, size boundaries, scope restrictions,
structured execution, and chained choices. Pinned formatting and all
four ratchets (lint, complexity, circular dependencies, test debt) pass
against `origin/main`.
- Size tests cover exactly 64 KiB, one byte over, UTF-8 encoding,
aggregate/inherited context, empty values, and no producer replay. Scope
tests stop at an offline translator boundary; no model calls are needed.
- Hosted verification for `2029d4fc1029cf4e9e34872a076f5943c88f5983`:
[build-ts run
36187276971](https://github.com/microsoft/TypeAgent/actions/runs/36187276971)
passed all six Windows/Linux/macOS × Node 22/24 jobs. All applicable PR
checks also pass, including Linux/Windows Shell & CLI smoke tests, shell
packaging on all three operating systems, .NET, CodeQL, formatting,
repository policy, and documentation generation. The fork-only
formatting check is skipped as expected.

The original Linux/macOS CI failures were inherited eval portability
issues, fixed separately in microsoft#3072 (`b47f8bacd`), whose six-job OS/Node
matrix is green. microsoft#3072 and microsoft#3077 remain in their own eval stack.

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Bind Run/Cancel approvals to admitted fixture actions and single-use structured interactions. Preserve clarification and terminal-stop rules; record separately authorized fresh evaluation allowance.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep common-files at 20 cases across seven candidates; add a separate 20-case lists corpus for C1-C4 with category-scoped policy, reset, readiness, oracles and reporting.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Fix filename clarification and private consent-context redaction. Enforce exact candidate tool routes with runtime metadata readiness, pending contract guards and audits. Add sanitized denial diagnostics and offline SDK regressions; preserve frozen measured results under prospective protocol 7.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Base automatically changed from georgengmsft-ghcp-eval-implementation to main September 30, 2026 04:33

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants