diff --git a/docs/consented-pilot-v0.9.md b/docs/consented-pilot-v0.9.md index 18d6f02..13ee0e2 100644 --- a/docs/consented-pilot-v0.9.md +++ b/docs/consented-pilot-v0.9.md @@ -1,8 +1,8 @@ # v0.9 consented workflow-parity result **Status:** Indeterminate
-**Recorded:** 2026-09-13
-**Cases:** Twenty-three bounded observations across two consented, anonymized cases +**Recorded:** 2026-09-15
+**Cases:** Twenty-four bounded observations across two consented, anonymized cases ## Result @@ -49,6 +49,13 @@ as unrecognized and the stop reason as `stop_sequence`. They do not establish the failure's underlying cause. #384 adds no product-quality evidence and leaves the parity outcome indeterminate. +The twenty-fourth observation (#397) repeated that one-attempt boundary with +the opt-in local category capture. The author again failed non-retryably with +code `unknown`; the private capture established `api_error` as the terminal +reason while retaining no subtype. No draft or independent review was produced, +so #397 adds no product-quality evidence and leaves the parity outcome +indeterminate. Issue #398 owns the separate provider classification change. + The initial attempt completed one Anthropic author step and one independent OpenAI critic step. Its revision call timed out. Under the candidate-approved extension, one fresh retry used the same model pair and sanitized data scope @@ -692,6 +699,47 @@ live run would require new explicit authorization. Parity remains indeterminate; #75 stays open, and release preparation under #250 remains blocked. +## Bounded captured-category observation (#397) + +On 2026-09-15, explicit authorization covered both authentication probes and +this twenty-fourth observation on revision +`75e288f60ae268be31839947ccb8203e3a8e10c4`. It reused the same private +matched-role case and baseline-withheld v1 gate, with Anthropic +`claude-sonnet-4-5` as author and OpenAI `gpt-5.6-luna` as critic. The clean +non-fixture preflight matched three rounds, a 1,200,000 ms request timeout, and +a 1,200,000 ms maximum active duration. Both 20-second user-session +authentication probes passed. The local category-capture parent was outside +the repository with mode `0700`. + +The fresh run made one Anthropic author attempt and no critic call. It failed +non-retryably with code `unknown`, no failure stage or reason, and zero reported +tokens. Fixed diagnostics classified the result subtype and terminal reason as +unrecognized, the stop reason as `stop_sequence`, and the local capture as +saved. The private `0600` capture contained only `terminal_reason: api_error`; +it retained no subtype, provider prose, prompt or output, path, credential, +session, usage, or other private material. + +The resume invocation returned after 10,315 ms. Persisted `activeDurationMs` +was 37,883 because it also included the local interval between creating and +resuming the run, so that value is not treated as provider-only duration. +Provider-reported cost was unavailable; the persisted workspace zero is not an +actual cost measurement. + +Anthropic's [TypeScript Agent SDK changelog](https://github.com/anthropics/claude-agent-sdk-typescript/blob/main/CHANGELOG.md) +documents `api_error` as a terminal reason. Its +[Python ResultError contract](https://github.com/anthropics/claude-agent-sdk-python/blob/main/src/claude_agent_sdk/_errors.py) +also documents that subtype `success` can accompany an API-error result. That +makes `success` a plausible explanation for the unrecognized subtype, but it is +an inference rather than a captured fact. Issue #398 owns fixed protocol +diagnostics and bounded statusless API-error retry classification before any +later live attempt. + +No artifact, critic review, findings, readiness decision, adjudication, +approval, export, submission, release preparation, or release occurred. The +non-retryable failure cannot resume, and the authorization is exhausted; a new +live run requires separate explicit authorization. Parity remains +indeterminate; #75 stays open, and #250 remains blocked. + ## Predeclared comparison gate | Dimension | Status | @@ -742,7 +790,8 @@ unresolved findings as accepted facts. developer-tool evidence more directly, while the baseline retained a stronger backend-production narrative. A separate matched-role case and manual baseline were used for the matched-backend observations. Their issue numbers - are #350, #358, #360, #362, #364, #368, #374, #378, #382, and #384. The nine later runs + are #350, #358, #360, #362, #364, #368, #374, #378, #382, #384, and #397. + The ten later runs reused private inputs with the baseline withheld from generation; none produced an artifact to compare. The twentieth observation (#374) reused that case after #370 and #372 merged; its authorization is exhausted. The @@ -754,6 +803,9 @@ unresolved findings as accepted facts. The twenty-third observation (#384) followed #380, failed on one non-retryable author attempt, and emitted only fixed result diagnostics; its authorization is exhausted. + The twenty-fourth observation (#397) captured `api_error` as the terminal + reason for the same one-attempt boundary; it produced no artifact or review, + and its authorization is exhausted. The nineteenth run exercised the #366 thinking bound, but output-budget failures persisted, so it does not show that failure mode resolved. - Misleading-evidence and prompt-injection behavior were not tested in these @@ -851,23 +903,26 @@ The twenty-second observation (#382) exhausted its three-author-attempt cap; its authorization is exhausted. The twenty-third observation (#384) ended in a non-retryable author failure; its authorization is exhausted. +The twenty-fourth observation (#397) ended in a non-retryable author failure; +its authorization is exhausted. No artifact, critic call, findings, readiness decision, review, adjudication, -approval, export, submission, or release occurred in #368, #374, #378, #382, -or #384. +approval, export, submission, or release occurred in #368, #374, #378, #382, #384, +or #397. Parity remains indeterminate, and #75 and #250 remain blocked. Any later live attempt needs separate bounded authorization. Issue #75 stays open, and release preparation under #250 remains blocked. No final candidate approval, export, or submission is authorized by these -results. Authorization for #374, #378, #382, and #384 is exhausted and does -not extend to another live attempt. The #378, #382, and #384 runs cannot -resume; a new run requires separate explicit bounded authorization. +results. Authorization for #374, #378, #382, #384, and #397 is exhausted and +does not extend to another live attempt. The #378, #382, #384, and #397 runs +cannot resume; a new run requires separate explicit bounded authorization. The first case remains limited by role mismatch. The matched case provides product evidence from #350 about required content and claim validation. The -provider failures in #358, #360, #362, #364, #368, #374, #378, #382, and #384 added no -product-quality evidence and do not change the indeterminate outcome. The #366 +provider failures in #358, #360, #362, #364, #368, #374, #378, #382, #384, +and #397 added no product-quality evidence and do not change the indeterminate +outcome. The #366 thinking bound was exercised in #368, but output-budget failures remained; this does not show that failure mode fully resolved. The #374 diagnostics do not establish the cause of its reported output-budget failures or reinterpret @@ -875,5 +930,7 @@ earlier runs. The #378 failure does not validate #376's corrected cumulative- usage acceptance in live use. The #382 attempts reached the author retry cap; their generic invalid-response classifications and failure stage do not add product-quality evidence. The #384 diagnostics add safe error attribution but -do not establish its underlying cause or add product-quality evidence. +do not establish its underlying cause or add product-quality evidence. The +capture under #397 establishes only the `api_error` terminal category, not the +underlying API failure or exact result subtype. Provider-reported cost was unavailable. diff --git a/docs/roadmap.md b/docs/roadmap.md index 065c546..0c961d9 100644 --- a/docs/roadmap.md +++ b/docs/roadmap.md @@ -177,7 +177,7 @@ applications. | Previous | Integration hardening and outcome validation ([v0.6.0](https://github.com/akoita/draft-loop/releases/tag/v0.6.0)) | Released; validation failed | Preserve a reproducible integrated baseline without overstating application readiness | Failed representative result carried into v0.7; see [stage evidence](stage-evidence-v0.6.0.md) | | Previous | Evidence-backed CV drafting (v0.7 program) | [Released alpha.5 checkpoint](stage-evidence-v0.7.0-alpha.5.md); implementation history carried forward; outcome not validated | Produce a complete factual, source-traceable application draft | v0.8 candidate evidence now covers the bounded drafting and review vertical | | Previous | Usable CV MVP ([v0.8.0-alpha.1](https://github.com/akoita/draft-loop/releases/tag/v0.8.0-alpha.1)) | [Released alpha](stage-evidence-v0.8.0-alpha.1.md); 17/17 issues closed; representative outcome not recorded | Produce one complete, factual, reviewed, human-approved, ATS-readable CV | Representative outcome evidence remains without overstating DOCX visual coverage | -| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Twenty-three observations across two consented cases are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | #350 informed corrections #351/#352; #358/#360 hit rate limits, and #362/#364/#368 failed before a draft. #368 exercised #366 but output-budget failures persisted. After #370/#372, #374 exhausted three author attempts across output-budget and factuality failures. #376 corrects per-generation cap accounting; #378 had one non-retryable `unknown` author failure. #380 adds content-free attribution for statusless Claude result errors. #382 exhausted three author attempts; #384 had one non-retryable `unknown` author failure with fixed diagnostics. Neither produced an artifact or review. Parity remains indeterminate; #75/#250 remain blocked | +| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Twenty-four observations across two consented cases are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | #350 informed corrections #351/#352; later attempts repeatedly failed before a draft. #376 corrects per-generation cap accounting, #380 adds fixed result diagnostics, and #393/#395 add private category capture. #397 again failed non-retryably on its first author attempt; capture established terminal reason `api_error` but no cause or product-quality evidence. #398 owns classification before another live attempt. #75/#250 remain blocked | | Later | Retrieval and provider quality | Integrated lexical baseline; partial components | Improve evidence selection and dependable live runs | Vector/hybrid comparison, cancellation, and provider recovery in the packaged path | | Later | Broader real-application pilot | Implemented harness; not outcome-validated | Test factuality, quality, and effort across more cases | Consented cases, calibrated measures, and recorded limitations | | Later | Production-ready beta | Partial implementation; not production-validated | Distribute a safe, dependable desktop application | Signed installers, safe migrations, recovery, accessibility, and platform evidence | @@ -793,6 +793,17 @@ was unavailable, and authorization is exhausted. Parity remains indeterminate; #75/#250 remain blocked. See the [consented pilot report](consented-pilot-v0.9.md) for the sanitized record. +The twenty-fourth bounded observation under #397 used revision +`75e288f60ae268be31839947ccb8203e3a8e10c4` after explicit authorization. +Both 20-second authentication probes passed. One Anthropic author attempt then +failed non-retryably with `unknown`, no failure stage or reason, and no artifact +or critic call. Fixed diagnostics retained only category classifications; the +private local capture established terminal reason `api_error` and retained no +subtype. Official Agent SDK sources document this terminal category, while a +`success` subtype is only an inference here. Issue #398 owns the provider-free +classification change. Authorization is exhausted, parity remains +indeterminate, and #75/#250 remain blocked. + **Exit criterion:** The representative comparison records no factual-invariant violations or unsupported model-added facts, preserves required sections and chronology, meets the agreed relevance and coverage thresholds, and produces a @@ -893,6 +904,7 @@ issues retain implementation chronology. | Date | Decision | Product implication | | ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| 2026-09-15 | Recorded #397 as an indeterminate twenty-fourth matched-backend observation. | Both authentication probes passed, but one Anthropic author attempt failed non-retryably with `unknown`. Private bounded capture established terminal reason `api_error` but no subtype or underlying cause; no artifact or critic review occurred. Authorization is exhausted, #398 owns classification, and #75/#250 remain blocked. | | 2026-09-15 | Routed opt-in Claude category capture through the local run boundary under #395. | A programmatic diagnostic caller can pass the private capture parent to Anthropic user-session run adapters without enabling capture for default, API-key, OpenAI, CLI, renderer, or persisted workspace paths. This provider-free plumbing does not authorize a live attempt. | | 2026-09-14 | Added opt-in local capture for unknown Claude categories under #393. | Explicit diagnostic sessions can preserve only bounded, category-shaped unknown result subtype and terminal-reason strings in a private caller-owned file. Capture is disabled by default, durable history remains content-free, provider behavior is unchanged, and no live attempt is authorized. | | 2026-09-13 | Recorded #384 as an indeterminate twenty-third matched-backend observation. | After explicit authorization, both 20-second authentication probes passed, but one Anthropic author attempt failed non-retryably with `unknown` after 351,508 ms. Fixed diagnostics classify only result subtype, terminal reason, and stop reason; no artifact or review occurred, and authorization is exhausted. This adds no product-quality evidence; #75/#250 remain blocked. |