Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
79 changes: 68 additions & 11 deletions docs/consented-pilot-v0.9.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# v0.9 consented workflow-parity result

**Status:** Indeterminate<br>
**Recorded:** 2026-09-13<br>
**Cases:** Twenty-three bounded observations across two consented, anonymized cases
**Recorded:** 2026-09-15<br>
**Cases:** Twenty-four bounded observations across two consented, anonymized cases

## Result

Expand Down Expand Up @@ -49,6 +49,13 @@ as unrecognized and the stop reason as `stop_sequence`. They do not establish
the failure's underlying cause. #384 adds no product-quality evidence and
leaves the parity outcome indeterminate.

The twenty-fourth observation (#397) repeated that one-attempt boundary with
the opt-in local category capture. The author again failed non-retryably with
code `unknown`; the private capture established `api_error` as the terminal
reason while retaining no subtype. No draft or independent review was produced,
so #397 adds no product-quality evidence and leaves the parity outcome
indeterminate. Issue #398 owns the separate provider classification change.

The initial attempt completed one Anthropic author step and one independent
OpenAI critic step. Its revision call timed out. Under the candidate-approved
extension, one fresh retry used the same model pair and sanitized data scope
Expand Down Expand Up @@ -692,6 +699,47 @@ live run would require new explicit authorization. Parity remains
indeterminate; #75 stays open, and release preparation under #250 remains
blocked.

## Bounded captured-category observation (#397)

On 2026-09-15, explicit authorization covered both authentication probes and
this twenty-fourth observation on revision
`75e288f60ae268be31839947ccb8203e3a8e10c4`. It reused the same private
matched-role case and baseline-withheld v1 gate, with Anthropic
`claude-sonnet-4-5` as author and OpenAI `gpt-5.6-luna` as critic. The clean
non-fixture preflight matched three rounds, a 1,200,000 ms request timeout, and
a 1,200,000 ms maximum active duration. Both 20-second user-session
authentication probes passed. The local category-capture parent was outside
the repository with mode `0700`.

The fresh run made one Anthropic author attempt and no critic call. It failed
non-retryably with code `unknown`, no failure stage or reason, and zero reported
tokens. Fixed diagnostics classified the result subtype and terminal reason as
unrecognized, the stop reason as `stop_sequence`, and the local capture as
saved. The private `0600` capture contained only `terminal_reason: api_error`;
it retained no subtype, provider prose, prompt or output, path, credential,
session, usage, or other private material.

The resume invocation returned after 10,315 ms. Persisted `activeDurationMs`
was 37,883 because it also included the local interval between creating and
resuming the run, so that value is not treated as provider-only duration.
Provider-reported cost was unavailable; the persisted workspace zero is not an
actual cost measurement.

Anthropic's [TypeScript Agent SDK changelog](https://github.com/anthropics/claude-agent-sdk-typescript/blob/main/CHANGELOG.md)
documents `api_error` as a terminal reason. Its
[Python ResultError contract](https://github.com/anthropics/claude-agent-sdk-python/blob/main/src/claude_agent_sdk/_errors.py)
also documents that subtype `success` can accompany an API-error result. That
makes `success` a plausible explanation for the unrecognized subtype, but it is
an inference rather than a captured fact. Issue #398 owns fixed protocol
diagnostics and bounded statusless API-error retry classification before any
later live attempt.

No artifact, critic review, findings, readiness decision, adjudication,
approval, export, submission, release preparation, or release occurred. The
non-retryable failure cannot resume, and the authorization is exhausted; a new
live run requires separate explicit authorization. Parity remains
indeterminate; #75 stays open, and #250 remains blocked.

## Predeclared comparison gate

| Dimension | Status |
Expand Down Expand Up @@ -742,7 +790,8 @@ unresolved findings as accepted facts.
developer-tool evidence more directly, while the baseline retained a stronger
backend-production narrative. A separate matched-role case and manual
baseline were used for the matched-backend observations. Their issue numbers
are #350, #358, #360, #362, #364, #368, #374, #378, #382, and #384. The nine later runs
are #350, #358, #360, #362, #364, #368, #374, #378, #382, #384, and #397.
The ten later runs
reused private inputs with the baseline withheld from generation; none
produced an artifact to compare. The twentieth observation (#374) reused
that case after #370 and #372 merged; its authorization is exhausted. The
Expand All @@ -754,6 +803,9 @@ unresolved findings as accepted facts.
The twenty-third observation (#384) followed #380, failed on one
non-retryable author attempt, and emitted only fixed result diagnostics; its
authorization is exhausted.
The twenty-fourth observation (#397) captured `api_error` as the terminal
reason for the same one-attempt boundary; it produced no artifact or review,
and its authorization is exhausted.
The nineteenth run exercised the #366 thinking bound, but output-budget
failures persisted, so it does not show that failure mode resolved.
- Misleading-evidence and prompt-injection behavior were not tested in these
Expand Down Expand Up @@ -851,29 +903,34 @@ The twenty-second observation (#382) exhausted its three-author-attempt cap;
its authorization is exhausted.
The twenty-third observation (#384) ended in a non-retryable author failure;
its authorization is exhausted.
The twenty-fourth observation (#397) ended in a non-retryable author failure;
its authorization is exhausted.

No artifact, critic call, findings, readiness decision, review, adjudication,
approval, export, submission, or release occurred in #368, #374, #378, #382,
or #384.
approval, export, submission, or release occurred in #368, #374, #378, #382, #384,
or #397.
Parity remains indeterminate, and #75 and #250 remain blocked. Any later live
attempt needs separate bounded authorization.

Issue #75 stays open, and release preparation under #250 remains blocked.
No final candidate approval, export, or submission is authorized by these
results. Authorization for #374, #378, #382, and #384 is exhausted and does
not extend to another live attempt. The #378, #382, and #384 runs cannot
resume; a new run requires separate explicit bounded authorization.
results. Authorization for #374, #378, #382, #384, and #397 is exhausted and
does not extend to another live attempt. The #378, #382, #384, and #397 runs
cannot resume; a new run requires separate explicit bounded authorization.

The first case remains limited by role mismatch. The matched case provides
product evidence from #350 about required content and claim validation. The
provider failures in #358, #360, #362, #364, #368, #374, #378, #382, and #384 added no
product-quality evidence and do not change the indeterminate outcome. The #366
provider failures in #358, #360, #362, #364, #368, #374, #378, #382, #384,
and #397 added no product-quality evidence and do not change the indeterminate
outcome. The #366
thinking bound was exercised in #368, but output-budget failures remained;
this does not show that failure mode fully resolved. The #374 diagnostics do
not establish the cause of its reported output-budget failures or reinterpret
earlier runs. The #378 failure does not validate #376's corrected cumulative-
usage acceptance in live use. The #382 attempts reached the author retry cap;
their generic invalid-response classifications and failure stage do not add
product-quality evidence. The #384 diagnostics add safe error attribution but
do not establish its underlying cause or add product-quality evidence.
do not establish its underlying cause or add product-quality evidence. The
capture under #397 establishes only the `api_error` terminal category, not the
underlying API failure or exact result subtype.
Provider-reported cost was unavailable.
14 changes: 13 additions & 1 deletion docs/roadmap.md
Original file line number Diff line number Diff line change
Expand Up @@ -177,7 +177,7 @@ applications.
| Previous | Integration hardening and outcome validation ([v0.6.0](https://github.com/akoita/draft-loop/releases/tag/v0.6.0)) | Released; validation failed | Preserve a reproducible integrated baseline without overstating application readiness | Failed representative result carried into v0.7; see [stage evidence](stage-evidence-v0.6.0.md) |
| Previous | Evidence-backed CV drafting (v0.7 program) | [Released alpha.5 checkpoint](stage-evidence-v0.7.0-alpha.5.md); implementation history carried forward; outcome not validated | Produce a complete factual, source-traceable application draft | v0.8 candidate evidence now covers the bounded drafting and review vertical |
| Previous | Usable CV MVP ([v0.8.0-alpha.1](https://github.com/akoita/draft-loop/releases/tag/v0.8.0-alpha.1)) | [Released alpha](stage-evidence-v0.8.0-alpha.1.md); 17/17 issues closed; representative outcome not recorded | Produce one complete, factual, reviewed, human-approved, ATS-readable CV | Representative outcome evidence remains without overstating DOCX visual coverage |
| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Twenty-three observations across two consented cases are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | #350 informed corrections #351/#352; #358/#360 hit rate limits, and #362/#364/#368 failed before a draft. #368 exercised #366 but output-budget failures persisted. After #370/#372, #374 exhausted three author attempts across output-budget and factuality failures. #376 corrects per-generation cap accounting; #378 had one non-retryable `unknown` author failure. #380 adds content-free attribution for statusless Claude result errors. #382 exhausted three author attempts; #384 had one non-retryable `unknown` author failure with fixed diagnostics. Neither produced an artifact or review. Parity remains indeterminate; #75/#250 remain blocked |
| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Twenty-four observations across two consented cases are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | #350 informed corrections #351/#352; later attempts repeatedly failed before a draft. #376 corrects per-generation cap accounting, #380 adds fixed result diagnostics, and #393/#395 add private category capture. #397 again failed non-retryably on its first author attempt; capture established terminal reason `api_error` but no cause or product-quality evidence. #398 owns classification before another live attempt. #75/#250 remain blocked |
| Later | Retrieval and provider quality | Integrated lexical baseline; partial components | Improve evidence selection and dependable live runs | Vector/hybrid comparison, cancellation, and provider recovery in the packaged path |
| Later | Broader real-application pilot | Implemented harness; not outcome-validated | Test factuality, quality, and effort across more cases | Consented cases, calibrated measures, and recorded limitations |
| Later | Production-ready beta | Partial implementation; not production-validated | Distribute a safe, dependable desktop application | Signed installers, safe migrations, recovery, accessibility, and platform evidence |
Expand Down Expand Up @@ -793,6 +793,17 @@ was unavailable, and authorization is exhausted. Parity remains
indeterminate; #75/#250 remain blocked. See the [consented pilot report](consented-pilot-v0.9.md)
for the sanitized record.

The twenty-fourth bounded observation under #397 used revision
`75e288f60ae268be31839947ccb8203e3a8e10c4` after explicit authorization.
Both 20-second authentication probes passed. One Anthropic author attempt then
failed non-retryably with `unknown`, no failure stage or reason, and no artifact
or critic call. Fixed diagnostics retained only category classifications; the
private local capture established terminal reason `api_error` and retained no
subtype. Official Agent SDK sources document this terminal category, while a
`success` subtype is only an inference here. Issue #398 owns the provider-free
classification change. Authorization is exhausted, parity remains
indeterminate, and #75/#250 remain blocked.

**Exit criterion:** The representative comparison records no factual-invariant
violations or unsupported model-added facts, preserves required sections and
chronology, meets the agreed relevance and coverage thresholds, and produces a
Expand Down Expand Up @@ -893,6 +904,7 @@ issues retain implementation chronology.

| Date | Decision | Product implication |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 2026-09-15 | Recorded #397 as an indeterminate twenty-fourth matched-backend observation. | Both authentication probes passed, but one Anthropic author attempt failed non-retryably with `unknown`. Private bounded capture established terminal reason `api_error` but no subtype or underlying cause; no artifact or critic review occurred. Authorization is exhausted, #398 owns classification, and #75/#250 remain blocked. |
| 2026-09-15 | Routed opt-in Claude category capture through the local run boundary under #395. | A programmatic diagnostic caller can pass the private capture parent to Anthropic user-session run adapters without enabling capture for default, API-key, OpenAI, CLI, renderer, or persisted workspace paths. This provider-free plumbing does not authorize a live attempt. |
| 2026-09-14 | Added opt-in local capture for unknown Claude categories under #393. | Explicit diagnostic sessions can preserve only bounded, category-shaped unknown result subtype and terminal-reason strings in a private caller-owned file. Capture is disabled by default, durable history remains content-free, provider behavior is unchanged, and no live attempt is authorized. |
| 2026-09-13 | Recorded #384 as an indeterminate twenty-third matched-backend observation. | After explicit authorization, both 20-second authentication probes passed, but one Anthropic author attempt failed non-retryably with `unknown` after 351,508 ms. Fixed diagnostics classify only result subtype, terminal reason, and stop reason; no artifact or review occurred, and authorization is exhausted. This adds no product-quality evidence; #75/#250 remain blocked. |
Expand Down