Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 57 additions & 11 deletions docs/consented-pilot-v0.9.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

**Status:** Indeterminate<br>
**Recorded:** 2026-09-15<br>
**Cases:** Twenty-four bounded observations across two consented, anonymized cases
**Cases:** Twenty-five bounded observations across two consented, anonymized cases

## Result

Expand Down Expand Up @@ -56,6 +56,14 @@ reason while retaining no subtype. No draft or independent review was produced,
so #397 adds no product-quality evidence and leaves the parity outcome
indeterminate. Issue #398 owns the separate provider classification change.

The twenty-fifth observation (#401) followed the merged #398 classification.
Both user-session authentication probes passed, but all three Anthropic author
responses failed local factual and substantive-coverage validation. Attempts
one and two were retryable; attempt three exhausted the author cap. No draft or
independent review was persisted, and no statusless API error occurred. The
classification branch from #398 was therefore not exercised. #401 adds no
product-quality evidence and leaves the parity outcome indeterminate.

The initial attempt completed one Anthropic author step and one independent
OpenAI critic step. Its revision call timed out. Under the candidate-approved
extension, one fresh retry used the same model pair and sanitized data scope
Expand Down Expand Up @@ -740,6 +748,36 @@ non-retryable failure cannot resume, and the authorization is exhausted; a new
live run requires separate explicit authorization. Parity remains
indeterminate; #75 stays open, and #250 remains blocked.

## Bounded post-classification observation (#401)

On 2026-09-15, explicit authorization covered both authentication probes and
this twenty-fifth observation on revision
`10bed45cf577e1d8bc2697571c86ae453ce1bce9`. It reused the same private
matched-role case and baseline-withheld v1 gate, with Anthropic
`claude-sonnet-4-5` as author and OpenAI `gpt-5.6-luna` as critic through their
authenticated user sessions. The clean non-fixture preflight matched three
rounds, a 1,200,000 ms request timeout, and a 1,200,000 ms maximum active
duration. Both 20-second authentication probes passed. The optional category
capture parent was outside the repository with mode `0700`.

The fresh run used all three Anthropic author attempts. Each response reached
local validation and failed with generic `invalid-response`, zero reported
tokens, and `factual-invariant-rejection` as both failure stage and reason.
Attempt one recorded one factual-invariant and seven substantive-coverage
diagnostics. Attempts two and three each recorded two factual-invariant and six
substantive-coverage diagnostics. The first two failures were retryable; the
third was non-retryable because it exhausted the three-attempt cap. Persisted
active duration was 598,833 ms. Provider-reported cost was unavailable; the
persisted workspace zero is not an actual cost measurement.

No statusless structured API error occurred, so this observation did not
exercise #398's retry classification. No unknown category was captured. No
artifact, critic call, findings, readiness decision, review, adjudication,
approval, export, submission, release preparation, or release occurred. The
authorization is exhausted; a later live run requires separate explicit
authorization. Parity remains indeterminate; #75 stays open, and #250 remains
blocked.

## Predeclared comparison gate

| Dimension | Status |
Expand Down Expand Up @@ -790,8 +828,8 @@ unresolved findings as accepted facts.
developer-tool evidence more directly, while the baseline retained a stronger
backend-production narrative. A separate matched-role case and manual
baseline were used for the matched-backend observations. Their issue numbers
are #350, #358, #360, #362, #364, #368, #374, #378, #382, #384, and #397.
The ten later runs
are #350, #358, #360, #362, #364, #368, #374, #378, #382, #384, #397, and
#401. The eleven later runs
reused private inputs with the baseline withheld from generation; none
produced an artifact to compare. The twentieth observation (#374) reused
that case after #370 and #372 merged; its authorization is exhausted. The
Expand All @@ -806,6 +844,10 @@ unresolved findings as accepted facts.
The twenty-fourth observation (#397) captured `api_error` as the terminal
reason for the same one-attempt boundary; it produced no artifact or review,
and its authorization is exhausted.
The twenty-fifth observation (#401) exhausted three author attempts on local
factual and substantive-coverage validation; it produced no artifact or
review, did not exercise #398's API-error branch, and its authorization is
exhausted.
The nineteenth run exercised the #366 thinking bound, but output-budget
failures persisted, so it does not show that failure mode resolved.
- Misleading-evidence and prompt-injection behavior were not tested in these
Expand Down Expand Up @@ -905,24 +947,26 @@ The twenty-third observation (#384) ended in a non-retryable author failure;
its authorization is exhausted.
The twenty-fourth observation (#397) ended in a non-retryable author failure;
its authorization is exhausted.
The twenty-fifth observation (#401) exhausted its three-author-attempt cap; its
authorization is exhausted.

No artifact, critic call, findings, readiness decision, review, adjudication,
approval, export, submission, or release occurred in #368, #374, #378, #382, #384,
or #397.
approval, export, submission, or release occurred in #368, #374, #378, #382,
the #384 and #397 observations, or #401.
Parity remains indeterminate, and #75 and #250 remain blocked. Any later live
attempt needs separate bounded authorization.

Issue #75 stays open, and release preparation under #250 remains blocked.
No final candidate approval, export, or submission is authorized by these
results. Authorization for #374, #378, #382, #384, and #397 is exhausted and
does not extend to another live attempt. The #378, #382, #384, and #397 runs
cannot resume; a new run requires separate explicit bounded authorization.
results. Authorization for #374, #378, #382, #384, #397, and #401 is exhausted
and does not extend to another live attempt. None of those runs can resume; a
new run requires separate explicit bounded authorization.

The first case remains limited by role mismatch. The matched case provides
product evidence from #350 about required content and claim validation. The
provider failures in #358, #360, #362, #364, #368, #374, #378, #382, #384,
and #397 added no product-quality evidence and do not change the indeterminate
outcome. The #366
and #397, plus the local-validation failures in #401, added no product-quality
evidence and do not change the indeterminate outcome. The #366
thinking bound was exercised in #368, but output-budget failures remained;
this does not show that failure mode fully resolved. The #374 diagnostics do
not establish the cause of its reported output-budget failures or reinterpret
Expand All @@ -932,5 +976,7 @@ their generic invalid-response classifications and failure stage do not add
product-quality evidence. The #384 diagnostics add safe error attribution but
do not establish its underlying cause or add product-quality evidence. The
capture under #397 establishes only the `api_error` terminal category, not the
underlying API failure or exact result subtype.
underlying API failure or exact result subtype. The #401 observation did not
encounter a statusless API error and therefore does not validate #398 in live
use.
Provider-reported cost was unavailable.
14 changes: 13 additions & 1 deletion docs/roadmap.md
Original file line number Diff line number Diff line change
Expand Up @@ -177,7 +177,7 @@ applications.
| Previous | Integration hardening and outcome validation ([v0.6.0](https://github.com/akoita/draft-loop/releases/tag/v0.6.0)) | Released; validation failed | Preserve a reproducible integrated baseline without overstating application readiness | Failed representative result carried into v0.7; see [stage evidence](stage-evidence-v0.6.0.md) |
| Previous | Evidence-backed CV drafting (v0.7 program) | [Released alpha.5 checkpoint](stage-evidence-v0.7.0-alpha.5.md); implementation history carried forward; outcome not validated | Produce a complete factual, source-traceable application draft | v0.8 candidate evidence now covers the bounded drafting and review vertical |
| Previous | Usable CV MVP ([v0.8.0-alpha.1](https://github.com/akoita/draft-loop/releases/tag/v0.8.0-alpha.1)) | [Released alpha](stage-evidence-v0.8.0-alpha.1.md); 17/17 issues closed; representative outcome not recorded | Produce one complete, factual, reviewed, human-approved, ATS-readable CV | Representative outcome evidence remains without overstating DOCX visual coverage |
| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Twenty-four observations across two consented cases are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | #350 informed corrections #351/#352; later attempts repeatedly failed before a draft. #376 corrects per-generation cap accounting, #380 adds fixed result diagnostics, and #393/#395 add private category capture. #397 established terminal reason `api_error`; #398 classifies a statusless structured occurrence as bounded and retryable without consuming provider prose. A later live observation still requires fresh authorization. #75/#250 remain blocked |
| Now | Workflow parity and release ([milestone v0.9.0](https://github.com/akoita/draft-loop/milestone/4)) | [Twenty-five observations across two consented cases are indeterminate](consented-pilot-v0.9.md); parity not validated | Demonstrate the complete application-grade workflow and publish evidence | #350 informed corrections #351/#352; later attempts repeatedly failed before a draft. #376 corrects per-generation cap accounting, #380 adds fixed result diagnostics, and #393/#395 add private category capture. #397 established terminal reason `api_error`, and #398 added its bounded statusless classification. #401 then exhausted three author attempts on local factual and coverage validation without an artifact or review. Any later live observation requires fresh authorization. #75/#250 remain blocked |
| Later | Retrieval and provider quality | Integrated lexical baseline; partial components | Improve evidence selection and dependable live runs | Vector/hybrid comparison, cancellation, and provider recovery in the packaged path |
| Later | Broader real-application pilot | Implemented harness; not outcome-validated | Test factuality, quality, and effort across more cases | Consented cases, calibrated measures, and recorded limitations |
| Later | Production-ready beta | Partial implementation; not production-validated | Distribute a safe, dependable desktop application | Signed installers, safe migrations, recovery, accessibility, and platform evidence |
Expand Down Expand Up @@ -808,6 +808,17 @@ Authorization is exhausted, parity remains indeterminate, and #75/#250 remain
blocked. A later live observation requires a fresh bounded issue and explicit
authorization.

The twenty-fifth bounded observation under #401 used revision
`10bed45cf577e1d8bc2697571c86ae453ce1bce9` after explicit authorization.
Both 20-second user-session authentication probes passed. All three Anthropic
author responses reached local validation but failed with
`factual-invariant-rejection`: the first two attempts were retryable, and the
third exhausted the cap. No statusless API error occurred, so the #398 branch
was not exercised. No artifact or critic review existed, provider-reported cost
was unavailable, and authorization is exhausted. Parity remains indeterminate;
issues #75/#250 remain blocked, and another live attempt requires a new bounded
issue and explicit authorization.

**Exit criterion:** The representative comparison records no factual-invariant
violations or unsupported model-added facts, preserves required sections and
chronology, meets the agreed relevance and coverage thresholds, and produces a
Expand Down Expand Up @@ -908,6 +919,7 @@ issues retain implementation chronology.

| Date | Decision | Product implication |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 2026-09-15 | Recorded #401 as an indeterminate twenty-fifth matched-backend observation. | Both user-session authentication probes passed, but three Anthropic author responses failed local factual and substantive-coverage validation before an artifact or critic review. The run did not exercise #398's API-error branch. Authorization is exhausted, and #75/#250 remain blocked. |
| 2026-09-15 | Classified statusless structured Claude API errors under #398. | The documented `api_error` terminal reason and error-result `success` subtype now produce fixed diagnostics. Without a finite numeric status, the failure is transient and retryable only within existing orchestration caps; numeric statuses retain precedence, provider prose remains excluded, and no live attempt is authorized. |
| 2026-09-15 | Recorded #397 as an indeterminate twenty-fourth matched-backend observation. | Both authentication probes passed, but one Anthropic author attempt failed non-retryably with `unknown`. Private bounded capture established terminal reason `api_error` but no subtype or underlying cause; no artifact or critic review occurred. Authorization is exhausted, #398 owns classification, and #75/#250 remain blocked. |
| 2026-09-15 | Routed opt-in Claude category capture through the local run boundary under #395. | A programmatic diagnostic caller can pass the private capture parent to Anthropic user-session run adapters without enabling capture for default, API-key, OpenAI, CLI, renderer, or persisted workspace paths. This provider-free plumbing does not authorize a live attempt. |
Expand Down