Conversation
…rence-drift Two packages/cli tests failed in the Full Suite shard on b4cddcb. Both were pre-existing drift that PR #36 (test files only) neither caused nor fixed. 1. task-retry.test.ts "clears the deadlock auto-pause" — expected 'todo', got 'in-review'. The fixture is a MERGE failure (all steps done, mergeRetries 4), and FN-9317 ("recover stalled in-review merges", 706c155) made that shape retry IN PLACE: clear status/error/auto-pause, reset mergeRetries, keep the card in review so approved work is not re-run. The `todo` assertion is pre-FN-9317 drift from FN-6173 (1c69ea7), which briefly sent CLI merge retries back to todo and was later reverted. Every sibling surface already asserts the in-place behavior for this exact shape: commands/__tests__/task.test.ts asserts moveTask is NOT called and logs "in-review merge retry, mergeRetries reset"; __tests__/extension.test.ts asserts details.newColumn === "in-review"; the dashboard route classifies it the same way. The product is correct, so the stale expectation is fixed and the test is renamed to state the contract it pins. Added the complementary execution-failure case (unfinished steps -> re-queue to the hold column with progress preserved), mirroring extension.test.ts, so the re-queue half of the deadlock auto-pause stays covered. 3 passed / 1 failed -> 4 passed. 2. skills-get.test.ts "prints a guide and version from the same built CLI entry point" — 'Test timed out in 5000ms'. The test did three SERIAL cold boots of the ~19 MB dist/bin.js ESM bundle inside one `it` (~1.0 s each locally, 2-3 s on a shared GitHub runner): the guide, --version, and the unknown-skill error path. Three serial boots need ~3.1 s locally against Vitest's 5000 ms default and exceed it on the runner. Nothing hung; the seam is what was wrong. Fix: run the two independent guide/version invocations concurrently via Promise.all so their boots overlap, and move the error-path boot into its own test. Wall clock drops from ~3.1 s to ~1.2 s (measured) with identical assertions and no timeout bump (check-no-test-timeout-appeasement.mjs stays green). Coverage unchanged: the guide must still come from the built entry point and its embedded version must still match that same built binary's --version. Gates: task-retry 4/4 pass, skills-get 6/6 pass, runTaskRetry suite 12/12 pass, cli tsc --noEmit clean, check-no-test-timeout-appeasement.mjs clean, check-changeset-format.mjs clean, eslint 0 errors (test files are eslint-ignored). No changeset: test-only change, no published behavior delta (AGENTS.md). Pipeline smoke tier is a SEPARATE pre-existing timing failure, not touched here.
ThreatCrush Security Scan4589 finding(s) HIGH/CRITICAL: 43 | MEDIUM: 4035 | LOW: 511
…and 4539 more. Full results in the Security tab. Snippets are redacted; ThreatCrush never prints matched credential material. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Continuação do #36
Depois do merge do #36 o shard 3 do Full Suite ainda falhava com exatamente 2 arquivos em
packages/cli(169 passed | 2 failed | 3 skipped). Ambas as falhas eram drift pré-existente que o #36 não causou nem corrigiu.1.
task-retry.test.ts— o produto estava certo, o teste estava erradoAssertionError: expected 'in-review' to be 'todo'O fixture é uma falha de merge (todos os steps
done,mergeRetries: 4). O commit706c15560(FN-9317, "recover stalled in-review merges") deliberadamente mudou essa forma para retry no lugar: limpastatus/error/auto-pause, zeramergeRetriese mantém o card em review para não reexecutar trabalho já aprovado. A expectativatodoé drift anterior ao FN-9317, vindo de1c69ea78b(FN-6173), que enviava retries de volta paratodoe depois foi revertido.Todas as outras superfícies já afirmavam o comportamento no lugar para essa mesma forma:
commands/__tests__/task.test.ts—moveTaskNÃO é chamado, e registra"in-review merge retry, mergeRetries reset"__tests__/extension.test.ts—details.newColumn === "in-review"register-task-workflow-routes.ts— classifica igualO teste do CLI era o único atrasado. Corrigi a expectativa, renomeei o teste para declarar o contrato que ele fixa, e adicionei o caso complementar (falha de execução, steps incompletos -> re-enfileira no hold preservando progresso), espelhando
extension.test.ts. Cobertura aumentou: 3 passed / 1 failed -> 4 passed.2.
skills-get.test.ts— não era hang, era formatoTest timed out in 5000msO teste fazia três boots seriais do bundle ESM de ~19 MB (
dist/bin.js) dentro de um únicoit: guide,--versione o caminho de erro. Cada boot custa ~1,0 s local e 2-3 s num runner compartilhado. Três boots seriais = ~3,1 s local contra o default de 5000 ms do Vitest, e estoura no CI. Nada travava.Correção na costura, não no orçamento: os dois spawns independentes (guide e
--version) rodam concorrentes viaPromise.all, e o boot do caminho de erro vai para o seu próprioit. Tempo cai de ~3,1 s para ~1,2 s com as mesmas asserções e sem aumentar timeout —check-no-test-timeout-appeasement.mjscontinua limpo. 6/6 passam.Verificação
task-retry4/4,skills-get6/6,runTaskRetry12/12tsc --noEmitlimpo,check-no-test-timeout-appeasement.mjslimpo,check-changeset-format.mjslimpo, eslint 0 errosO que NÃO está neste PR
Pipeline smoke tiercontinua vermelho, e é outra coisa: orçamento, não asserção. Medido localmente, o projetoengine-pipeline-smokeleva 160,1 s contraPIPELINE_SMOKE_DURATION_BUDGET_MS = 175_000— só ~9% de folga, e o wrapper orça a lane inteira numa única invocação. O arquivo mais lento sozinho (pipeline-resilience.pipeline.test.ts, 157,0 s) quase consome o orçamento inteiro dos 3 workers. Num runner do GitHub isso estoura o watchdog. O comentário do código diz que os 175 s são "a regression detector, not a knob for hiding overruns", então subir o orçamento é exatamente o remendo que a política do repo proíbe. A correção legítima é reduzir o custo por cenário ou aumentar workers/shardar a lane.