Skip to content

fix: keep runs with dropped batches open and honour Retry-After (4.1.62) - #259

Merged
gibiw merged 2 commits into
mainfrom
fix/dropped-batches-and-retry-after
Aug 27, 2026
Merged

gibiw merged 2 commits into
mainfrom
fix/dropped-batches-and-retry-after

Conversation

@gibiw

@gibiw gibiw commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Releases as 4.1.62. Fixes two gaps in the TestOps reporter, both variants of the qase-python#504 failure mode: the report looks trustworthy while data is missing.

Do not complete a run whose batches were dropped

uploadBatch already counted a permanently failed batch, but completeTestRun called client.completeTestRun() unconditionally — so a run over partial data was marked complete and CI stayed green over missing results.

  • Completion is skipped when statFailedBatches > 0. The run is left open, which is a visible signal; a completed run is not.
  • The error names how many results were lost, not just batches (uploadBatch already knows batch.size() at the point of failure).
  • Enabling the public report link is skipped for the same reason.
  • The dropped-batch message was raised from WARN to ERROR, since it means lost data.

Honour Retry-After

Qase answers HTTP 429 with Retry-After at roughly 60 seconds, while the computed backoff ladder (1s + 3s + 9s) was exhausted long before the limit cleared — so retries did not survive rate limiting in large parallel CI runs, the most common real-world trigger. Nothing read the header.

  • A numeric Retry-After replaces the computed delay, read case-insensitively off ApiException (v1 and v2), capped at 120s so an implausible value cannot park an upload thread for the whole build.
  • An HTTP-date value falls back to the computed backoff rather than guessing at clock skew. Same for garbage, empty, 0 and negative values.
  • computeUploadTimeout now reserves the worst-case retry budget per submitted batch, so awaitTermination cannot expire while a thread is correctly waiting out a rate limit. awaitTermination is a maximum wait, not a fixed pause, so this only takes effect in the pathological case.

Version

Version bumped 4.1.61 → 4.1.62 across all 12 poms, with a matching changelog section. 4.1.61 is already released (tag qase-java-v4.1.61), so the entries were moved out of it rather than appended to a shipped release.

Tests

  • TestopsReporterFailedBatchTest (new, 5 tests): a failed batch means no completeTestRun and no public report; a clean run and an empty run still complete; the loss is logged at ERROR.
  • RetryHelperRetryAfterTest (new, 11 tests): header parsing (numeric, case, whitespace, absent, HTTP-date, garbage, non-positive, v2 exception, cap) plus two timing tests — Retry-After: 3 waits ≥ 3000ms, and without the header the wait stays at the computed 1000ms backoff.
  • Both suites were verified to fail against the pre-fix implementation.
  • UploadSummaryLogTest.summaryCountsFailedBatches expected Test run N completed as its third INFO on a failed batch — exactly the behaviour being removed. Its assertion now covers the two logs that are its actual subject (LOGS-04) and adds an explicit verify(never()).completeTestRun(...).
  • Full suite: 338 tests, 0 failures, all 12 modules green.

Out of scope

this.results is still cleared before the batch is submitted to the executor — the same pattern behind #504, but mitigated here because retries run inside uploadBatch and a permanent failure is now counted and blocks completion rather than being swallowed. Restructuring batching or the upload executor is a larger change.

gibiw added 2 commits August 27, 2026 12:48
A batch that permanently failed to upload was counted and logged, but the
run was still marked complete, so a report over partial data looked
trustworthy and CI stayed green over missing results — the qase-python#504
symptom. completeTestRun now skips completion (and the public report link)
when any batch was dropped, reporting how many results were lost, and the
dropped-batch message is logged at ERROR since it means lost data.

RetryHelper now honours the Retry-After response header. Qase answers HTTP
429 with Retry-After at roughly 60 seconds, while the computed backoff
ladder (1s + 3s + 9s) was exhausted long before the rate limit cleared, so
retries did not survive rate limiting in large parallel CI runs. A numeric
value replaces the computed delay, capped at 120 seconds; an HTTP-date value
falls back to the computed backoff rather than guessing at clock skew.
computeUploadTimeout reserves the worst-case retry budget per submitted
batch so awaitTermination cannot expire while a thread is correctly waiting
out a rate limit.
4.1.61 was already released (tag qase-java-v4.1.61), so the fixes move into
a new 4.1.62 section instead of being appended to a shipped release.
@gibiw gibiw changed the title fix: keep runs with dropped batches open and honour Retry-After fix: keep runs with dropped batches open and honour Retry-After (4.1.62) Aug 27, 2026
@gibiw
gibiw merged commit 16bbd2e into main Aug 27, 2026
14 checks passed
@gibiw
gibiw deleted the fix/dropped-batches-and-retry-after branch August 27, 2026 13:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants