fix: keep runs with dropped batches open and honour Retry-After (4.1.62) - #259
Merged
Merged
Conversation
A batch that permanently failed to upload was counted and logged, but the run was still marked complete, so a report over partial data looked trustworthy and CI stayed green over missing results — the qase-python#504 symptom. completeTestRun now skips completion (and the public report link) when any batch was dropped, reporting how many results were lost, and the dropped-batch message is logged at ERROR since it means lost data. RetryHelper now honours the Retry-After response header. Qase answers HTTP 429 with Retry-After at roughly 60 seconds, while the computed backoff ladder (1s + 3s + 9s) was exhausted long before the rate limit cleared, so retries did not survive rate limiting in large parallel CI runs. A numeric value replaces the computed delay, capped at 120 seconds; an HTTP-date value falls back to the computed backoff rather than guessing at clock skew. computeUploadTimeout reserves the worst-case retry budget per submitted batch so awaitTermination cannot expire while a thread is correctly waiting out a rate limit.
4.1.61 was already released (tag qase-java-v4.1.61), so the fixes move into a new 4.1.62 section instead of being appended to a shipped release.
nismangulov
approved these changes
Aug 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Releases as 4.1.62. Fixes two gaps in the TestOps reporter, both variants of the
qase-python#504failure mode: the report looks trustworthy while data is missing.Do not complete a run whose batches were dropped
uploadBatchalready counted a permanently failed batch, butcompleteTestRuncalledclient.completeTestRun()unconditionally — so a run over partial data was marked complete and CI stayed green over missing results.statFailedBatches > 0. The run is left open, which is a visible signal; a completed run is not.uploadBatchalready knowsbatch.size()at the point of failure).WARNtoERROR, since it means lost data.Honour
Retry-AfterQase answers HTTP 429 with
Retry-Afterat roughly 60 seconds, while the computed backoff ladder (1s + 3s + 9s) was exhausted long before the limit cleared — so retries did not survive rate limiting in large parallel CI runs, the most common real-world trigger. Nothing read the header.Retry-Afterreplaces the computed delay, read case-insensitively offApiException(v1 and v2), capped at 120s so an implausible value cannot park an upload thread for the whole build.0and negative values.computeUploadTimeoutnow reserves the worst-case retry budget per submitted batch, soawaitTerminationcannot expire while a thread is correctly waiting out a rate limit.awaitTerminationis a maximum wait, not a fixed pause, so this only takes effect in the pathological case.Version
Version bumped
4.1.61→4.1.62across all 12 poms, with a matching changelog section. 4.1.61 is already released (tagqase-java-v4.1.61), so the entries were moved out of it rather than appended to a shipped release.Tests
TestopsReporterFailedBatchTest(new, 5 tests): a failed batch means nocompleteTestRunand no public report; a clean run and an empty run still complete; the loss is logged atERROR.RetryHelperRetryAfterTest(new, 11 tests): header parsing (numeric, case, whitespace, absent, HTTP-date, garbage, non-positive, v2 exception, cap) plus two timing tests —Retry-After: 3waits ≥ 3000ms, and without the header the wait stays at the computed 1000ms backoff.UploadSummaryLogTest.summaryCountsFailedBatchesexpectedTest run N completedas its third INFO on a failed batch — exactly the behaviour being removed. Its assertion now covers the two logs that are its actual subject (LOGS-04) and adds an explicitverify(never()).completeTestRun(...).Out of scope
this.resultsis still cleared before the batch is submitted to the executor — the same pattern behind #504, but mitigated here because retries run insideuploadBatchand a permanent failure is now counted and blocks completion rather than being swallowed. Restructuring batching or the upload executor is a larger change.