Repository navigation
Delay verify-replay retries without holding a SQL thread - #6236
emelialei88 wants to merge 1 commit into
Conversation
1662eee to
be60598
Compare
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
analyze_partial_index_off_generated [failed with core dumped] **quarantined**
analyze [failed with core dumped] **quarantined**
consumer_non_atomic_default_consumer_generated **quarantined**
tunables
sc_downgrade [timeout] **quarantined**
reco-ddlk-sql [timeout] **quarantined**
skipscan [timeout] **quarantined**
be60598 to
096b5be
Compare
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Success ✓.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_truncate_multiddl_generated [db unavailable at finish] **quarantined**
consumer_non_atomic_default_consumer_generated **quarantined**
tunables
sc_downgrade [timeout] **quarantined**
Signed-off-by: Emelia Lei <wlei29@bloomberg.net>
096b5be to
c7f2c5e
Compare
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Success ✓.
Regression testing: 2/730 tests failed ⚠.
The first 10 failing tests are:
comdb2sys_queueodh_generated
consumer_non_atomic_default_consumer_generated **quarantined**
🎯 The "Why" (Intent)
When a transaction loses the commit-time verify check, it is replayed immediately. Under contention every loser retries in lockstep and collides again, so replays climb into the dozens and latency accumulates into seconds. Only distributed transactions had a backoff, and it sleeps on the SQL thread, starving unrelated clients when the pool is small.
🛠️ The "What" (Critical Changes)
verify_retry_poll_after(default 3) immediate replays, a verify-failed txn is parked on a timer for a random[0, verify_retry_poll)ms (default 50), then re-queued. The SQL thread is released while it waits.verify_retry_poll 0restores lockstep retries.tests/verify_retry_poll.test: a benchmark (not pass/fail).📊 Benchmark
4-node cluster, 32 writers × 200 upserts on one row, SQL pool of 8 threads. Bystander = an unrelated write to another table during the storm (~45 ms idle, mostly connection setup). "On-thread" sleeps on the SQL worker; "parked" is this PR. Both pause from the 2nd retry on (threshold 1, see below), so only the thread use differs.
Threshold: how many retries go back immediately before pausing (
verify_retry_poll_after)The original attempt is not a retry. With threshold N, retries 1..N are re-queued immediately and the pause starts before retry N+1:
Parked, 50 ms max wait, pool 48, 32 writers. Each cell is total time / total replays.
Each extra immediate retry mostly collides again: replays grow with the threshold everywhere, and time never improves beyond noise (±0.6 s). Threshold 1 is fastest or tied. The default here is 3; the data favours 1, pending a re-run on a larger cluster.
With writers spread over 2000 rows (little conflict), every setting gives 2.2–2.6 s and under 1k replays. With a 48-thread pool, the bystander is unaffected either way.