Skip to content

test(fdv2): stop the recovery test spinning into an OutOfMemoryError - #404

Merged
abelonogov-ld merged 3 commits into
mainfrom
andrey/fix-fdv2-datasource-test-oom
Sep 25, 2026
Merged

abelonogov-ld merged 3 commits into
mainfrom
andrey/fix-fdv2-datasource-test-oom

Conversation

@abelonogov-ld

@abelonogov-ld abelonogov-ld commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

What failed

FDv2DataSourceTest.fallbackAndRecoveryTasksWellBehaved failed on CI with java.lang.OutOfMemoryError (seen on the tier-1 event-durability branch, but the test and the cause are on main):

com.launchdarkly.sdk.android.FDv2DataSourceTest > fallbackAndRecoveryTasksWellBehaved FAILED
    java.lang.OutOfMemoryError at FDv2DataSourceTest.java:71
    java.lang.OutOfMemoryError at Arrays.java:3481

Why

The test's synchronizer factories returned the same MockQueuedSynchronizer instance on every call. After fallback, the recovery timer closes the active synchronizer and builds the primary again. That returns the already-closed first mock, and a closed mock answers every next() with SHUTDOWN immediately. On Android SHUTDOWN advances to the next synchronizer, which is the other closed mock, and SourceManager wraps the index. So the orchestrator spun between two closed mocks with nothing blocking it. Each pass logs "Synchronizer '…' is starting." into LogCaptureRule and creates and cancels the condition timers.

The test then waits up to 10 s for a third changeset. No closed mock can deliver it, and awaitApplyCount returns silently on timeout, so the spin ran for the whole wait. With a temporary probe I measured about 630,000 factory calls and 3.8 million captured log lines in that window, with only 2 changesets applied. It passes on a developer machine with enough heap and fails on CI.

This is a test bug, not an SDK bug. A real factory builds a new synchronizer each time.

Fix

  • Each factory call builds a new MockQueuedSynchronizer, as a real factory does. Recovery now gets a working primary that delivers the third changeset.
  • Assert getApplyCount() >= 3 after the await. If the shared-instance pattern comes back, the test now fails with an AssertionError (checked by reverting the fix) instead of passing while it spins.
  • orchestrationLogging_recovery_logsInfo had the same shape and gets the same change. It stopped as soon as the log line appeared, so it spun only briefly.
  • Separate commit: FlagTest uses Long.valueOf/Integer.valueOf instead of the deprecated boxing constructors, which removes the three test-compile deprecation warnings from the same CI log.

Verification

  • fallbackAndRecoveryTasksWellBehaved now takes about 3.0 s, down from the full 10 s wait.
  • FDv2DataSourceTest passed 5/5 runs; the full unit suite passes (730 tests).

Note

Overview
Fixes CI OutOfMemoryError in FDv2DataSourceTest.fallbackAndRecoveryTasksWellBehaved by correcting how mock synchronizer factories behave.

The fallback/recovery tests used factories that returned the same MockQueuedSynchronizer instances. After recovery closes a synchronizer, reusing it makes every next() return SHUTDOWN immediately, so the orchestrator can spin between two closed mocks (logging and timer churn) while the test waits for a third changeset that never arrives.

Changes: factory lambdas now construct a fresh MockQueuedSynchronizer on each call (matching real factories), in both fallbackAndRecoveryTasksWellBehaved and orchestrationLogging_recovery_logsInfo. The recovery test also asserts getApplyCount() >= 3 after the wait so a regression fails fast instead of timing out while spinning.

Separate cleanup: FlagTest replaces deprecated new Integer / new Long with Integer.valueOf / Long.valueOf in assertions (compile warnings only).

Reviewed by Cursor Bugbot for commit 486b53a. Bugbot is set up for automated code reviews on this repo. Configure here.

abelonogov-ld and others added 2 commits September 23, 2026 11:13
… tests

fallbackAndRecoveryTasksWellBehaved handed the same MockQueuedSynchronizer
back from each factory call. Recovery closes the active synchronizer and
builds the primary again; the closed mock answers every next() with
SHUTDOWN immediately, which advances to the other closed mock, and the
orchestrator spun between the two until the test stopped it. The third
changeset the test waits for could never arrive, so it spun for the whole
ten-second await: locally about 630,000 factory calls and 3.8 million
captured log lines, and on CI an OutOfMemoryError.

Each factory call now builds a new synchronizer, as a real factory does,
and the test asserts the third changeset arrives rather than letting the
await time out silently. It finishes in about three seconds.
orchestrationLogging_recovery_logsInfo had the same shape and gets the
same change.

Co-authored-by: Cursor <cursoragent@cursor.com>
…agTest

Co-authored-by: Cursor <cursoragent@cursor.com>
@abelonogov-ld
abelonogov-ld requested a review from a team as a code owner September 23, 2026 18:14
@abelonogov-ld
abelonogov-ld merged commit 78db8b6 into main Sep 25, 2026
7 checks passed
@abelonogov-ld
abelonogov-ld deleted the andrey/fix-fdv2-datasource-test-oom branch September 25, 2026 14:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants