feat(bridgeservice): sync status for all syncers with is_halted (#1861) - #1877
Conversation
Add a public IsActive(ctx) method to L1InfoTreeSync mirroring bridgesync.BridgeSync.IsActive, so callers can tell whether the syncer's processor is currently halted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
…erfaces Extend L1InfoTreeSyncer with IsActive and GetLastProcessedBlock, and Claimer with GetLastProcessedBlock, so the sync-status and health handlers can report halt state and progress for these syncers. Regenerate mocks accordingly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
…ypes Add IsHalted and Error to NetworkSyncInfo and L2GERSyncInfo, and a new SyncerSyncInfo type used by three new SyncStatus entries (l1_info_tree_info, claim_l1_info, claim_l2_info). Add IsHalted to ComponentHealth and matching new keys (l1_info_tree, claim_l1, claim_l2) to HealthCheckDetails. All additive; no existing JSON tag is renamed or removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
Compute every configured syncer independently in computeSyncStatus, so one component's failure no longer blanks the others or turns the whole request into a 500. GetSyncStatusHandler now always answers 200 with a per-entry, redacted error instead. Populate the three new entries (l1_info_tree_info, claim_l1_info, claim_l2_info) and is_halted on every entry. Derive /health details from the same result, adding l1_info_tree, claim_l1 and claim_l2 alongside is_halted, and aggregate sync_status from halt/error state across all six components. Also fix cmd/run_bridgeservice.go so that a nil concrete syncer is converted to an untyped nil interface before being wired into BridgeService: passing a nil concrete pointer through directly created a non-nil interface holding a typed nil, which made both endpoints panic (and answer 500) on any instance without a claim or l1infotree syncer configured. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
…dling Add table-driven tests proving the new contract: all six syncers reporting when healthy, each halt-capable syncer reporting is_halted independently, per-component errors staying isolated with a 200 response, unconfigured syncers omitted or defaulted, concurrent computation, panic propagation, and error redaction across every component. Add a real-processor halt test for l1infotreesync using new export_test.go helpers. Add a bridgetracker regression test confirming isNetworkSynced still treats a component error as not synced. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
…erate swagger Update the sync-status and health-check sections with the new entries, is_halted semantics, the not-configured-vs-halted rule, per-entry error and the always-200 behaviour. Regenerate swagger from the updated handler annotations, including dropping the now-unused @failure 500 on GetSyncStatusHandler. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
Add TestBridgeServiceSyncStatusAllSyncers, which polls /bridge/v1/sync-status and /health against a live env, asserts every running syncer's entry (including the new l1_info_tree/claim_l1/claim_l2 ones), checks that last_processed_block advances, and asserts consistency between the two endpoints. Extend TestBridgeServiceHealthSyncStatus for the new details keys. Tighten checks.go's connectivity pre-check to inspect per-entry errors now that sync-status always answers 200. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
GetContractDepositCount, l1infotreesync's GetLastProcessedBlock, and l2gersync's GetLastProcessedBlock all ignored the caller's context, so a hung L1 RPC call or a stuck DB read could outlive the sync-status/health endpoint timeouts even though those timeouts were already threaded down to these getters. - bridgesync: pass ctx via bind.CallOpts to DepositCount instead of nil. - l1infotreesync: use db.QueryRowContext instead of db.QueryRow. - l2gersync: replace meddler.QueryRow (no ctx-aware variant) with database.QueryRowContext + Scan, dropping the now-unused BlockNum helper type. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo
|
Claude finished @arnaubennassar's task in 5m 54s —— View job Review:
|
🔄 Changes Summary
/bridge/v1/sync-statusentries:l1_info_tree_info,claim_l1_infoandclaim_l2_info, so all six syncers used by a bridge-service instance (bridgesync L1/L2, l2gersync, l1infotreesync, claimsync L1/L2) are now reported, not just the first three.is_haltedto every sync-status entry (l1_info,l2_info,l2_ger_infoand the three new ones). It isfalsefor syncers with no halt state, andfalse(alongsideis_active:false) for a syncer that isn't configured on this instance — the two cases are distinguished by which entries are present, per the docs.error(redacted via the existingaggkitcommon.RedactErrorhelper from fix(bridgetracker): redact error URLs + gate bridgeservicefinder auto-registration #1862: URLs and hosts are stripped) instead of failing the whole request on the first component error./healthgains matchingdetailskeysl1_info_tree,claim_l1andclaim_l2, plusis_haltedon everydetails.*entry./healthis derived from the exact same sync-status computation as before (pure function over the result), so the two endpoints can never drift.computeSyncStatusnow computes all six components concurrently instead of sequentially, and a panic in a single component's goroutine is logged with its original stack before being re-raised.L1InfoTreeSync.IsActive(ctx) booltol1infotreesync, mirroringbridgesync.BridgeSync.IsActive.cmd/run_bridgeservice.gowas passing nil concrete syncer pointers into interface parameters, which produces a non-nil interface holding a typed nil. On any instance without a claim or l1infotree syncer configured, both/bridge/v1/sync-statusand/healthpanicked and returned 500. Nil concrete pointers are now converted to untyped nil interfaces before being wired in.bridgesync.GetContractDepositCount,l1infotreesync.GetLastProcessedBlockandl2gersync.GetLastProcessedBlockto honour the caller'sctx: the bridgesync getter now passesctxviabind.CallOpts, and the two SQLite getters useQueryRowContextinstead ofQueryRow/meddler.QueryRow. The/bridge/v1/sync-statusread timeout and/healthcompute timeout now actually bound these calls, which previously ignoredctx. Also removes the unusedl2gersync.BlockNumhelper type.GET /bridge/v1/sync-statusnow always answers 200, with a per-entryerrorstring, instead of returning 500 on the first failing component.GET /healthdetails.*.errortext is now redacted (hosts/URLs stripped) instead of a raw error. Clients that previously treated a 500 from/bridge/v1/sync-statusas their failure signal must switch to inspecting each entry'serror/is_haltedfield; a 500 is no longer produced by this endpoint for component failures.📋 Config Updates
🔌 API Updates
🔌 Bridge service API
GET /bridge/v1/sync-status: addedl1_info_tree_info,claim_l1_info,claim_l2_info; addedis_haltedto every entry; addederrorto every entry; the endpoint now always responds 200 instead of 500 on a component failure. Not a breaking interface change apart from the 200-vs-500 status behaviour described above — no field was removed or renamed.GET /health: addeddetails.l1_info_tree,details.claim_l1,details.claim_l2; addedis_haltedto everydetails.*entry; error text indetails.*.erroris now redacted. Still always returns HTTP 200. Not a breaking interface change.🔌 Proxy API
🔌 Others API
✅ Testing
🤖 Automatic: New unit tests in
bridgeservice/bridge_test.go(TestSyncStatusAllSyncers,TestHealthFromSyncStatus,TestSyncStatusRedactsErrors,TestComputeSyncStatus_ComponentsRunConcurrently,TestComputeSyncStatusPanicPropagates), a real halted-processor test inl1infotreesync/bridgeservice_halt_test.go(TestBridgeServiceReportsL1InfoTreeHalt), thecmd/run_bridgeservice_test.gonil-wiring regression tests, and abridgetracker/sources/activity_test.goregression test (TestIsNetworkSynced) confirming an erroring network still reports not-synced. New e2e testTestBridgeServiceSyncStatusAllSyncersintest/e2e/sync_status_all_syncers_test.go, plus an extendedTestBridgeServiceHealthSyncStatus.Regression tests for the ctx fix, verified to fail on the pre-fix code:
TestGetContractDepositCount_HonoursContextDeadlineinbridgesync(a blockedeth_call), andTestGetLastProcessedBlock_HonoursContextin bothl1infotreesyncandl2gersync(an already-cancelled ctx).🖱️ Manual:
golangci-lintv2.4.0 reported 0 issues.make test-unitpassed. E2E run (AGGKIT_E2E_ENV=anvil-2chains make test-e2e TEST_RUN='^(TestBridgeServiceSyncStatusAllSyncers|TestBridgeServiceHealthSyncStatus)$') passed for both L2A and L2B:--- PASS: TestBridgeServiceHealthSyncStatus (3.51s),--- PASS: TestBridgeServiceSyncStatusAllSyncers (9.03s).Before (develop today):
/bridge/v1/sync-statusreturns onlyl1_info,l2_infoandl2_ger_info— nol1_info_tree_info,claim_l1_infoorclaim_l2_info, and nois_haltedanywhere. A single failing component (e.g. an RPC blip on l1infotreesync) causes the whole endpoint to answer HTTP 500 instead of returning what it could compute.After (this branch), real payload from the e2e env:
GET /bridge/v1/sync-status{ "l1_info": {"contract_deposit_count":2,"synchronized_deposit_count":2,"is_synced":true,"is_active":true,"is_halted":false}, "l2_info": {"contract_deposit_count":0,"synchronized_deposit_count":0,"is_synced":true,"is_active":true,"is_halted":false}, "l2_ger_info": {"is_active":true,"last_processed_block":100,"is_halted":false}, "l1_info_tree_info": {"is_active":true,"is_halted":false,"last_processed_block":329}, "claim_l1_info": {"is_active":true,"is_halted":false,"last_processed_block":329}, "claim_l2_info": {"is_active":true,"is_halted":false,"last_processed_block":234} }GET /health{ "status": "ok", "time": "2026-09-29T07:33:26.425307489Z", "version": "v0.11.0-rc10-13-gaf2f650d", "sync_status": "done", "details": { "l1": {"is_active":true,"is_synced":true,"is_halted":false}, "l2": {"is_active":true,"is_synced":true,"is_halted":false}, "l2_ger": {"is_active":true,"is_halted":false}, "l1_info_tree": {"is_active":true,"is_halted":false}, "claim_l1": {"is_active":true,"is_halted":false}, "claim_l2": {"is_active":true,"is_halted":false} } }🐞 Issues
🔗 Related PRs
RedactErrorhelper this PR reuses for per-entry error redaction)📝 Notes
is_halted), with no reason or halted block, because those strings can carry DB/RPC internals — this follows the same redaction policy as fix(bridgetracker): redact error URLs + gate bridgeservicefinder auto-registration #1862.pendingin/health's aggregation: they have no in-service "caught up" signal the way the L1/L2 bridge syncers do (claimsync syncs on demand viaSetNextRequiredBlock). They only ever feed the error bucket ordone.🤖 Generated with Claude Code
https://claude.ai/code/session_01GRC3mRecjKCrgAHRcDnWuo