Skip to content

fix: recover inbound replies across worker handoffs and update dependencies - #8

Closed
CPU-JIA wants to merge 5 commits into
SMNETSTUDIO:masterfrom
CPU-JIA:fix/atomic-bot-leases
Closed

CPU-JIA wants to merge 5 commits into
SMNETSTUDIO:masterfrom
CPU-JIA:fix/atomic-bot-leases

Conversation

@CPU-JIA

@CPU-JIA CPU-JIA commented Oct 1, 2026 •

Copy link
Copy Markdown

During a bot handoff, a delayed node can overwrite/delete its successor's lease. Independently, accepted messages live only in the receiving process: rebalance removes its client, and restart loses queued replies after the poll cursor has already advanced. This PR fixes those paths, restores the tests omitted by Linux shell glob expansion, and updates the dependencies flagged by the production audit.

Changes

  • Make lease renewal, individual/batch release, and overclaim cleanup atomic with owner-checked Redis Lua. Propagate pipeline errors. Fence cursor commits against a changed owner or deleted bot.
  • Persist inbound payloads before advancing the poll cursor. Atomically deduplicate queued/completed messages and apply capacity limits without marking rejected work as seen.
  • Use per-bot/per-peer FIFO processing claims: different peers run concurrently, an in-flight old owner can finish after polling handoff, and a successor recovers expired work. Graceful shutdown and OTA wait up to 25 seconds for active replies; unacknowledged payloads remain in Redis.
  • Save routing/generation decisions and per-part send checkpoints. Confirmed parts are skipped on retry, and ambiguous sends reuse a deterministic client_id. Keep the reply plan stable across runtime configuration changes.
  • Retain exhausted work after three processing attempts. Add super-admin-only inspection/retry endpoints, return only failure metadata, and audit manual retries. Bot deletion clears pending/failed payloads and invalidates processing claims.
  • Quote every package's test glob so Node discovers top-level and nested tests. Previously Linux ran only 40 of the core package's 152 tests. Run Redis integration tests in CI on Node 22 and 24.
  • Update Fastify 5.12.3 → 5.12.5 and both installed fast-uri majors (3.1.7 → 3.1.8, 4.1.4 → 4.2.1). This clears four audit records for the already-public GHSA-hrr3-gc8f-f4qj, GHSA-jvvf-x445-j334 and GHSA-4mh8-r7rc-xpvc advisories. The application does not currently enable HTTP/2/trailers or use fast-uri as its outbound URL authorization boundary; the advisories are not a claim of an exploitable production route here.

Validation

  • The four lease handoff regression cases fail against the original implementation and pass with the atomic operations; seven real-Redis lease cases cover normal and error behavior.
  • Twenty additional cases cover durable queue acceptance/capacity, FIFO handoff, crash recovery, stale-token fencing, payload cleanup, cursor commits, worker rebalance, partial-send retry, graceful stop, partial poll batches, admin authorization/redaction, and dependency regressions.
  • Reproduced encoded-host/mailto inconsistencies using the old URI packages and the uncaught HTTP/2 trailer exception using old Fastify. Patched regression cases and ordinary URI/server logging controls pass.
  • pnpm -r typecheck: passes on local Node 24.19.0 / pnpm 11.15.0.
  • pnpm -r test with dedicated Redis databases: 666 passed, 0 failed, 0 skipped.
  • pnpm audit --prod --registry=https://registry.npmjs.org --json: 0 vulnerabilities at review time. Existing advisory-only CI policy is unchanged.
  • git diff --check: passes.
  • Fork Linux CI at 73a05f9: Node 22 and Node 24 each pass type checks and all 666 tests, with zero failures/skips; audit reports no known vulnerabilities. Counts verified from the job logs.
  • Local application startup (Node 24 + dedicated Redis): /health, /health/ready, / and /app return 200; unauthenticated access to the inbox admin endpoint returns 401. Workers and external services were disabled for this smoke check.
  • Exact Dockerfile build remains unverified: Docker Hub base-image downloads stalled for over 20 minutes and the local build was stopped. This was a download limitation, not a passing container build.
  • Upstream workflow currently reports action_required; the same commit is fully checked by the fork CI above.

Operations and limits

The queue requires one shared, non-cluster Redis database with Lua and the used data-structure/TIME commands enabled. Configure persistence, backups, and no eviction. Payloads contain message text, context tokens, media credentials, and reply checkpoints; protect them like existing bot credentials. INBOX_MAX_LEN now caps retained jobs per bot, including failed jobs. Failed jobs do not expire automatically; manual retry appends to that peer's queue while retaining checkpoints. See the updated Docker, runbook and admin API docs.

Drain old in-memory work and upgrade all nodes together; mixed old/new workers retain unsafe paths. Old binaries do not consume the new queue, so drain it before rollback. Completion dedup retains the existing ten-minute window.

This provides recoverable, bounded at-least-once processing, not end-to-end exactly-once delivery. A process can die after an external operation succeeds but before its checkpoint is saved. Stable client_id reduces ambiguity, but actual iLink deduplication has not been verified; generation/P2P state transitions can also repeat in that window. No live WeChat/model calls, Lightsail deployment or cross-region load test was performed.

@CPU-JIA CPU-JIA closed this by deleting the head repository Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant