Skip to content

fix(experiment): durable batched log queue with session/seq ids - #669

Open
kcarnold wants to merge 4 commits into
mainfrom
claude/brave-curie-reb89q
Open

kcarnold wants to merge 4 commits into
mainfrom
claude/brave-curie-reb89q

Conversation

@kcarnold

@kcarnold kcarnold commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Why

Before this change, each experiment log event was a separate POST that stopped retrying about 300ms after the first try. A Wi‑Fi blip or a redeploy during the writing task could silently lose chat messages and AI events, and those can't be recovered from anything else.

What

  • lib/logging.ts: log() now timestamps each event and assigns it a sessionId (new per page load) and a seq number, then adds it to a queue. The queue uploads in batches about once per second, one request at a time. Events leave the queue only after the server confirms them; failed uploads retry with backoff for as long as the page is open.
  • Surviving unloads: when the tab is hidden or unloaded (visibilitychange/pagehide), the queue is saved to localStorage, and the next page load uploads it. It is not written on every event. pagehide also tries a keepalive upload of up to 60KB. beforeunload shows a prompt if events are still queued.
  • Page transitions wait for logs to be saved: redirect(url) / logThenRedirect() wait, retrying indefinitely, until the queue is empty. All study page transitions now go through them, replacing the duplicated log-then-window.location code. While a transition is blocked and uploads are failing, a single LogStatusBanner on the study page explains the wait. The page moves on by itself once the upload succeeds. Batches the server rejected (4xx) don't block, since waiting can't recover them.
  • /api/log: takes an array of entries and groups them by username, since a queue recovered from localStorage on a shared machine can hold another participant's events. It skips invalid entries instead of rejecting the batch, so one bad entry can't block the queue. The client drops a batch on a 4xx rather than retrying it.
  • Document snapshots are still full snapshots taken on every keystroke. docs/logging.md recommends compressing the log files on disk.

Analysis impact

Deduplicate on (sessionId, seq), since duplicates are expected. Sort by timestamp. A gap in seq within a session means lost events. See experiment/docs/logging.md.

Testing

  • New __tests__/lib/logging.test.ts: batching and seq numbering, retry after a 5xx, dropping a batch on a 4xx, hand-off to the next page load via localStorage, flushing before a redirect, and a redirect that stays blocked while uploads fail and reports its status.
  • Called the route handler directly with a mixed batch: valid entries were written per user, the invalid username was skipped, and a non-array body returned 400.
  • tsc passes. ESLint reports no new issues; the existing error in ScreenSizeCheck.tsx is unchanged.
  • Not yet tested in a real browser (unload prompt, keepalive upload during pagehide, banner appearance).

🤖 Generated with Claude Code

https://claude.ai/code/session_018SvPCoGPWdKzNdUZVCoW29

Log events are now timestamped and numbered at log time, queued, and
uploaded in batches that leave the queue only on server ack. Unsent events
survive unloads via localStorage (written on hide/pagehide only) and are
uploaded by the next page load. Page transitions go through redirect(),
which drains the queue first. /api/log accepts batches and skips invalid
entries rather than rejecting the batch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018SvPCoGPWdKzNdUZVCoW29
Comment thread experiment/lib/logging.ts Fixed
redirect() now waits indefinitely for the log queue to drain instead of
timing out after 5s. A single LogStatusBanner on the study page explains
the wait when uploads are failing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018SvPCoGPWdKzNdUZVCoW29
Works in non-secure contexts too, so no Math.random fallback is needed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018SvPCoGPWdKzNdUZVCoW29
@kcarnold

Copy link
Copy Markdown
Contributor Author

@neh8 can you test that this addresses the too-fast-to-log issue that you were seeing?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants