Skip to content

Ops: email operator error alerts to support_filemill@keywind.cc #7

Description

@brocla

Backlog: Observability & ops

Problem

When something goes wrong (systemic job failure, delivery failure, panic, a stuck worker loop, a crash), nobody is notified. An unattended background service fails silently.

Proposed solution

Email operator alerts to support_filemill@keywind.cc (replaces the earlier support@keywind.cc and support@mill.keywind.cc). It's deliberately on keywind.cc, whose mail goes through Cloudflare Email Routing, not mill.keywind.cc, which is Mailgun's receiving domain. So each alert is exactly one Mailgun send and never re-enters Mailgun's inbound routing. No extra setup: keywind.cc's Cloudflare catch-all rule already delivers it. Requirement: one test send before relying on it (also shows whether alerts land in spam). The full design is in ERROR-ALERTING-PLAN.md, revised 2026-09-11 against the current code (supervisor, boot start, retention sweeps, sheets-link delivery).

Key points:

  • New leaf package internal/alert: Reporter interface (no-op default), Mailer (satisfied by mailgun.Service.SendAlert) and Ledger (satisfied by store.Store). A package of its own because mailgun imports app and both need to report.
  • Alert only on systemic failures. Split the two kinds of "job failed" in App.execute: a valid result.json with success:false is the sender's problem and never alerts. A timeout, crash, missing/invalid result, missing transformer, or success:true with a nonzero exit is systemic.
  • Mandatory throttle, persisted in SQLite: 15-minute cooldown per category, global caps of 10 per hour and 20 per day (rolling 24h), suppressed counts in the next email. The daily cap is set by the Mailgun Free plan: 100 sends a day, shared with replies, and Mailgun rejects further sends past that. The hourly cap alone would allow 240 a day and lock out replies. Reaching the cap sends one final notice, which counts within the 20. Persisted because the supervisor restarts a crashing worker every ≤120s, and an in-memory throttle would reset each time.
  • Crashes are reported by the next process. The supervisor passes FILEMILL_PREVIOUS_EXIT / FILEMILL_RAPID_RESTARTS to the restarted worker, which sends one restart alert, folded together with the count of jobs store.Open marks interrupted.
  • Report never blocks: a buffered queue drained by a goroutine. A failed alert send is logged and dropped, never re-reported.

Alert sites

Intake 500s · wrong route (store(notify=)) · systemic job failures · job panics · reply send failing ≥5 min · MarkEmailDelivered failing after a successful send (immediate; resends every second, see #6) · sheets-link publish failures · orphaned Drive files · job-claim errors ≥1 min · both retention sweeps · restart after crash.

Deliberately not alerting: 401/400 webhook noise, unrouted/disallowed/benign mail.

What this can't cover

Startup fatal()s (no reporter yet), a failing Mailgun send (it's the alert channel), and a machine that's off. These are the heartbeat's job, #5.

Phases (one PR each; tests with fakes written first)

  1. internal/alert core: throttle, queue, persisted Ledger. No behavior change.
  2. Job taxonomy split in execute, plus panic recovery.
  3. Mailgun sites and main.go wiring. send takes its subject verbatim; add SendAlert and alert_recipient config. Alerting goes live.
  4. Crash reporting across restarts: supervisor env vars and the interrupted-job count; recover in the delivery and sweep loops.
  5. Live verification against support_filemill@keywind.cc (after a test send).

About 3 days in total.

Decisions (defaults)

Config

alert_recipient: support_filemill@keywind.cc   # in the gitignored config/email.yaml; empty = disabled
# alert_cooldown_minutes: 15
# alert_max_per_hour: 10
# alert_max_per_day: 20      # Mailgun Free plan budget; revisit on a paid plan

Activity

  1. changed the title [-]Ops: email operator error alerts to support@keywind.cc[/-] [+]Ops: email operator error alerts to support@mill.keywind.cc[/+] on Sep 11, 2026
  2. changed the title [-]Ops: email operator error alerts to support@mill.keywind.cc[/-] [+]Ops: email operator error alerts to support_filemill@keywind.cc[/+] on Sep 11, 2026
  3. brocla commented on Sep 12, 2026

    @brocla
    OwnerAuthor

    Done, merged in #29 (tagged v0.3.0) and running in production since 2026-09-11.

    What shipped

    • internal/alert: a Reporter with a throttle in front of the Mailgun sender - 15 minutes per category, 10 an hour, 20 a day over a rolling 24h - draining a queue on its own goroutine so Report never blocks. Throttle state lives in SQLite, so a crash-looping worker cannot spend the Free plan budget. If the database fails, the throttle falls back to an in-memory copy instead of going silent, since a sick database is exactly what is worth hearing about.
    • Alert sites per the plan: systemic job failures (crash, timeout, missing or invalid result.json, missing transformer, a result contradicting its exit code), panics, intake 500s, the store(notify=) route warning, replies or sheets-link publishes failing for 5 minutes, MarkEmailDelivered failing (immediately), orphaned Drive files, job-claim errors, both retention sweeps, and a restart after a crash. A transformer that rejected its input through the contract stays silent, as intended.
    • Crash reporting across restarts: the supervisor passes the previous exit code and rapid-restart count to each relaunch, folded together with the count of jobs left running.
    • filemill alert-test sends one test alert, to prove the channel before relying on it.

    Verified in production, not just in tests

    A test send; one restart email from a killed worker with four further restarts suppressed inside the cooldown; 21 probe jobs producing exactly 2 emails and 19 suppressions, the second carrying the suppressed count. Those four suppressed restarts were recorded by one build and reported 2.5 hours later by another, after a binary swap - which is the persisted ledger doing what it was built for.

    The bad-Mailgun-domain check from phase 5 was skipped deliberately: it delays real senders' replies, and that path is covered by tests and by the four live alerts above.

    Fixed along the way

    • The reply resend storm (see Reliability: at-least-once delivery can send a duplicate reply (#6) #6): a reply that was sent but could not be marked delivered was re-sent every second, which would flood the sender and spend the day's Mailgun budget in minutes. Only the mark is retried now. Reliability: at-least-once delivery can send a duplicate reply (#6) #6's remaining crash window is documented and accepted; a per-submission retry cap and dead-letter state is still worth its own issue.
    • store.Open marked every running job interrupted, so submit or jobs get run while the worker was mid-job made the delivery loop send that sender a premature reply. Only the starting continuous worker does this now.
    • Two misleading sender-facing messages: a result claiming success with a nonzero exit no longer repeats the transformer's success text, and an interrupted job now asks the sender to send the file again instead of leaving the line blank.

    Still not covered, as the plan said from the start: startup fatal()s, a Mailgun outage (it is the alert channel), and a machine that is off. Those remain #5's job.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions