Repository navigation
Ops: email operator error alerts to support_filemill@keywind.cc #7
Description
Activity
- changed the title
[-]Ops: email operator error alerts to support@keywind.cc[/-][+]Ops: email operator error alerts to support@mill.keywind.cc[/+]on Sep 11, 2026 - changed the title
[-]Ops: email operator error alerts to support@mill.keywind.cc[/-][+]Ops: email operator error alerts to support_filemill@keywind.cc[/+]on Sep 11, 2026 - added 6 commits that reference this issue
on Sep 11, 2026 Done, merged in #29 (tagged v0.3.0) and running in production since 2026-09-11.
What shipped
internal/alert: aReporterwith a throttle in front of the Mailgun sender - 15 minutes per category, 10 an hour, 20 a day over a rolling 24h - draining a queue on its own goroutine soReportnever blocks. Throttle state lives in SQLite, so a crash-looping worker cannot spend the Free plan budget. If the database fails, the throttle falls back to an in-memory copy instead of going silent, since a sick database is exactly what is worth hearing about.- Alert sites per the plan: systemic job failures (crash, timeout, missing or invalid
result.json, missing transformer, a result contradicting its exit code), panics, intake 500s, thestore(notify=)route warning, replies or sheets-link publishes failing for 5 minutes,MarkEmailDeliveredfailing (immediately), orphaned Drive files, job-claim errors, both retention sweeps, and a restart after a crash. A transformer that rejected its input through the contract stays silent, as intended. - Crash reporting across restarts: the supervisor passes the previous exit code and rapid-restart count to each relaunch, folded together with the count of jobs left running.
filemill alert-testsends one test alert, to prove the channel before relying on it.
Verified in production, not just in tests
A test send; one
restartemail from a killed worker with four further restarts suppressed inside the cooldown; 21 probe jobs producing exactly 2 emails and 19 suppressions, the second carrying the suppressed count. Those four suppressed restarts were recorded by one build and reported 2.5 hours later by another, after a binary swap - which is the persisted ledger doing what it was built for.The bad-Mailgun-domain check from phase 5 was skipped deliberately: it delays real senders' replies, and that path is covered by tests and by the four live alerts above.
Fixed along the way
- The reply resend storm (see Reliability: at-least-once delivery can send a duplicate reply (#6) #6): a reply that was sent but could not be marked delivered was re-sent every second, which would flood the sender and spend the day's Mailgun budget in minutes. Only the mark is retried now. Reliability: at-least-once delivery can send a duplicate reply (#6) #6's remaining crash window is documented and accepted; a per-submission retry cap and dead-letter state is still worth its own issue.
store.Openmarked every running jobinterrupted, sosubmitorjobs getrun while the worker was mid-job made the delivery loop send that sender a premature reply. Only the starting continuous worker does this now.- Two misleading sender-facing messages: a result claiming success with a nonzero exit no longer repeats the transformer's success text, and an interrupted job now asks the sender to send the file again instead of leaving the line blank.
Still not covered, as the plan said from the start: startup
fatal()s, a Mailgun outage (it is the alert channel), and a machine that is off. Those remain #5's job.
Backlog: Observability & ops
Problem
When something goes wrong (systemic job failure, delivery failure, panic, a stuck worker loop, a crash), nobody is notified. An unattended background service fails silently.
Proposed solution
Email operator alerts to
support_filemill@keywind.cc(replaces the earliersupport@keywind.ccandsupport@mill.keywind.cc). It's deliberately onkeywind.cc, whose mail goes through Cloudflare Email Routing, notmill.keywind.cc, which is Mailgun's receiving domain. So each alert is exactly one Mailgun send and never re-enters Mailgun's inbound routing. No extra setup:keywind.cc's Cloudflare catch-all rule already delivers it. Requirement: one test send before relying on it (also shows whether alerts land in spam). The full design is inERROR-ALERTING-PLAN.md, revised 2026-09-11 against the current code (supervisor, boot start, retention sweeps, sheets-link delivery).Key points:
internal/alert:Reporterinterface (no-op default),Mailer(satisfied bymailgun.Service.SendAlert) andLedger(satisfied bystore.Store). A package of its own becausemailgunimportsappand both need to report.App.execute: a validresult.jsonwithsuccess:falseis the sender's problem and never alerts. A timeout, crash, missing/invalid result, missing transformer, orsuccess:truewith a nonzero exit is systemic.FILEMILL_PREVIOUS_EXIT/FILEMILL_RAPID_RESTARTSto the restarted worker, which sends onerestartalert, folded together with the count of jobsstore.Openmarksinterrupted.Alert sites
Intake 500s · wrong route (
store(notify=)) · systemic job failures · job panics · reply send failing ≥5 min ·MarkEmailDeliveredfailing after a successful send (immediate; resends every second, see #6) · sheets-link publish failures · orphaned Drive files · job-claim errors ≥1 min · both retention sweeps · restart after crash.Deliberately not alerting: 401/400 webhook noise, unrouted/disallowed/benign mail.
What this can't cover
Startup
fatal()s (no reporter yet), a failing Mailgun send (it's the alert channel), and a machine that's off. These are the heartbeat's job, #5.Phases (one PR each; tests with fakes written first)
internal/alertcore: throttle, queue, persistedLedger. No behavior change.execute, plus panic recovery.main.gowiring.sendtakes its subject verbatim; addSendAlertandalert_recipientconfig. Alerting goes live.recoverin the delivery and sweep loops.support_filemill@keywind.cc(after a test send).About 3 days in total.
Decisions (defaults)
REPLY_FROM.Config