Changed project or global settings produce a settings:updated Activity Log entry, except engine-owned churn (engineLastActiveAt, engineActiveSinceMs, postgresMigrationInboxMessageSentAt, and secretsSyncPassphraseConfigured). Secret-bearing values are redacted or summarized. Workflow setting values, including requirePrApproval, never appear in this payload; use the configuration revision history for those changes. See Who changed this setting, and when for both audit mechanisms and their query contracts.
Fusion provides two complementary audit surfaces: append-only configuration revisions for project settings and the Activity Log for concise operational events. The related rebase controls are documented in Worktree and pre-merge rebase settings.
GET /api/config/revisions reads project-setting revisions from the PostgreSQL async revision store. It accepts only configKind=project-settings (or an omitted configKind), returns newest first by createdAt and sequence, and supports limit (1–500; default 100) plus offset; the response includes hasMore. This route is unavailable without the PostgreSQL async layer.
POST /api/config/revisions/:revisionId/rollback applies a selected historical revision and records that action as a new, forward-only rollback revision rather than deleting history. The configuration_revisions table stores project-partitioned records including changed_by, before, after, diffs, source (mutation or rollback), and the database-assigned sequence that orders same-timestamp writes. GLOBAL_CONFIGURATION_OWNER_ID is the reserved owner for user-global settings history, preventing it from being attributed to whichever project happened to issue the write.
GET /api/activity returns newest-first events. Its query parameters are limit (default 100, maximum 1000), since (an ISO timestamp), and type. Valid types are task:created, task:moved, task:updated, task:deleted, task:merged, task:failed, and settings:updated. DELETE /api/activity clears the activity log for maintenance. settings:updated entries are generated by the Task Store listener using the generic diff policy in packages/core/src/task-store/settings-activity.ts.
GET /api/agent-activity is a separate project-scoped history for agent-attributed task, gate, approval, and state transitions; it does not expand the Activity Log type contract. Its inspectable wire, cursor, continuation, and retention contract is Agent activity API contract. Live dashboards receive durable agent:activity SSE frames; reconnecting consumers can close any bounded-tail truncation gap through this route. Payload metadata is IDs/counts/outcomes-only, and only roster-proven agent attribution represents an org-map node.
- Attribution: revision
changedByrecords the actor supplied to the settings mutation. Older investigations may encounter historical{ kind: "human", id: "local-user" }records, but current settings-operation defaults use the system actor when a caller supplies none; treat historical attribution according to the stored record. - Revision history: the API is paginated rather than capped to one unpageable 100-row window. Request later pages with
offset; concurrent appends can make offset pages overlap, so preserve revision IDs when collecting an incident timeline. - Activity coverage: the former four-key Activity Log allowlist is removed. All changed project/global settings are recorded except the four churn or secret-configuration exclusions above; values for sensitive keys are redacted. Workflow settings remain outside the
settings:updatedpayload.
Engine and core subsystem loggers (createLogger) expose a debug() level for routine diagnostics. It is off by default so the TUI log pane and engine stderr show state changes rather than repeated resting-state chatter.
| Severity | Use it for |
|---|---|
debug() |
Repeated poll/sweep lines with unchanged state, expected skips and no-ops, per-item progress, expected-and-handled failures (fallback/retry/optional dependency), and diagnostics already recorded as run-audit events. |
log() |
A state transition or operator-visible event. |
warn() |
Handled degradation that needs eventual operator attention. |
error() |
An unrecovered failure requiring operator action; do not use it when code recovered or scheduled a retry. |
Most new diagnostic sites should therefore start at debug() and only be promoted when they meet a higher-severity rule. debug() output is enabled per subsystem with the logger prefix:
FUSION_DEBUG=scheduler # one subsystem
FUSION_DEBUG=scheduler,merger # several
FUSION_DEBUG=1 # everything (also: true, all, *)The variable is re-read per call, so it can be toggled on a long-lived process without recreating loggers. Debug lines emit under the info severity marker and render like any other info line.
Currently debug-gated classes include local/default routing, capacity and re-entrancy skips, poll/sweep no-actions, per-step success/progress, optional integration probes, successful verification bookkeeping, per-session agent setup (agent-session runtime/fallback resolution, planning mode, stuck-detector track bookkeeping), per-skill selection chatter and intentional skill-exclusion notices ([skills] info: … disabled by project execution settings), expected-missing PROMPT.md seed reads (ENOENT), token-cache metrics JSON, duplicate runtime Specifying … echoes, zero-count recovery summaries, mission "no linked feature" skips, event-driven "triggering scheduling" echoes, auto-claim snapshot invalidation, assignment-trigger skip guards (ephemeral/disabled/active-run), runtime Scheduled/Started executing echoes of executor Starting, worktree warm-reuse, runtime-env injection counts, baseCommitSha capture, parse-steps reconcile diagnostics, graph node worktree re-acquire, ephemeral already-owned worker skip, embedded-postgres rejoin "already running", and fn_run_verification non-timeout command-fail detail (failed done lines stay at info). Each agent session instead emits one normal-level [skills] summary with the available count, resolved forced skills, and any unavailable forced requests. State-changing recovery and dispatch outcomes remain visible (Starting/Specifying/Worktree created/Auto-merge merged/column moves/slow hold-release).
Dashboard server code uses the core logger only: import { createLogger } from "@fusion/core";. Do not import an engine logger, use a relative cross-package logger path, or add a dashboard-local logger implementation.
packages/engine/src/__tests__/log-severity-manifest.ts and package contract tests pin individual demotions and the no-bare-console rule. This makes an accidental severity reversion a CI failure instead of an operator-visible log flood.
Executor, heartbeat, and planning runs emit one goal-injection diagnostic with outcome applied, no-goals, or disabled-or-failed.
- Run-audit event:
prompt:goal-injection(databasedomain, target lane) with metadata{ lane, outcome, goalCount, goalIds, provenanceGoalIds, truncated, reason?, errorClass?, runId?, agentId?, taskId? }. - Goal anchoring events also persist
metadata.goalIds(alongside existing count/tool fields):goal:injection-applied/goal:injection-skipped→{ lane, count, goalIds, truncated?, reason? }goal:retrieval-invoked→{ toolName, count, goalIds, notFound }
- Run cited-goals read path:
GET /api/agents/:id/runs/:runId/cited-goalsreturns{ runId, taskId?, injectedGoalIds, retrievedGoalIds, citedGoalIds }aggregated fromgoal:*+prompt:goal-injectionrun-audit events. - Task log (executor lane with
taskId):[goal-injection] <outcome> count=<n> ids=<json-array> provenance=<json-array> truncated=<bool> .... goalIds/goalCountdescribe the active goals injected into the prompt;provenanceGoalIdsadditively records mission-derived task provenance and does not affect prompt selection.- Guardrail: diagnostics persist goal IDs/counts only; never prompt text, goal titles, or goal descriptions.
Agent reflection generation emits one run-audit event for every AgentReflectionService.generateReflection attempt, covering manual dashboard requests, executor/post-task tools, heartbeat tools, and self-improve callers from the shared service seam.
- Run-audit events (
databasedomain, targetagentId):reflection:generatedmetadata:{ agentId, trigger, taskId?, reflectionId, tasksCompleted?, tasksFailed?, avgDurationMs?, commonErrorCount, insightCount, suggestedImprovementCount }.reflection:skippedmetadata:{ agentId, trigger, taskId?, reason: "no-history" | "not-completed" }.reflection:failedmetadata:{ agentId, trigger, taskId?, errorClass }.
- Trigger taxonomy is preserved from
ReflectionTrigger:manual,periodic,post-task, anduser-requested. - The events use synthetic run context with phase
reflectionand source equal to the trigger so they correlate in the run-audit stream without requiring caller-specific wiring. - Guardrail: reflection diagnostics persist IDs, counts, reasons, and error classes only; never prompt text, reflection summaries, insight strings, suggested-improvement text, or free-form trigger details.
AgentReflectionService.captureTaskPerformance is a deterministic, non-LLM counterpart to generateReflection: it runs once per completed task at the executor completion seam (TaskExecutor.signalTaskComplete), guarded by reflectionService presence, settings.reflectionEnabled, and an assigned agent id. It never calls the model provider and persists a compact structured post-task ReflectionMetrics record — duration, packages/files touched, verification command(s) + file-scoped-vs-broader classification, and retry/rework count — sourced only from the completed Task record. Fields whose source is unavailable are omitted, never fabricated.
- Run-audit event (
databasedomain, targetagentId):reflection:capturedmetadata:{ agentId, trigger: "post-task", taskId?, reflectionId, retryReworkCount?, filesTouchedCount?, packagesTouchedCount?, verificationFileScoped?, durationMs? }.- Skipped/failed captures reuse the existing
reflection:skipped(reason: "no-history" | "not-completed") andreflection:failed(errorClass) event types above.
- Guardrail: capture telemetry stays ids/counts/outcomes-only —
verificationScopeReasonfree-text, the deterministic one-linesummary, and any prompt/reflection prose never reach run-audit metadata (they are stored only in theReflectionMetrics/AgentReflectionrecord itself). - Capture is best-effort and fire-and-forget: a capture failure never blocks or fails task completion, and an in-memory per-taskId guard prevents duplicate captures across the executor's several completion call sites (fresh completion, duplicate in-review re-entry, auto-recovery, paused-after-completion finalize, retry-completed).
The dashboard insight router runs stale-run recovery sweeps for project_insight_runs rows stuck in pending/running without a live controller owner.
- Recovery writes
terminalCause: "orphaned_active_run_recovered"and lifecycle failure metadata (failureClass: "non_retryable",retryable: false). - Recovery appends both
warningandstatus_changedevents onproject_insight_run_eventswithmetadata.recovery = "orphaned_active_run". metadata.recoverySourceindicates where recovery occurred:startup,periodic,drive_by, ormanual.
Self-healing now runs surface-dependency-blocked-todos during both startup recovery and periodic maintenance.
- Normal path emits a workflow insight titled
Backlog health: dependency-blocked todos YYYY-MM-DD. - Fallback path (insight store unavailable) writes a per-task log entry prefixed with
[dependency-blocked-todo]against the top blocker task. - Reporter summary warnings include group count, total blocked Todo count, and top blocker IDs.
Operator interpretation:
ageBucket: "fresh"→ expected dependency queueing.ageBucket: "aging"→ review blocker progress.ageBucket: "stale"→ emerging stall; escalate/unblock blocker.
When a Fusion-owned Windows embedded cluster reports the exact 0xC0000142 backend DLL-initialization failure followed by PostgreSQL's shutdown chain, the existing startup/System diagnostic sink records:
detected Windows DLL initialization shutdown; attempting one owned-cluster recoveryWindows owned-cluster recovery completed; existing pools may reconnect, orWindows DLL initialization recovery failed after one retry; restart Fusion and inspect the System log
The recovery budget is one per lifecycle and applies only to a post-readiness cluster Fusion started. It never restarts a joined cluster. On the terminal message, restart Fusion; if it repeats, retain the System log and bundled-runtime version for support rather than deleting the data directory.
The process supervisor logs when it registers a supervised child, starts teardown, expires the grace window, escalates to SIGKILL, or observes a natural child exit.
spawned pid=<pid> pgid=<pgid|n/a> command=<cmd>— child registered for parent-death supervision.terminating pid=<pid> pgid=<pgid|n/a> reason=<reason>— teardown cascade started.grace expired for pid=<pid>; escalating to SIGKILL— child ignored the grace window.sent SIGKILL to pid=<pid> pgid=<pgid|n/a>— hard-kill escalation sent.maxLifetime exceeded for pid=<pid> after <ms>ms— lifetime watchdog fired.child pid=<pid> exited naturally code=<n|null> signal=<n|null>— child deregistered after exit.
surface-in-review-stalls- Log prefix:
In-review stall surfaced [ - Purpose: reason-driven in-review stall detector (
merge-blocker, retry exhaustion, no-worktree, transient merge-status orphaning).
- Log prefix:
surface-in-review-stalled- Log prefix:
In-review stalled surfaced [in-review-stalled]: quiet ... - Purpose: time-quiet detector for unpaused in-review tasks beyond
inReviewStalledThresholdMs. - Non-overlap: skipped when reason-driven
In-review stall surfaced [is fresh, and skipped for paused tasks (owned by stale-paused-review).
- Log prefix:
surface-stale-paused-reviews- Log prefix:
Stale paused review surfaced [stale-paused-review]: paused ... - Purpose: paused in-review backlog-health detector gated by
stalePausedReviewThresholdMs.
- Log prefix:
- FN-5335 backward-move annotations
- Log prefix shape:
[<stage-name>] <taskId>: triple-proof not satisfied — no action (operator-decides) - Representative stage names:
no-progress-no-task-done,partial-progress-no-task-done,stale-incomplete-review,ghost-review,missing-worktree-review,stuck-merge-deadlock,finalize-no-op-review,reclaim-pr-conflict,reclaim-self-owned-branch-conflict,auto-rebound-paused-scope-decay.
- Log prefix shape:
Time-based stuck detection floors activity timestamps using settings.engineActiveSinceMs plus settings.engineActivationGraceMs (default 300000). The runtime stamps engineActiveSinceMs on startup and each unpause transition so engine pause/downtime does not count as quiet time.
A silent or repetitive session is disposed and the same task is re-dispatched after the old execution lock unwinds. Recovery preserves the current column, workflow node, current step, worktree, branch, and completed-step progress. Repeated silence remains automatically recoverable; no stuck-kill budget, terminal failure, decompose instruction, or human approval hold is emitted. User-paused tasks remain untouched.
FN-5941 adds a convergence backstop for repeated todo↔in-progress churn.
- Suppressed false-positive branch-conflict recovery emits
task:reclaim-self-owned-branch-conflict-no-actionwith liveness metadata (taskId,branch,worktree,checkedOutBy,executionStartedAt,executionAgeMs,graceMs,liveWorktreeBoundBranch,reason). In-progress-limbo and stuck-budget terminal recovery events were removed by FN-217. - Scheduler settle-window diagnostic:
Task <id> was engine-requeued <age>ms ago — waiting <settleMs>ms settle window before redispatch. - Terminal audit event:
task:dispatch-oscillation-terminalizedwith{ taskId, cycleCount, windowMs, lastMoveSource }. - Outcome: task stays in
todo, is auto-paused withpausedReason: "dispatch-oscillation", and requires operator unpause/forward progress to reset the counter.
FN-5346 adds a same-task stale-binding reconcile marker before worktree removal:
[FN-5346] <taskId>: dropped stale self-owned activeSessionRegistry entry before removeWorktree at <worktreePath>- Follow-up task log entry:
Cleared stale self-owned active-session entry before remove
Engine stop now aborts in-flight executor AI sessions before the runtime drain wait.
- Executor summary log:
[executor] abortAllInFlight: aborted N task surface(s) — engine stop - Runtime warning when in-flight work still exists after configured post-abort drain:
[runtime-stop] post-abort drain timeout reached with N tasks still in-flight
Use these together to distinguish expected immediate session teardown from genuinely stuck cleanup surfaces that outlive the configured runtimeStopDrainMs window.
Direct-report stale decisions in HeartbeatMonitor.buildReportsHealthSection() now emit a structured log when an agent is marked **stale**.
- Log shape:
[reports-health] stale report <agentId> intervalSource=<source> staleThresholdMs=<n> heartbeatAgeMs=<n> intervalSourcevalues:runtimeConfig— interval came from cached per-agent runtime configpersisted-agent— cache was missing/sparse; interval came from persistedgetAgent()rowmonitor-default— no per-agent interval available; monitor default interval used
staleThresholdMsis the computed stale threshold (max(1.5 × interval, 5m floor))heartbeatAgeMsis the report's current heartbeat age at classification time- Healthy reports do not emit this diagnostic; only stale decisions do
Dashboard Phase 1 resume instrumentation adds observation-only client/server traces for refetch/reconnect attribution. It does not change visibility/pageshow/SSE behavior; FN-5392 consumes this data for fixes.
- Client event shape (
ResumeEvent):{ ts, view, trigger, projectId?, gapMs?, replayAttempted, replayFromEventId?, lastEventId?, sseChannel?, reason?, detail? }. - Trigger taxonomy:
visibility,focus,pageshow,sse-error,sse-reconnect,sse-open,remount,route-active,route-inactive,project-context-change. The browser and diagnostics route both import this accepted vocabulary frompackages/dashboard/src/shared/resume-triggers.ts;focusis accepted for tab-return diagnostics. - Sources:
sse-bus(pageshow, visiblevisibilitychange,openChannel,forceReconnect, EventSourceerror)- Hooks:
useTasks(visibility,focus,sse-reconnect),useChatRooms(sse-reconnect),useChat(sse-open,project-context-change) - Components:
BoardandChatViewmount/unmount route markers (remount/route-active/route-inactive), andtaskActivityFeedwhen Feed becomes visible or its stream reconnects
- Access paths:
- Client ring (500):
window.__fusionDebug.resumeInstrumentation.get()/.clear() - Server ring (5000, in-memory):
GET /api/diagnostics/resume-events?limit=&since=&view=returns{ events, droppedSinceLastRead }
- Client ring (500):
- Client batching: POST
/api/diagnostics/resume-eventsin idle batches (<=25per POST). - Disable knob:
window.__fusionDebug.resumeInstrumentation.setEnabled(false).
FN-5415 extends this coverage across remaining board/data visibility hooks: useNodes, useMeshState, useProjects, and useManagedDockerNodes. Each now emits trigger: "visibility" with reason: "debounced-refresh" when refresh is taken and reason: "debounce-skipped" (including detail.timeSinceLastRefreshMs) when suppressed by debounce. This completes board/data-hook resume-correlation coverage needed for FN-5392 Phase 2 remediation analysis.
FN-5416 extends resume-correlation coverage to stream-focused hooks and their primary route shells:
- Hooks
usePrChecksStream:remount,visibilityuseDevServerLogs:project-context-change,sse-open,sse-reconnectuseResearch:sse-open,sse-reconnectuseBackgroundSessions:sse-open,sse-reconnectuseAgentLogs:project-context-change,sse-open,sse-reconnecton/api/tasks/:id/logs/stream(live tail via SSE; historical reads are backed by.fusion/tasks/{ID}/agent-log.jsonl)
- Route shells
DevServerView:remount/route-active/route-inactiveResearchView:remount/route-active/route-inactive
Fusion merge cleanup treats a narrow class of temporary merge/post-merge worktree removal failures as non-fatal only after Git admin state proves there is no registered worktree leak.
- Applies to Fusion-created temp merge paths such as
fusion-ai-merge-*and post-merge paths such aspost-merge-*duringmerger-cleanup/merger-post-mergeremoval. - Trigger shape:
git worktree remove --force <path>fails with validation text such asfatal: validation failed, cannot remove working tree: '<path>/.git' is not a .git file. - Recovery proof: Fusion runs
git worktree prune, then inspectsgit worktree list --porcelain. - Harmless classification: if the target path is absent from porcelain after prune, the merger logs that cleanup remove failed but no registered worktree remains. If a directory still exists, Fusion reports it as residue for operator inspection; it does not delete arbitrary
/var/folderscontent. - Leak classification: if the target path is still present in porcelain after prune, the cleanup failure remains visible as a real registered-worktree leak.
Operator verification command:
git worktree list --porcelain | grep -F "<temp-worktree-path>"No output means Git no longer registers that temp path; matching worktree <temp-worktree-path> output means the leak is still registered and needs operator cleanup.
Hold-release summaries include prefetch, IR-resolution, and evaluation accumulators; measured sweep-attributable read counts; scanned-task and held-candidate counts; released/held totals; unevaluatedCount; and budgetOverrunMs. A budget-truncated sweep warns even when it stops in the preamble. Resolver reads are measured by a delegating counting facade (and direct sweep reads at their call sites), not cache-size inference. PostgreSQL health probes return a degraded timeout reason when the pool is saturated. Migration-state probe timeouts remain advisory after database and task-ID integrity checks succeed.
The dashboard HTTP server exposes GET /metrics — a plain-text Prometheus exposition endpoint (text/plain; version=0.0.4; charset=utf-8) serving the system / runtime / Fusion-domain measurements that earlier CPU and UI-responsiveness diagnoses had to collect by hand (curl /api/health for event-loop latency, ps for child-process cadence, psql for query rate, RAM-usage sampling for RSS). A curl /metrics returns the same numbers a Prometheus/Grafana scrape would consume; no OTLP collector or prom-client dependency is involved — the module serializes Prometheus text in-process.
- Public, outside
/api: the route is mounted at the app level (before the SPA catch-all andexpress.static), so it returns Prometheus text rather thanindex.html. Daemon bearer-token auth only protects/api/*;/metricsis intentionally unauthenticated. The body carries numeric values plus low-cardinality string label values — no secrets, no prose, no request payloads, no run-audit telemetry. The label values do include registered project identifiers (fusion_domain_project_running_agents{project="…"}) and board column names, so any client that can reach the port can enumerate open project ids. Bind the port to a trusted network when that disclosure is not acceptable (CodeRabbit Major review fix 2026-08-18-11:53: the earlier "numeric gauges only" wording did not match the emitted body). - Non-blocking by construction: the
/metricshandler renders synchronously from pre-read snapshots. It performs zero awaited I/O — a scrape completes in O(metric count) work and can never itself starve the event loop or trigger an on-demand DB/ps query. All sampling happens on pre-read tick timers (see cadence below). - Run-audit blackline (FN-7158/FN-7528): no metric content, relabeled names, timestamps, or numeric snapshots are written to the run-audit. The endpoint computes on scrape from in-process state; nothing here emits a documented run-audit event.
| Metric | Type | Cadence source | Meaning |
|---|---|---|---|
fusion_system_request_count_total |
counter | live request pipeline | Requests served through the latency recorder |
fusion_system_request_latency_ms{quantile="p50|p95|max"} |
gauge | live request pipeline | Histogram over the recent served-request ring |
fusion_system_request_latency_bucket{le="…"} |
gauge | live request pipeline | Cumulative bucket counts over the ring |
fusion_system_last_request_age_ms |
gauge | live request pipeline | ms since the last served request — grows during event-loop starvation (the freeze indicator) |
fusion_system_process_rss_bytes |
gauge | ~5s tick | RSS of the serving process |
fusion_system_process_heap_used_bytes / …_heap_total_bytes |
gauge | ~5s tick | Heap usage of the serving process |
fusion_system_cpu_user_seconds_total / …_system_seconds_total |
counter | ~5s tick | CPU time consumed by the serving process |
fusion_system_child_process_spawn_total |
counter | spawn hook | Cumulative child_process spawn/fork/execFile/exec invocations |
fusion_system_child_process_spawn_total_by_kind{kind="…"} |
counter | spawn hook | Per-kind cumulative spawn counts |
fusion_system_git_child_processes |
gauge | ~15s ps |
Live git children of the serving process (best-effort) |
fusion_domain_postgres_queries_per_second |
gauge | ~5s tick | Derived PG xact rate from pg_stat_database deltas (best-effort) |
fusion_domain_projects_total / …_active / …_idle |
gauge | ~5s tick | Registered open project split by running-agent activity |
fusion_domain_project_running_agents{project="…"} |
gauge | ~5s tick | Running agents per registered project |
fusion_domain_board_tasks{column="…"} |
gauge | ~5s tick | Tasks per board column across registered projects |
- The request-latency recorder is an Express middleware mounted before route handlers, so it measures the live serving path (including
GET /api/health), not a synthetic probe. - Process CPU/memory gauges update every ~5s; the git-subprocess gauge every ~15s; the PG rate and domain gauges every ~5s. All timers are
unref()'d so a running sampler never keeps the process alive. The process and git arms invoke the same sampler, so they share ONE in-flight guard key: at the default 5s/15s cadence the arms coincide every 15 s and the coinciding tick skips the duplicatepsprobe instead of double-probing (CodeRabbit Major review fix 2026-08-18-11:53). - Sampler start is wired into the server's listen override and stop into the close handler (co-located with the OTLP exporter lifecycle), and runs in both headless and non-headless modes. Both start and stop are idempotent and never break server startup/shutdown; both calls are wrapped in try/catch (same pattern as the OTLP exporter) so a failure is logged and can never skip the remaining close handlers.
- Spawn-count hook: on start,
child_process.spawn,fork,execFile, andexecare wrapped with an atomic counter that delegates to the original via.apply, so child spawning (includingsuperviseSpawn/runCommandAsync/execFileAsync) is never broken. On stop the original functions are restored exactly. The hook is idempotent (starting twice never stacks a second wrap).
- Git-subprocess gauge: a single-level
ps -o comm= --ppid <pid>scan every ~15s counts livegitchildren of the serving process. It never recurses and never scans the whole process tree. Whenpsis unavailable (Windows, non-POSIX, missing procfs), the gauge degrades to0rather than throwing. - PG query-rate sampler: reads cumulative
pg_stat_databasexact_commit/xact_rollback deltas PER DATABASE from the store's live async layer on the tick, normalized to a per-second rate. It is best-effort: on a privilege-fenced PG, transient pool error, or absent async layer it keeps the last-known rate (or0on the first invalid sample) rather than throwing or hammering the DB. A failed-probe gap invalidates the retained baseline — the first success after the gap re-baselines and keeps the last-known rate, so a stats reset landing inside the gap can never produce a cross-epoch rate — and a backward delta on ANY single database is treated as a stats reset even when the cross-database sum stays positive. The baseline is also marked stale on stop: after a dashboard stop/restart the first success re-baselines and keeps the last-known rate, so a stats reset during the stop gap can never emit a cross-epoch rate either (Greptile P1 review fix 2026-08-18-11:53). Embedded PostgreSQL reads may be operator-only depending on context. - Domain gauges come only from already-open project stores via
countRunningAgentsInStore/listRegisteredProjectStores/store.listTasks({ slim: true }). Empty/undefined/duplicate project and empty-column states produce well-formed0-valued or absent metric lines, never malformed output; the sampler never opens a store or starts an engine to answer a scrape. - Value safety: non-finite or non-numeric values are coerced to
0so a single bad sample cannot abort the whole body; invalid metric/label names are sanitized to the permitted Prometheus character set.
- Endpoint acceptance (
packages/dashboard/src/routes/__tests__/metrics-endpoint.test.ts): drivesGET /metricsthrough the real server creator and an independent exposition-text parser (packages/dashboard/src/__tests__/prometheus-text-parse.ts) to prove the served body is well-formed Prometheus text covering all five measurement gaps and is NOT the pre-RUFU-081 SPAindex.htmlfallback, that a scrape writes no run-audit row, and that repeat scrapes render a fresh, bounded snapshot. - Sampler acceptance (
packages/dashboard/src/metrics/__tests__/metrics-samplers-acceptance.test.ts): exercises the orchestratorrender()end-to-end — synchronous pre-read render (no on-demand DB/ps on a scrape), all five family gaps as finite gauges, and the spawn-count hook incrementing on a real child process with the wrapper restored infinally. - Parser unit cases (
packages/dashboard/src/__tests__/prometheus-text-parse.test.ts): gauge/counter_total/NaN/Inf/labeled families and non-exposition-text rejection.