Files
Fenris/docs/adr/0005-failure-detection-and-recovery.md

4.7 KiB

5. Failure and recovery: visible degradation, never fabrication

Status

Accepted — resolves Define failure and recovery behavior on the Wayfinder map.

Context

The observation store (ADR 0001) and the service lifecycle (ADR 0003) settled single-writer transactions, bounded single failed runs, freshness grading (fresh / missed / stale), and absent-hour classification through power-on-hours evidence. Left open by the failure ticket: behavior per failure class — malformed observations inside the store, missed observations, store faults (unreadable, corrupt, or newer-schema database), and repeated collector failures — and how the habit record re-anchors after the store itself is lost.

Decision

  1. Malformed observations — refuse at the write boundary. The collector validates every row it would write against the store's domain invariants (hour seconds sum to 3600, non-negative DUW delta within a controller segment, coverage consistent with sample count). A violating run writes nothing for that run, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Readers (TUI, status) defensively exclude and count malformed rows as a contributing fact, but under a single trusted writer they should never see one. Store invariant: everything persisted is well-formed.
  2. Missed observations — never backfill. Fenris never interpolates, estimates, or fabricates an hour. Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. ADR 0003's power-on-hours classification is the only inference admitted.
  3. Store faults — degrade, never recreate over. An unreadable or corrupt database is a store fault: readers surface a "observation store unreadable" fact with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside, the next run starts a fresh store, and if the legacy import never completed, the still-present history.jsonl is re-imported. No built-in destructive command exists.
  4. Newer schema — readers refuse symmetrically. The TUI and status detect a user_version newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal in ADR 0001 and the forward-only upgrade rule of ADR 0004.
  5. Repeated collector failures — flat cadence, no escalation. The timer's retry is the recovery path; the settled freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff and no notification machinery; a persistent failure reads as stale exactly like any other gap.
  6. Drive-reported anomalies — facts, not alerts. critical_warning, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and status; no alerting or notification surface exists. Fenris observes and projects; it does not alarm. The projection is unaffected: endurance math consumes writes, not warnings.
  7. Orphaned samples — the collector re-anchors observed fact. When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent: only the sanctioned path of ADR 0003 records a user_disabled close. Coverage semantics stay intact without requiring a re-run of fenris-monitor enable after recovery.

Consequences

  • Validation lives at one boundary — the collector — so the store's contract is "everything in it is well-formed" and readers only defend against the impossible.
  • No synthetic data can ever enter the habit record; confidence categories can be trusted to reflect real evidence.
  • Store-fault recovery can lose history; the mitigation is the one-file backup story of ADR 0001, kept human-sanctioned so loss is never silent.
  • Period bookkeeping splits by epistemics: the helper records intent, the collector records observed fact.
  • Fenris stays fully local and silent: no notification, escalation, or alerting machinery anywhere.