docs(adr): 0005 failure and recovery — refuse bad writes, never backfill, degrade store faults; glossary term for store fault

This commit is contained in:
xavierk
2026-08-31 20:14:23 +05:30
parent ffa6f22777
commit d45d931a0d
2 changed files with 31 additions and 0 deletions
+4
View File
@@ -28,6 +28,10 @@ _Avoid_: Daemon uptime, calibration window
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline. The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic) _Avoid_: Data directory, history.jsonl, the database (generic)
**Store fault**:
The condition where the observation store is present but cannot be read or trusted — unreadable, corrupt, or written by a newer Fenris — degrading every view that depends on it rather than crashing or guessing.
_Avoid_: Database error, corruption, broken data
**Hour observation**: **Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage. One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry _Avoid_: Hourly record, hourly.jsonl entry
@@ -0,0 +1,27 @@
# 5. Failure and recovery: visible degradation, never fabrication
## Status
Accepted — resolves [Define failure and recovery behavior](https://git.bongbetic.com/xavierk/Fenris/issues/10) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The observation store ([ADR 0001](0001-observation-store-sqlite.md)) and the service lifecycle ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) settled single-writer transactions, bounded single failed runs, freshness grading (fresh / missed / stale), and absent-hour classification through power-on-hours evidence. Left open by the [failure ticket](https://git.bongbetic.com/xavierk/Fenris/issues/10): behavior per failure class — malformed observations inside the store, missed observations, store faults (unreadable, corrupt, or newer-schema database), and repeated collector failures — and how the habit record re-anchors after the store itself is lost.
## Decision
1. **Malformed observations — refuse at the write boundary.** The collector validates every row it would write against the store's domain invariants (hour seconds sum to 3600, non-negative DUW delta within a controller segment, coverage consistent with sample count). A violating run writes nothing for that run, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact, but under a single trusted writer they should never see one. Store invariant: everything persisted is well-formed.
2. **Missed observations — never backfill.** Fenris never interpolates, estimates, or fabricates an hour. Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s power-on-hours classification is the only inference admitted.
3. **Store faults — degrade, never recreate over.** An unreadable or corrupt database is a store fault: readers surface a "observation store unreadable" fact with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside, the next run starts a fresh store, and if the legacy import never completed, the still-present `history.jsonl` is re-imported. No built-in destructive command exists.
4. **Newer schema — readers refuse symmetrically.** The TUI and `status` detect a `user_version` newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal in [ADR 0001](0001-observation-store-sqlite.md) and the forward-only upgrade rule of [ADR 0004](0004-install-upgrade-removal-lifecycle.md).
5. **Repeated collector failures — flat cadence, no escalation.** The timer's retry is the recovery path; the settled freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff and no notification machinery; a persistent failure reads as stale exactly like any other gap.
6. **Drive-reported anomalies — facts, not alerts.** `critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status`; no alerting or notification surface exists. Fenris observes and projects; it does not alarm. The projection is unaffected: endurance math consumes writes, not warnings.
7. **Orphaned samples — the collector re-anchors observed fact.** When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent: only the sanctioned path of [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) records a `user_disabled` close. Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
## Consequences
- Validation lives at one boundary — the collector — so the store's contract is "everything in it is well-formed" and readers only defend against the impossible.
- No synthetic data can ever enter the habit record; confidence categories can be trusted to reflect real evidence.
- Store-fault recovery can lose history; the mitigation is the one-file backup story of [ADR 0001](0001-observation-store-sqlite.md), kept human-sanctioned so loss is never silent.
- Period bookkeeping splits by epistemics: the helper records intent, the collector records observed fact.
- Fenris stays fully local and silent: no notification, escalation, or alerting machinery anywhere.