Define failure and recovery behavior #10

Closed
opened 2026-08-31 09:19:06 +00:00 by xavierk · 1 comment
Owner

Parent map: Chart Fenris’s persistent TUI monitoring redesign

Question

How does Fenris detect, degrade, and recover when things go wrong? Define behavior per failure class: corrupt or malformed observations in the observation store, missed observations (gaps in the habit record), observation-store failures (unreadable or corrupt SQLite database, unknown newer schema), and collector/service failures. The observation store is settled (ADR 0001); this ticket follows the lifecycle decision so service-failure handling can build on the chosen timer/service arrangement.

Parent map: [Chart Fenris’s persistent TUI monitoring redesign](https://git.bongbetic.com/xavierk/Fenris/issues/1) ## Question How does Fenris detect, degrade, and recover when things go wrong? Define behavior per failure class: corrupt or malformed observations in the observation store, missed observations (gaps in the habit record), observation-store failures (unreadable or corrupt SQLite database, unknown newer schema), and collector/service failures. The observation store is settled ([ADR 0001](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/adr/0001-observation-store-sqlite.md)); this ticket follows the lifecycle decision so service-failure handling can build on the chosen timer/service arrangement.
xavierk added this to the Wayfinder: Fenris persistent TUI monitoring redesign milestone 2026-08-31 09:19:06 +00:00
xavierk added the wayfinder:grilling label 2026-08-31 09:19:06 +00:00
xavierk added a new dependency 2026-08-31 09:19:28 +00:00
xavierk added a new dependency 2026-08-31 11:45:44 +00:00
xavierk self-assigned this 2026-08-31 13:33:06 +00:00
Author
Owner

Resolution

Failure behavior is defined per class in ADR 0005 — "visible degradation, never fabrication":

  1. Malformed observations: write-time validation in the collector; an invariant-violating run writes nothing, journals the refused row, and fails visibly. Readers exclude-and-count defensively only.
  2. Missed observations: never backfill or interpolate; gaps stay unknown seconds and degrade through coverage, freshness facts, and confidence categories; recovery is the next successful run.
  3. Store faults (unreadable/corrupt DB): readers surface a store-fault fact and nothing dependent; the collector bounded-fails and never recreates over an existing file; recovery is human-sanctioned (move aside → fresh store → legacy re-import if never completed). Glossary term Store fault added to CONTEXT.md.
  4. Newer schema: TUI and status refuse symmetrically with the collector ("observation store written by a newer Fenris").
  5. Repeated collector failures: flat cadence, no backoff or escalation; settled freshness grading makes degradation visible.
  6. Drive-reported anomalies: facts, not alerts; no notification machinery; projection unaffected.
  7. Orphaned samples: a run finding no open monitoring period opens one at the run moment (observed fact, never backdated); only the sanctioned path records intent.

No new tickets surfaced; the map's spec-assembly fog stays until the sibling tickets (Verify the controller identity that segments observation history, Define endurance-baseline provenance and validation) close.

## Resolution Failure behavior is defined per class in [ADR 0005](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/adr/0005-failure-detection-and-recovery.md) — "visible degradation, never fabrication": 1. **Malformed observations**: write-time validation in the collector; an invariant-violating run writes nothing, journals the refused row, and fails visibly. Readers exclude-and-count defensively only. 2. **Missed observations**: never backfill or interpolate; gaps stay unknown seconds and degrade through coverage, freshness facts, and confidence categories; recovery is the next successful run. 3. **Store faults** (unreadable/corrupt DB): readers surface a store-fault fact and nothing dependent; the collector bounded-fails and never recreates over an existing file; recovery is human-sanctioned (move aside → fresh store → legacy re-import if never completed). Glossary term **Store fault** added to CONTEXT.md. 4. **Newer schema**: TUI and status refuse symmetrically with the collector ("observation store written by a newer Fenris"). 5. **Repeated collector failures**: flat cadence, no backoff or escalation; settled freshness grading makes degradation visible. 6. **Drive-reported anomalies**: facts, not alerts; no notification machinery; projection unaffected. 7. **Orphaned samples**: a run finding no open monitoring period opens one at the run moment (observed fact, never backdated); only the sanctioned path records intent. No new tickets surfaced; the map's spec-assembly fog stays until the sibling tickets ([Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11), [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12)) close.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/Fenris#10