Files
Fenris/docs/adr/0001-observation-store-sqlite.md

5.2 KiB

1. Observation store: a single SQLite database

Status

Accepted — resolves Define the persistent observation store and legacy migration on the Wayfinder map.

Context

Fenris today persists full SMART samples to an append-only data/history.jsonl beside a derived data/hourly.jsonl, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI (lifecycle research), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence (endurance research). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.

Decision

  1. Substrate: one SQLite database in WAL mode at /var/lib/fenris/observations.db. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
  2. Access: the database is root-owned and group-readable through the fenris read group created by packaging; the TUI opens it read-only. No /run snapshot or export layer.
  3. Entities:
    • samples — recent raw SMART samples: timestamp, controller identity, raw data_units_written/data_units_read integers, percentage_used, available_spare, media_errors, power_on_hours, power_cycles, unsafe_shutdowns, temperature, critical_warning.
    • hour_observations — one row per UTC hour: the usage-habit split (seconds_active, seconds_idle, seconds_powered_off, seconds_unknown), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
    • day_aggregates — one row per UTC day; the habit-evidence grain.
    • monitoring_periods — started_at, ended_at (NULL = open), end_cause enum (user_disabled, migrated, …). Powered-off time stays inside a period; deliberately disabled time does not.
    • controller_segments — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
    • endurance_baseline — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
    • Projections are not stored; they are recomputed on read. There is no separate latest-status table.
  4. Day boundary: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
  5. Retention: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
  6. Migration (first new-version collection run):
    1. If the database already carries the legacy-import marker, do nothing.
    2. history.jsonl is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore hourly.jsonl as derived data (diff and log mismatches; do not trust).
    3. One implicit monitoring_periods row opens at the first legacy sample and closes with end_cause = migrated at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
    4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
    5. Only after commit are legacy files renamed to *.migrated (never deleted).
    6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
  7. Projection inputs: the endurance_baseline table lives in the database and is edited via the CLI; /etc/fenris/ holds only operational configuration.
  8. Versioning: PRAGMA user_version plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
  9. Collector health: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.

Consequences

  • Backups and state migration are copying one file (plus its WAL sidecars).
  • SQLite becomes a runtime dependency of both the collector and the TUI (Python sqlite3 stdlib suffices; no server).
  • The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
  • Legacy checkout-relative data/ files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
  • The active/idle/powered-off classification contract with the projection model is the hour_observations column set, keeping storage and model decisions separable.