43 lines
5.2 KiB
Markdown
43 lines
5.2 KiB
Markdown
# 1. Observation store: a single SQLite database
|
|
|
|
## Status
|
|
|
|
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
|
|
|
|
## Context
|
|
|
|
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
|
|
|
|
## Decision
|
|
|
|
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
|
|
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
|
|
3. **Entities**:
|
|
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
|
|
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
|
|
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
|
|
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
|
|
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
|
|
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
|
|
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
|
|
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
|
|
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
|
|
6. **Migration** (first new-version collection run):
|
|
1. If the database already carries the legacy-import marker, do nothing.
|
|
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
|
|
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
|
|
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
|
|
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
|
|
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
|
|
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
|
|
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
|
|
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
|
|
|
|
## Consequences
|
|
|
|
- Backups and state migration are copying one file (plus its WAL sidecars).
|
|
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
|
|
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
|
|
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
|
|
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
|