6.4 KiB
1. Observation store: a single SQLite database
Status
Accepted — resolves Define the persistent observation store and legacy migration on the Wayfinder map.
Amended by Define endurance-baseline provenance and validation: the endurance_baseline field set and validation contract — derived verification, entry-time unprivileged sysfs validation, and read-time controller-segment applicability.
Context
Fenris today persists full SMART samples to an append-only data/history.jsonl beside a derived data/hourly.jsonl, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI (lifecycle research), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence (endurance research). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
Decision
- Substrate: one SQLite database in WAL mode at
/var/lib/fenris/observations.db. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions. - Access: the database is root-owned and group-readable through the
fenrisread group created by packaging; the TUI opens it read-only. No/runsnapshot or export layer. - Entities:
samples— recent raw SMART samples: timestamp, controller identity, rawdata_units_written/data_units_readintegers,percentage_used,available_spare,media_errors,power_on_hours,power_cycles,unsafe_shutdowns, temperature,critical_warning.hour_observations— one row per UTC hour: the usage-habit split (seconds_active,seconds_idle,seconds_powered_off,seconds_unknown), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.day_aggregates— one row per UTC day; the habit-evidence grain.monitoring_periods—started_at,ended_at(NULL = open),end_causeenum (user_disabled,migrated, …). Powered-off time stays inside a period; deliberately disabled time does not.controller_segments— boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.endurance_baseline— one active row, replaced on edit (Define endurance-baseline provenance and validation): the rated-TBW value in bytes (E_rated = entered_TBW × 10¹²) plus mandatory provenance — source URL, document revision, entry date, model string, nominal capacity — and frozen validation facts (detected model, detected capacity bytes,validated_bymachine/user,validated_at). Verification is derived at read — complete provenance and a drive match (machine or attested), never a stored boolean; incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier. Entry validation is an unprivileged live sysfs read of the configured device (normalized model containment with an interactive confirm recorded asvalidated_by = user; capacity within ±1%); at projection time applicability is a model match against the current controller segment, and a mismatch is retained — never auto-deleted — leaving the projection Unavailable.- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
- Day boundary: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
- Retention: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
- Migration (first new-version collection run):
- If the database already carries the legacy-import marker, do nothing.
history.jsonlis the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignorehourly.jsonlas derived data (diff and log mismatches; do not trust).- One implicit
monitoring_periodsrow opens at the first legacy sample and closes withend_cause = migratedat the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred). - The import is a single transaction: interruption leaves the database fully pre- or post-migration.
- Only after commit are legacy files renamed to
*.migrated(never deleted). - Malformed legacy lines are quarantined with a logged count, never silently dropped.
- Projection inputs: the
endurance_baselinetable lives in the database and is edited via the CLI;/etc/fenris/holds only operational configuration. - Versioning:
PRAGMA user_versionplus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version. - Collector health: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python
sqlite3stdlib suffices; no server). - The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative
data/files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail. - The active/idle/powered-off classification contract with the projection model is the
hour_observationscolumn set, keeping storage and model decisions separable.