Migration runbook at docs/install/migrate-from-makeinstall.md covers the mandatory remove-then-install path, why over-install is forbidden, no-move continuity guarantees, and reset-to-dormant expectations. Acceptance criteria MG-1 through MG-4 added to the install criteria section. Containerized tests verify no-move continuity: existing group makes sysusers a no-op, existing store dir makes tmpfiles a no-op, hand-edited config survives as a non-database file, and store schema is caught up by the upgrade-path migration. Dead code from a prior merge removed. Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
23 KiB
23 KiB
Acceptance criteria: the Fenris redesign
Status: Accepted — resolves Define cross-cutting acceptance criteria on the Wayfinder map. These criteria are the accepted definition of done for the finished redesign; the implementation-ready specification assembles them with ADRs 0001–0006 at handoff.
Framework
- Canonical term: acceptance criterion — one testable behavioral statement. "Behavioral gate" is avoided as a synonym. Wording follows the repository glossary (
CONTEXT.md). - Evidence classes — every criterion carries exactly one:
- A — automated test (unit/integration, fixture-driven).
- P — scripted system probe on a host with systemd, polkit, and the configured NVMe device.
- M — manual checklist, reserved for interactions a fixture cannot capture (live polkit agent prompts, TUI keyboard feel).
- Whatever can be automated must be; M only where automation cannot reach.
- Traceability-only: every number and behavior cites the ADR or ticket that fixed it. Nothing undecided enters here; new demands become new tickets, never criteria.
- Organization: criteria are grouped by subsystem, with cross-cutting invariants spanning them. Coverage spans all decided areas — observation store and migration, collector lifecycle and privileges, projection and confidence, controller identity, the Panes TUI, failure paths, and installation lifecycle.
- Placeholders: none remain. SLOT-B was filled by Choose the collector's NVMe acquisition path as AC-1–AC-5; SLOT-A was filled by Decide how degraded identity affects projection confidence as PR-15, PR-16, and ID-4.
- Test-plan boundary: Given/When/Then test specs are derived by the implementer at implementation time. This effort produces criteria only.
Cross-cutting invariants
- CI-1 (A; ADR 0002 §§6–8, ADR 0003 §10) Exhaustive state matrix: from a synthetic observation store, the TUI and
fenris statusrender every realizable combination of confidence state (Unavailable, Limited, Supported) × freshness grade (fresh, missed, stale, empty store) × endurance-baseline tier (verified override, unverified override, implied, none) exactly as the ADR 0002 rule table and ADR 0003 freshness constants dictate — headline number present only when the rules allow it, contributing facts always, never a percentage. - CI-2 (P/A; ADR 0003 §§4–9) TUI/CLI parity: every TUI action has a CLI twin — pause (
fenris monitor pause), resume (fenris monitor resume), collect-now (fenris sample), baseline set/clear, and the status fact set — with identical outcomes and wording. - CI-3 Prohibition set (A = test, P = probe; each cites its clause):
- No code path outside
fenris-collectinterrogates the device (ADR 0003 §7). - No
/run/fenriscoordination surface or export layer exists anywhere; state lives in the observation store and coordination in systemd (ADR 0003 §1, ADR 0001 §2). - Polkit authorizes exactly one binary,
fenris-monitor, undercom.bongbetic.fenris.monitorauth_admin(ADR 0003 §5, ADR 0004 §3). - No absent hour is ever interpolated, estimated, or fabricated (ADR 0005 §2).
- No alerting, notification, or escalation machinery exists anywhere (ADR 0005 §§5–6).
/etc/fenris/fenris.confholds exactly one key — the device selector (ADR 0003 §3).- No synthetic or capacity-derived baseline is ever created, including for legacy history (ADR 0002, Consequences).
- Readers never partially interpret a newer-schema store (ADR 0001 §8, ADR 0005 §4).
- Projections are never stored; always recomputed on read (ADR 0001 §3, ADR 0002 §13).
- No code path outside
- CI-4 (A; ADR 0002 §§11–12) User-facing language: TUI and status render the endurance research's required wording and six disclosures as adopted; zero-rate and unavailable cases use their exact phrasing; the scenario range is the only spread shown anywhere.
Observation store and legacy migration (ADR 0001)
- ST-1 (P) One SQLite database in WAL mode at
/var/lib/fenris/observations.db, root-owned and group-readable through thefenrisread group; the TUI opens it read-only. - ST-2 (A) An unprivileged reader querying during a collector write sees a consistent snapshot.
- ST-3 (A) The schema carries
samples,hour_observations,day_aggregates,monitoring_periods,controller_segments, andendurance_baselinewith the ADR 0001 column sets as amended by Define endurance-baseline provenance and validation and Decide controller-segment metadata columns. - ST-4 (A) Hours and days are UTC-bounded; day derivation from hours is monotonic; no 23- or 25-hour days exist.
- ST-5 (A) Raw samples are pruned opportunistically to 14 days; hour observations and day aggregates are retained indefinitely.
- ST-6 (P) Legacy import is one transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
- ST-7 (A) Import is idempotent: a second run no-ops on the legacy-import marker.
- ST-8 (A) Legacy files are renamed
*.migratedonly after commit and never deleted. - ST-9 (A) Malformed legacy lines are quarantined with a logged count, never silently dropped.
- ST-10 (A)
hourly.jsonlis never trusted: mismatches against derived data are diffed and logged. - ST-11 (A) Migration opens one implicit monitoring period at the first legacy sample, closed
end_cause = migratedat the migration moment; pre-migration hours carry an unknown activity split except directly evidenced facts. - ST-12 (A) Schema versioning:
PRAGMA user_versionwith ordered, per-step transactional migrations; the collector refuses an unknown newer version.
Collector lifecycle and privilege boundaries (ADR 0003)
- LC-1 (P) Exactly two system units exist —
fenris-collect.timer(timers.target) andfenris-collect.service(Type=oneshot, root,ExecStart=/usr/libexec/fenris/fenris-collect); the TUI and CLI are ordinary unprivileged processes and never units. - LC-2 (P) Timer defaults ship as
OnBootSec=2min,OnUnitInactiveSec=5min,AccuracySec=30s,Persistent=no,TimeoutStartSec=90s; cadence changes are documented drop-ins and no interval key exists in configuration. - LC-3 (P) A hung device interrogation fails visibly within
TimeoutStartSec=90sas a bounded failed run retried next interval. - LC-4 (A/P)
/etc/fenris/fenris.confholds exactly the device selector (stable/dev/disk/by-id/…path; raw nodes warned), re-read every run; an invalid selector is a bounded failed run surfaced asconfiguration error: <reason>instatusand the TUI. - LC-5 (P) Two privileged binaries ship at
/usr/libexec/fenris/fenris-collectand/usr/libexec/fenris/fenris-monitor; the unprivilegedfenriswrapper opens the TUI with no arguments. - LC-6 (P+M) Pause =
fenris-monitor disable --nowasks for confirmation; Resume =enable --nowdoes not; both perform the systemctl operation and period bookkeeping in one step under polkitcom.bongbetic.fenris.monitor(auth_admin), failing cleanly with the printed root equivalent where no polkit agent exists. (M covers the live agent prompt.) - LC-7 (A) The period-row idempotent matrix of ADR 0003 §6 holds exactly: first-ever enable opens; resume with an open period changes nothing; resume without one opens anew; pause with an open period closes
user_disabled; pause otherwise no-ops; a raw systemctl stop/disable never recordsuser_disabled. - LC-8 (P)
fenris sampleand the TUI's collect-now route throughfenris-monitor→systemctl start fenris-collect.service, block until exit, and report the outcome (freshness line or journal hint) synchronously; the TUI never samples in-process. - LC-9 (A/P) CLI compatibility:
statusis a read-only composition (projection facts, enabled/active, last collect outcome,journalctlhint on failure or staleness) that never auto-samples and never prompts;sampleis retained via the helper;--deviceis rejected with a pointer to the configuration file;start,stop, andrunare rejected with one-line migration pointers;fenris.shis not shipped and is removed from the repository; the README maps its five menu options to successors. - LC-10 (A) Freshness constants are defined once and shared by TUI and CLI: fresh = newest sample within 2× cadence +
AccuracySec+ 60 s; missed between that and 48 h; stale ≥ 48 h; an empty store reads "no observations yet" with an enable hint; freshness derives from the newest sample timestamp, never a stored health flag.
Projection contract and confidence (ADR 0002)
- PR-1 (A) Exactly one projection, from the precedence-chosen baseline (verified override → unverified override → implied → unavailable); Percentage Used renders as a vendor-wear context line, with a note when it disagrees with the observed write rate by more than a factor of 2; the PU-slope regression and
capacity × 600synthesis are gone. - PR-2 (A) The headline rate is the sustained-regime rate (regime DUW bytes ÷ in-period wall-clock seconds), default regime = full history capped at 90 days; the 7/28/90-day scenario range is computed independently and shows only covered horizons, with no placeholders.
- PR-3 (A) Habit change: trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean for 3 consecutive days starts a new regime at the first divergence day, adopted automatically and labeled "usage habit changed N days ago"; a regime younger than 7 days caps confidence at Limited evidence.
- PR-4 (A) Hour classification uses the named constants: powered-off below 90% of power-on-hours span; active at ≥ 256 MiB DUW; idle below it while powered on and sampled; unknown otherwise; disabled time is wall-clock outside monitoring periods, never an hour state.
- PR-5 (A) The denominator is wall-clock seconds inside monitoring periods including powered-off and unknown time; disabled periods are excluded from numerator and denominator; unexplained gaps keep the aggregate counter delta, remain as unknown seconds, and reduce coverage.
- PR-6 (A) Warming up until 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage; the projection still renders with its facts while warming; every Unavailable condition renders no lifespan number.
- PR-7 (A) A newest day aggregate older than 48 h drops confidence one level and is shown as a contributing fact.
- PR-8 (A) The confidence rule table of ADR 0002 §8 holds verbatim, rendering state plus contributing facts and never a percentage.
- PR-9 (A) Segment breaks: a DUW decrease with unchanged identity keeps prior day aggregates as habit evidence with the projection Unavailable until re-warm; a controller-identity change quarantines prior history from projection entirely.
- PR-10 (A) The implied baseline is eligible only after ≥ 2 Percentage-Used increments within the current controller segment; until then, Unavailable with "vendor wear estimate too coarse to imply endurance".
- PR-11 (A) Zero rate renders "no finite projection from this history" — never infinity or zero; no statistical confidence interval appears anywhere.
- PR-12 (A) The projection contract hands the TUI exactly: confidence state, contributing facts, headline remaining time when one exists, scenario range, Percentage-Used context line, disclosure text — recomputed on read, never stored.
- PR-13 (A) Baseline provenance and validation per Define endurance-baseline provenance and validation: mandatory provenance (URL, revision, entry date, model, nominal capacity); one active row replaced on edit; verification derived at read (machine match or recorded attestation), never a stored boolean; incomplete provenance stores only behind explicit acknowledgment as the unverified tier; entry-time unprivileged sysfs validation (normalized model containment; capacity within ±1%; interactive confirm recorded as
validated_by = user); read-time applicability is a model match against the current controller segment, with a mismatch retained — never auto-deleted — leaving the projection Unavailable. - PR-14 (P)
baseline set/baseline clearpersist through the polkit-guardedfenris-monitorverb after CLI-side validation. - PR-15 (A) An identity-degraded controller segment (blank identity key — every rung of the key ladder empty) caps projection confidence at Limited evidence, with the contributing fact "controller identity unavailable — replacement detection relies on write-counter continuity only" rendered in every state; the cap combines idempotently with the 48-hour staleness drop, and ephemeral markers (model "Linux", non-pcie transport) never render as confidence facts (Decide how degraded identity affects projection confidence; ADR 0002 §§8–9 as amended).
- PR-16 (A) Identity-change semantics extend to blank keys verbatim: any visible change of the recorded identity key — including to or from a blank key — quarantines prior history from projection as a controller-identity change, while equal blank keys continue the segment segmented by DUW monotonicity alone (ADR 0002 §9 as amended).
- PR-17 (A) Projection arithmetic is exactly
E_rated = entered_TBW × 10¹²bytes,E_implied = 100 · W_t / pcomputed only for1 ≤ p ≤ 254(Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable), andprojected = max(E_baseline − W_t, 0) / rateforrate > 0(ADR 0002 §2).
Controller identity (Verify the controller identity that segments observation history, Decide controller-segment metadata columns; ADR 0001 §3 as amended)
- ID-1 (A) The controller-segment identity key is the normalized kernel-exposed subsystem NQN, with the kernel composite then model|serial as fallbacks; FR is metadata only; identity change and DUW decrease act as independent axes.
- ID-2 (A) Segments freeze a fully nullable metadata snapshot at open — normalized
subnqn/sn/mn/frplusvid/ssvid/transportandidentity_degraded— immutable thereafter, withcntlidexcluded. - ID-3 (A) Legacy history imports under a labeled model-scoped legacy identity (mn-only segments).
- ID-4 (A)
identity_degradedis set at segment open exactly when the identity key is blank; keys from the kernel-composite ormodel|serialrungs are not degraded (Decide how degraded identity affects projection confidence).
Panes TUI (Prototype the TUI information architecture, Evaluate Python TUI frameworks; ADR 0003 §§8, 10; ADR 0004 §10)
- TUI-1 (A) Variant A "Panes": one dense keyboard-first screen; confidence rendered as evidence (state + contributing facts); boot enablement, runtime activity, last collect outcome, and freshness displayed as four separate facts.
- TUI-2 (M) Pause/resume asymmetry and polkit tty passthrough work in a live terminal: pause confirms, resume does not, and the platform agent prompts without breaking the TUI.
- TUI-3 (P) Textual runs on Python 3.9+, gated at install time, never a runtime crash.
- TUI-4 (A) The Panes screen layout is normative: a full-width headline band (lifespan headline or its no-projection wording, confidence state with contributing facts, scenario range); a usage-history pane on the left (write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus legend, habit-split bar with active/idle/powered-off/unknown shares); a drive-health and settings pane on the right (health facts, vendor-wear context line, read-only settings with the endurance baseline and its provenance label); a full-width service strip at the bottom (the four separate service facts, the monitoring-period line, the action legend). Production bindings are
ppause (asks),rresume (does not),ccollect now,ddisclosures,qquit (Prototype the TUI information architecture); the prototype branch is visual reference only.
Failure and recovery (ADR 0005)
- FL-1 (A) The collector validates every row against the store invariants (hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count); a violating run writes nothing, logs the refused row, and fails visibly.
- FL-2 (A) Readers defensively exclude and count malformed rows as a contributing fact.
- FL-3 (A) No backfill ever: gaps remain unknown seconds; degradation flows only through coverage, freshness facts, and confidence categories.
- FL-4 (P/A) A store fault surfaces "observation store unreadable" with a journal hint and suppresses everything else store-dependent; the collector treats it as a bounded failed run and never recreates or overwrites the file; recovery is the documented human-sanctioned move-aside (with
history.jsonlre-import if the legacy import never completed); no built-in destructive command exists. - FL-5 (A) A newer-schema store renders "observation store written by a newer Fenris — upgrade Fenris" in TUI and status, with no partial interpretation.
- FL-6 (A/P) Repeated collector failures retry at flat cadence with no backoff or notification; persistence reads as stale exactly like any other gap.
- FL-7 (A)
critical_warning, media errors, and unsafe shutdowns render as ordinary facts in TUI and status and never affect the projection. - FL-8 (A) A collection run finding no open monitoring period opens one at the run moment, never backdated.
Installation, upgrade, and removal (ADR 0004)
- IN-1 (P)
sudo make installbuilds a wheel from the checkout and installs pinned dependencies into the dedicated venv at/opt/fenris, with a/usr/local/bin/fenriswrapper; after install nothing references the checkout. - IN-2 (P) The installer records every placed file in an explicit manifest consumed by upgrade and uninstall.
- IN-3 (P) The installer never enables or starts units: a fresh install is dormant (units disabled, nothing running, no monitoring period); the only opt-in is the sanctioned toggle —
fenris monitor resume [--now]or the first-run TUI prompt — enabling the timer and opening the first period in one step. - IN-4 (P) Install-time legacy import detects
./data/history.jsonl(or an explicit path), runs the idempotent single-transaction import, and reports imported counts;fenris import <path>remains available. - IN-5 (P)
sudo make upgradeinstalls into the same venv, syncs units and polkit against the manifest (daemon-reload; timer restarted only if unit contents changed and it is active), never kills an in-flight collection run, then applies forward-only schema migrations;/var/lib/fenrisis never rebuilt. - IN-6 (P) Before migrations,
observations.dbis snapshotted to a one-generation.bak; rollback is reinstall-previous plus restore; automatic schema downgrade does not exist. - IN-7 (P)
make uninstallperforms the sanctioned disable first (open period closesuser_disabled), then removes venv, helpers, units, polkit policy, and wrapper while keeping/etc/fenrisand the observation store;make purgeadditionally removes configuration and store. - IN-8 (P) Dependencies are exact pins in a committed lockfile installed by both install and upgrade; refreshing pins is an explicit
make update-depsstep, never an install side effect. - IN-9 (P) The installer verifies
python3 ≥ 3.9and fails cleanly otherwise;/var/lib/fenrisis created with root-written group-read permissions; the database file is created lazily by the first write. - IN-10 (P) Installed artifacts sit only at their fixed locations — units in
/etc/systemd/system, helpers in/usr/libexec/fenris, polkit policy under/usr/share/polkit-1/actions/, configuration at/etc/fenris, observation store under/var/lib/fenris— and every placed file is recorded in the manifest (ADR 0004 §2; ADR 0003 §4).
Migration from make-install systems (ADR 0007 §10, spec §9)
- MG-1 (M) The migration runbook is published in the install docs (
docs/install/migrate-from-makeinstall.md): mandatory remove-then-install steps, why over-install is forbidden (stale admin-directory units silently shadow vendor units; the local wrapper shadows the package wrapper), no-move continuity, and the reset-to-dormant expectation (the user opts back in with the sanctioned resume). - MG-2 (A) The install guard is verified across the matrix: either make-install marker (the legacy placement manifest, or a unit file under the admin unit directory) causes an abort with a runbook pointer — never auto-clean. Tested by
test_migration_guardon all four targets (Debian 12, Ubuntu 22.04, Ubuntu 24.04, Fedora 40). - MG-3 (A) No-move continuity is verified in a container seeded with a make-install-shaped system: existing group makes sysusers a no-op, existing store directory makes tmpfiles a no-op, the hand-written configuration survives as a non-database file (package default lands beside it), and the store schema is caught up by the upgrade-path migration. Tested by
test_no_move_continuity_debandtest_no_move_continuity_rpm. - MG-4 (M) The migration costs at most one short sample gap, honestly recorded in the endurance timeline:
make uninstall's sanctioned disable closes the open perioduser_disabled; after migration the user opts back in withfenris monitor resume.
Collector acquisition path (ADR 0006)
- AC-1 (P) Each collection run acquires counters and thermal evidence solely from
smartctl -a -j <device>and controller identity (subnqn,sn,mn,fr,transport) solely from sysfs; no other acquisition path exists anywhere in the codebase. - AC-2 (A) Identity normalization is applied exactly once, at write time — trailing spaces and newlines stripped, no case folding, empty-after-strip stored blank — so padded and unpadded renderings of the same field yield byte-identical stored values.
- AC-3 (A) Any acquisition failure — missing binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the whole collection run; no partial sample (identity without counters, or counters without identity) is ever written; the miss surfaces through ADR 0005 freshness, never as degraded identity.
- AC-4 (P)
vid/ssvidare read from the PCI sysfs node when present and stored null otherwise; they are segment metadata only, never key components. - AC-5 (P)
make installverifiessmartctland fails cleanly otherwise; the acquisition path adds no Python dependency and no OS package beyond smartmontools (ADR 0004 §9).