docs(spec): acceptance criteria for the redesign — subsystem gates, A/P/M evidence classes, ADR traceability; state-matrix and parity gates, prohibition set

This commit is contained in:
xavierk
2026-08-31 22:13:20 +05:30
parent 305bdd4779
commit bf83ac5481
+116
View File
@@ -0,0 +1,116 @@
# Acceptance criteria: the Fenris redesign
Status: Accepted — resolves [Define cross-cutting acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/issues/13) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). These criteria are the accepted definition of done for the finished redesign; the implementation-ready specification assembles them with ADRs 0001–0005 at handoff.
## Framework
- **Canonical term**: *acceptance criterion* — one testable behavioral statement. "Behavioral gate" is avoided as a synonym. Wording follows the repository glossary (`CONTEXT.md`).
- **Evidence classes** — every criterion carries exactly one:
- **A** — automated test (unit/integration, fixture-driven).
- **P** — scripted system probe on a host with systemd, polkit, and the configured NVMe device.
- **M** — manual checklist, reserved for interactions a fixture cannot capture (live polkit agent prompts, TUI keyboard feel).
- Whatever can be automated must be; **M** only where automation cannot reach.
- **Traceability-only**: every number and behavior cites the ADR or ticket that fixed it. Nothing undecided enters here; new demands become new tickets, never criteria.
- **Organization**: criteria are grouped by subsystem, with cross-cutting invariants spanning them. Coverage spans all decided areas — observation store and migration, collector lifecycle and privileges, projection and confidence, controller identity, the Panes TUI, failure paths, and installation lifecycle.
- **Placeholders**: SLOT-A and SLOT-B are reserved for the open decisions [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15) and [Choose the collector's NVMe acquisition path](https://git.bongbetic.com/xavierk/Fenris/issues/16); their resolutions fill them.
- **Test-plan boundary**: Given/When/Then test specs are derived by the implementer at implementation time. This effort produces criteria only.
## Cross-cutting invariants
- **CI-1** (A; ADR 0002 §§6–8, ADR 0003 §10) *Exhaustive state matrix*: from a synthetic observation store, the TUI and `fenris status` render every realizable combination of confidence state (Unavailable, Limited, Supported) × freshness grade (fresh, missed, stale, empty store) × endurance-baseline tier (verified override, unverified override, implied, none) exactly as the ADR 0002 rule table and ADR 0003 freshness constants dictate — headline number present only when the rules allow it, contributing facts always, never a percentage.
- **CI-2** (P/A; ADR 0003 §§4–9) *TUI/CLI parity*: every TUI action has a CLI twin — pause (`fenris monitor pause`), resume (`fenris monitor resume`), collect-now (`fenris sample`), baseline set/clear, and the status fact set — with identical outcomes and wording.
- **CI-3** *Prohibition set* (A = test, P = probe; each cites its clause):
- No code path outside `fenris-collect` interrogates the device (ADR 0003 §7).
- Polkit authorizes exactly one binary, `fenris-monitor`, under `com.bongbetic.fenris.monitor` `auth_admin` (ADR 0003 §5, ADR 0004 §3).
- No absent hour is ever interpolated, estimated, or fabricated (ADR 0005 §2).
- No alerting, notification, or escalation machinery exists anywhere (ADR 0005 §§5–6).
- `/etc/fenris/fenris.conf` holds exactly one key — the device selector (ADR 0003 §3).
- No synthetic or capacity-derived baseline is ever created, including for legacy history (ADR 0002, Consequences).
- Readers never partially interpret a newer-schema store (ADR 0001 §8, ADR 0005 §4).
- Projections are never stored; always recomputed on read (ADR 0001 §3, ADR 0002 §13).
- **CI-4** (A; ADR 0002 §§11–12) *User-facing language*: TUI and status render the endurance research's required wording and six disclosures as adopted; zero-rate and unavailable cases use their exact phrasing; the scenario range is the only spread shown anywhere.
## Observation store and legacy migration (ADR 0001)
- **ST-1** (P) One SQLite database in WAL mode at `/var/lib/fenris/observations.db`, root-owned and group-readable through the `fenris` read group; the TUI opens it read-only.
- **ST-2** (A) An unprivileged reader querying during a collector write sees a consistent snapshot.
- **ST-3** (A) The schema carries `samples`, `hour_observations`, `day_aggregates`, `monitoring_periods`, `controller_segments`, and `endurance_baseline` with the ADR 0001 column sets as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) and [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14).
- **ST-4** (A) Hours and days are UTC-bounded; day derivation from hours is monotonic; no 23- or 25-hour days exist.
- **ST-5** (A) Raw samples are pruned opportunistically to 14 days; hour observations and day aggregates are retained indefinitely.
- **ST-6** (P) Legacy import is one transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
- **ST-7** (A) Import is idempotent: a second run no-ops on the legacy-import marker.
- **ST-8** (A) Legacy files are renamed `*.migrated` only after commit and never deleted.
- **ST-9** (A) Malformed legacy lines are quarantined with a logged count, never silently dropped.
- **ST-10** (A) `hourly.jsonl` is never trusted: mismatches against derived data are diffed and logged.
- **ST-11** (A) Migration opens one implicit monitoring period at the first legacy sample, closed `end_cause = migrated` at the migration moment; pre-migration hours carry an unknown activity split except directly evidenced facts.
- **ST-12** (A) Schema versioning: `PRAGMA user_version` with ordered, per-step transactional migrations; the collector refuses an unknown newer version.
## Collector lifecycle and privilege boundaries (ADR 0003)
- **LC-1** (P) Exactly two system units exist — `fenris-collect.timer` (`timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`); the TUI and CLI are ordinary unprivileged processes and never units.
- **LC-2** (P) Timer defaults ship as `OnBootSec=2min`, `OnUnitInactiveSec=5min`, `AccuracySec=30s`, `Persistent=no`, `TimeoutStartSec=90s`; cadence changes are documented drop-ins and no interval key exists in configuration.
- **LC-3** (P) A hung device interrogation fails visibly within `TimeoutStartSec=90s` as a bounded failed run retried next interval.
- **LC-4** (A/P) `/etc/fenris/fenris.conf` holds exactly the device selector (stable `/dev/disk/by-id/…` path; raw nodes warned), re-read every run; an invalid selector is a bounded failed run surfaced as `configuration error: <reason>` in `status` and the TUI.
- **LC-5** (P) Two privileged binaries ship at `/usr/libexec/fenris/fenris-collect` and `/usr/libexec/fenris/fenris-monitor`; the unprivileged `fenris` wrapper opens the TUI with no arguments.
- **LC-6** (P+M) Pause = `fenris-monitor disable --now` asks for confirmation; Resume = `enable --now` does not; both perform the systemctl operation and period bookkeeping in one step under polkit `com.bongbetic.fenris.monitor` (`auth_admin`), failing cleanly with the printed root equivalent where no polkit agent exists. (M covers the live agent prompt.)
- **LC-7** (A) The period-row idempotent matrix of ADR 0003 §6 holds exactly: first-ever enable opens; resume with an open period changes nothing; resume without one opens anew; pause with an open period closes `user_disabled`; pause otherwise no-ops; a raw systemctl stop/disable never records `user_disabled`.
- **LC-8** (P) `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, block until exit, and report the outcome (freshness line or journal hint) synchronously; the TUI never samples in-process.
- **LC-9** (A/P) CLI compatibility: `status` is a read-only composition (projection facts, enabled/active, last collect outcome, `journalctl` hint on failure or staleness) that never auto-samples and never prompts; `sample` is retained via the helper; `--device` is rejected with a pointer to the configuration file; `start`, `stop`, and `run` are rejected with one-line migration pointers; `fenris.sh` is not shipped and is removed from the repository; the README maps its five menu options to successors.
- **LC-10** (A) Freshness constants are defined once and shared by TUI and CLI: fresh = newest sample within 2× cadence + `AccuracySec` + 60 s; missed between that and 48 h; stale ≥ 48 h; an empty store reads "no observations yet" with an enable hint; freshness derives from the newest sample timestamp, never a stored health flag.
## Projection contract and confidence (ADR 0002)
- **PR-1** (A) Exactly one projection, from the precedence-chosen baseline (verified override → unverified override → implied → unavailable); Percentage Used renders as a vendor-wear context line, with a note when it disagrees with the observed write rate by more than a factor of 2; the PU-slope regression and `capacity × 600` synthesis are gone.
- **PR-2** (A) The headline rate is the sustained-regime rate (regime DUW bytes ÷ in-period wall-clock seconds), default regime = full history capped at 90 days; the 7/28/90-day scenario range is computed independently and shows only covered horizons, with no placeholders.
- **PR-3** (A) Habit change: trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean for 3 consecutive days starts a new regime at the first divergence day, adopted automatically and labeled "usage habit changed N days ago"; a regime younger than 7 days caps confidence at Limited evidence.
- **PR-4** (A) Hour classification uses the named constants: powered-off below 90% of power-on-hours span; active at ≥ 256 MiB DUW; idle below it while powered on and sampled; unknown otherwise; disabled time is wall-clock outside monitoring periods, never an hour state.
- **PR-5** (A) The denominator is wall-clock seconds inside monitoring periods including powered-off and unknown time; disabled periods are excluded from numerator and denominator; unexplained gaps keep the aggregate counter delta, remain as unknown seconds, and reduce coverage.
- **PR-6** (A) Warming up until 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage; the projection still renders with its facts while warming; every Unavailable condition renders no lifespan number.
- **PR-7** (A) A newest day aggregate older than 48 h drops confidence one level and is shown as a contributing fact.
- **PR-8** (A) The confidence rule table of ADR 0002 §8 holds verbatim, rendering state plus contributing facts and never a percentage.
- **PR-9** (A) Segment breaks: a DUW decrease with unchanged identity keeps prior day aggregates as habit evidence with the projection Unavailable until re-warm; a controller-identity change quarantines prior history from projection entirely.
- **PR-10** (A) The implied baseline is eligible only after ≥ 2 Percentage-Used increments within the current controller segment; until then, Unavailable with "vendor wear estimate too coarse to imply endurance".
- **PR-11** (A) Zero rate renders "no finite projection from this history" — never infinity or zero; no statistical confidence interval appears anywhere.
- **PR-12** (A) The projection contract hands the TUI exactly: confidence state, contributing facts, headline remaining time when one exists, scenario range, Percentage-Used context line, disclosure text — recomputed on read, never stored.
- **PR-13** (A) Baseline provenance and validation per [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12): mandatory provenance (URL, revision, entry date, model, nominal capacity); one active row replaced on edit; verification derived at read (machine match or recorded attestation), never a stored boolean; incomplete provenance stores only behind explicit acknowledgment as the unverified tier; entry-time unprivileged sysfs validation (normalized model containment; capacity within ±1%; interactive confirm recorded as `validated_by = user`); read-time applicability is a model match against the current controller segment, with a mismatch retained — never auto-deleted — leaving the projection Unavailable.
- **PR-14** (P) `baseline set` / `baseline clear` persist through the polkit-guarded `fenris-monitor` verb after CLI-side validation.
- **SLOT-A** — pending [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15): how a blank or degraded identity key moves confidence.
## Controller identity ([Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11), [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14); ADR 0001 §3 as amended)
- **ID-1** (A) The controller-segment identity key is the normalized kernel-exposed subsystem NQN, with the kernel composite then model|serial as fallbacks; FR is metadata only; identity change and DUW decrease act as independent axes.
- **ID-2** (A) Segments freeze a fully nullable metadata snapshot at open — normalized `subnqn`/`sn`/`mn`/`fr` plus `vid`/`ssvid`/`transport` and `identity_degraded` — immutable thereafter, with `cntlid` excluded.
- **ID-3** (A) Legacy history imports under a labeled model-scoped legacy identity (mn-only segments).
## Panes TUI ([Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3), [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6); ADR 0003 §§8, 10; ADR 0004 §10)
- **TUI-1** (A) Variant A "Panes": one dense keyboard-first screen; confidence rendered as evidence (state + contributing facts); boot enablement, runtime activity, last collect outcome, and freshness displayed as four separate facts.
- **TUI-2** (M) Pause/resume asymmetry and polkit tty passthrough work in a live terminal: pause confirms, resume does not, and the platform agent prompts without breaking the TUI.
- **TUI-3** (P) Textual runs on Python 3.9+, gated at install time, never a runtime crash.
## Failure and recovery (ADR 0005)
- **FL-1** (A) The collector validates every row against the store invariants (hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count); a violating run writes nothing, logs the refused row, and fails visibly.
- **FL-2** (A) Readers defensively exclude and count malformed rows as a contributing fact.
- **FL-3** (A) No backfill ever: gaps remain unknown seconds; degradation flows only through coverage, freshness facts, and confidence categories.
- **FL-4** (P/A) A store fault surfaces "observation store unreadable" with a journal hint and suppresses everything else store-dependent; the collector treats it as a bounded failed run and never recreates or overwrites the file; recovery is the documented human-sanctioned move-aside (with `history.jsonl` re-import if the legacy import never completed); no built-in destructive command exists.
- **FL-5** (A) A newer-schema store renders "observation store written by a newer Fenris — upgrade Fenris" in TUI and status, with no partial interpretation.
- **FL-6** (A/P) Repeated collector failures retry at flat cadence with no backoff or notification; persistence reads as stale exactly like any other gap.
- **FL-7** (A) `critical_warning`, media errors, and unsafe shutdowns render as ordinary facts in TUI and status and never affect the projection.
- **FL-8** (A) A collection run finding no open monitoring period opens one at the run moment, never backdated.
## Installation, upgrade, and removal (ADR 0004)
- **IN-1** (P) `sudo make install` builds a wheel from the checkout and installs pinned dependencies into the dedicated venv at `/opt/fenris`, with a `/usr/local/bin/fenris` wrapper; after install nothing references the checkout.
- **IN-2** (P) The installer records every placed file in an explicit manifest consumed by upgrade and uninstall.
- **IN-3** (P) The installer never enables or starts units: a fresh install is dormant (units disabled, nothing running, no monitoring period); the only opt-in is the sanctioned toggle — `fenris monitor resume [--now]` or the first-run TUI prompt — enabling the timer and opening the first period in one step.
- **IN-4** (P) Install-time legacy import detects `./data/history.jsonl` (or an explicit path), runs the idempotent single-transaction import, and reports imported counts; `fenris import <path>` remains available.
- **IN-5** (P) `sudo make upgrade` installs into the same venv, syncs units and polkit against the manifest (`daemon-reload`; timer restarted only if unit contents changed and it is active), never kills an in-flight collection run, then applies forward-only schema migrations; `/var/lib/fenris` is never rebuilt.
- **IN-6** (P) Before migrations, `observations.db` is snapshotted to a one-generation `.bak`; rollback is reinstall-previous plus restore; automatic schema downgrade does not exist.
- **IN-7** (P) `make uninstall` performs the sanctioned disable first (open period closes `user_disabled`), then removes venv, helpers, units, polkit policy, and wrapper while keeping `/etc/fenris` and the observation store; `make purge` additionally removes configuration and store.
- **IN-8** (P) Dependencies are exact pins in a committed lockfile installed by both install and upgrade; refreshing pins is an explicit `make update-deps` step, never an install side effect.
- **IN-9** (P) The installer verifies `python3 ≥ 3.9` and fails cleanly otherwise; `/var/lib/fenris` is created with root-written group-read permissions; the database file is created lazily by the first write.
## Collector acquisition path
- **SLOT-B** — pending [Choose the collector's NVMe acquisition path](https://git.bongbetic.com/xavierk/Fenris/issues/16): dependency-weight, lockfile, privilege, and normalization-consistency criteria for the chosen path.