Fenris today is a checkout-resident daemon (fenris.py + fenris.sh menu) that polls the drive, appends JSONL files beside the code, and serves an HTML dashboard. As a drive owner I can't trust what it tells me: the projection is a trailing-24-hour rate against endurance guessed from Percentage Used or synthesized from capacity, with no notion of evidence quality, drive replacement, or changed habits. As an administrator I can't manage it: state lives wherever the checkout sits, elevation means passwordless sudo for smartctl, and nothing survives moving or deleting the checkout. And the dashboard needs a browser and a port — there is no keyboard-first, terminal-native way to watch the drive.
Solution
A systemd timer drives a short-lived privileged collector that interrogates one configured NVMe drive every five minutes and persists a compact observation history into a single SQLite observation store. An unprivileged, keyboard-first Panes TUI (plus a headless status CLI twin) recomputes — on every read — one usage-adjusted theoretical lifespan from the best available endurance baseline, with projection confidence rendered as categorical evidence (state plus contributing facts, never a percentage), a 7/28/90-day scenario range, and the six endurance disclosures. Monitoring pause/resume is a sanctioned, polkit-authorized toggle that records intent as monitoring-period rows; legacy history imports in one interruption-safe transaction; install is a dormant, manifest-tracked, venv-delivered system package with exact dependency pins.
User Stories
As a drive owner, I want every collection run persisted to a durable observation store, so that my observation history survives reboots, restarts, and upgrades.
As a drive owner, I want the observation store consistent to read while the collector writes, so that opening the TUI never blocks or corrupts.
As a drive owner, I want hour observations and day aggregates retained indefinitely, so that long-term usage-habit evidence accumulates.
As a drive owner, I want raw samples pruned after 14 days, so that the store stays compact without losing habit evidence.
As a drive owner, I want hours and days bounded by UTC, so that day boundaries never shift with daylight-saving time.
As a drive owner, I want my legacy history.jsonl observations imported in one interruption-safe transaction, so that nothing is lost or duplicated when I move to the redesign.
As a drive owner, I want a second import run to no-op on the legacy-import marker, so that migration is idempotent.
As a drive owner, I want malformed legacy lines quarantined with a logged count, so that nothing is silently dropped.
As a drive owner, I want legacy files renamed .migrated rather than deleted after commit, so that I keep a recovery trail.
As a drive owner, I want the legacy hourly.jsonl treated as derived and never trusted, so that mismatches are diffed and logged rather than imported as truth.
As a drive owner, I want the schema versioned with ordered, transactional, forward-only migrations, so that upgrades never destroy the observation store.
As a drive owner, I want a store written by a newer Fenris refused with "observation store written by a newer Fenris — upgrade Fenris", so that no partial interpretation garbles my history.
As a drive owner, I want collection to run on a five-minute systemd timer with no resident daemon, so that observation continues without a long-lived process to babysit.
As a drive owner, I want each collection run to acquire counters and thermal evidence from smartctl and controller identity from sysfs in one all-or-nothing step, so that partial samples never enter the store.
As a drive owner, I want identity normalized exactly once at write time, so that an implementation change can never split my drive's own history.
As an administrator, I want a hung device interrogation to fail visibly within 90 seconds as a bounded failed run, so that one stuck run can't wedge the system.
As an administrator, I want the device selector as the single configuration key, so that changing the monitored drive is a one-key edit re-read every run.
As an administrator, I want an invalid device selector surfaced as a configuration error fact in status and the TUI, so that misconfiguration is visible rather than mysterious.
As a drive owner, I want failed collection runs retried at the flat cadence with freshness visibly degrading, so that recovery is simply the next successful run.
As a drive owner, I want exactly one usage-adjusted theoretical lifespan computed from the precedence-chosen endurance baseline, so that I'm never shown two competing numbers.
As a drive owner, I want Percentage Used rendered as a vendor-wear context line with a note when it disagrees with my observed write rate, so that vendor wear is visible but never a second projection.
As a drive owner, I want my rated-TBW override recorded with mandatory provenance and validated against the detected drive, so that the projection rests on evidence I can audit.
As a drive owner, I want to knowingly store an unverified override when provenance is incomplete, so that honesty about provenance never blocks a projection I asked for.
As a drive owner, I want an implied-from-vendor-wear baseline used only after at least two Percentage-Used increments in the current controller segment, so that quantization noise never masquerades as endurance.
As a drive owner, I want a baseline that no longer matches my drive retained rather than auto-deleted, so that my records are never silently destroyed.
As a drive owner, I want the headline rate taken from my most recent sustained regime, so that the projection reflects how I actually use the drive now.
As a drive owner, I want a detected habit change adopted automatically and labeled "usage habit changed N days ago", so that a new regime replaces a stale one without my intervention.
As a drive owner, I want a 7/28/90-day scenario range showing only horizons my history covers, so that I see horizon sensitivity instead of false precision.
As a drive owner, I want projection confidence rendered as a state plus contributing facts and never a percentage, so that I can judge the evidence myself.
As a drive owner, I want a warming-up indication while evidence accumulates, so that I understand why the projection is soft and see it improve.
As a drive owner, I want a newest day aggregate older than 48 hours to drop confidence one level as a shown fact, so that an abandoned monitor never looks current.
As a drive owner, I want a zero write rate reported as "no finite projection from this history", so that infinity or zero is never displayed.
As a drive owner, I want the six endurance disclosures always available in TUI and status, so that I'm never misled about what the number means.
As a drive owner, I want observation history segmented by controller identity, so that a drive replacement never contaminates my projection.
As a drive owner, I want a write-counter decrease with unchanged identity to quarantine nothing but re-warm, so that a controller reset doesn't erase my history.
As a drive owner, I want a degraded-identity segment capped at Limited evidence with its explaining fact, so that replacement-detection blindness is disclosed, not hidden.
As a drive owner, I want segment metadata frozen at segment open, so that diagnostics always reflect the identity that opened the segment.
As a drive owner, I want my legacy history imported under a labeled model-scoped legacy identity, so that it never blends with post-redesign identity.
As a drive owner, I want pause to ask for confirmation and explain that paused time is excluded from my usage habit, so that deliberate disables are never accidental.
As a drive owner, I want resume to act immediately without confirmation, so that getting back to monitoring is frictionless.
As a drive owner, I want pause/resume to perform the systemctl operation and the monitoring-period bookkeeping in one polkit-authorized step, so that intent and service state never diverge.
As an administrator, I want a raw systemctl stop or disable to register as an unexplained gap rather than a deliberate disable, so that only Fenris's own control path records intent.
As a drive owner, I want collect-now on demand with a synchronous outcome, so that I can refresh without waiting for the timer and without the TUI touching the device.
As a drive owner, I want status as a read-only composition that never auto-samples and never prompts, so that checking state never mutates anything.
As a drive owner, I want retired commands (start, stop, run, --device) rejected with one-line migration pointers, so that old habits fail loudly instead of silently changing meaning.
As a drive owner, I want boot enablement, runtime activity, last collect outcome, and freshness displayed as four separate facts, so that service state is unambiguous at a glance.
As a drive owner, I want one dense keyboard-first Panes screen with no page navigation, so that everything is visible at once.
As a drive owner, I want a write-history sparkline with habit-change and unexplained-gap markers, so that changes and holes stand out.
As a drive owner, I want a habit-split bar with active/idle/powered-off/unknown shares, so that my usage habit is glanceable.
As a drive owner, I want drive health and read-only settings in a side pane, so that facts and configuration are clearly separated.
As a drive owner, I want polkit prompts to appear on my real terminal when acting from the TUI, so that elevation works without breaking the interface.
As a drive owner, I want an empty store greeted with "no observations yet" and an enable hint, so that a first run is self-explanatory.
As a drive owner, I want an invariant-violating collection run to write nothing and fail visibly, so that the observation store only ever holds well-formed data.
As a drive owner, I want readers to defensively exclude and count malformed rows as a fact, so that even the impossible is visible.
As a drive owner, I want gaps never backfilled — no interpolation, estimation, or fabrication — so that no synthetic hour ever enters my history.
As a drive owner, I want a store fault to surface "observation store unreadable" with a journal hint and suppress everything store-dependent, so that corruption is visible without silent guessing.
As an administrator, I want store-fault recovery to be a documented, human-sanctioned move-aside with re-import if the legacy import never completed, so that nothing destructive is automated.
As a drive owner, I want critical warnings, media errors, and unsafe shutdowns rendered as ordinary facts that never affect the projection, so that I'm informed without alarm machinery.
As a drive owner, I want a collection run that finds no open monitoring period to open one at the run moment, so that observed fact is recorded without pretending intent.
As an administrator, I want make install to deliver a fully dormant system, so that nothing starts monitoring without my explicit opt-in.
As an administrator, I want install-time gates for Python ≥ 3.9 and smartctl, so that missing prerequisites fail cleanly instead of crashing at runtime.
As an administrator, I want every placed file recorded in an explicit manifest, so that upgrade and uninstall are exact.
As an administrator, I want upgrades to never kill an in-flight collection run and never rebuild the observation store, so that history is safe across versions.
As an administrator, I want a one-generation database backup before migrations, so that rollback is possible.
As an administrator, I want uninstall to perform the sanctioned disable first and keep my configuration and observation history, so that removal is deliberate but reversible — and purge available when I truly want everything gone.
As an administrator, I want exact dependency pins installed from a committed lockfile, so that installs and upgrades are reproducible.
As a drive owner, I want the README to map the retired menu's five options to their successors, so that migrating my habits is easy.
Implementation Decisions
All decisions below are settled; the redesign specification restates each normatively with a criteria block, and ADRs 0001–0006 hold the rationale. Nothing here should be re-litigated during implementation.
Architecture. The checkout-resident daemon, HTTP dashboard, PID file, and menu script are replaced by: a timer-driven oneshot collector (the only code path that interrogates the device), a privileged fixed-operation helper mediating enable/disable/collect-trigger/baseline-persistence under one polkit action (auth_admin), and an unprivileged wrapper exposing the Panes TUI (no arguments) and the CLI subcommands. No /run coordination surface; state lives in the observation store, coordination in systemd, failures in the journal.
Observation store. One SQLite database in WAL mode at /var/lib/fenris/observations.db, root-owned and group-readable through a fenris read group; readers open read-only. Six entities: samples (14-day raw retention), hour_observations (UTC-hour usage-habit split, write/read deltas, thermal min/avg/max, sample count, coverage flag), day_aggregates (the habit-evidence grain), monitoring_periods (started_at, nullable ended_at, end_cause enum user_disabled/migrated/…), controller_segments (identity key plus frozen nullable metadata snapshot: normalized subnqn/sn/mn/fr, vid/ssvid/transport, identity_degraded; cntlid excluded), and endurance_baseline (one active row, replaced on edit). Projections are never stored; no latest-status table; no stored health flag. Schema versioning via PRAGMA user_version with ordered per-step transactional migrations; unknown newer version is refused by collector and readers alike.
Acquisition. Hard pin, no fallback: counters and thermal evidence solely from smartctl -a -j, controller identity solely from sysfs (/sys/class/nvme/<ctrl>/); vid/ssvid from the PCI node when present, else null. Any acquisition failure fails the whole run — a partial sample is never written. Identity normalization happens exactly once at write time: strip trailing spaces and newlines, no case folding, empty-after-strip stored blank. make install verifies smartctl and adds no dependency beyond smartmontools.
Controller identity. Segment key ladder: normalized kernel-exposed subsystem NQN → kernel composite → model|serial; FR is metadata only. Identity change and DUW decrease are independent axes: an identity change (including any to-or-from-blank change) quarantines prior history from projection; a DUW decrease with unchanged identity re-warms only; equal blank keys continue a segment by DUW monotonicity. A blank key marks the segment identity_degraded (set exactly when the key is blank), capping confidence at Limited with the fixed fact "controller identity unavailable — replacement detection relies on write-counter continuity only". Legacy history imports under a labeled mn-only legacy identity.
Hour/day derivation. Named constants, no configuration surface: powered-off when the hour's power-on-hours delta is below 90 % of its wall-clock span; active at ≥ 256 MiB DUW; idle when powered on, sampled, below active; unknown otherwise (machine-off and collector failure are indistinguishable by design). Denominator is wall-clock seconds inside monitoring periods, including powered-off and unknown time; disabled time is not an hour state and is excluded from numerator and denominator. Unexplained gaps keep the aggregate counter delta as unknown seconds and reduce coverage. No hour is ever interpolated, estimated, or fabricated.
Projection and confidence. One projection from the precedence-chosen baseline (verified override → unverified override → implied → unavailable), with the arithmetic fixed as:
Default regime: full observation history capped at 90 days. Habit change: trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean for 3 consecutive days → new regime at the first divergence day, auto-adopted, labeled; a regime younger than 7 days caps at Limited. Warming until 14 distinct UTC day aggregates with ≤ 2 below 50 % coverage, rendering facts meanwhile. Confidence rule table (verbatim from ADR 0002 §8 as amended):
Unavailable — no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
Supported — verified baseline and ≥ 14 qualifying days and coverage ≥ 80 % and fresh (< 48 h) and 7/28/90 rates within a factor of 2 and no single day ≥ 50 % of trailing 28-day bytes and regime ≥ 7 days old and current segment's identity key not degraded.
Limited — every other case with a baseline and a positive rate; failing facts shown.
Staleness (newest day aggregate ≥ 48 h old) drops confidence one level. Zero rate renders "no finite projection from this history". The scenario range (7/28/90, computed independently, covered horizons only) is the only spread shown; no statistical interval anywhere. The projection contract hands the views exactly: confidence state, contributing facts, headline when one exists, scenario range, Percentage-Used context line, disclosure text — recomputed on read.
Endurance baseline provenance. Mandatory provenance (source URL, document revision, entry date, model string, nominal capacity); verification derived at read (machine match or recorded user attestation), never a stored boolean; incomplete provenance stores only behind an explicit unverified acknowledgment; entry-time unprivileged sysfs validation (normalized model containment with interactive confirm recorded as validated_by = user; capacity within ± 1 %); read-time applicability is a model match against the current controller segment — a mismatch is retained, never auto-deleted, and leaves the projection Unavailable. Persistence goes through the polkit-guarded helper verb after CLI-side validation.
Panes TUI. Textual on Python 3.9+ (install-time gate). One dense keyboard-first screen: full-width headline band (lifespan headline or its no-projection wording with regime line; confidence state plus contributing facts; scenario range); usage-history pane left (sparkline with ▲ habit-change and ? unexplained-gap markers plus legend; habit-split bar); drive-health + settings pane right (health facts, vendor-wear context line, read-only settings with baseline and provenance label); full-width service strip bottom (four separate service facts, monitoring-period line, action legend). Bindings: p pause (asks), r resume (doesn't), c collect now, d disclosures, q quit. Privileged actions run as terminal-attached subprocesses so the platform polkit agent prompts on the real terminal (validated live in the prototype). The prototype is visual reference only; the layout above is normative.
Service lifecycle and sanctioned toggle. Two units only: fenris-collect.timer (timers.target; OnBootSec=2min, OnUnitInactiveSec=5min, AccuracySec=30s, Persistent=no) and fenris-collect.service (Type=oneshot, root, 90 s timeout). /etc/fenris/fenris.conf holds exactly one key — the device selector (stable by-id path preferred, raw nodes warned) — re-read every run. Period-row idempotent matrix (from ADR 0003 §6):
Situation
Effect
First-ever enable
Opens a period at the enable moment
Resume with an open period (raw stop intervened)
No row changes; gap stays inside as unknown seconds
Resume with no open period
Opens a new row at the resume moment
Pause with an open period
Closes it user_disabled at the pause moment
Pause otherwise
No-op
Raw systemctl stop/disable
Unexplained gap, never user_disabled
fenris sample and collect-now route through the helper → systemctl start the oneshot, block, and report the outcome synchronously. Freshness constants shared by TUI and CLI: fresh within 2 × cadence + AccuracySec + 60 s; missed until 48 h; stale at ≥ 48 h; empty store reads "no observations yet" with an enable hint; grade derives from the newest sample timestamp, never a stored flag. CLI twins exist for every TUI action with identical outcomes and wording; status is a read-only composition with a journal hint on failure or staleness; retired commands are rejected with pointers; the menu script is removed from the repository and the README maps its five options to successors.
Failure and recovery. Visible degradation, never fabrication: the collector validates every row against store invariants (hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count) and a violating run writes nothing, logs the refused row, and fails visibly; readers defensively exclude and count malformed rows; never backfill; a store fault degrades every dependent view and the collector never recreates or overwrites an existing file — recovery is a documented human-sanctioned move-aside (with legacy re-import if the import never completed); newer-schema stores are refused symmetrically; repeated failures retry at flat cadence with no backoff and no notification machinery; drive-reported anomalies render as ordinary facts and never affect the projection; a run finding no open monitoring period opens one at the run moment, never backdated.
Installation.make install builds a wheel into a dedicated /opt/fenris venv with exact lockfile pins, places units/helpers/polkit policy at their fixed locations recorded in an explicit manifest, verifies Python ≥ 3.9 and smartctl, creates /var/lib/fenris with root-written group-read permissions (database created lazily by first write), detects and imports legacy history, and never enables or starts anything — opt-in is the sanctioned toggle or the first-run TUI prompt. make upgrade installs into the same venv, syncs units and polkit against the manifest (daemon-reload; timer restarted only if contents changed and active), never kills an in-flight run, snapshots the database to a one-generation .bak, then applies forward-only migrations; /var/lib/fenris is never rebuilt. make uninstall performs the sanctioned disable first, then removes artifacts while keeping configuration and store; make purge removes those too. Refreshing pins is an explicit make update-deps step, never an install side effect.
Testing Decisions
What makes a good test here: assert external behavior at the agreed seams — store contents after a collection run, the projection-contract tuple, rendered status text, headless TUI snapshots, CLI exit outcomes — never internal helpers. The acceptance criteria register (evidence classes A/P/M) is the definition of done; every A-class criterion maps to a test at one of the two seam directions below.
One pivot seam — the observation store database file, confirmed with the owner, tested from both directions:
Write side: a collection run is exercised as a function of (smartctl-JSON fixture, sysfs fixture tree, existing store, config fixture, injected clock) → (resulting store contents, run outcome). Covers acquisition and normalization, store schema, migration (idempotency, interruption, quarantine, rename-after-commit, hourly.jsonl distrust), retention pruning, versioning refusal, write-boundary validation, hour/day derivation — no device or privileges needed.
Read side: the projection and both views are exercised as a function of (synthetic store, injected clock) → (projection-contract tuple, rendered status text, headless TUI snapshot). The flagship is the exhaustive state matrix: every realizable combination of confidence state × freshness grade × baseline tier rendered exactly per the rule table and freshness constants — headline number only when allowed, contributing facts always, never a percentage.
Beyond the seam: P-class scripted system probes on a host with systemd, polkit, and the configured drive (units, timer behavior, toggle paths, install/upgrade/uninstall, acquisition against the real device); M-class manual checks reserved for what fixtures cannot capture (live polkit agent prompt, pause/resume keyboard feel, tty passthrough).
Prior art: no legacy test suite exists. The prototype's headless Textual smoke test (SVG captures of every variant × state, all checks green) is the pattern for read-side TUI testing; its scenario fixtures (steady, warming, habit-changed, stale, no-baseline, paused) seed the synthetic-store builders.
Given/When/Then test specs are derived at implementation time from the criteria register; this issue fixes the seams, not the test list.
Out of Scope
Retaining the HTML dashboard, HTTP server, or any HTTP API.
Non-systemd operating systems or init systems; Windows/macOS.
Simultaneous monitoring of multiple drives.
Network telemetry or automatic vendor-data fetching of any kind.
Alerting, notification, or escalation machinery.
Statistical confidence intervals or percentage-based confidence.
The TUI prototype (branch prototype/tui-information-architecture) remains the visual reference; the tty-passthrough mechanism was validated live under Textual 8.x on Python 3.13.
Python floor 3.9 is an install-time gate; the venv at /opt/fenris carries exact pins so the runtime floor is the lockfile's.
Implements the redesign whose decisions are settled and assembled: [docs/spec/fenris-redesign.md](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/spec/fenris-redesign.md) (operative contracts) · [docs/spec/acceptance-criteria.md](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/spec/acceptance-criteria.md) (definition of done) · ADRs 0001–0006 (rationale) · [`CONTEXT.md`](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/CONTEXT.md) (glossary). Terminology throughout follows the glossary.
## Problem Statement
Fenris today is a checkout-resident daemon (`fenris.py` + `fenris.sh` menu) that polls the drive, appends JSONL files beside the code, and serves an HTML dashboard. As a drive owner I can't trust what it tells me: the projection is a trailing-24-hour rate against endurance guessed from Percentage Used or synthesized from capacity, with no notion of evidence quality, drive replacement, or changed habits. As an administrator I can't manage it: state lives wherever the checkout sits, elevation means passwordless sudo for smartctl, and nothing survives moving or deleting the checkout. And the dashboard needs a browser and a port — there is no keyboard-first, terminal-native way to watch the drive.
## Solution
A systemd timer drives a short-lived privileged collector that interrogates one configured NVMe drive every five minutes and persists a compact observation history into a single SQLite observation store. An unprivileged, keyboard-first Panes TUI (plus a headless `status` CLI twin) recomputes — on every read — one usage-adjusted theoretical lifespan from the best available endurance baseline, with projection confidence rendered as categorical evidence (state plus contributing facts, never a percentage), a 7/28/90-day scenario range, and the six endurance disclosures. Monitoring pause/resume is a sanctioned, polkit-authorized toggle that records intent as monitoring-period rows; legacy history imports in one interruption-safe transaction; install is a dormant, manifest-tracked, venv-delivered system package with exact dependency pins.
## User Stories
1. As a drive owner, I want every collection run persisted to a durable observation store, so that my observation history survives reboots, restarts, and upgrades.
2. As a drive owner, I want the observation store consistent to read while the collector writes, so that opening the TUI never blocks or corrupts.
3. As a drive owner, I want hour observations and day aggregates retained indefinitely, so that long-term usage-habit evidence accumulates.
4. As a drive owner, I want raw samples pruned after 14 days, so that the store stays compact without losing habit evidence.
5. As a drive owner, I want hours and days bounded by UTC, so that day boundaries never shift with daylight-saving time.
6. As a drive owner, I want my legacy `history.jsonl` observations imported in one interruption-safe transaction, so that nothing is lost or duplicated when I move to the redesign.
7. As a drive owner, I want a second import run to no-op on the legacy-import marker, so that migration is idempotent.
8. As a drive owner, I want malformed legacy lines quarantined with a logged count, so that nothing is silently dropped.
9. As a drive owner, I want legacy files renamed `.migrated` rather than deleted after commit, so that I keep a recovery trail.
10. As a drive owner, I want the legacy `hourly.jsonl` treated as derived and never trusted, so that mismatches are diffed and logged rather than imported as truth.
11. As a drive owner, I want the schema versioned with ordered, transactional, forward-only migrations, so that upgrades never destroy the observation store.
12. As a drive owner, I want a store written by a newer Fenris refused with "observation store written by a newer Fenris — upgrade Fenris", so that no partial interpretation garbles my history.
13. As a drive owner, I want collection to run on a five-minute systemd timer with no resident daemon, so that observation continues without a long-lived process to babysit.
14. As a drive owner, I want each collection run to acquire counters and thermal evidence from smartctl and controller identity from sysfs in one all-or-nothing step, so that partial samples never enter the store.
15. As a drive owner, I want identity normalized exactly once at write time, so that an implementation change can never split my drive's own history.
16. As an administrator, I want a hung device interrogation to fail visibly within 90 seconds as a bounded failed run, so that one stuck run can't wedge the system.
17. As an administrator, I want the device selector as the single configuration key, so that changing the monitored drive is a one-key edit re-read every run.
18. As an administrator, I want an invalid device selector surfaced as a `configuration error` fact in status and the TUI, so that misconfiguration is visible rather than mysterious.
19. As a drive owner, I want failed collection runs retried at the flat cadence with freshness visibly degrading, so that recovery is simply the next successful run.
20. As a drive owner, I want exactly one usage-adjusted theoretical lifespan computed from the precedence-chosen endurance baseline, so that I'm never shown two competing numbers.
21. As a drive owner, I want Percentage Used rendered as a vendor-wear context line with a note when it disagrees with my observed write rate, so that vendor wear is visible but never a second projection.
22. As a drive owner, I want my rated-TBW override recorded with mandatory provenance and validated against the detected drive, so that the projection rests on evidence I can audit.
23. As a drive owner, I want to knowingly store an unverified override when provenance is incomplete, so that honesty about provenance never blocks a projection I asked for.
24. As a drive owner, I want an implied-from-vendor-wear baseline used only after at least two Percentage-Used increments in the current controller segment, so that quantization noise never masquerades as endurance.
25. As a drive owner, I want a baseline that no longer matches my drive retained rather than auto-deleted, so that my records are never silently destroyed.
26. As a drive owner, I want the headline rate taken from my most recent sustained regime, so that the projection reflects how I actually use the drive now.
27. As a drive owner, I want a detected habit change adopted automatically and labeled "usage habit changed N days ago", so that a new regime replaces a stale one without my intervention.
28. As a drive owner, I want a 7/28/90-day scenario range showing only horizons my history covers, so that I see horizon sensitivity instead of false precision.
29. As a drive owner, I want projection confidence rendered as a state plus contributing facts and never a percentage, so that I can judge the evidence myself.
30. As a drive owner, I want a warming-up indication while evidence accumulates, so that I understand why the projection is soft and see it improve.
31. As a drive owner, I want a newest day aggregate older than 48 hours to drop confidence one level as a shown fact, so that an abandoned monitor never looks current.
32. As a drive owner, I want a zero write rate reported as "no finite projection from this history", so that infinity or zero is never displayed.
33. As a drive owner, I want the six endurance disclosures always available in TUI and status, so that I'm never misled about what the number means.
34. As a drive owner, I want observation history segmented by controller identity, so that a drive replacement never contaminates my projection.
35. As a drive owner, I want a write-counter decrease with unchanged identity to quarantine nothing but re-warm, so that a controller reset doesn't erase my history.
36. As a drive owner, I want a degraded-identity segment capped at Limited evidence with its explaining fact, so that replacement-detection blindness is disclosed, not hidden.
37. As a drive owner, I want segment metadata frozen at segment open, so that diagnostics always reflect the identity that opened the segment.
38. As a drive owner, I want my legacy history imported under a labeled model-scoped legacy identity, so that it never blends with post-redesign identity.
39. As a drive owner, I want pause to ask for confirmation and explain that paused time is excluded from my usage habit, so that deliberate disables are never accidental.
40. As a drive owner, I want resume to act immediately without confirmation, so that getting back to monitoring is frictionless.
41. As a drive owner, I want pause/resume to perform the systemctl operation and the monitoring-period bookkeeping in one polkit-authorized step, so that intent and service state never diverge.
42. As an administrator, I want a raw systemctl stop or disable to register as an unexplained gap rather than a deliberate disable, so that only Fenris's own control path records intent.
43. As a drive owner, I want collect-now on demand with a synchronous outcome, so that I can refresh without waiting for the timer and without the TUI touching the device.
44. As a drive owner, I want `status` as a read-only composition that never auto-samples and never prompts, so that checking state never mutates anything.
45. As a drive owner, I want retired commands (`start`, `stop`, `run`, `--device`) rejected with one-line migration pointers, so that old habits fail loudly instead of silently changing meaning.
46. As a drive owner, I want boot enablement, runtime activity, last collect outcome, and freshness displayed as four separate facts, so that service state is unambiguous at a glance.
47. As a drive owner, I want one dense keyboard-first Panes screen with no page navigation, so that everything is visible at once.
48. As a drive owner, I want a write-history sparkline with habit-change and unexplained-gap markers, so that changes and holes stand out.
49. As a drive owner, I want a habit-split bar with active/idle/powered-off/unknown shares, so that my usage habit is glanceable.
50. As a drive owner, I want drive health and read-only settings in a side pane, so that facts and configuration are clearly separated.
51. As a drive owner, I want polkit prompts to appear on my real terminal when acting from the TUI, so that elevation works without breaking the interface.
52. As a drive owner, I want an empty store greeted with "no observations yet" and an enable hint, so that a first run is self-explanatory.
53. As a drive owner, I want an invariant-violating collection run to write nothing and fail visibly, so that the observation store only ever holds well-formed data.
54. As a drive owner, I want readers to defensively exclude and count malformed rows as a fact, so that even the impossible is visible.
55. As a drive owner, I want gaps never backfilled — no interpolation, estimation, or fabrication — so that no synthetic hour ever enters my history.
56. As a drive owner, I want a store fault to surface "observation store unreadable" with a journal hint and suppress everything store-dependent, so that corruption is visible without silent guessing.
57. As an administrator, I want store-fault recovery to be a documented, human-sanctioned move-aside with re-import if the legacy import never completed, so that nothing destructive is automated.
58. As a drive owner, I want critical warnings, media errors, and unsafe shutdowns rendered as ordinary facts that never affect the projection, so that I'm informed without alarm machinery.
59. As a drive owner, I want a collection run that finds no open monitoring period to open one at the run moment, so that observed fact is recorded without pretending intent.
60. As an administrator, I want `make install` to deliver a fully dormant system, so that nothing starts monitoring without my explicit opt-in.
61. As an administrator, I want install-time gates for Python ≥ 3.9 and smartctl, so that missing prerequisites fail cleanly instead of crashing at runtime.
62. As an administrator, I want every placed file recorded in an explicit manifest, so that upgrade and uninstall are exact.
63. As an administrator, I want upgrades to never kill an in-flight collection run and never rebuild the observation store, so that history is safe across versions.
64. As an administrator, I want a one-generation database backup before migrations, so that rollback is possible.
65. As an administrator, I want uninstall to perform the sanctioned disable first and keep my configuration and observation history, so that removal is deliberate but reversible — and purge available when I truly want everything gone.
66. As an administrator, I want exact dependency pins installed from a committed lockfile, so that installs and upgrades are reproducible.
67. As a drive owner, I want the README to map the retired menu's five options to their successors, so that migrating my habits is easy.
## Implementation Decisions
All decisions below are settled; the [redesign specification](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/spec/fenris-redesign.md) restates each normatively with a criteria block, and ADRs 0001–0006 hold the rationale. Nothing here should be re-litigated during implementation.
**Architecture.** The checkout-resident daemon, HTTP dashboard, PID file, and menu script are replaced by: a timer-driven oneshot collector (the only code path that interrogates the device), a privileged fixed-operation helper mediating enable/disable/collect-trigger/baseline-persistence under one polkit action (`auth_admin`), and an unprivileged wrapper exposing the Panes TUI (no arguments) and the CLI subcommands. No `/run` coordination surface; state lives in the observation store, coordination in systemd, failures in the journal.
**Observation store.** One SQLite database in WAL mode at `/var/lib/fenris/observations.db`, root-owned and group-readable through a `fenris` read group; readers open read-only. Six entities: `samples` (14-day raw retention), `hour_observations` (UTC-hour usage-habit split, write/read deltas, thermal min/avg/max, sample count, coverage flag), `day_aggregates` (the habit-evidence grain), `monitoring_periods` (`started_at`, nullable `ended_at`, `end_cause` enum `user_disabled`/`migrated`/…), `controller_segments` (identity key plus frozen nullable metadata snapshot: normalized `subnqn`/`sn`/`mn`/`fr`, `vid`/`ssvid`/`transport`, `identity_degraded`; `cntlid` excluded), and `endurance_baseline` (one active row, replaced on edit). Projections are never stored; no latest-status table; no stored health flag. Schema versioning via `PRAGMA user_version` with ordered per-step transactional migrations; unknown newer version is refused by collector and readers alike.
**Acquisition.** Hard pin, no fallback: counters and thermal evidence solely from `smartctl -a -j`, controller identity solely from sysfs (`/sys/class/nvme/<ctrl>/`); `vid`/`ssvid` from the PCI node when present, else null. Any acquisition failure fails the whole run — a partial sample is never written. Identity normalization happens exactly once at write time: strip trailing spaces and newlines, no case folding, empty-after-strip stored blank. `make install` verifies smartctl and adds no dependency beyond smartmontools.
**Controller identity.** Segment key ladder: normalized kernel-exposed subsystem NQN → kernel composite → `model|serial`; FR is metadata only. Identity change and DUW decrease are independent axes: an identity change (including any to-or-from-blank change) quarantines prior history from projection; a DUW decrease with unchanged identity re-warms only; equal blank keys continue a segment by DUW monotonicity. A blank key marks the segment `identity_degraded` (set exactly when the key is blank), capping confidence at Limited with the fixed fact "controller identity unavailable — replacement detection relies on write-counter continuity only". Legacy history imports under a labeled mn-only legacy identity.
**Hour/day derivation.** Named constants, no configuration surface: powered-off when the hour's power-on-hours delta is below 90 % of its wall-clock span; active at ≥ 256 MiB DUW; idle when powered on, sampled, below active; unknown otherwise (machine-off and collector failure are indistinguishable by design). Denominator is wall-clock seconds inside monitoring periods, including powered-off and unknown time; disabled time is not an hour state and is excluded from numerator and denominator. Unexplained gaps keep the aggregate counter delta as unknown seconds and reduce coverage. No hour is ever interpolated, estimated, or fabricated.
**Projection and confidence.** One projection from the precedence-chosen baseline (verified override → unverified override → implied → unavailable), with the arithmetic fixed as:
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
Default regime: full observation history capped at 90 days. Habit change: trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean for 3 consecutive days → new regime at the first divergence day, auto-adopted, labeled; a regime younger than 7 days caps at Limited. Warming until 14 distinct UTC day aggregates with ≤ 2 below 50 % coverage, rendering facts meanwhile. Confidence rule table (verbatim from ADR 0002 §8 as amended):
- **Unavailable** — no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported** — verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80 % **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 **and** no single day ≥ 50 % of trailing 28-day bytes **and** regime ≥ 7 days old **and** current segment's identity key not degraded.
- **Limited** — every other case with a baseline and a positive rate; failing facts shown.
Staleness (newest day aggregate ≥ 48 h old) drops confidence one level. Zero rate renders "no finite projection from this history". The scenario range (7/28/90, computed independently, covered horizons only) is the only spread shown; no statistical interval anywhere. The projection contract hands the views exactly: confidence state, contributing facts, headline when one exists, scenario range, Percentage-Used context line, disclosure text — recomputed on read.
**Endurance baseline provenance.** Mandatory provenance (source URL, document revision, entry date, model string, nominal capacity); verification derived at read (machine match or recorded user attestation), never a stored boolean; incomplete provenance stores only behind an explicit unverified acknowledgment; entry-time unprivileged sysfs validation (normalized model containment with interactive confirm recorded as `validated_by = user`; capacity within ± 1 %); read-time applicability is a model match against the current controller segment — a mismatch is retained, never auto-deleted, and leaves the projection Unavailable. Persistence goes through the polkit-guarded helper verb after CLI-side validation.
**Panes TUI.** Textual on Python 3.9+ (install-time gate). One dense keyboard-first screen: full-width headline band (lifespan headline or its no-projection wording with regime line; confidence state plus contributing facts; scenario range); usage-history pane left (sparkline with ▲ habit-change and ? unexplained-gap markers plus legend; habit-split bar); drive-health + settings pane right (health facts, vendor-wear context line, read-only settings with baseline and provenance label); full-width service strip bottom (four separate service facts, monitoring-period line, action legend). Bindings: `p` pause (asks), `r` resume (doesn't), `c` collect now, `d` disclosures, `q` quit. Privileged actions run as terminal-attached subprocesses so the platform polkit agent prompts on the real terminal (validated live in the prototype). The [prototype](https://git.bongbetic.com/xavierk/Fenris/src/branch/prototype/tui-information-architecture/prototype/tui-ia) is visual reference only; the layout above is normative.
**Service lifecycle and sanctioned toggle.** Two units only: `fenris-collect.timer` (timers.target; `OnBootSec=2min`, `OnUnitInactiveSec=5min`, `AccuracySec=30s`, `Persistent=no`) and `fenris-collect.service` (`Type=oneshot`, root, 90 s timeout). `/etc/fenris/fenris.conf` holds exactly one key — the device selector (stable by-id path preferred, raw nodes warned) — re-read every run. Period-row idempotent matrix (from ADR 0003 §6):
| Situation | Effect |
|---|---|
| First-ever enable | Opens a period at the enable moment |
| Resume with an open period (raw stop intervened) | No row changes; gap stays inside as unknown seconds |
| Resume with no open period | Opens a new row at the resume moment |
| Pause with an open period | Closes it `user_disabled` at the pause moment |
| Pause otherwise | No-op |
| Raw systemctl stop/disable | Unexplained gap, never `user_disabled` |
`fenris sample` and collect-now route through the helper → `systemctl start` the oneshot, block, and report the outcome synchronously. Freshness constants shared by TUI and CLI: fresh within 2 × cadence + AccuracySec + 60 s; missed until 48 h; stale at ≥ 48 h; empty store reads "no observations yet" with an enable hint; grade derives from the newest sample timestamp, never a stored flag. CLI twins exist for every TUI action with identical outcomes and wording; `status` is a read-only composition with a journal hint on failure or staleness; retired commands are rejected with pointers; the menu script is removed from the repository and the README maps its five options to successors.
**Failure and recovery.** Visible degradation, never fabrication: the collector validates every row against store invariants (hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count) and a violating run writes nothing, logs the refused row, and fails visibly; readers defensively exclude and count malformed rows; never backfill; a store fault degrades every dependent view and the collector never recreates or overwrites an existing file — recovery is a documented human-sanctioned move-aside (with legacy re-import if the import never completed); newer-schema stores are refused symmetrically; repeated failures retry at flat cadence with no backoff and no notification machinery; drive-reported anomalies render as ordinary facts and never affect the projection; a run finding no open monitoring period opens one at the run moment, never backdated.
**Installation.** `make install` builds a wheel into a dedicated `/opt/fenris` venv with exact lockfile pins, places units/helpers/polkit policy at their fixed locations recorded in an explicit manifest, verifies Python ≥ 3.9 and smartctl, creates `/var/lib/fenris` with root-written group-read permissions (database created lazily by first write), detects and imports legacy history, and never enables or starts anything — opt-in is the sanctioned toggle or the first-run TUI prompt. `make upgrade` installs into the same venv, syncs units and polkit against the manifest (daemon-reload; timer restarted only if contents changed and active), never kills an in-flight run, snapshots the database to a one-generation `.bak`, then applies forward-only migrations; `/var/lib/fenris` is never rebuilt. `make uninstall` performs the sanctioned disable first, then removes artifacts while keeping configuration and store; `make purge` removes those too. Refreshing pins is an explicit `make update-deps` step, never an install side effect.
## Testing Decisions
- **What makes a good test here:** assert external behavior at the agreed seams — store contents after a collection run, the projection-contract tuple, rendered `status` text, headless TUI snapshots, CLI exit outcomes — never internal helpers. The [acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/spec/acceptance-criteria.md) register (evidence classes A/P/M) is the definition of done; every A-class criterion maps to a test at one of the two seam directions below.
- **One pivot seam — the observation store database file**, confirmed with the owner, tested from both directions:
- **Write side:** a collection run is exercised as a function of (smartctl-JSON fixture, sysfs fixture tree, existing store, config fixture, injected clock) → (resulting store contents, run outcome). Covers acquisition and normalization, store schema, migration (idempotency, interruption, quarantine, rename-after-commit, `hourly.jsonl` distrust), retention pruning, versioning refusal, write-boundary validation, hour/day derivation — no device or privileges needed.
- **Read side:** the projection and both views are exercised as a function of (synthetic store, injected clock) → (projection-contract tuple, rendered `status` text, headless TUI snapshot). The flagship is the exhaustive state matrix: every realizable combination of confidence state × freshness grade × baseline tier rendered exactly per the rule table and freshness constants — headline number only when allowed, contributing facts always, never a percentage.
- **Beyond the seam:** P-class scripted system probes on a host with systemd, polkit, and the configured drive (units, timer behavior, toggle paths, install/upgrade/uninstall, acquisition against the real device); M-class manual checks reserved for what fixtures cannot capture (live polkit agent prompt, pause/resume keyboard feel, tty passthrough).
- **Prior art:** no legacy test suite exists. The prototype's headless Textual smoke test (SVG captures of every variant × state, all checks green) is the pattern for read-side TUI testing; its scenario fixtures (steady, warming, habit-changed, stale, no-baseline, paused) seed the synthetic-store builders.
- Given/When/Then test specs are derived at implementation time from the criteria register; this issue fixes the seams, not the test list.
## Out of Scope
- Retaining the HTML dashboard, HTTP server, or any HTTP API.
- Non-systemd operating systems or init systems; Windows/macOS.
- Simultaneous monitoring of multiple drives.
- Network telemetry or automatic vendor-data fetching of any kind.
- Alerting, notification, or escalation machinery.
- Statistical confidence intervals or percentage-based confidence.
- Automatic schema downgrade; built-in destructive recovery commands.
- Re-opening settled decisions: behavioral changes to the redesign specification or ADRs become new tickets, not implementation liberties.
- Tuning the documented constants (thresholds, cadence, retention) beyond their fixed values.
- Anything the [map's Out of scope](https://git.bongbetic.com/xavierk/Fenris/issues/1) already ruled out.
## Further Notes
- Canonical reading order for the implementer: [`CONTEXT.md`](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/CONTEXT.md) (glossary) → [redesign specification](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/spec/fenris-redesign.md) (operative contracts, normative constants index, traceability appendix) → [acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/docs/spec/acceptance-criteria.md) (definition of done) → ADRs 0001–0006 (rationale).
- The Wayfinder map [Chart Fenris's persistent TUI monitoring redesign](https://git.bongbetic.com/xavierk/Fenris/issues/1) is closed; its decisions live on in the ADRs and the two spec documents. This issue is the implementation handoff.
- The TUI prototype (branch `prototype/tui-information-architecture`) remains the visual reference; the tty-passthrough mechanism was validated live under Textual 8.x on Python 3.13.
- Python floor 3.9 is an install-time gate; the venv at `/opt/fenris` carries exact pins so the runtime floor is the lockfile's.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Implements the redesign whose decisions are settled and assembled: docs/spec/fenris-redesign.md (operative contracts) · docs/spec/acceptance-criteria.md (definition of done) · ADRs 0001–0006 (rationale) ·
CONTEXT.md(glossary). Terminology throughout follows the glossary.Problem Statement
Fenris today is a checkout-resident daemon (
fenris.py+fenris.shmenu) that polls the drive, appends JSONL files beside the code, and serves an HTML dashboard. As a drive owner I can't trust what it tells me: the projection is a trailing-24-hour rate against endurance guessed from Percentage Used or synthesized from capacity, with no notion of evidence quality, drive replacement, or changed habits. As an administrator I can't manage it: state lives wherever the checkout sits, elevation means passwordless sudo for smartctl, and nothing survives moving or deleting the checkout. And the dashboard needs a browser and a port — there is no keyboard-first, terminal-native way to watch the drive.Solution
A systemd timer drives a short-lived privileged collector that interrogates one configured NVMe drive every five minutes and persists a compact observation history into a single SQLite observation store. An unprivileged, keyboard-first Panes TUI (plus a headless
statusCLI twin) recomputes — on every read — one usage-adjusted theoretical lifespan from the best available endurance baseline, with projection confidence rendered as categorical evidence (state plus contributing facts, never a percentage), a 7/28/90-day scenario range, and the six endurance disclosures. Monitoring pause/resume is a sanctioned, polkit-authorized toggle that records intent as monitoring-period rows; legacy history imports in one interruption-safe transaction; install is a dormant, manifest-tracked, venv-delivered system package with exact dependency pins.User Stories
history.jsonlobservations imported in one interruption-safe transaction, so that nothing is lost or duplicated when I move to the redesign..migratedrather than deleted after commit, so that I keep a recovery trail.hourly.jsonltreated as derived and never trusted, so that mismatches are diffed and logged rather than imported as truth.configuration errorfact in status and the TUI, so that misconfiguration is visible rather than mysterious.statusas a read-only composition that never auto-samples and never prompts, so that checking state never mutates anything.start,stop,run,--device) rejected with one-line migration pointers, so that old habits fail loudly instead of silently changing meaning.make installto deliver a fully dormant system, so that nothing starts monitoring without my explicit opt-in.Implementation Decisions
All decisions below are settled; the redesign specification restates each normatively with a criteria block, and ADRs 0001–0006 hold the rationale. Nothing here should be re-litigated during implementation.
Architecture. The checkout-resident daemon, HTTP dashboard, PID file, and menu script are replaced by: a timer-driven oneshot collector (the only code path that interrogates the device), a privileged fixed-operation helper mediating enable/disable/collect-trigger/baseline-persistence under one polkit action (
auth_admin), and an unprivileged wrapper exposing the Panes TUI (no arguments) and the CLI subcommands. No/runcoordination surface; state lives in the observation store, coordination in systemd, failures in the journal.Observation store. One SQLite database in WAL mode at
/var/lib/fenris/observations.db, root-owned and group-readable through afenrisread group; readers open read-only. Six entities:samples(14-day raw retention),hour_observations(UTC-hour usage-habit split, write/read deltas, thermal min/avg/max, sample count, coverage flag),day_aggregates(the habit-evidence grain),monitoring_periods(started_at, nullableended_at,end_causeenumuser_disabled/migrated/…),controller_segments(identity key plus frozen nullable metadata snapshot: normalizedsubnqn/sn/mn/fr,vid/ssvid/transport,identity_degraded;cntlidexcluded), andendurance_baseline(one active row, replaced on edit). Projections are never stored; no latest-status table; no stored health flag. Schema versioning viaPRAGMA user_versionwith ordered per-step transactional migrations; unknown newer version is refused by collector and readers alike.Acquisition. Hard pin, no fallback: counters and thermal evidence solely from
smartctl -a -j, controller identity solely from sysfs (/sys/class/nvme/<ctrl>/);vid/ssvidfrom the PCI node when present, else null. Any acquisition failure fails the whole run — a partial sample is never written. Identity normalization happens exactly once at write time: strip trailing spaces and newlines, no case folding, empty-after-strip stored blank.make installverifies smartctl and adds no dependency beyond smartmontools.Controller identity. Segment key ladder: normalized kernel-exposed subsystem NQN → kernel composite →
model|serial; FR is metadata only. Identity change and DUW decrease are independent axes: an identity change (including any to-or-from-blank change) quarantines prior history from projection; a DUW decrease with unchanged identity re-warms only; equal blank keys continue a segment by DUW monotonicity. A blank key marks the segmentidentity_degraded(set exactly when the key is blank), capping confidence at Limited with the fixed fact "controller identity unavailable — replacement detection relies on write-counter continuity only". Legacy history imports under a labeled mn-only legacy identity.Hour/day derivation. Named constants, no configuration surface: powered-off when the hour's power-on-hours delta is below 90 % of its wall-clock span; active at ≥ 256 MiB DUW; idle when powered on, sampled, below active; unknown otherwise (machine-off and collector failure are indistinguishable by design). Denominator is wall-clock seconds inside monitoring periods, including powered-off and unknown time; disabled time is not an hour state and is excluded from numerator and denominator. Unexplained gaps keep the aggregate counter delta as unknown seconds and reduce coverage. No hour is ever interpolated, estimated, or fabricated.
Projection and confidence. One projection from the precedence-chosen baseline (verified override → unverified override → implied → unavailable), with the arithmetic fixed as:
Default regime: full observation history capped at 90 days. Habit change: trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean for 3 consecutive days → new regime at the first divergence day, auto-adopted, labeled; a regime younger than 7 days caps at Limited. Warming until 14 distinct UTC day aggregates with ≤ 2 below 50 % coverage, rendering facts meanwhile. Confidence rule table (verbatim from ADR 0002 §8 as amended):
Staleness (newest day aggregate ≥ 48 h old) drops confidence one level. Zero rate renders "no finite projection from this history". The scenario range (7/28/90, computed independently, covered horizons only) is the only spread shown; no statistical interval anywhere. The projection contract hands the views exactly: confidence state, contributing facts, headline when one exists, scenario range, Percentage-Used context line, disclosure text — recomputed on read.
Endurance baseline provenance. Mandatory provenance (source URL, document revision, entry date, model string, nominal capacity); verification derived at read (machine match or recorded user attestation), never a stored boolean; incomplete provenance stores only behind an explicit unverified acknowledgment; entry-time unprivileged sysfs validation (normalized model containment with interactive confirm recorded as
validated_by = user; capacity within ± 1 %); read-time applicability is a model match against the current controller segment — a mismatch is retained, never auto-deleted, and leaves the projection Unavailable. Persistence goes through the polkit-guarded helper verb after CLI-side validation.Panes TUI. Textual on Python 3.9+ (install-time gate). One dense keyboard-first screen: full-width headline band (lifespan headline or its no-projection wording with regime line; confidence state plus contributing facts; scenario range); usage-history pane left (sparkline with ▲ habit-change and ? unexplained-gap markers plus legend; habit-split bar); drive-health + settings pane right (health facts, vendor-wear context line, read-only settings with baseline and provenance label); full-width service strip bottom (four separate service facts, monitoring-period line, action legend). Bindings:
ppause (asks),rresume (doesn't),ccollect now,ddisclosures,qquit. Privileged actions run as terminal-attached subprocesses so the platform polkit agent prompts on the real terminal (validated live in the prototype). The prototype is visual reference only; the layout above is normative.Service lifecycle and sanctioned toggle. Two units only:
fenris-collect.timer(timers.target;OnBootSec=2min,OnUnitInactiveSec=5min,AccuracySec=30s,Persistent=no) andfenris-collect.service(Type=oneshot, root, 90 s timeout)./etc/fenris/fenris.confholds exactly one key — the device selector (stable by-id path preferred, raw nodes warned) — re-read every run. Period-row idempotent matrix (from ADR 0003 §6):user_disabledat the pause momentuser_disabledfenris sampleand collect-now route through the helper →systemctl startthe oneshot, block, and report the outcome synchronously. Freshness constants shared by TUI and CLI: fresh within 2 × cadence + AccuracySec + 60 s; missed until 48 h; stale at ≥ 48 h; empty store reads "no observations yet" with an enable hint; grade derives from the newest sample timestamp, never a stored flag. CLI twins exist for every TUI action with identical outcomes and wording;statusis a read-only composition with a journal hint on failure or staleness; retired commands are rejected with pointers; the menu script is removed from the repository and the README maps its five options to successors.Failure and recovery. Visible degradation, never fabrication: the collector validates every row against store invariants (hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count) and a violating run writes nothing, logs the refused row, and fails visibly; readers defensively exclude and count malformed rows; never backfill; a store fault degrades every dependent view and the collector never recreates or overwrites an existing file — recovery is a documented human-sanctioned move-aside (with legacy re-import if the import never completed); newer-schema stores are refused symmetrically; repeated failures retry at flat cadence with no backoff and no notification machinery; drive-reported anomalies render as ordinary facts and never affect the projection; a run finding no open monitoring period opens one at the run moment, never backdated.
Installation.
make installbuilds a wheel into a dedicated/opt/fenrisvenv with exact lockfile pins, places units/helpers/polkit policy at their fixed locations recorded in an explicit manifest, verifies Python ≥ 3.9 and smartctl, creates/var/lib/fenriswith root-written group-read permissions (database created lazily by first write), detects and imports legacy history, and never enables or starts anything — opt-in is the sanctioned toggle or the first-run TUI prompt.make upgradeinstalls into the same venv, syncs units and polkit against the manifest (daemon-reload; timer restarted only if contents changed and active), never kills an in-flight run, snapshots the database to a one-generation.bak, then applies forward-only migrations;/var/lib/fenrisis never rebuilt.make uninstallperforms the sanctioned disable first, then removes artifacts while keeping configuration and store;make purgeremoves those too. Refreshing pins is an explicitmake update-depsstep, never an install side effect.Testing Decisions
statustext, headless TUI snapshots, CLI exit outcomes — never internal helpers. The acceptance criteria register (evidence classes A/P/M) is the definition of done; every A-class criterion maps to a test at one of the two seam directions below.hourly.jsonldistrust), retention pruning, versioning refusal, write-boundary validation, hour/day derivation — no device or privileges needed.statustext, headless TUI snapshot). The flagship is the exhaustive state matrix: every realizable combination of confidence state × freshness grade × baseline tier rendered exactly per the rule table and freshness constants — headline number only when allowed, contributing facts always, never a percentage.Out of Scope
Further Notes
CONTEXT.md(glossary) → redesign specification (operative contracts, normative constants index, traceability appendix) → acceptance criteria (definition of done) → ADRs 0001–0006 (rationale).prototype/tui-information-architecture) remains the visual reference; the tty-passthrough mechanism was validated live under Textual 8.x on Python 3.13./opt/fenriscarries exact pins so the runtime floor is the lockfile's.