Files
Fenris/docs/spec/fenris-redesign.md
T

52 KiB
Raw Permalink Blame History

Fenris redesign specification

Status: implementation-ready. Assembled by Write the Fenris redesign specification and close the map, executing the assembly decision Assemble the implementation-ready specification (all seven recommendations accepted) on the Wayfinder map Chart Fenris's persistent TUI monitoring redesign.

Canonical roles. ADRs 0001–0006 are the immutable rationale records — the why. Acceptance criteria are the single register of testable statements — the definition of done. This document normatively restates every operative contract — the what — so an implementer never needs Wayfinder-ticket access: schema column sets, constants, rule tables, unit and CLI definitions, and the Panes TUI layout. Nothing here overrides an ADR or restates a criterion as a criterion.

How to read this document

  • Binding language. Must, exactly, and never are normative. Terminology follows the glossary in CONTEXT.md: observation history, usage-adjusted theoretical lifespan, projection confidence, monitoring period, observation store, hour observation, day aggregate, controller segment, degraded identity, endurance baseline, verified override, unverified override, sustained regime, habit change, scenario range, coverage, collection run, deliberate disable, store fault.
  • Ordering. Sections follow data flow: system context → collector acquisition → observation store → controller identity & segmentation → hour/day derivation → projection & confidence → Panes TUI → service lifecycle & sanctioned toggle → failure & recovery → installation. Each section opens with its ADR links and criterion-ID block.
  • Implementation boundary. This specification plans the redesign; it does not implement it. The complete handoff is this document + the criteria register + ADRs 0001–0006 + the glossary. Given/When/Then test specs are derived by the implementer at implementation time.

Normative constants index

Every constant is defined once, in the section named below; other sections cite, never redefine. All are named constants in code, not configuration.

Constant Value Defined in
Collection cadence (default) 5 min (OnUnitInactiveSec) §8.2
First-boot delay 2 min (OnBootSec) §8.2
Timer accuracy window 30 s (AccuracySec) §8.2
Collection-run timeout 90 s (TimeoutStartSec) §8.2
Fresh threshold newest sample within 2 × cadence + AccuracySec + 60 s §8.9
Missed → stale boundary 48 h §8.9, §6.7
Powered-off hour threshold power-on-hours delta < 90 % of the hour's wall-clock span §5.1
Active hour threshold DUW delta ≥ 256 MiB in the hour §5.1
Raw-sample retention 14 days §3.4
Warming gate 14 distinct UTC day aggregates, ≤ 2 below 50 % coverage §6.6
Supported coverage floor 80 % §6.7
Horizon agreement 7/28/90-day rates within a factor of 2 §6.7
Burst guard no single day ≥ 50 % of trailing 28-day bytes §6.7
Young-regime cap regime < 7 days old → Limited §6.4
Habit-change trigger trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean, 3 consecutive days §6.4
Regime span cap (default) full observation history capped at 90 days §6.4
Scenario horizons 7 / 28 / 90 days §6.5
Implied-baseline eligibility ≥ 2 Percentage-Used increments within the current controller segment §6.3
Rated-TBW conversion E_rated = entered_TBW × 10¹² bytes §6.3
Implied-baseline validity window 1 ≤ p ≤ 254 §6.3
Wear-disagreement note vendor wear vs. observed write rate by more than a factor of 2 §6.1
Capacity validation tolerance ± 1 % §6.2

1. System context

ADRs: 0001, 0003, 0006. Criteria: CI-3, LC-1, LC-5, ST-1.

1.1 Scope

Fenris observes one configured NVMe drive's real-world use and translates the observation history into a usage-adjusted theoretical lifespan. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer, persistent compact observation storage, and categorical projection confidence reflecting the length, completeness, and stability of real usage history.

Standing constraints, binding on every section:

  • Linux with systemd and polkit only; no other init system is supported.
  • Exactly one configured NVMe drive — the device named by /etc/fenris/fenris.conf (§8.3).
  • Fully local: no telemetry, no network fetching, no automatic vendor-data retrieval.
  • The HTML dashboard and HTTP server are gone; nothing of the daemonization, PID files, or /run state survives.
  • CLI status and sample are retained (§8.8).

1.2 Components and privilege boundaries

Component Privilege Path Role
fenris-collect.service root oneshot unit /usr/libexec/fenris/fenris-collect The only code path that interrogates the device and writes the observation store.
fenris-collect.timer system timer — Schedules collection runs; WantedBy=timers.target.
fenris-monitor root helper /usr/libexec/fenris/fenris-monitor Fixed privileged operations: enable/disable (optional --now), the collect trigger, monitoring-period bookkeeping, and baseline persistence. The only binary polkit authorizes.
fenris unprivileged /usr/local/bin/fenris Human entry point: no arguments opens the TUI; subcommands are the CLI (§8.8). Never a unit.
Observation store root-written, group-read /var/lib/fenris/observations.db Single SQLite database in WAL mode (§3). The TUI and status open it read-only.
Configuration world-readable /etc/fenris/fenris.conf Exactly one key: the device selector (§8.3).

The TUI and CLI are ordinary unprivileged processes. Elevation is exclusively polkit, exclusively for fenris-monitor (§8.5). There is no /run/fenris coordination surface and no export layer: systemd serializes collection runs, the observation store holds state, failures go to the journal.

1.3 Data flow

  1. The timer fires; fenris-collect.service runs fenris-collect.
  2. The collector acquires counters and thermal evidence from smartctl -a -j and controller identity from sysfs (§2), normalizes identity exactly once (§2.3), and either fails the whole run or writes one complete sample.
  3. The collector derives and validates hour observations and day aggregates, advances controller segmentation and period bookkeeping, prunes raw samples, and commits (§3–§5, §9.1).
  4. Readers — the TUI and fenris status — open the store read-only and recompute the projection on every read (§6); nothing derived is ever stored (§3.7).

Control flow is separate: the human drives the TUI/CLI; privileged operations route through fenris-monitor under polkit to systemctl; period rows record intent (only the sanctioned path), while the collector records observed fact (§8.5–§8.6, §9.8).

1.4 Cross-cutting prohibitions

These are operative contracts; each is restated in its home section and gated by the criteria block CI-3:

  1. No code path outside fenris-collect interrogates the device (§2.1, §8.7).
  2. Polkit authorizes exactly one binary, fenris-monitor, under com.bongbetic.fenris.monitor auth_admin (§8.5).
  3. No /run/fenris coordination surface or export layer exists (§1.2).
  4. No absent hour is ever interpolated, estimated, or fabricated (§5.3, §9.3).
  5. No alerting, notification, or escalation machinery exists anywhere (§9.6–§9.7).
  6. /etc/fenris/fenris.conf holds exactly one key — the device selector (§8.3).
  7. No synthetic or capacity-derived baseline is ever created, including for legacy history (§6.1, §3.5).
  8. Readers never partially interpret a newer-schema store (§3.6, §9.5).
  9. Projections are never stored; always recomputed on read (§3.7, §6.10).

2. Collector acquisition

ADR: 0006. Criteria: AC-1–AC-5; miss absorption per 0005 §5.

2.1 Channels — the hard pin

Every collection run acquires exactly two ways:

  • Counters and thermal evidence — solely from smartctl -a -j <device>: data_units_written, data_units_read, percentage_used, available_spare, media_errors, power_on_hours, power_cycles, unsafe_shutdowns, temperature, critical_warning — consumed as-is (smartmontools already trims the strings it copies).
  • Controller identity — solely from sysfs (/sys/class/nvme/<ctrl>/): subnqn, sn, mn, fr, transport.

No other acquisition path exists anywhere in the codebase. There is no fallback: libnvme bindings and the nvme CLI JSON interface are excluded (ADR 0006, Considered options).

2.2 All-or-nothing runs

Any acquisition failure — missing smartctl binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the whole collection run. A partial sample (identity without counters, or counters without identity) is never written: a transient read failure must never push a healthy drive down the degraded-identity path (§4). The miss surfaces through freshness grading (§8.9) and the flat retry cadence (§9.6), never as degraded identity.

2.3 Identity normalization — once, at write time

One collector-side function normalizes every identity field, applied exactly once at write time:

  • strip trailing spaces and newlines;
  • no case folding;
  • empty-after-strip is stored blank.

Padded and unpadded renderings of the same field therefore yield byte-identical stored values — a collector implementation change can never split a drive's own history. A future acquisition-path change must deliver byte-identical normalized identity values, or the change itself forces a controller-segment boundary.

2.4 Segment metadata sourcing

transport comes from the NVMe class sysfs directory. vid/ssvid come from the PCI node (/sys/class/nvme/<ctrl>/device/{vendor,subsystem_vendor}) when present and are stored null otherwise. Both are segment metadata only (§4.2), never key components.

2.5 Prerequisites

make install verifies smartctl is present and fails cleanly otherwise (§10.1). The acquisition path adds no Python dependency and no OS package beyond smartmontools; the dependency lockfile (§10.5) is untouched by this section.


3. Observation store

ADR: 0001 as amended by Define endurance-baseline provenance and validation and Decide controller-segment metadata columns. Criteria: ST-1–ST-12; FL-5.

3.1 Substrate and access

  • One SQLite database in WAL mode at /var/lib/fenris/observations.db. An unprivileged reader querying during a collector write sees a consistent snapshot.
  • The database is root-owned and group-readable through the fenris read group created by packaging; the TUI and status open it read-only. No /run snapshot, no export layer.
  • /var/lib/fenris is created by the installer with root-written group-read permissions; the database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet (§7.6, §10.1).
  • Migration, schema changes, prune, and import are each single transactions — a killed timer run can never leave partial state.

3.2 Entities and column sets

The schema carries exactly six entities:

samples — recent raw samples (14-day retention, §3.4): timestamp (UTC); the normalized controller-identity fields captured at acquisition (§2.3); raw integer data_units_written, data_units_read; percentage_used; available_spare; media_errors; power_on_hours; power_cycles; unsafe_shutdowns; temperature; critical_warning.

hour_observations — one row per UTC hour: the usage-habit split seconds_active, seconds_idle, seconds_powered_off, seconds_unknown (summing to 3600, §5.1); DUW/DUR deltas; temperature min/avg/max; sample count; coverage flag. Classification thresholds belong to the projection model (§5.1), not the store.

day_aggregates — one row per UTC day, the habit-evidence grain: each day row carries, at minimum, the day's activity-split sums, write deltas, and coverage share — the inputs the evidence gates of §6.6 consume — derived monotonically from its hour rows.

monitoring_periods — started_at; ended_at (NULL = open); end_cause enum (user_disabled, migrated, …). Powered-off time stays inside a period; deliberately disabled time does not (§5.2, §8.6).

controller_segments — spans of unchanged controller identity and monotonic counters; write deltas are never computed across a segment boundary. Columns: the identity key (§4.1) and the frozen metadata snapshot of §4.2, plus the segment's span bounds.

endurance_baseline — one active row, replaced on edit (§6.2): the rated-TBW value in bytes (E_rated = entered_TBW × 10¹²); mandatory provenance — source URL, document revision, entry date, model string, nominal capacity; frozen validation facts — detected model, detected capacity bytes, validated_by (machine/user), validated_at.

3.3 Time model

Hours and days are UTC-bounded. Day derivation from hour rows is monotonic; DST-ambiguous 23- or 25-hour days never exist in the store.

3.4 Retention

Raw samples are pruned opportunistically by the collector to 14 days. Hour observations and day aggregates are retained indefinitely.

3.5 Legacy migration

The migration procedure, invoked from the entry points below, is idempotent and interruption-safe:

  1. If the store already carries the legacy-import marker, do nothing.
  2. history.jsonl is the sole authority: import raw samples and derive hour observations and day aggregates from them.
  3. hourly.jsonl is never trusted as input: mismatches against derived data are diffed and logged.
  4. Open one implicit monitoring_periods row at the first legacy sample, closed end_cause = migrated at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts — a sample present means powered on; a DUW delta means writes occurred.
  5. The import is a single transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
  6. Only after commit are legacy files renamed *.migrated — never deleted.
  7. Malformed legacy lines are quarantined with a logged count, never silently dropped.

No synthetic or capacity-derived baseline is ever created for legacy history (§1.4–7). Entry points: the installer's import detection at ./data/history.jsonl (or an explicit path) (§10.1); fenris import <path> for later finds (§8.8); and the collector's first new-version run, which performs this same procedure (ADR 0001 §6).

3.6 Schema versioning

PRAGMA user_version plus ordered migration steps in code, each in its own transaction. The collector refuses to run against an unknown newer version; readers refuse symmetrically with the exact wording of §9.5 and never partially interpret. Reconciliation note: ADR 0004 §6 describes upgrade-time migrations as "governed by a schema_version table" — the operative mechanism is this section's user_version (ADR 0001 §8, criterion ST-12); there is one version authority, not two.

3.7 Nothing derived is stored

Projections are not stored; there is no separate latest-status table and no stored health flag. The freshest sample timestamp is the store's own staleness signal (§8.9). The baseline lives in the database (§6.2); /etc/fenris/ holds only operational configuration (§8.3).


4. Controller identity and segmentation

Decisions: Verify the controller identity that segments observation history, Decide controller-segment metadata columns, Decide how degraded identity affects projection confidence. ADRs: 0001 §3 (as amended), 0002 §§8–9 (as amended). Criteria: ID-1–ID-4, PR-9, PR-15, PR-16.

4.1 Identity key ladder

The controller-segment identity key is the normalized, kernel-exposed subsystem NQN (subnqn), with fallbacks, in order:

  1. kernel-exposed subsystem NQN;
  2. the kernel composite;
  3. model|serial.

fr (firmware revision) is metadata only — it may go stale after a mid-segment firmware update. Identity change and DUW decrease are independent axes (§4.3).

4.2 Frozen metadata snapshot

Each segment freezes, at open, a fully nullable metadata snapshot — immutable thereafter: normalized subnqn, sn, mn, fr, plus vid, ssvid, transport, and the identity_degraded flag. All columns are nullable so incompleteness stays explicit: legacy-imported segments carry mn with NULLs (§4.4); degraded segments carry whatever was observed. These are human diagnostics, never key components. cntlid is excluded — it distinguishes controllers within one subsystem, out of scope for a single-drive monitor.

4.3 Segmentation axes

  • DUW decrease, unchanged identity — a segment boundary within the same drive. Prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.8).
  • Identity-key change — quarantines prior history from projection entirely: it describes a different drive (§6.8).
  • Degraded identity — a segment whose identity key is blank (every rung of the ladder empty). identity_degraded is set at segment open exactly when the key is blank; keys from the kernel-composite or model|serial rungs are not degraded. Blank-key semantics extend identity-change rules verbatim: any visible change of the recorded key — including to or from blank — is a controller-identity change and quarantines; equal blank keys continue the segment, segmented by DUW monotonicity alone. Even a degraded→healthy transition quarantines, so the projection window only ever spans segments sharing one key (§6.8).
  • Ephemeral markers (model "Linux", non-pcie transport) are segment metadata, never confidence facts.

The confidence consequence of degraded identity — capped at Limited with its fixed contributing fact — is §6.7's rule.

4.4 Legacy identity

Legacy history imports under a labeled, model-scoped legacy identity (mn-only segments), so it never blends with the post-redesign identity of the same physical drive.


5. Hour and day derivation

ADRs: 0002 §§4–6; 0001 §3; 0003 §2 (power-on-hours evidence). Criteria: PR-4–PR-6, ST-4, FL-3.

5.1 Hour classification

Each UTC hour is classified by named constants, in this order of evidence:

  • Powered-off — the hour's power-on-hours delta is below 90 % of its wall-clock span.
  • Active — DUW delta ≥ 256 MiB in the hour.
  • Idle — powered on, sampled, below the active threshold.
  • Unknown — everything else: unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable by design), or inconsistent counters.

There is no configuration surface for these thresholds; they are documented constants in one projection module.

5.2 Denominator and disabled time

The projection denominator is wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods — wall-clock outside monitoring periods — are excluded from numerator and denominator. Disabled time is not an hour state.

5.3 Gaps and coverage — never backfill

No absent hour is ever interpolated, estimated, or fabricated. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage. Power-on-hours classification (§5.1) is the only inference admitted. Coverage is the share of wall-clock seconds inside monitoring periods whose classification is known rather than unknown — a first-class displayed fact (§6.10, §7.3).

5.4 Day aggregates

One row per UTC day, derived monotonically from hour rows (§3.2–§3.3) — the grain at which usage-habit evidence is judged (§6.6).


6. Projection and confidence

ADRs: 0002 as amended by Decide how degraded identity affects projection confidence; baseline per Define endurance-baseline provenance and validation. Criteria: PR-1–PR-17, CI-4.

6.1 One projection; baseline precedence

Exactly one usage-adjusted theoretical lifespan is computed, against the endurance baseline chosen by precedence:

  1. Verified override — a rated-TBW override with complete provenance whose applicability to the detected drive was confirmed by machine match or explicit user attestation;
  2. Unverified override — a rated-TBW override knowingly stored with incomplete provenance; always presented as user-supplied, never as verified;
  3. Implied baseline — derived from vendor wear (§6.3), eligible only per §6.3's gate;
  4. otherwise the projection is Unavailable.

Percentage Used is context, never a second projection: it renders as a vendor-wear context line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The legacy PU-slope regression and capacity × 600 synthesis are gone; no synthetic or capacity-derived baseline is ever created (§1.4–7).

6.2 Endurance baseline: provenance and validation

The baseline lives in the observation store's endurance_baseline table (§3.2) and is edited via the CLI (§8.8) — /etc/fenris/ holds no baseline.

  • Mandatory provenance: source URL, document revision, entry date, model string, nominal capacity.
  • One active row, replaced on edit.
  • Verification is derived at read — complete provenance and a drive match (machine or attested) — never a stored boolean.
  • Unverified tier: incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier.
  • Entry-time validation (unprivileged, live sysfs read of the configured device): normalized model containment, with an interactive confirm recorded as validated_by = user; nominal capacity within ± 1 %.
  • Read-time applicability: a model match against the current controller segment (§4). A mismatch is retained — never auto-deleted — and leaves the projection Unavailable.
  • Persistence goes through the polkit-guarded fenris-monitor verb after CLI-side validation (§8.5).

6.3 Arithmetic

rate          = regime DUW delta bytes / in-period wall-clock seconds
projected     = max(E_baseline − W_t, 0) / rate      (rate > 0)
E_rated       = entered_TBW × 10¹² bytes
E_implied     = 100 · W_t / p                         (1 ≤ p ≤ 254)
  • E_rated is exact: rated TBW converts to bytes by × 10¹².
  • E_implied is computed only for 1 ≤ p ≤ 254; Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable. The implied baseline is labeled implied from vendor wear estimate and shown with few significant digits.
  • Implied-baseline eligibility: the implied tier is used only after ≥ 2 Percentage-Used increments within the current controller segment; until then the projection is Unavailable with the fixed phrase "vendor wear estimate too coarse to imply endurance".

6.4 Sustained regime and habit change

The headline rate is the sustained-regime rate: regime DUW bytes ÷ in-period wall-clock seconds. The default regime is the full observation history capped at 90 days.

A habit change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence, is adopted automatically, and is labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.

6.5 Scenario range

The 7-, 28-, and 90-day rates are computed independently of the regime and shown as the scenario range. Only horizons the history actually covers appear — no placeholders. The scenario range is the only spread shown anywhere (§6.9).

6.6 Minimum evidence

Warming up until 14 distinct UTC day aggregates of which at most 2 fall below 50 % coverage. The projection still renders while warming up, labeled with its facts (e.g. "warming up: N of 14 qualifying days"). Every Unavailable condition renders no lifespan number.

6.7 Confidence rule table

Confidence renders as state plus contributing facts, never a percentage. Three states:

  • Unavailable — no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
  • Supported — verified baseline and ≥ 14 qualifying days and coverage ≥ 80 % and fresh (< 48 h) and 7/28/90 rates within a factor of 2 across existing horizons and no single day ≥ 50 % of trailing 28-day bytes and regime ≥ 7 days old and the current controller segment's identity key is not degraded.
  • Limited — every other case with a baseline and a positive rate; the failing facts are shown.

Staleness: a newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.

Degraded identity: a controller segment whose identity key is blank (§4.3) caps confidence at Limited evidence, with the contributing fact "controller identity unavailable — replacement detection relies on write-counter continuity only" rendered in every state. Supported is unreachable while the current segment is degraded. The cap combines idempotently with the staleness drop (both land at Limited).

6.8 Segment-break effects

  • DUW decrease, unchanged identity: prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.6).
  • Controller-identity change — including any to-or-from-blank key change (§4.3): prior history is quarantined from projection entirely.
  • Since even degraded→healthy transitions quarantine, the projection window only ever spans segments sharing one key; no cross-segment propagation rule is needed.

6.9 Zero rate and uncertainty

Zero rate renders "no finite projection from this history" — never infinity, never zero. No statistical confidence interval appears anywhere; the scenario range is the only spread.

6.10 The projection contract

The projection function hands the TUI and status exactly: the confidence state; the contributing facts — including the degraded-identity fact when the current segment's key is blank; the headline remaining time when one exists; the scenario range; the Percentage-Used context line; the disclosure text (§6.11). Recomputed on read, never stored.

6.11 User-facing language

Adopted from the endurance research as fixed by ADR 0002 §12; rendered identically by TUI and status.

Headline wording (equivalent phrasing required):

Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.

Fixed phrases (exact): no finite projection from this history (zero rate); vendor wear estimate too coarse to imply endurance (§6.3); usage habit changed N days ago (§6.4); controller identity unavailable — replacement detection relies on write-counter continuity only (§6.7); observation store unreadable (§9.4); observation store written by a newer Fenris — upgrade Fenris (§9.5); no observations yet with an enable hint (§8.9); configuration error: ⟨reason⟩ (§8.3).

Confidence rendering: state plus contributing facts, in the research's evidence style, e.g.

Supported evidence · verified manufacturer TBW · 42 calendar days · 96 % interval coverage · 6 weekly cycles · recent and 28-day rates agree

Never "82 % confidence" or "95 % accurate".

The six disclosures (verbatim, always available — TUI disclosures view and status):

  1. This is an endurance projection, not a predicted hardware-failure date.
  2. Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated.
  3. Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold.
  4. DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes.
  5. Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
  6. Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.

7. Panes TUI

Decisions: Prototype the TUI information architecture (Variant A adopted), Evaluate Python TUI frameworks (Textual). ADRs: 0003 §§8, 10; 0004 §10. Criteria: TUI-1–TUI-4, CI-1, CI-2, CI-4. The prototype is visual reference only; this section is normative.

7.1 Framework and floor

The TUI is built on Textual. It runs on Python 3.9+, gated at install time (§10.1) — never a runtime crash. The tty-passthrough mechanism below was validated live under Textual on a real terminal (prototype decision).

7.2 Layout — one dense keyboard-first screen

Variant A Panes: everything on one screen, no page navigation. The screen is a grid of four regions:

  1. Headline band — full width, top: the lifespan headline (or its no-projection wording) with its regime line ("if current habits continue · sustained regime: N days at R GB/day"); the confidence state with contributing facts; the scenario range.
  2. Usage-history pane — left, wider column: the write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus their legend; the habit-split bar with active/idle/powered-off/unknown shares.
  3. Drive-health and settings pane — right, narrower column: health facts (model, temperature, spare, media errors, unsafe shutdowns, power-on hours, power cycles, capacity); the vendor-wear context line (Percentage Used · total written of rated — context, not a second projection); a read-only settings view (device selector, cadence with drop-in pointer, raw retention, endurance baseline value with its provenance label). Edits happen via CLI / drop-ins, not in the TUI.
  4. Service strip — full width, bottom: the four separate service facts (§7.3), the monitoring-period line, the action legend.

Exact proportions, glyphs, and borders follow the prototype's validated arrangement as visual reference; the region arrangement, contents, and bindings above are normative.

7.3 Content contracts per region

  • Confidence is evidence: state plus contributing facts, never a percentage (§6.7); the headline band renders the §6.10 contract in full, including the disclosure affordance (d).
  • Four separate service facts, always: boot enablement (enabled/disabled) · runtime activity (timer active/inactive) · last collect outcome (ok/FAILED, age, reason) · freshness (fresh/missed/stale with newest-sample age, §8.9). They are never merged into one "service status".
  • Monitoring-period line: open-since / closed with end cause; deliberate-disable count where nonzero.
  • The scenario range shows only covered horizons (§6.5); the vendor-wear context line carries the >2× disagreement note when it applies (§6.1).

7.4 Keybindings and asymmetry

Production bindings:

Key Action
p Pause — asks for confirmation (y pause · n cancel), stating that paused time is excluded from the usage habit while powered-off time would still count.
r Resume — no confirmation (benign; friction invites raw-systemctl escapes).
c Collect now — synchronous outcome (§8.7), no confirmation.
d Disclosures — the six disclosures of §6.11.
q Quit.

No bare start/stop exists anywhere; no page navigation keys exist (variant switching was prototype-only). Framework defaults apply for focus and scrolling otherwise.

7.5 Privileged actions and tty passthrough

Pause, resume, collect-now, and baseline operations run through fenris-monitor as a terminal-attached subprocess: the TUI suspends, the platform polkit agent prompts on the real terminal, and control returns cleanly with the outcome reflected in the service facts. Where no polkit agent exists the operation fails cleanly with the printed root equivalent (§8.5).

7.6 State rendering obligations

  • From a synthetic observation store, the TUI renders every realizable combination of confidence state × freshness grade × baseline tier exactly as the §6.7 rule table and §8.9 constants dictate — headline number only when the rules allow it, contributing facts always, never a percentage (criterion CI-1).
  • Empty store: "no observations yet" with an enable hint; the first-run prompt is an opt-in that enables the timer and opens the first period in one step (§10.1, dormant install).
  • A configuration error: ⟨reason⟩ fact renders when the device selector is invalid (§8.3); a store fault suppresses everything store-dependent (§9.4); a newer schema renders its fixed phrase (§9.5).
  • Every TUI action has a CLI twin with identical outcomes and wording (§8.8, CI-2).

8. Service lifecycle and sanctioned toggle

ADR: 0003 as amended by Define endurance-baseline provenance and validation. Criteria: LC-1–LC-10, CI-2, CI-3.

8.1 Units

Exactly two system units exist:

  • fenris-collect.timer — WantedBy=timers.target.
  • fenris-collect.service — Type=oneshot, root, ExecStart=/usr/libexec/fenris/fenris-collect; no listener, no UI code.

The TUI and CLI are ordinary unprivileged processes and never units.

8.2 Cadence

Shipped defaults: OnBootSec=2min, OnUnitInactiveSec=5min (measured from run completion; drift accepted because hours are the evidence grain), AccuracySec=30s, Persistent=no (no suspend catch-up — absent hours classify through power-on-hours evidence, §5.1), TimeoutStartSec=90s so a hung device interrogation fails visibly as a bounded failed run retried next interval. Cadence changes are documented drop-ins on the timer unit (systemctl edit + daemon-reload); no interval key exists in configuration.

8.3 Configuration

/etc/fenris/fenris.conf holds exactly one key: the device selector, a stable /dev/disk/by-id/… path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run — there is no reload path. An invalid selector is a bounded failed run (journal + failed unit result, retried next interval); status and the TUI also read the world-readable file directly and surface configuration error: ⟨reason⟩.

8.4 Entry points

Two privileged binaries — /usr/libexec/fenris/fenris-collect (device interrogation and store writes; the unit's ExecStart) and /usr/libexec/fenris/fenris-monitor (fixed operations enable/disable with optional --now, the collect trigger, monitoring-period bookkeeping, and baseline set/baseline clear persistence for the CLI-validated baseline; the only binary the polkit policy authorizes). One unprivileged fenris wrapper (§1.2). Root invokes the helpers directly; unprivileged users go through polkit.

8.5 Sanctioned toggle and polkit

  • Pause = fenris-monitor disable --now; Resume = enable --now. Both perform the systemctl operation and the monitoring-period bookkeeping in one step. The human-facing twins fenris monitor pause / fenris monitor resume map to these and always act immediately; pause asks for confirmation in both TUI and CLI, resume does not (§7.4).
  • Polkit action com.bongbetic.fenris.monitor (auth_admin) covers the toggle and the collect trigger and baseline persistence — authorizing exactly the one binary fenris-monitor.
  • Where no polkit agent exists the operation fails cleanly and prints the root equivalent.
  • This is the only sanctioned control path: a raw systemctl stop/disable never records user_disabled — only the sanctioned path records intent (§8.6).

8.6 Period-row idempotent matrix

Situation Effect on monitoring_periods
First-ever enable Opens a period at the enable moment (hours before the first successful sample are unknown-but-inside — correct when the device errors).
Resume with an open period (a raw systemctl stop intervened) No row changes; the gap remains inside as unknown seconds.
Resume with no open period Opens a new row at the resume moment.
Pause with an open period Closes it user_disabled at the pause moment.
Pause otherwise No-op.
Raw systemctl stop/disable outside the helper An unexplained gap, never user_disabled.

8.7 On-demand collection

fenris sample and the TUI's collect-now route through fenris-monitor → systemctl start fenris-collect.service, which blocks until the oneshot exits; the outcome (freshness line or journal hint) is reported synchronously. No confirmation is required. No code path outside fenris-collect touches the device; the TUI never samples in-process.

8.8 CLI surface

Command Behavior
fenris (no arguments) Opens the TUI (§7).
fenris status Read-only composition of the observation store and allow-listed systemctl show properties: projection facts, enabled/active, last collect outcome, and a journalctl -u fenris-collect.service hint on failure or staleness. Never auto-samples, never prompts.
fenris sample On-demand collection via the helper path (§8.7).
fenris monitor pause / resume The sanctioned toggle (§8.5), pause asking confirmation.
fenris baseline set / clear CLI-side validation (§6.2), then polkit-guarded persistence.
fenris import ⟨path⟩ The idempotent single-transaction legacy import (§3.5).
--device Rejected with a pointer to the configuration file.
start, stop, run Rejected with one-line migration pointers — never aliased (an alias would silently change meaning).

fenris.sh is retired: not shipped, removed from the repository; the README maps its five menu options to their successors. Headless administration has full parity: every TUI action has a CLI twin (pause, resume, collect-now, baseline set/clear, the status fact set) with identical outcomes and wording.

8.9 Freshness grading

Constants defined once, consumed by TUI and CLI alike; the grade derives from the newest sample timestamp, never a stored flag:

  • fresh — newest sample within 2 × cadence + AccuracySec + 60 s (11.5 min at default cadence);
  • missed — between that and 48 h (a contributing fact);
  • stale — ≥ 48 h, matching the §6.7 evidence gate;
  • empty store — "no observations yet" with an enable hint.

9. Failure and recovery

ADR: 0005. Criteria: FL-1–FL-8.

The posture: visible degradation, never fabrication.

9.1 Write-boundary validation

The collector validates every row it would write against the store's domain invariants: hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count. A violating run writes nothing, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Store invariant: everything persisted is well-formed.

9.2 Reader defense

Readers (TUI, status) defensively exclude and count malformed rows as a contributing fact. Under a single trusted writer they should never see one.

9.3 No backfill, ever

Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. Power-on-hours classification (§5.1) is the only inference admitted.

9.4 Store faults — degrade, never recreate over

An unreadable or corrupt database is a store fault: readers surface "observation store unreadable" with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside; the next run starts a fresh store; if the legacy import never completed, the still-present history.jsonl is re-imported (§3.5). No built-in destructive command exists.

9.5 Newer schema — symmetric refusal

The TUI and status detect a user_version newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal (§3.6) and the forward-only upgrade rule (§10.2).

9.6 Repeated collector failures — flat cadence, no escalation

The timer's retry is the recovery path; the freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff, no notification machinery; a persistent failure reads as stale exactly like any other gap.

9.7 Drive-reported anomalies — facts, not alerts

critical_warning, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and status (§7.2); no alerting or notification surface exists. The projection is unaffected: endurance math consumes writes, not warnings.

9.8 Orphaned samples — the collector re-anchors observed fact

When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent — only the sanctioned path records a user_disabled close (§8.5). Coverage semantics stay intact without requiring a re-run of fenris-monitor enable after recovery.


10. Installation

ADR: 0004. Criteria: IN-1–IN-10.

10.1 Install

sudo make install:

  1. Builds a wheel from the checkout and installs it, with pinned dependencies (§10.5), into the dedicated Fenris-owned venv at /opt/fenris; a /usr/local/bin/fenris wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only — after install, nothing references it.
  2. Verifies python3 ≥ 3.9 and smartctl presence, failing cleanly otherwise (never a runtime crash).
  3. Creates /var/lib/fenris with root-written group-read permissions and the fenris read group; the database file is created lazily by the first write (§3.1).
  4. Places units in /etc/systemd/system, helpers in /usr/libexec/fenris, polkit policy under /usr/share/polkit-1/actions/ — recording every placed file in an explicit manifest consumed by upgrade and uninstall (§10.6).
  5. Never enables or starts units. A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle — fenris monitor resume or the first-run TUI prompt — enabling the timer and opening the first period in one step.
  6. Detects ./data/history.jsonl beside the source (or accepts an explicit path), runs the idempotent single-transaction import (§3.5), and reports imported counts.

10.2 Upgrade

sudo make upgrade installs the new wheel into the same venv, syncs units and polkit against the manifest (daemon-reload; restart the timer only if unit contents changed and it is active — safe with Persistent=no), leaves timer state untouched, and never kills an in-flight collection run: a running oneshot finishes on its mapped interpreter; at worst one old-code run completes to the store. It then applies forward-only observation-store schema migrations (§3.6). /var/lib/fenris is never rebuilt.

10.3 Rollback

Best-effort by design: before migrations run, the installer snapshots observations.db to a one-generation observations.db.bak; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade does not exist.

10.4 Uninstall and purge

  • make uninstall first performs the sanctioned disable (fenris-monitor disable --now) so an open monitoring period closes user_disabled — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, keeping /etc/fenris and the observation store. Journal entries age out naturally.
  • make purge additionally removes configuration and store.
  • Reinstall after uninstall resumes from the preserved observation store; only purge erases history.

10.5 Dependencies

Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (make update-deps, committed), never a side effect of installing. The acquisition path adds no Python dependency and no OS package beyond smartmontools (§2.5).

10.6 Placement manifest

Installed artifacts sit only at their fixed locations — units in /etc/systemd/system, helpers in /usr/libexec/fenris, polkit policy under /usr/share/polkit-1/actions/, configuration at /etc/fenris, observation store under /var/lib/fenris, venv at /opt/fenris, wrapper at /usr/local/bin/fenris — and every placed file is recorded in the manifest (criterion IN-10).


Appendix A: Traceability matrix

Built as assembly's first step (assembly decision, recommendation 5). Two-way: every ADR section maps to at least one criterion ID; every criterion cites its ADR or ticket.

A.1 ADR section → criteria

ADR section Criteria
0001 §1 Substrate ST-1, ST-2
0001 §2 Access ST-1, CI-3 (/run)
0001 §3 Entities (incl. #12/#14 amendments) ST-3, ST-5, PR-13, ID-2, CI-3 (no stored projections)
0001 §4 Day boundary ST-4
0001 §5 Retention ST-5
0001 §6 Migration ST-6, ST-7, ST-8, ST-9, ST-10, ST-11
0001 §7 Projection inputs ST-3, PR-14, CI-3 (one-key config)
0001 §8 Versioning ST-12, FL-5
0001 §9 Collector health LC-10
0002 §1 One projection PR-1
0002 §2 Rate/formulas/regime/scenario PR-2, PR-17
0002 §3 Habit change PR-3
0002 §4 Hour classification PR-4
0002 §5 Denominator PR-5
0002 §6 Minimum evidence PR-6
0002 §7 Staleness PR-7
0002 §8 Confidence table (incl. #15 amendment) PR-8, PR-15, CI-1, CI-4
0002 §9 Segment breaks (incl. #15 amendment) PR-9, PR-16
0002 §10 Implied eligibility PR-10
0002 §11 Uncertainty PR-11
0002 §12 Language CI-4
0002 §13 Contract PR-12
0003 §1 Units LC-1, CI-3 (/run)
0003 §2 Cadence LC-2, LC-3
0003 §3 Configuration LC-4, CI-3
0003 §4 Entry points LC-5, IN-10
0003 §5 Sanctioned toggle LC-6, CI-2, CI-3
0003 §6 Period rows LC-7
0003 §7 On-demand collection LC-8, CI-2, CI-3
0003 §8 TUI controls TUI-1, TUI-2, TUI-4
0003 §9 CLI compatibility LC-9, CI-2
0003 §10 Freshness constants LC-10, CI-1, CI-2
0004 §1 Delivery IN-1
0004 §2 Layout and manifest IN-2, IN-10
0004 §3 Privilege IN-3, CI-3
0004 §4 Dormant install IN-3
0004 §5 Legacy import IN-4
0004 §6 Upgrade IN-5
0004 §7 Rollback IN-6
0004 §8 Removal IN-7
0004 §9 Dependencies IN-8, AC-5
0004 §10 Scaffolding and floor IN-9, TUI-3
0005 §1 Malformed observations FL-1, FL-2
0005 §2 Missed observations FL-3, CI-3
0005 §3 Store faults FL-4
0005 §4 Newer schema FL-5
0005 §5 Repeated failures FL-6
0005 §6 Drive-reported anomalies FL-7, CI-3
0005 §7 Orphaned samples FL-8
0006 §1 Pin AC-1
0006 §2 Hard pin, no fallback AC-3
0006 §3 Normalization AC-2
0006 §4 Segment metadata sourcing AC-4
0006 §5 Prerequisites AC-5

A.2 Ticket decisions → criteria

Ticket Criteria
Evaluate Python TUI frameworks TUI-3
Prototype the TUI information architecture TUI-1, TUI-2, TUI-4
Verify the controller identity that segments observation history ID-1, ID-3
Define endurance-baseline provenance and validation PR-13, PR-14, ST-3
Decide controller-segment metadata columns ID-2
Decide how degraded identity affects projection confidence PR-15, PR-16, ID-4
Define cross-cutting acceptance criteria the register itself
Assemble the implementation-ready specification / Write the Fenris redesign specification and close the map this document and this appendix

A.3 Criterion → source

Every criterion carries its citation inline in the register: CI-1–CI-4 (ADR 0002 §§6–8, 0003 §§1–10, 0001 §§2–3/8, 0005 §§2/4–6); ST-1–ST-12 (ADR 0001, with ST-3 amended by tickets #12/#14); LC-1–LC-10 (ADR 0003); PR-1–PR-12 (ADR 0002), PR-13–PR-14 (ticket #12), PR-15–PR-16 (ticket #15, ADR 0002 §§8–9 as amended), PR-17 (ADR 0002 §2); ID-1–ID-4 (tickets #11/#14/#15, ADR 0001 §3 as amended); TUI-1–TUI-4 (tickets #3/#6, ADR 0003 §§8/10, ADR 0004 §10); FL-1–FL-8 (ADR 0005); IN-1–IN-10 (ADR 0004, with IN-10 also citing ADR 0003 §4); AC-1–AC-5 (ADR 0006).

A.4 Assembly result

  • Every ADR 0001–0006 section maps to at least one criterion — the A.1 table is complete; no orphan sections.
  • Every criterion cites its ADR or ticket — verified in the register; no orphan criteria.
  • Decided-but-uncitered gaps found and filled inline in the register during assembly: CI-3 bullet (no /run coordination surface; ADR 0003 §1, ADR 0001 §2), PR-17 (projection arithmetic; ADR 0002 §2), TUI-4 (normative Panes layout and bindings; ticket #3), IN-10 (fixed artifact placement; ADR 0004 §2, ADR 0003 §4).
  • No genuinely undecided behavior remained — no blocking ticket was raised.
  • One reconciliation: ADR 0004 §6's "schema_version table" wording resolves to ADR 0001 §8's PRAGMA user_version as the single version authority (§3.6); criterion ST-12 already fixed the mechanism.