- Change default collection cadence from 5 minutes to 3 minutes (CADENCE_DEFAULT_S=180, systemd OnUnitInactiveSec=3min, runit CADENCE=180) - Update freshness threshold to 450s (2×180 + AccuracySec + 60) - Add _query_live_graph_data(): queries raw samples from the last 3 hours and computes interval byte deltas with actual timestamps - Add LiveActivityGraph widget: vertical bar chart of interval volumes with read/write toggle (w key), arrow key inspection, and click support - Wire live graph into TUI layout (full-width row between daily graph and drive health), refresh cycle, and CSS grid - Replace t theme binding with t today/live binding; theme selection via preferences file - Add w binding for read/write toggle on live graph - Update action legend, help screen, and grid layout for new live-activity row - Add 20 tests covering cadence constants, live query, widget rendering, toggle, and TUI integration - Update all cadence documentation (README, ADR 0003, acceptance criteria LC-2, fenris-redesign constants table, CHANGELOG)
52 KiB
Fenris redesign specification
Status: implementation-ready. Assembled by Write the Fenris redesign specification and close the map, executing the assembly decision Assemble the implementation-ready specification (all seven recommendations accepted) on the Wayfinder map Chart Fenris's persistent TUI monitoring redesign.
Canonical roles. ADRs 0001–0006 are the immutable rationale records — the why. Acceptance criteria are the single register of testable statements — the definition of done. This document normatively restates every operative contract — the what — so an implementer never needs Wayfinder-ticket access: schema column sets, constants, rule tables, unit and CLI definitions, and the Panes TUI layout. Nothing here overrides an ADR or restates a criterion as a criterion.
How to read this document
- Binding language. Must, exactly, and never are normative. Terminology follows the glossary in
CONTEXT.md: observation history, usage-adjusted theoretical lifespan, projection confidence, monitoring period, observation store, hour observation, day aggregate, controller segment, degraded identity, endurance baseline, verified override, unverified override, sustained regime, habit change, scenario range, coverage, collection run, deliberate disable, store fault. - Ordering. Sections follow data flow: system context → collector acquisition → observation store → controller identity & segmentation → hour/day derivation → projection & confidence → Panes TUI → service lifecycle & sanctioned toggle → failure & recovery → installation. Each section opens with its ADR links and criterion-ID block.
- Implementation boundary. This specification plans the redesign; it does not implement it. The complete handoff is this document + the criteria register + ADRs 0001–0006 + the glossary. Given/When/Then test specs are derived by the implementer at implementation time.
Normative constants index
Every constant is defined once, in the section named below; other sections cite, never redefine. All are named constants in code, not configuration.
| Constant | Value | Defined in |
|---|---|---|
| Collection cadence (default) | 3 min (OnUnitInactiveSec) |
§8.2 |
| First-boot delay | 2 min (OnBootSec) |
§8.2 |
| Timer accuracy window | 30 s (AccuracySec) |
§8.2 |
| Collection-run timeout | 90 s (TimeoutStartSec) |
§8.2 |
| Fresh threshold | newest sample within 2 × cadence + AccuracySec + 60 s |
§8.9 |
| Missed → stale boundary | 48 h | §8.9, §6.7 |
| Powered-off hour threshold | power-on-hours delta < 90 % of the hour's wall-clock span | §5.1 |
| Active hour threshold | DUW delta ≥ 256 MiB in the hour | §5.1 |
| Raw-sample retention | 14 days | §3.4 |
| Warming gate | 14 distinct UTC day aggregates, ≤ 2 below 50 % coverage | §6.6 |
| Supported coverage floor | 80 % | §6.7 |
| Horizon agreement | 7/28/90-day rates within a factor of 2 | §6.7 |
| Burst guard | no single day ≥ 50 % of trailing 28-day bytes | §6.7 |
| Young-regime cap | regime < 7 days old → Limited | §6.4 |
| Habit-change trigger | trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean, 3 consecutive days | §6.4 |
| Regime span cap (default) | full observation history capped at 90 days | §6.4 |
| Scenario horizons | 7 / 28 / 90 days | §6.5 |
| Implied-baseline eligibility | ≥ 2 Percentage-Used increments within the current controller segment | §6.3 |
| Rated-TBW conversion | E_rated = entered_TBW × 10¹² bytes |
§6.3 |
| Implied-baseline validity window | 1 ≤ p ≤ 254 | §6.3 |
| Wear-disagreement note | vendor wear vs. observed write rate by more than a factor of 2 | §6.1 |
| Capacity validation tolerance | ± 1 % | §6.2 |
1. System context
ADRs: 0001, 0003, 0006. Criteria: CI-3, LC-1, LC-5, ST-1.
1.1 Scope
Fenris observes one configured NVMe drive's real-world use and translates the observation history into a usage-adjusted theoretical lifespan. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer, persistent compact observation storage, and categorical projection confidence reflecting the length, completeness, and stability of real usage history.
Standing constraints, binding on every section:
- Linux with systemd and polkit only; no other init system is supported.
- Exactly one configured NVMe drive — the device named by
/etc/fenris/fenris.conf(§8.3). - Fully local: no telemetry, no network fetching, no automatic vendor-data retrieval.
- The HTML dashboard and HTTP server are gone; nothing of the daemonization, PID files, or
/runstate survives. - CLI
statusandsampleare retained (§8.8).
1.2 Components and privilege boundaries
| Component | Privilege | Path | Role |
|---|---|---|---|
fenris-collect.service |
root oneshot unit | /usr/libexec/fenris/fenris-collect |
The only code path that interrogates the device and writes the observation store. |
fenris-collect.timer |
system timer | — | Schedules collection runs; WantedBy=timers.target. |
fenris-monitor |
root helper | /usr/libexec/fenris/fenris-monitor |
Fixed privileged operations: enable/disable (optional --now), the collect trigger, monitoring-period bookkeeping, and baseline persistence. The only binary polkit authorizes. |
fenris |
unprivileged | /usr/local/bin/fenris |
Human entry point: no arguments opens the TUI; subcommands are the CLI (§8.8). Never a unit. |
| Observation store | root-written, group-read | /var/lib/fenris/observations.db |
Single SQLite database in WAL mode (§3). The TUI and status open it read-only. |
| Configuration | world-readable | /etc/fenris/fenris.conf |
Exactly one key: the device selector (§8.3). |
The TUI and CLI are ordinary unprivileged processes. Elevation is exclusively polkit, exclusively for fenris-monitor (§8.5). There is no /run/fenris coordination surface and no export layer: systemd serializes collection runs, the observation store holds state, failures go to the journal.
1.3 Data flow
- The timer fires;
fenris-collect.servicerunsfenris-collect. - The collector acquires counters and thermal evidence from
smartctl -a -jand controller identity from sysfs (§2), normalizes identity exactly once (§2.3), and either fails the whole run or writes one complete sample. - The collector derives and validates hour observations and day aggregates, advances controller segmentation and period bookkeeping, prunes raw samples, and commits (§3–§5, §9.1).
- Readers — the TUI and
fenris status— open the store read-only and recompute the projection on every read (§6); nothing derived is ever stored (§3.7).
Control flow is separate: the human drives the TUI/CLI; privileged operations route through fenris-monitor under polkit to systemctl; period rows record intent (only the sanctioned path), while the collector records observed fact (§8.5–§8.6, §9.8).
1.4 Cross-cutting prohibitions
These are operative contracts; each is restated in its home section and gated by the criteria block CI-3:
- No code path outside
fenris-collectinterrogates the device (§2.1, §8.7). - Polkit authorizes exactly one binary,
fenris-monitor, undercom.bongbetic.fenris.monitorauth_admin(§8.5). - No
/run/fenriscoordination surface or export layer exists (§1.2). - No absent hour is ever interpolated, estimated, or fabricated (§5.3, §9.3).
- No alerting, notification, or escalation machinery exists anywhere (§9.6–§9.7).
/etc/fenris/fenris.confholds exactly one key — the device selector (§8.3).- No synthetic or capacity-derived baseline is ever created, including for legacy history (§6.1, §3.5).
- Readers never partially interpret a newer-schema store (§3.6, §9.5).
- Projections are never stored; always recomputed on read (§3.7, §6.10).
2. Collector acquisition
ADR: 0006. Criteria: AC-1–AC-5; miss absorption per 0005 §5.
2.1 Channels — the hard pin
Every collection run acquires exactly two ways:
- Counters and thermal evidence — solely from
smartctl -a -j <device>:data_units_written,data_units_read,percentage_used,available_spare,media_errors,power_on_hours,power_cycles,unsafe_shutdowns, temperature,critical_warning— consumed as-is (smartmontools already trims the strings it copies). - Controller identity — solely from sysfs (
/sys/class/nvme/<ctrl>/):subnqn,sn,mn,fr,transport.
No other acquisition path exists anywhere in the codebase. There is no fallback: libnvme bindings and the nvme CLI JSON interface are excluded (ADR 0006, Considered options).
2.2 All-or-nothing runs
Any acquisition failure — missing smartctl binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the whole collection run. A partial sample (identity without counters, or counters without identity) is never written: a transient read failure must never push a healthy drive down the degraded-identity path (§4). The miss surfaces through freshness grading (§8.9) and the flat retry cadence (§9.6), never as degraded identity.
2.3 Identity normalization — once, at write time
One collector-side function normalizes every identity field, applied exactly once at write time:
- strip trailing spaces and newlines;
- no case folding;
- empty-after-strip is stored blank.
Padded and unpadded renderings of the same field therefore yield byte-identical stored values — a collector implementation change can never split a drive's own history. A future acquisition-path change must deliver byte-identical normalized identity values, or the change itself forces a controller-segment boundary.
2.4 Segment metadata sourcing
transport comes from the NVMe class sysfs directory. vid/ssvid come from the PCI node (/sys/class/nvme/<ctrl>/device/{vendor,subsystem_vendor}) when present and are stored null otherwise. Both are segment metadata only (§4.2), never key components.
2.5 Prerequisites
make install verifies smartctl is present and fails cleanly otherwise (§10.1). The acquisition path adds no Python dependency and no OS package beyond smartmontools; the dependency lockfile (§10.5) is untouched by this section.
3. Observation store
ADR: 0001 as amended by Define endurance-baseline provenance and validation and Decide controller-segment metadata columns. Criteria: ST-1–ST-12; FL-5.
3.1 Substrate and access
- One SQLite database in WAL mode at
/var/lib/fenris/observations.db. An unprivileged reader querying during a collector write sees a consistent snapshot. - The database is root-owned and group-readable through the
fenrisread group created by packaging; the TUI andstatusopen it read-only. No/runsnapshot, no export layer. /var/lib/fenrisis created by the installer with root-written group-read permissions; the database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet (§7.6, §10.1).- Migration, schema changes, prune, and import are each single transactions — a killed timer run can never leave partial state.
3.2 Entities and column sets
The schema carries exactly six entities:
samples — recent raw samples (14-day retention, §3.4): timestamp (UTC); the normalized controller-identity fields captured at acquisition (§2.3); raw integer data_units_written, data_units_read; percentage_used; available_spare; media_errors; power_on_hours; power_cycles; unsafe_shutdowns; temperature; critical_warning.
hour_observations — one row per UTC hour: the usage-habit split seconds_active, seconds_idle, seconds_powered_off, seconds_unknown (summing to 3600, §5.1); DUW/DUR deltas; temperature min/avg/max; sample count; coverage flag. Classification thresholds belong to the projection model (§5.1), not the store.
day_aggregates — one row per UTC day, the habit-evidence grain: each day row carries, at minimum, the day's activity-split sums, write deltas, and coverage share — the inputs the evidence gates of §6.6 consume — derived monotonically from its hour rows.
monitoring_periods — started_at; ended_at (NULL = open); end_cause enum (user_disabled, migrated, …). Powered-off time stays inside a period; deliberately disabled time does not (§5.2, §8.6).
controller_segments — spans of unchanged controller identity and monotonic counters; write deltas are never computed across a segment boundary. Columns: the identity key (§4.1) and the frozen metadata snapshot of §4.2, plus the segment's span bounds.
endurance_baseline — one active row, replaced on edit (§6.2): the rated-TBW value in bytes (E_rated = entered_TBW × 10¹²); mandatory provenance — source URL, document revision, entry date, model string, nominal capacity; frozen validation facts — detected model, detected capacity bytes, validated_by (machine/user), validated_at.
3.3 Time model
Hours and days are UTC-bounded. Day derivation from hour rows is monotonic; DST-ambiguous 23- or 25-hour days never exist in the store.
3.4 Retention
Raw samples are pruned opportunistically by the collector to 14 days. Hour observations and day aggregates are retained indefinitely.
3.5 Legacy migration
The migration procedure, invoked from the entry points below, is idempotent and interruption-safe:
- If the store already carries the legacy-import marker, do nothing.
history.jsonlis the sole authority: import raw samples and derive hour observations and day aggregates from them.hourly.jsonlis never trusted as input: mismatches against derived data are diffed and logged.- Open one implicit
monitoring_periodsrow at the first legacy sample, closedend_cause = migratedat the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts — a sample present means powered on; a DUW delta means writes occurred. - The import is a single transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
- Only after commit are legacy files renamed
*.migrated— never deleted. - Malformed legacy lines are quarantined with a logged count, never silently dropped.
No synthetic or capacity-derived baseline is ever created for legacy history (§1.4–7). Entry points: the installer's import detection at ./data/history.jsonl (or an explicit path) (§10.1); fenris import <path> for later finds (§8.8); and the collector's first new-version run, which performs this same procedure (ADR 0001 §6).
3.6 Schema versioning
PRAGMA user_version plus ordered migration steps in code, each in its own transaction. The collector refuses to run against an unknown newer version; readers refuse symmetrically with the exact wording of §9.5 and never partially interpret. Reconciliation note: ADR 0004 §6 describes upgrade-time migrations as "governed by a schema_version table" — the operative mechanism is this section's user_version (ADR 0001 §8, criterion ST-12); there is one version authority, not two.
3.7 Nothing derived is stored
Projections are not stored; there is no separate latest-status table and no stored health flag. The freshest sample timestamp is the store's own staleness signal (§8.9). The baseline lives in the database (§6.2); /etc/fenris/ holds only operational configuration (§8.3).
4. Controller identity and segmentation
Decisions: Verify the controller identity that segments observation history, Decide controller-segment metadata columns, Decide how degraded identity affects projection confidence. ADRs: 0001 §3 (as amended), 0002 §§8–9 (as amended). Criteria: ID-1–ID-4, PR-9, PR-15, PR-16.
4.1 Identity key ladder
The controller-segment identity key is the normalized, kernel-exposed subsystem NQN (subnqn), with fallbacks, in order:
- kernel-exposed subsystem NQN;
- the kernel composite;
model|serial.
fr (firmware revision) is metadata only — it may go stale after a mid-segment firmware update. Identity change and DUW decrease are independent axes (§4.3).
4.2 Frozen metadata snapshot
Each segment freezes, at open, a fully nullable metadata snapshot — immutable thereafter: normalized subnqn, sn, mn, fr, plus vid, ssvid, transport, and the identity_degraded flag. All columns are nullable so incompleteness stays explicit: legacy-imported segments carry mn with NULLs (§4.4); degraded segments carry whatever was observed. These are human diagnostics, never key components. cntlid is excluded — it distinguishes controllers within one subsystem, out of scope for a single-drive monitor.
4.3 Segmentation axes
- DUW decrease, unchanged identity — a segment boundary within the same drive. Prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.8).
- Identity-key change — quarantines prior history from projection entirely: it describes a different drive (§6.8).
- Degraded identity — a segment whose identity key is blank (every rung of the ladder empty).
identity_degradedis set at segment open exactly when the key is blank; keys from the kernel-composite ormodel|serialrungs are not degraded. Blank-key semantics extend identity-change rules verbatim: any visible change of the recorded key — including to or from blank — is a controller-identity change and quarantines; equal blank keys continue the segment, segmented by DUW monotonicity alone. Even a degraded→healthy transition quarantines, so the projection window only ever spans segments sharing one key (§6.8). - Ephemeral markers (model "Linux", non-pcie transport) are segment metadata, never confidence facts.
The confidence consequence of degraded identity — capped at Limited with its fixed contributing fact — is §6.7's rule.
4.4 Legacy identity
Legacy history imports under a labeled, model-scoped legacy identity (mn-only segments), so it never blends with the post-redesign identity of the same physical drive.
5. Hour and day derivation
ADRs: 0002 §§4–6; 0001 §3; 0003 §2 (power-on-hours evidence). Criteria: PR-4–PR-6, ST-4, FL-3.
5.1 Hour classification
Each UTC hour is classified by named constants, in this order of evidence:
- Powered-off — the hour's power-on-hours delta is below 90 % of its wall-clock span.
- Active — DUW delta ≥ 256 MiB in the hour.
- Idle — powered on, sampled, below the active threshold.
- Unknown — everything else: unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable by design), or inconsistent counters.
There is no configuration surface for these thresholds; they are documented constants in one projection module.
5.2 Denominator and disabled time
The projection denominator is wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods — wall-clock outside monitoring periods — are excluded from numerator and denominator. Disabled time is not an hour state.
5.3 Gaps and coverage — never backfill
No absent hour is ever interpolated, estimated, or fabricated. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage. Power-on-hours classification (§5.1) is the only inference admitted. Coverage is the share of wall-clock seconds inside monitoring periods whose classification is known rather than unknown — a first-class displayed fact (§6.10, §7.3).
5.4 Day aggregates
One row per UTC day, derived monotonically from hour rows (§3.2–§3.3) — the grain at which usage-habit evidence is judged (§6.6).
6. Projection and confidence
ADRs: 0002 as amended by Decide how degraded identity affects projection confidence; baseline per Define endurance-baseline provenance and validation. Criteria: PR-1–PR-17, CI-4.
6.1 One projection; baseline precedence
Exactly one usage-adjusted theoretical lifespan is computed, against the endurance baseline chosen by precedence:
- Verified override — a rated-TBW override with complete provenance whose applicability to the detected drive was confirmed by machine match or explicit user attestation;
- Unverified override — a rated-TBW override knowingly stored with incomplete provenance; always presented as user-supplied, never as verified;
- Implied baseline — derived from vendor wear (§6.3), eligible only per §6.3's gate;
- otherwise the projection is Unavailable.
Percentage Used is context, never a second projection: it renders as a vendor-wear context line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The legacy PU-slope regression and capacity × 600 synthesis are gone; no synthetic or capacity-derived baseline is ever created (§1.4–7).
6.2 Endurance baseline: provenance and validation
The baseline lives in the observation store's endurance_baseline table (§3.2) and is edited via the CLI (§8.8) — /etc/fenris/ holds no baseline.
- Mandatory provenance: source URL, document revision, entry date, model string, nominal capacity.
- One active row, replaced on edit.
- Verification is derived at read — complete provenance and a drive match (machine or attested) — never a stored boolean.
- Unverified tier: incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier.
- Entry-time validation (unprivileged, live sysfs read of the configured device): normalized model containment, with an interactive confirm recorded as
validated_by = user; nominal capacity within ± 1 %. - Read-time applicability: a model match against the current controller segment (§4). A mismatch is retained — never auto-deleted — and leaves the projection Unavailable.
- Persistence goes through the polkit-guarded
fenris-monitorverb after CLI-side validation (§8.5).
6.3 Arithmetic
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
E_ratedis exact: rated TBW converts to bytes by × 10¹².E_impliedis computed only for1 ≤ p ≤ 254; Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable. The implied baseline is labeled implied from vendor wear estimate and shown with few significant digits.- Implied-baseline eligibility: the implied tier is used only after ≥ 2 Percentage-Used increments within the current controller segment; until then the projection is Unavailable with the fixed phrase "vendor wear estimate too coarse to imply endurance".
6.4 Sustained regime and habit change
The headline rate is the sustained-regime rate: regime DUW bytes ÷ in-period wall-clock seconds. The default regime is the full observation history capped at 90 days.
A habit change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence, is adopted automatically, and is labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
6.5 Scenario range
The 7-, 28-, and 90-day rates are computed independently of the regime and shown as the scenario range. Only horizons the history actually covers appear — no placeholders. The scenario range is the only spread shown anywhere (§6.9).
6.6 Minimum evidence
Warming up until 14 distinct UTC day aggregates of which at most 2 fall below 50 % coverage. The projection still renders while warming up, labeled with its facts (e.g. "warming up: N of 14 qualifying days"). Every Unavailable condition renders no lifespan number.
6.7 Confidence rule table
Confidence renders as state plus contributing facts, never a percentage. Three states:
- Unavailable — no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- Supported — verified baseline and ≥ 14 qualifying days and coverage ≥ 80 % and fresh (< 48 h) and 7/28/90 rates within a factor of 2 across existing horizons and no single day ≥ 50 % of trailing 28-day bytes and regime ≥ 7 days old and the current controller segment's identity key is not degraded.
- Limited — every other case with a baseline and a positive rate; the failing facts are shown.
Staleness: a newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
Degraded identity: a controller segment whose identity key is blank (§4.3) caps confidence at Limited evidence, with the contributing fact "controller identity unavailable — replacement detection relies on write-counter continuity only" rendered in every state. Supported is unreachable while the current segment is degraded. The cap combines idempotently with the staleness drop (both land at Limited).
6.8 Segment-break effects
- DUW decrease, unchanged identity: prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.6).
- Controller-identity change — including any to-or-from-blank key change (§4.3): prior history is quarantined from projection entirely.
- Since even degraded→healthy transitions quarantine, the projection window only ever spans segments sharing one key; no cross-segment propagation rule is needed.
6.9 Zero rate and uncertainty
Zero rate renders "no finite projection from this history" — never infinity, never zero. No statistical confidence interval appears anywhere; the scenario range is the only spread.
6.10 The projection contract
The projection function hands the TUI and status exactly: the confidence state; the contributing facts — including the degraded-identity fact when the current segment's key is blank; the headline remaining time when one exists; the scenario range; the Percentage-Used context line; the disclosure text (§6.11). Recomputed on read, never stored.
6.11 User-facing language
Adopted from the endurance research as fixed by ADR 0002 §12; rendered identically by TUI and status.
Headline wording (equivalent phrasing required):
Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
Fixed phrases (exact): no finite projection from this history (zero rate); vendor wear estimate too coarse to imply endurance (§6.3); usage habit changed N days ago (§6.4); controller identity unavailable — replacement detection relies on write-counter continuity only (§6.7); observation store unreadable (§9.4); observation store written by a newer Fenris — upgrade Fenris (§9.5); no observations yet with an enable hint (§8.9); configuration error: ⟨reason⟩ (§8.3).
Confidence rendering: state plus contributing facts, in the research's evidence style, e.g.
Supported evidence · verified manufacturer TBW · 42 calendar days · 96 % interval coverage · 6 weekly cycles · recent and 28-day rates agree
Never "82 % confidence" or "95 % accurate".
The six disclosures (verbatim, always available — TUI disclosures view and status):
- This is an endurance projection, not a predicted hardware-failure date.
- Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated.
- Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold.
- DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes.
- Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
- Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
7. Panes TUI
Decisions: Prototype the TUI information architecture (Variant A adopted), Evaluate Python TUI frameworks (Textual). ADRs: 0003 §§8, 10; 0004 §10. Criteria: TUI-1–TUI-4, CI-1, CI-2, CI-4. The prototype is visual reference only; this section is normative.
7.1 Framework and floor
The TUI is built on Textual. It runs on Python 3.9+, gated at install time (§10.1) — never a runtime crash. The tty-passthrough mechanism below was validated live under Textual on a real terminal (prototype decision).
7.2 Layout — one dense keyboard-first screen
Variant A Panes: everything on one screen, no page navigation. The screen is a grid of four regions:
- Headline band — full width, top: the lifespan headline (or its no-projection wording) with its regime line ("if current habits continue · sustained regime: N days at R GB/day"); the confidence state with contributing facts; the scenario range.
- Usage-history pane — left, wider column: the write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus their legend; the habit-split bar with active/idle/powered-off/unknown shares.
- Drive-health and settings pane — right, narrower column: health facts (model, temperature, spare, media errors, unsafe shutdowns, power-on hours, power cycles, capacity); the vendor-wear context line (Percentage Used · total written of rated — context, not a second projection); a read-only settings view (device selector, cadence with drop-in pointer, raw retention, endurance baseline value with its provenance label). Edits happen via CLI / drop-ins, not in the TUI.
- Service strip — full width, bottom: the four separate service facts (§7.3), the monitoring-period line, the action legend.
Exact proportions, glyphs, and borders follow the prototype's validated arrangement as visual reference; the region arrangement, contents, and bindings above are normative.
7.3 Content contracts per region
- Confidence is evidence: state plus contributing facts, never a percentage (§6.7); the headline band renders the §6.10 contract in full, including the disclosure affordance (
d). - Four separate service facts, always: boot enablement (enabled/disabled) · runtime activity (timer active/inactive) · last collect outcome (ok/FAILED, age, reason) · freshness (fresh/missed/stale with newest-sample age, §8.9). They are never merged into one "service status".
- Monitoring-period line: open-since / closed with end cause; deliberate-disable count where nonzero.
- The scenario range shows only covered horizons (§6.5); the vendor-wear context line carries the >2× disagreement note when it applies (§6.1).
7.4 Keybindings and asymmetry
Production bindings:
| Key | Action |
|---|---|
p |
Pause — asks for confirmation (y pause · n cancel), stating that paused time is excluded from the usage habit while powered-off time would still count. |
r |
Resume — no confirmation (benign; friction invites raw-systemctl escapes). |
c |
Collect now — synchronous outcome (§8.7), no confirmation. |
d |
Disclosures — the six disclosures of §6.11. |
q |
Quit. |
No bare start/stop exists anywhere; no page navigation keys exist (variant switching was prototype-only). Framework defaults apply for focus and scrolling otherwise.
7.5 Privileged actions and tty passthrough
Pause, resume, collect-now, and baseline operations run through fenris-monitor as a terminal-attached subprocess: the TUI suspends, the platform polkit agent prompts on the real terminal, and control returns cleanly with the outcome reflected in the service facts. Where no polkit agent exists the operation fails cleanly with the printed root equivalent (§8.5).
7.6 State rendering obligations
- From a synthetic observation store, the TUI renders every realizable combination of confidence state × freshness grade × baseline tier exactly as the §6.7 rule table and §8.9 constants dictate — headline number only when the rules allow it, contributing facts always, never a percentage (criterion CI-1).
- Empty store: "no observations yet" with an enable hint; the first-run prompt is an opt-in that enables the timer and opens the first period in one step (§10.1, dormant install).
- A
configuration error: ⟨reason⟩fact renders when the device selector is invalid (§8.3); a store fault suppresses everything store-dependent (§9.4); a newer schema renders its fixed phrase (§9.5). - Every TUI action has a CLI twin with identical outcomes and wording (§8.8, CI-2).
8. Service lifecycle and sanctioned toggle
ADR: 0003 as amended by Define endurance-baseline provenance and validation. Criteria: LC-1–LC-10, CI-2, CI-3.
8.1 Units
Exactly two system units exist:
fenris-collect.timer—WantedBy=timers.target.fenris-collect.service—Type=oneshot, root,ExecStart=/usr/libexec/fenris/fenris-collect; no listener, no UI code.
The TUI and CLI are ordinary unprivileged processes and never units.
8.2 Cadence
Shipped defaults: OnBootSec=2min, OnUnitInactiveSec=5min (measured from run completion; drift accepted because hours are the evidence grain), AccuracySec=30s, Persistent=no (no suspend catch-up — absent hours classify through power-on-hours evidence, §5.1), TimeoutStartSec=90s so a hung device interrogation fails visibly as a bounded failed run retried next interval. Cadence changes are documented drop-ins on the timer unit (systemctl edit + daemon-reload); no interval key exists in configuration.
8.3 Configuration
/etc/fenris/fenris.conf holds exactly one key: the device selector, a stable /dev/disk/by-id/… path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run — there is no reload path. An invalid selector is a bounded failed run (journal + failed unit result, retried next interval); status and the TUI also read the world-readable file directly and surface configuration error: ⟨reason⟩.
8.4 Entry points
Two privileged binaries — /usr/libexec/fenris/fenris-collect (device interrogation and store writes; the unit's ExecStart) and /usr/libexec/fenris/fenris-monitor (fixed operations enable/disable with optional --now, the collect trigger, monitoring-period bookkeeping, and baseline set/baseline clear persistence for the CLI-validated baseline; the only binary the polkit policy authorizes). One unprivileged fenris wrapper (§1.2). Root invokes the helpers directly; unprivileged users go through polkit.
8.5 Sanctioned toggle and polkit
- Pause =
fenris-monitor disable --now; Resume =enable --now. Both perform the systemctl operation and the monitoring-period bookkeeping in one step. The human-facing twinsfenris monitor pause/fenris monitor resumemap to these and always act immediately; pause asks for confirmation in both TUI and CLI, resume does not (§7.4). - Polkit action
com.bongbetic.fenris.monitor(auth_admin) covers the toggle and the collect trigger and baseline persistence — authorizing exactly the one binaryfenris-monitor. - Where no polkit agent exists the operation fails cleanly and prints the root equivalent.
- This is the only sanctioned control path: a raw
systemctl stop/disablenever recordsuser_disabled— only the sanctioned path records intent (§8.6).
8.6 Period-row idempotent matrix
| Situation | Effect on monitoring_periods |
|---|---|
| First-ever enable | Opens a period at the enable moment (hours before the first successful sample are unknown-but-inside — correct when the device errors). |
Resume with an open period (a raw systemctl stop intervened) |
No row changes; the gap remains inside as unknown seconds. |
| Resume with no open period | Opens a new row at the resume moment. |
| Pause with an open period | Closes it user_disabled at the pause moment. |
| Pause otherwise | No-op. |
Raw systemctl stop/disable outside the helper |
An unexplained gap, never user_disabled. |
8.7 On-demand collection
fenris sample and the TUI's collect-now route through fenris-monitor → systemctl start fenris-collect.service, which blocks until the oneshot exits; the outcome (freshness line or journal hint) is reported synchronously. No confirmation is required. No code path outside fenris-collect touches the device; the TUI never samples in-process.
8.8 CLI surface
| Command | Behavior |
|---|---|
fenris (no arguments) |
Opens the TUI (§7). |
fenris status |
Read-only composition of the observation store and allow-listed systemctl show properties: projection facts, enabled/active, last collect outcome, and a journalctl -u fenris-collect.service hint on failure or staleness. Never auto-samples, never prompts. |
fenris sample |
On-demand collection via the helper path (§8.7). |
fenris monitor pause / resume |
The sanctioned toggle (§8.5), pause asking confirmation. |
fenris baseline set / clear |
CLI-side validation (§6.2), then polkit-guarded persistence. |
fenris import ⟨path⟩ |
The idempotent single-transaction legacy import (§3.5). |
--device |
Rejected with a pointer to the configuration file. |
start, stop, run |
Rejected with one-line migration pointers — never aliased (an alias would silently change meaning). |
fenris.sh is retired: not shipped, removed from the repository; the README maps its five menu options to their successors. Headless administration has full parity: every TUI action has a CLI twin (pause, resume, collect-now, baseline set/clear, the status fact set) with identical outcomes and wording.
8.9 Freshness grading
Constants defined once, consumed by TUI and CLI alike; the grade derives from the newest sample timestamp, never a stored flag:
- fresh — newest sample within 2 × cadence +
AccuracySec+ 60 s (7.5 min at default cadence); - missed — between that and 48 h (a contributing fact);
- stale — ≥ 48 h, matching the §6.7 evidence gate;
- empty store — "no observations yet" with an enable hint.
9. Failure and recovery
ADR: 0005. Criteria: FL-1–FL-8.
The posture: visible degradation, never fabrication.
9.1 Write-boundary validation
The collector validates every row it would write against the store's domain invariants: hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count. A violating run writes nothing, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Store invariant: everything persisted is well-formed.
9.2 Reader defense
Readers (TUI, status) defensively exclude and count malformed rows as a contributing fact. Under a single trusted writer they should never see one.
9.3 No backfill, ever
Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. Power-on-hours classification (§5.1) is the only inference admitted.
9.4 Store faults — degrade, never recreate over
An unreadable or corrupt database is a store fault: readers surface "observation store unreadable" with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside; the next run starts a fresh store; if the legacy import never completed, the still-present history.jsonl is re-imported (§3.5). No built-in destructive command exists.
9.5 Newer schema — symmetric refusal
The TUI and status detect a user_version newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal (§3.6) and the forward-only upgrade rule (§10.2).
9.6 Repeated collector failures — flat cadence, no escalation
The timer's retry is the recovery path; the freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff, no notification machinery; a persistent failure reads as stale exactly like any other gap.
9.7 Drive-reported anomalies — facts, not alerts
critical_warning, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and status (§7.2); no alerting or notification surface exists. The projection is unaffected: endurance math consumes writes, not warnings.
9.8 Orphaned samples — the collector re-anchors observed fact
When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent — only the sanctioned path records a user_disabled close (§8.5). Coverage semantics stay intact without requiring a re-run of fenris-monitor enable after recovery.
10. Installation
ADR: 0004. Criteria: IN-1–IN-10.
10.1 Install
sudo make install:
- Builds a wheel from the checkout and installs it, with pinned dependencies (§10.5), into the dedicated Fenris-owned venv at
/opt/fenris; a/usr/local/bin/fenriswrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only — after install, nothing references it. - Verifies
python3 ≥ 3.9andsmartctlpresence, failing cleanly otherwise (never a runtime crash). - Creates
/var/lib/fenriswith root-written group-read permissions and thefenrisread group; the database file is created lazily by the first write (§3.1). - Places units in
/etc/systemd/system, helpers in/usr/libexec/fenris, polkit policy under/usr/share/polkit-1/actions/— recording every placed file in an explicit manifest consumed by upgrade and uninstall (§10.6). - Never enables or starts units. A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle —
fenris monitor resumeor the first-run TUI prompt — enabling the timer and opening the first period in one step. - Detects
./data/history.jsonlbeside the source (or accepts an explicit path), runs the idempotent single-transaction import (§3.5), and reports imported counts.
10.2 Upgrade
sudo make upgrade installs the new wheel into the same venv, syncs units and polkit against the manifest (daemon-reload; restart the timer only if unit contents changed and it is active — safe with Persistent=no), leaves timer state untouched, and never kills an in-flight collection run: a running oneshot finishes on its mapped interpreter; at worst one old-code run completes to the store. It then applies forward-only observation-store schema migrations (§3.6). /var/lib/fenris is never rebuilt.
10.3 Rollback
Best-effort by design: before migrations run, the installer snapshots observations.db to a one-generation observations.db.bak; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade does not exist.
10.4 Uninstall and purge
make uninstallfirst performs the sanctioned disable (fenris-monitor disable --now) so an open monitoring period closesuser_disabled— removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, keeping/etc/fenrisand the observation store. Journal entries age out naturally.make purgeadditionally removes configuration and store.- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
10.5 Dependencies
Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (make update-deps, committed), never a side effect of installing. The acquisition path adds no Python dependency and no OS package beyond smartmontools (§2.5).
10.6 Placement manifest
Installed artifacts sit only at their fixed locations — units in /etc/systemd/system, helpers in /usr/libexec/fenris, polkit policy under /usr/share/polkit-1/actions/, configuration at /etc/fenris, observation store under /var/lib/fenris, venv at /opt/fenris, wrapper at /usr/local/bin/fenris — and every placed file is recorded in the manifest (criterion IN-10).
Appendix A: Traceability matrix
Built as assembly's first step (assembly decision, recommendation 5). Two-way: every ADR section maps to at least one criterion ID; every criterion cites its ADR or ticket.
A.1 ADR section → criteria
| ADR section | Criteria |
|---|---|
| 0001 §1 Substrate | ST-1, ST-2 |
| 0001 §2 Access | ST-1, CI-3 (/run) |
| 0001 §3 Entities (incl. #12/#14 amendments) | ST-3, ST-5, PR-13, ID-2, CI-3 (no stored projections) |
| 0001 §4 Day boundary | ST-4 |
| 0001 §5 Retention | ST-5 |
| 0001 §6 Migration | ST-6, ST-7, ST-8, ST-9, ST-10, ST-11 |
| 0001 §7 Projection inputs | ST-3, PR-14, CI-3 (one-key config) |
| 0001 §8 Versioning | ST-12, FL-5 |
| 0001 §9 Collector health | LC-10 |
| 0002 §1 One projection | PR-1 |
| 0002 §2 Rate/formulas/regime/scenario | PR-2, PR-17 |
| 0002 §3 Habit change | PR-3 |
| 0002 §4 Hour classification | PR-4 |
| 0002 §5 Denominator | PR-5 |
| 0002 §6 Minimum evidence | PR-6 |
| 0002 §7 Staleness | PR-7 |
| 0002 §8 Confidence table (incl. #15 amendment) | PR-8, PR-15, CI-1, CI-4 |
| 0002 §9 Segment breaks (incl. #15 amendment) | PR-9, PR-16 |
| 0002 §10 Implied eligibility | PR-10 |
| 0002 §11 Uncertainty | PR-11 |
| 0002 §12 Language | CI-4 |
| 0002 §13 Contract | PR-12 |
| 0003 §1 Units | LC-1, CI-3 (/run) |
| 0003 §2 Cadence | LC-2, LC-3 |
| 0003 §3 Configuration | LC-4, CI-3 |
| 0003 §4 Entry points | LC-5, IN-10 |
| 0003 §5 Sanctioned toggle | LC-6, CI-2, CI-3 |
| 0003 §6 Period rows | LC-7 |
| 0003 §7 On-demand collection | LC-8, CI-2, CI-3 |
| 0003 §8 TUI controls | TUI-1, TUI-2, TUI-4 |
| 0003 §9 CLI compatibility | LC-9, CI-2 |
| 0003 §10 Freshness constants | LC-10, CI-1, CI-2 |
| 0004 §1 Delivery | IN-1 |
| 0004 §2 Layout and manifest | IN-2, IN-10 |
| 0004 §3 Privilege | IN-3, CI-3 |
| 0004 §4 Dormant install | IN-3 |
| 0004 §5 Legacy import | IN-4 |
| 0004 §6 Upgrade | IN-5 |
| 0004 §7 Rollback | IN-6 |
| 0004 §8 Removal | IN-7 |
| 0004 §9 Dependencies | IN-8, AC-5 |
| 0004 §10 Scaffolding and floor | IN-9, TUI-3 |
| 0005 §1 Malformed observations | FL-1, FL-2 |
| 0005 §2 Missed observations | FL-3, CI-3 |
| 0005 §3 Store faults | FL-4 |
| 0005 §4 Newer schema | FL-5 |
| 0005 §5 Repeated failures | FL-6 |
| 0005 §6 Drive-reported anomalies | FL-7, CI-3 |
| 0005 §7 Orphaned samples | FL-8 |
| 0006 §1 Pin | AC-1 |
| 0006 §2 Hard pin, no fallback | AC-3 |
| 0006 §3 Normalization | AC-2 |
| 0006 §4 Segment metadata sourcing | AC-4 |
| 0006 §5 Prerequisites | AC-5 |
A.2 Ticket decisions → criteria
| Ticket | Criteria |
|---|---|
| Evaluate Python TUI frameworks | TUI-3 |
| Prototype the TUI information architecture | TUI-1, TUI-2, TUI-4 |
| Verify the controller identity that segments observation history | ID-1, ID-3 |
| Define endurance-baseline provenance and validation | PR-13, PR-14, ST-3 |
| Decide controller-segment metadata columns | ID-2 |
| Decide how degraded identity affects projection confidence | PR-15, PR-16, ID-4 |
| Define cross-cutting acceptance criteria | the register itself |
| Assemble the implementation-ready specification / Write the Fenris redesign specification and close the map | this document and this appendix |
A.3 Criterion → source
Every criterion carries its citation inline in the register: CI-1–CI-4 (ADR 0002 §§6–8, 0003 §§1–10, 0001 §§2–3/8, 0005 §§2/4–6); ST-1–ST-12 (ADR 0001, with ST-3 amended by tickets #12/#14); LC-1–LC-10 (ADR 0003); PR-1–PR-12 (ADR 0002), PR-13–PR-14 (ticket #12), PR-15–PR-16 (ticket #15, ADR 0002 §§8–9 as amended), PR-17 (ADR 0002 §2); ID-1–ID-4 (tickets #11/#14/#15, ADR 0001 §3 as amended); TUI-1–TUI-4 (tickets #3/#6, ADR 0003 §§8/10, ADR 0004 §10); FL-1–FL-8 (ADR 0005); IN-1–IN-10 (ADR 0004, with IN-10 also citing ADR 0003 §4); AC-1–AC-5 (ADR 0006).
A.4 Assembly result
- Every ADR 0001–0006 section maps to at least one criterion — the A.1 table is complete; no orphan sections.
- Every criterion cites its ADR or ticket — verified in the register; no orphan criteria.
- Decided-but-uncitered gaps found and filled inline in the register during assembly: CI-3 bullet (no
/runcoordination surface; ADR 0003 §1, ADR 0001 §2), PR-17 (projection arithmetic; ADR 0002 §2), TUI-4 (normative Panes layout and bindings; ticket #3), IN-10 (fixed artifact placement; ADR 0004 §2, ADR 0003 §4). - No genuinely undecided behavior remained — no blocking ticket was raised.
- One reconciliation: ADR 0004 §6's "
schema_versiontable" wording resolves to ADR 0001 §8'sPRAGMA user_versionas the single version authority (§3.6); criterion ST-12 already fixed the mechanism.