Compare commits

..
Author SHA1 Message Date
xavierk b2bfbfea4f docs: research NVMe endurance signals 2026-08-31 13:50:07 +05:30
28 changed files with 127 additions and 4009 deletions
-69
View File
@@ -1,69 +0,0 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Endurance baseline**:
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
-42
View File
@@ -1,42 +0,0 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -1,49 +0,0 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -1,31 +0,0 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger and monitoring-period bookkeeping; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
+127
View File
@@ -0,0 +1,127 @@
# NVMe endurance signals and projection constraints
Research for [Verify NVMe endurance signals and projection constraints](https://git.bongbetic.com/xavierk/Fenris/issues/5).
## Decision
Fenris can defensibly project **when a host-write endurance baseline would be consumed if the observed usage habit continues**. It cannot predict SSD failure.
Use lifetime **Data Units Written (DUW)** as the write counter, a model-and-capacity-specific **rated-TBW override** as the preferred baseline, and wall-clock observation history as the rate denominator. Keep **Percentage Used** as a separate manufacturer wear signal; only use it for a coarse, explicitly labeled implied baseline when no rated baseline exists. Treat **Power On Hours** as context, not elapsed calendar time. Express projection confidence as categorical evidence backed by visible facts, never as an accuracy percentage.
This decision removes two unsupported assumptions in the current calculation: inferring endurance as `capacity × 600` when Percentage Used is zero, and presenting `DUW / Percentage Used` as non-estimated endurance ([current Fenris calculation](https://git.bongbetic.com/xavierk/Fenris/src/commit/91db519148a6359f6968b758a76ce30b480aaae0/fenris.py#L261-L270)).
## What the NVMe signals support
| Signal | Standard semantics and precision | Defensible Fenris use |
|---|---|---|
| **Percentage Used** | An unsigned one-byte, vendor-specific estimate based on actual use and the manufacturer's prediction of NVM life. `100` means estimated endurance consumed but may not mean subsystem failure; values may exceed 100, and values above 254 are represented as 255. It is updated once per power-on hour while the controller is not asleep ([NVM Express Base Specification 2.0e, SMART / Health Information](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); [official libnvme field documentation](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)). | Preserve the raw integer. Do not clamp at 100; render 255 as `≥255%`, not an exact value. Do not call 100 a failure point. Zero is too coarse to establish zero wear or infer a baseline. |
| **Data Units Written** | A 128-bit cumulative count of host-written 512-byte data units, excluding metadata, reported in thousands and rounded upward. For the NVM command set, Write logical blocks count; Write Uncorrectable and Write Zeroes do not. Zero means the counter is not reported ([official libnvme structure and semantics](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L15-L23), [field definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)). | Store the raw integer and derive `reported_host_bytes = DUW × 512,000`. Call it **reported host writes**, not physical NAND writes or exact bytes. Treat zero as unsupported/ambiguous unless later positive samples prove support. |
| **Power On Hours** | Integer power-on hours; the controller may omit time powered in a non-operational power state ([official libnvme field definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L159-L162)). | Display as drive context and use changes as a diagnostic. Do not use it as exact active time, idle time, powered-off time, or the denominator of a calendar-life projection. |
The SMART / Health log describes controller-level lifetime information; Fenris should therefore bind a history segment to a stable controller identity and avoid implying filesystem-level or physical-NAND-write precision ([NVM Express Base Specification 2.0e](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
### DUW quantization
Let `q = 512,000 bytes`, raw counter `U_t`, and reported cumulative host writes `W_t = qU_t`. Since each cumulative endpoint is rounded upward, the reported interval delta is:
```text
ΔW = q(U_b - U_a)
```
Its error from endpoint quantization alone is less than one quantum: `|ΔW - actual interval writes| < 512,000 bytes`. Thus a zero hourly delta does not prove no writes below that resolution, and “exact bytes written” is not defensible. This bound follows directly from the standard's upward-rounded cumulative representation ([official libnvme DUW definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)).
## Endurance baseline precedence
Use these sources in order:
1. **Verified rated-TBW override** for the exact manufacturer, model, and capacity, with source URL and document revision.
2. **Unverified manual override**, visibly labeled as user-supplied.
3. **Implied endurance from Percentage Used**, visibly labeled as a coarse heuristic.
4. Otherwise, **projection unavailable**. Never synthesize TBW from capacity alone.
Manufacturer TBW values are model- and capacity-specific: Samsung, for example, rates the 1 TB 990 PRO at 600 TBW and the 2 TB model at 1,200 TBW, and states that its warranty is limited by the stated period or TBW, whichever comes first ([Samsung 990 PRO data sheet, pp. 3–4](https://download.semiconductor.samsung.com/resources/data-sheet/Samsung_NVMe_SSD_990_PRO_Datasheet_Rev.1.0.pdf)). Samsung's warranty treats crossing TBW as a warranty-limit condition, not as a predicted failure event ([Samsung SSD Limited Warranty, sections A–B](https://download.semiconductor.samsung.com/resources/warranty/SAMSUNG_SSD_Limited_Warranty_English_US.pdf)). Fenris must therefore call rated TBW an endurance/warranty baseline rather than physical end of life.
Store an override as bytes plus provenance. If the input is labeled TBW, define it explicitly as decimal terabytes:
```text
E_rated = entered_TBW × 10^12 bytes
R_rated = max(E_rated - W_t, 0)
```
Keep rated-budget consumption and manufacturer Percentage Used separate; disagreement is useful evidence, not a reason to blend them into a fabricated wear percentage.
### Implied endurance constraints
Only for `1 ≤ p ≤ 254`:
```text
E_implied = 100W_t / p
R_implied = max(E_implied - W_t, 0)
```
This assumes the vendor's Percentage Used estimate is proportional to host writes, which NVMe does **not** require: the field is explicitly vendor-specific and based on the manufacturer's life prediction ([NVM Express Base Specification 2.0e](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); [libnvme documentation](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)). Do not compute it for 0 or saturated 255. Label it **implied from vendor wear estimate**, show few significant digits, and do not promote it until multiple wear increments make the estimate less dominated by one-percentage-point quantization. At low values it is intrinsically unstable: changing `p` from 1 to 2 halves the result.
## Usage-adjusted projection
For a selected valid wall-clock history interval from `a` to `b`:
```text
rate = q(U_b - U_a) / elapsed_wall_clock_seconds
projected_seconds = remaining_baseline_bytes / rate (rate > 0)
```
If the rate is zero, report **no finite projection from this history**, not infinity. Required wording should be equivalent to:
> Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
Use wall-clock elapsed time because powered-off and idle periods are part of the observed usage habit, whereas Power On Hours may exclude non-operational powered states ([libnvme Power On Hours definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L159-L162)). Deliberately disabled monitoring must be excluded or marked unknown by lifecycle records; a SMART counter pair can recover aggregate writes across a collector gap but cannot reveal when within that gap the writes occurred.
## History and changing habits
Persist interval observations rather than pretending every delta belongs to a clock-hour bucket:
- Compute deltas only within one controller-identity segment and only when DUW is monotonic. A decrease is a segment boundary or data fault, never a delta to clamp to zero.
- Preserve both endpoints, elapsed wall time, counter delta, and gap/monitoring state. Writes across a gap cannot be assigned exactly to individual hours; any proportional allocation must be labeled estimated.
- Use daily aggregates for habit evidence and retain hourly intervals for display/diagnostics. Hourly samples are time-dependent, so raw sample count is not independent evidence; NIST warns that autocorrelation can invalidate standard `s/√N` uncertainty calculations and other statistical conclusions ([NIST Autocorrelation Plot](https://www.itl.nist.gov/div898/handbook/eda/section3/eda331.htm)).
- Compare descriptive recent, medium, and longer horizons (for example 7, 28, and 90 days) and expose their projection spread as a **scenario range**, not a statistical confidence interval.
- Flag a changing habit when recent and earlier daily-rate windows diverge materially for a sustained period. Prefer the recent sustained regime for the headline projection while retaining the older regime as comparison. This is necessary because a stationary time series has stable mean, variance, and autocorrelation structure; trend, changing variance, and seasonality violate that assumption ([NIST Stationarity](https://www.itl.nist.gov/div898/handbook/pmc/section4/pmc442.htm)).
- Before treating days as interchangeable, account for weekly or other periodic patterns; NIST describes seasonality as regular periodic behavior that must be addressed in a time-series model ([NIST Seasonality](https://www.itl.nist.gov/div898/handbook/pmc/section4/pmc443.htm)).
Exact horizon lengths and change thresholds are product guardrails to validate later, not statistically guaranteed constants.
## Projection confidence
Use four categorical states:
- **Unavailable** — no applicable baseline, unsupported DUW, identity/counter discontinuity, or no positive usable rate.
- **Warming up** — too little history to represent ordinary usage cycles.
- **Limited evidence** — implied or unverified baseline, short/incomplete history, substantial horizon spread, stale observations, or a recent habit change.
- **Supported evidence** — verified baseline, multiple representative usage cycles, good interval coverage, current observations, and stable rates across relevant horizons.
Always show the contributing facts, for example:
> Supported evidence · verified manufacturer TBW · 42 calendar days · 96% interval coverage · 6 weekly cycles · recent and 28-day rates agree
Do not display “82% confidence” or “95% accurate.” NIST defines confidence level through the long-run coverage of an interval procedure, not as the probability that this particular estimate is correct ([NIST Confidence Limits](https://www.itl.nist.gov/div898/handbook/eda/section3/eda352.htm)). A future statistical rate interval would cover rate-estimation uncertainty only; it would not validate the endurance baseline or guarantee that habits remain unchanged.
## Required disclosures
1. This is an endurance projection, not a predicted hardware-failure date.
2. Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated ([NVMe definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)).
3. Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold ([Samsung warranty](https://download.semiconductor.samsung.com/resources/warranty/SAMSUNG_SSD_Limited_Warranty_English_US.pdf)).
4. DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes ([NVMe definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)).
5. Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
6. Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
## Newly surfaced questions
Carry these to the next Wayfinder session rather than expanding this ticket:
- Should rated-budget and manufacturer Percentage Used projections appear side by side when they disagree?
- Which provenance fields are mandatory for a TBW override: URL, revision, model, capacity, region, and entry date?
- What exact warming-up, coverage, horizon, and changing-habit thresholds should the projection-model specification adopt?
- How should powered-off periods, deliberately disabled monitoring, and unexplained gaps be represented separately?
- Which stable controller identity prevents observation history from crossing a drive replacement?
- Should a detected recent regime automatically replace the long-term rate or require acknowledgement?
- Should the TUI expose multi-horizon scenarios only, or also a model-based statistical rate interval?
- How should unsupported DUW, saturated Percentage Used, and counter discontinuities appear in the TUI?
-2
View File
@@ -1,2 +0,0 @@
.venv/
__pycache__/
-46
View File
@@ -1,46 +0,0 @@
# Fenris TUI information-architecture PROTOTYPE (throwaway)
**This is throwaway code answering [ticket #3](https://git.bongbetic.com/xavierk/Fenris/issues/3).** It is not the redesign, reads nothing real, and never ships. Branch: `prototype/tui-information-architecture`.
## Question
What screen hierarchy, navigation, and action model makes Fenris's projection contract (ADR 0002 §13), the four separate service facts (ADR 0003 §8), warming-up, unexplained gaps, and changing habits understandable in a keyboard-first terminal?
## Run (one command)
```sh
./run
```
(creates `.venv` and installs `textual` on first use)
## What to flip through
**Variants (← / →)** — three structurally different answers, not restylings:
| Key | Variant | Idea |
|-----|---------|------|
| A | **Panes** | everything on one dense screen, btop-style; no navigation, panes are zones |
| B | **Pages** | persistent three-fact header (lifespan · confidence · freshness) + pages 1–5 |
| C | **Ledger** | one scrolling document in reading order, headline sentence first |
**States (s)** — same variants, six shapes of the contract:
1. steady · Supported (with one unexplained 3-hour gap)
2. warming up · Limited (11 of 14 days)
3. habit changed · Limited (regime 6 days old, scenario spread visible)
4. stale · Supported→Limited (last collect FAILED, 61 h old)
5. no baseline · Unavailable (wear too coarse to imply endurance)
6. paused · Limited (period closed by deliberate disable)
**Actions** — `p` pause (asks confirmation) · `r` resume (doesn't) · `c` collect now (synchronous outcome). Each suspends the TUI and runs `polkit_stub.py` on the real terminal: this validates the tty passthrough ADR 0003 requires for the polkit prompt. Results land in the tty log (variant B service page; every variant's log is the same list).
## What to react to
- Which variant's hierarchy matches how you think about the drive? (Mixing — "header from B, density of A" — is a valid answer and the point.)
- Are the four service facts separable at a glance?
- Do confidence states + contributing facts read as evidence, not as a percentage?
- Is the pause confirmation the right amount of friction?
- Did the polkit tty stub actually prompt in your terminal? (That's the mechanism check.)
`screenshots/` holds headless captures (`smoke_test.py`) of each variant at 80×24 and 140×40.
-29
View File
@@ -1,29 +0,0 @@
#!/usr/bin/env python3
"""STUB polkit-agent stand-in for the Fenris TUI prototype (throwaway).
Runs attached to the real terminal while the Textual app is suspended, exactly
where the platform polkit agent would prompt for `com.bongbetic.fenris.monitor`.
Accepts any password; the point is validating tty passthrough, not auth.
"""
import getpass
import sys
import time
op = sys.argv[1] if len(sys.argv) > 1 else "unknown"
print("=" * 56)
print(" polkit STUB · com.bongbetic.fenris.monitor")
print(f" operation: {op}")
print(" Authentication required to manage Fenris monitoring")
print("=" * 56)
try:
getpass.getpass(" password (anything works): ")
except (EOFError, KeyboardInterrupt):
print("\n(cancelled — operation not performed)")
sys.exit(1)
time.sleep(0.6) # pretend systemctl + monitoring-period bookkeeping
print(f" fenris-monitor {op}: done")
try:
input(" [press Enter to return to the TUI] ")
except EOFError:
pass
sys.exit(0)
-1
View File
@@ -1 +0,0 @@
textual>=0.60
-14
View File
@@ -1,14 +0,0 @@
#!/usr/bin/env bash
# PROTOTYPE runner — throwaway, see README.md
set -euo pipefail
cd "$(dirname "$0")"
if [ ! -x .venv/bin/python ]; then
if command -v uv >/dev/null 2>&1; then
uv venv -q .venv
uv pip install -q --python .venv/bin/python -r requirements.txt
else
python3 -m venv .venv
.venv/bin/pip -q install -r requirements.txt
fi
fi
exec .venv/bin/python tui_prototype.py
File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 61 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 40 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 56 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 56 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 43 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 42 KiB

-67
View File
@@ -1,67 +0,0 @@
# Headless smoke test for the prototype: drives every variant × state through
# Textual's test pilot, exports SVG screenshots, and asserts contract strings render.
import asyncio, inspect, os, sys
os.environ["FENRIS_PROTOTYPE_NO_TTY"] = "1"
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from tui_prototype import FenrisPrototypeApp, VARIANTS
OUT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "screenshots")
os.makedirs(OUT, exist_ok=True)
async def snap(app, name):
r = app.export_screenshot()
if inspect.isawaitable(r):
r = await r
with open(os.path.join(OUT, name + ".svg"), "w") as f:
f.write(r)
async def main():
checks = []
for size in [(80, 24), (140, 40)]:
app = FenrisPrototypeApp()
async with app.run_test(size=size) as pilot:
await pilot.pause()
for i, (key, _name) in enumerate(VARIANTS):
if i:
await pilot.press("right"); await pilot.pause()
tag = f"v{key}_{size[0]}x{size[1]}"
await snap(app, tag + "_steady")
# rotate states on variant B
await pilot.press("left"); await pilot.pause() # back to A
await pilot.press("right"); await pilot.pause() # B
for st in ["warming", "changed", "stale", "nobaseline", "paused", "steady"]:
await pilot.press("s"); await pilot.pause()
await snap(app, f"vB_{size[0]}x{size[1]}_{st}")
# pages on B
for k in ["1", "2", "3", "4", "5"]:
await pilot.press(k); await pilot.pause()
body = app.query_one("#vB-service").content
checks.append(("service facts", "boot enablement" in str(body) and "freshness" in str(body)))
# pause flow: confirm modal, y, stub skipped headless -> state becomes paused? (skip branch)
await pilot.press("4"); await pilot.pause()
await pilot.press("p"); await pilot.pause()
await pilot.press("n"); await pilot.pause() # cancel, no state change
head = str(app.query_one("#vB-overview").content)
checks.append(("confidence renders", "Projection confidence" in head))
# disclosure modal
await pilot.press("d"); await pilot.pause()
await pilot.press("escape"); await pilot.pause()
steady = FenrisPrototypeApp()
async with steady.run_test(size=(140, 40)) as pilot:
await pilot.pause()
doc = str(steady.query_one("#vC-doc").content)
checks += [
("headline", "Usage-adjusted theoretical lifespan" in doc),
("scenario range", "Scenario range" in doc),
("wear line", "Vendor wear" in doc),
("gap marker", "unexplained gap" in doc),
]
failed = [n for n, ok in checks if not ok]
print("CHECKS:", "all ok" if not failed else f"FAILED: {failed}")
for n, ok in checks:
print(f" {'ok ' if ok else 'FAIL'} {n}")
if failed:
sys.exit(1)
asyncio.run(main())
-573
View File
@@ -1,573 +0,0 @@
# PROTOTYPE (THROWAWAY) — Fenris TUI information-architecture prototype.
# Question (ticket #3): what screen hierarchy, navigation, and action model makes
# the projection contract (ADR 0002 §13) and the service facts (ADR 0003 §8)
# understandable in a keyboard-first terminal?
# Plan: three structurally different variants (A Panes / B Pages / C Ledger),
# switchable live with ←/→, plus a scenario rotator (s) that drives the same
# variants through warming-up / habit-change / stale / paused / no-baseline states.
# Data is synthetic but modeled on the real drive (Micron 2400 512GB, ~101 TB
# written, 50 % used). Nothing here reads or writes the observation store.
from __future__ import annotations
import os
import random
import subprocess
import sys
from pathlib import Path
from textual.app import App, ComposeResult
from textual.binding import Binding
from textual.containers import Horizontal, Vertical, VerticalScroll
from textual.screen import ModalScreen
from textual.widgets import Static
STUB = Path(__file__).with_name("polkit_stub.py")
# ---------------------------------------------------------------- fake data
DRIVE = {
"model": "Micron_2400_MTFDKBA512QFM",
"capacity": "512 GB",
"percentage_used": 50,
"written_tb": 101.1,
"temp": 38,
"spare": 100,
"media_errors": 0,
"power_on_hours": 11126,
"power_cycles": 5188,
"unsafe_shutdowns": 142,
}
BASELINE = {"tb": 220.0, "source": "Micron 2400 datasheet (PROTOTYPE placeholder provenance)"}
def _years(gb_per_day: float) -> str:
days = (BASELINE["tb"] - DRIVE["written_tb"]) * 1000.0 / gb_per_day
if days >= 365.25:
return f"~{days / 365.25:.1f} years"
return f"~{days:.0f} days"
def _days_hist(n: int, rate: float, seed: int) -> list[float]:
rng = random.Random(seed)
return [max(2.0, rate + rng.gauss(0, rate * 0.18)) for _ in range(n)]
def _sparkline(vals: list[float], width: int = 40) -> str:
if not vals:
return ""
mx = max(vals) or 1.0
blocks = " ▁▂▃▄▅▆▇█"
step = max(1, len(vals) // width or 1)
picked = vals[-width * step:][::step][-width:]
return "".join(blocks[min(len(blocks) - 1, int(v / mx * (len(blocks) - 1)) + (1 if v > 0 else 0))] for v in picked)
def _habit_bar(active: float, idle: float, off: float, unknown: float, width: int = 44) -> str:
total = active + idle + off + unknown or 1.0
segs = [("active", active, "green"), ("idle", idle, "yellow"), ("off", off, "cyan"), ("?", unknown, "magenta")]
out = []
for label, v, color in segs:
n = max(1 if v else 0, round(v / total * width))
out.append((f"[{color}]{label[0] * n}[/{color}]"))
legend = f" active {active / total:.0%} · idle {idle / total:.0%} · powered-off {off / total:.0%} · unknown {unknown / total:.0%}"
return "".join(out) + "\n " + legend
def build_scenarios() -> list[dict]:
"""Six states of the ADR-0002 §13 contract + ADR-0003 service facts."""
hist_steady = _days_hist(34, 55, seed=7)
hist_warm = _days_hist(11, 62, seed=11)
hist_changed = _days_hist(35, 70, seed=3) + _days_hist(6, 130, seed=4)
hist_stale = _days_hist(34, 55, seed=7)
hist_nobase = _days_hist(26, 48, seed=5)
hist_paused = _days_hist(28, 51, seed=9)
def svc(enabled, timer, outcome, freshness, period):
return {"enabled": enabled, "timer": timer, "outcome": outcome, "freshness": freshness, "period": period}
return [
{
"key": "steady",
"name": "steady · Supported",
"days": 34,
"history": hist_steady,
"rate": 55,
"headline": _years(55),
"confidence": "Supported",
"facts": [
"34 qualifying days (≥ 14), coverage 92 %",
"7- and 28-day rates within a factor of 2",
"no single day ≥ 50 % of trailing 28-day writes",
"1 day with 3 unknown hours — unexplained gap inside the period",
],
"horizons": [("last 7 days", _years(48)), ("last 28 days", _years(57))],
"habit": (0.34, 0.52, 0.10, 0.04),
"gap_days": {-9},
"habit_change": None,
"service": svc(True, True, "ok · 3 min ago (5 min cadence)", "fresh · newest sample 3 min old", "open since Aug 3 · deliberate disables: 0"),
},
{
"key": "warming",
"name": "warming up · Limited",
"days": 11,
"history": hist_warm,
"rate": 62,
"headline": _years(62),
"confidence": "Limited",
"facts": [
"warming up: 11 of 14 qualifying days",
"coverage 84 %",
],
"horizons": [("last 7 days", _years(66))],
"habit": (0.38, 0.46, 0.12, 0.04),
"gap_days": set(),
"habit_change": None,
"service": svc(True, True, "ok · 2 min ago", "fresh · newest sample 2 min old", "open since Aug 20"),
},
{
"key": "changed",
"name": "habit changed · Limited",
"days": 41,
"history": hist_changed,
"rate": 130,
"headline": _years(130),
"confidence": "Limited",
"facts": [
"usage habit changed 6 days ago — new regime adopted",
"regime 6 days old (young — Limited evidence)",
"coverage 88 %",
"vendor wear line disagrees ×2.1 with observed write rate",
],
"horizons": [("last 7 days", _years(128)), ("last 28 days", _years(71))],
"habit": (0.47, 0.41, 0.08, 0.04),
"gap_days": set(),
"habit_change": 6,
"service": svc(True, True, "ok · 4 min ago", "fresh · newest sample 4 min old", "open since Jul 15 · habit change noted Aug 26"),
},
{
"key": "stale",
"name": "stale · Supported→Limited",
"days": 34,
"history": hist_stale,
"rate": 55,
"headline": _years(55),
"confidence": "Limited",
"facts": [
"newest evidence 61 h old (missed — older than 48 h)",
"34 qualifying days, coverage 92 %",
],
"horizons": [("last 7 days", _years(48)), ("last 28 days", _years(57))],
"habit": (0.34, 0.52, 0.10, 0.04),
"gap_days": {-9},
"habit_change": None,
"service": svc(True, True, "FAILED · exit 1 · 61 h ago (device busy)", "missed · newest sample 61 h old", "open since Aug 3 · gap is unknown time inside the period"),
},
{
"key": "nobaseline",
"name": "no baseline · Unavailable",
"days": 26,
"history": hist_nobase,
"rate": 48,
"headline": None,
"confidence": "Unavailable",
"facts": [
"no verified rated-TBW override on record",
"vendor wear estimate too coarse to imply endurance (1 of ≥ 2 Percentage Used increments)",
],
"horizons": [("last 7 days", "48 GB/day (no baseline to project)"), ("last 28 days", "44 GB/day (no baseline to project)")],
"habit": (0.31, 0.55, 0.10, 0.04),
"gap_days": set(),
"habit_change": None,
"service": svc(True, True, "ok · 3 min ago", "fresh · newest sample 3 min old", "open since Aug 8"),
},
{
"key": "paused",
"name": "paused · Limited",
"days": 28,
"history": hist_paused,
"rate": 51,
"headline": _years(51),
"confidence": "Limited",
"facts": [
"no open monitoring period — paused 2 days ago",
"paused time is excluded from the usage habit by your choice",
],
"horizons": [("last 7 days (pre-pause)", _years(53))],
"habit": (0.33, 0.51, 0.12, 0.04),
"gap_days": set(),
"habit_change": None,
"service": svc(False, False, "ok · 2 d ago (period closed by pause)", "stale · monitoring paused 2 days ago", "closed 2 days ago · end cause: deliberate disable"),
},
]
# ------------------------------------------------------------- renderers
def headline_block(sc: dict) -> str:
if sc["headline"]:
return (
f"[bold]Usage-adjusted theoretical lifespan: [white]{sc['headline']}[/white][/bold]\n"
f" if current habits continue · sustained regime: {sc['days'] if not sc['habit_change'] else sc['habit_change']} days at {sc['rate']} GB/day"
)
return "[bold]Usage-adjusted theoretical lifespan: [red]no projection from this history yet[/red][/bold]\n " + "\n ".join(sc["facts"][:2])
def confidence_block(sc: dict) -> str:
color = {"Supported": "green", "Limited": "yellow", "Unavailable": "red"}[sc["confidence"]]
lines = [f"[bold]Projection confidence: [{color}]{sc['confidence']}[/{color}][/bold]"]
lines += [f" · {f}" for f in sc["facts"]]
return "\n".join(lines)
def horizon_block(sc: dict) -> str:
rows = [f" {label:<28} → [cyan]{value}[/cyan]" for label, value in sc["horizons"]]
return "[bold]Scenario range[/bold] (same endurance, other horizons)\n" + "\n".join(rows)
def wear_line(sc: dict) -> str:
return f"Vendor wear: {DRIVE['percentage_used']} % used · {DRIVE['written_tb']} TB of {BASELINE['tb']:.0f} TB rated (context, not a second projection)"
def disclosure_lines() -> list[str]:
return [
"Rated endurance is a vendor guarantee boundary, not a predicted failure date.",
"Powered-off time counts toward the projection while monitoring is enabled; deliberately paused time does not.",
"The scenario range is a spread of horizons, not a statistical interval.",
f"Baseline provenance: {BASELINE['source']}.",
]
def history_block(sc: dict, width: int = 60) -> str:
vals = sc["history"]
spark = _sparkline(vals, width)
marks = [" "] * len(spark)
if sc["habit_change"]:
idx = len(spark) - max(1, round(sc["habit_change"] / max(1, len(vals) // width or 1)))
if 0 <= idx < len(marks):
marks[idx] = "▲"
for g in sc["gap_days"]:
idx = len(spark) + g - 1
if 0 <= idx < len(marks) and marks[idx] == " ":
marks[idx] = "?"
head = f"[bold]Usage history[/bold] · {sc['days']} days · {min(vals):.0f}–{max(vals):.0f} GB/day"
bar = f"[green]{spark}[/green]"
markline = "".join(marks)
a, i, o, u = sc["habit"]
return head + "\n " + bar + "\n " + markline + " ▲ habit change · ? unexplained gap\n " + _habit_bar(a, i, o, u)
def health_block() -> str:
d = DRIVE
rows = [
f"[bold]Drive health[/bold] · {d['model']}",
f" temperature {d['temp']} °C · spare {d['spare']} %",
f" media errors {d['media_errors']} · unsafe shutdowns {d['unsafe_shutdowns']}",
f" power-on {d['power_on_hours']:,} h · {d['power_cycles']:,} cycles · {d['capacity']}",
]
return "\n".join(rows)
def service_block(sc: dict) -> str:
s = sc["service"]
en = "[green]enabled[/green]" if s["enabled"] else "[red]disabled[/red]"
tm = "[green]timer active[/green]" if s["timer"] else "[red]timer inactive[/red]"
return "\n".join([
"[bold]Service[/bold] (four separate facts)",
f" boot enablement: {en}",
f" runtime activity: {tm}",
f" last collect outcome: {s['outcome']}",
f" freshness: {s['freshness']}",
f" monitoring period: {s['period']}",
])
def settings_block() -> str:
return "\n".join([
"[bold]Settings[/bold] (read view · edit via CLI / drop-ins)",
" device: /dev/disk/by-id/nvme-Micron_2400_MTFDKBA512QFM_2341ABCD",
" cadence: every 5 min (systemd drop-in to change) · raw retention 14 d",
f" endurance baseline: {BASELINE['tb']:.0f} TB rated — {BASELINE['source']}",
])
def actions_legend(paused: bool) -> str:
resume = "[bold green]r resume[/bold green]" if paused else "r resume"
pause = "[bold yellow]p pause[/bold yellow]" if not paused else "p pause"
return f"{pause} (asks) · {resume} · c collect now · s state · ←/→ variant · d disclosures · q quit"
VARIANTS = [
("A", "Panes — one dense screen"),
("B", "Pages — persistent header + tabbed body"),
("C", "Ledger — scrolling narrative document"),
]
# ----------------------------------------------------------------- screens
class ConfirmPause(ModalScreen[bool]):
"""Pause asks for confirmation (ADR 0003 §8)."""
BINDINGS = [
Binding("y", "yes", "Pause"),
Binding("n", "no", "Cancel"),
Binding("escape", "no", "Cancel", show=False),
]
def compose(self) -> ComposeResult:
yield Static(
"[bold]Pause monitoring?[/bold]\n\n"
"This closes the current monitoring period.\n"
"Paused time is [bold]excluded[/bold] from your usage habit\n"
"(powered-off time would still count).\n\n"
"[dim]y pause · n cancel[/dim]",
id="confirm-text",
)
def action_yes(self) -> None:
self.dismiss(True)
def action_no(self) -> None:
self.dismiss(False)
class Disclosures(ModalScreen):
BINDINGS = [Binding("escape", "close", "Close"), Binding("d", "close", "Close")]
def compose(self) -> ComposeResult:
yield VerticalScroll(Static("\n".join(["[bold]Disclosures[/bold]"] + [f" · {d}" for d in disclosure_lines()]) + "\n\n[dim]esc to close[/dim]", id="disc-text"), id="disc-wrap")
def action_close(self) -> None:
self.dismiss()
# -------------------------------------------------------------------- app
class FenrisPrototypeApp(App):
TITLE = "Fenris — PROTOTYPE (throwaway)"
SUB_TITLE = "TUI information architecture · ticket #3"
BINDINGS = [
Binding("left", "prev_variant", "‹ variant", show=False),
Binding("right", "next_variant", "variant ›", show=False),
Binding("s", "cycle_state", "state", show=False),
Binding("p", "pause", "pause", show=False),
Binding("r", "resume", "resume", show=False),
Binding("c", "collect", "collect now", show=False),
Binding("d", "disclose", "disclosures", show=False),
Binding("q", "quit", "quit", show=False),
Binding("1", "page('overview')", "overview", show=False),
Binding("2", "page('history')", "history", show=False),
Binding("3", "page('drive')", "drive", show=False),
Binding("4", "page('service')", "service", show=False),
Binding("5", "page('settings')", "settings", show=False),
Binding("j", "scroll_down", "scroll down", show=False),
Binding("k", "scroll_up", "scroll up", show=False),
]
CSS = """
#vhost { height: 1fr; }
#vA { layout: grid; grid-size: 2 3; grid-columns: 3fr 2fr; grid-rows: 8 1fr 7; height: 1fr; }
#vA-head, #vA-bottom { column-span: 2; }
#vA-history { overflow-y: auto; }
.pane { border: round #555555; padding: 0 1; }
#vB-head { height: 6; border-bottom: thick #555555; padding: 0 1; }
#vB-body { height: 1fr; padding: 0 1; }
#vB-legend { height: 3; }
.page { height: 1fr; padding: 0 1; }
#vC { height: 1fr; padding: 0 2; }
#switchbar { height: 3; dock: bottom; background: $boost; }
#sb-variant { width: 1fr; text-style: reverse; }
#sb-state { width: 1fr; }
#sb-actions { width: 2fr; }
#confirm-text { padding: 1 2; }
#disc-wrap { padding: 1 2; height: auto; max-height: 80%; }
"""
def __init__(self) -> None:
super().__init__()
self.scenarios = build_scenarios()
self.scenario_idx = 0
self.variant_idx = 0
self.b_page = "overview"
self.tty_log: list[str] = []
# ---- composition
def compose(self) -> ComposeResult:
with Vertical(id="vhost"):
with Vertical(id="vA"):
yield Static("", id="vA-head", classes="pane")
yield Static("", id="vA-history", classes="pane")
yield Static("", id="vA-health", classes="pane")
yield Static("", id="vA-bottom", classes="pane")
with Vertical(id="vB"):
yield Static("", id="vB-head")
with Vertical(id="vB-body"):
yield Static("", id="vB-overview", classes="page")
yield Static("", id="vB-history", classes="page")
yield Static("", id="vB-drive", classes="page")
yield Static("", id="vB-service", classes="page")
yield Static("", id="vB-settings", classes="page")
yield Static("", id="vB-legend")
with VerticalScroll(id="vC"):
yield Static("", id="vC-doc")
with Horizontal(id="switchbar"):
yield Static("", id="sb-variant")
yield Static("", id="sb-state")
yield Static("", id="sb-actions")
def on_mount(self) -> None:
for vid in ("vA-head", "vA-history", "vA-health", "vA-bottom"):
w = self.query_one(f"#{vid}", Static)
w.border_title = {"vA-head": "headline", "vA-history": "usage history", "vA-health": "drive", "vA-bottom": "service + actions"}[vid]
self.render_all()
# ---- helpers
@property
def sc(self) -> dict:
return self.scenarios[self.scenario_idx]
def w(self, vid: str) -> Static:
return self.query_one(f"#{vid}", Static)
def render_all(self) -> None:
sc = self.sc
paused = not sc["service"]["enabled"]
key, name = VARIANTS[self.variant_idx]
# variant A: everything on one dense screen
self.w("vA-head").update(headline_block(sc) + "\n" + confidence_block(sc))
self.w("vA-history").update(history_block(sc) + "\n" + horizon_block(sc))
self.w("vA-health").update(health_block() + "\n\n" + settings_block())
self.w("vA-bottom").update(service_block(sc) + "\n " + actions_legend(paused))
# variant B: persistent header, tabbed pages
head = "\n".join([
headline_block(sc),
confidence_block(sc).split("\n")[0] + f" · {sc['confidence']}",
f"freshness: {sc['service']['freshness']} · wear: {DRIVE['percentage_used']} %",
])
self.w("vB-head").update(head)
self.w("vB-overview").update(confidence_block(sc) + "\n\n" + horizon_block(sc) + "\n\n" + wear_line(sc))
self.w("vB-history").update(history_block(sc, 70))
self.w("vB-drive").update(health_block() + "\n\n" + wear_line(sc))
self.w("vB-service").update(service_block(sc) + "\n\n " + actions_legend(paused) + "\n\n tty log:\n" + ("\n".join(self.tty_log) if self.tty_log else " (no privileged action taken yet)"))
self.w("vB-settings").update(settings_block() + "\n\n" + "\n".join(" · " + d for d in disclosure_lines()))
self.w("vB-legend").update(f"pages: 1 overview · 2 history · 3 drive · 4 service · 5 settings [now: {self.b_page}]")
for p in ("overview", "history", "drive", "service", "settings"):
self.w(f"vB-{p}").styles.display = "block" if p == self.b_page else "none"
# variant C: one scrolling document in reading order
doc = "\n\n".join([
"[dim]═" * 70 + "[/dim]",
headline_block(sc),
confidence_block(sc),
horizon_block(sc),
wear_line(sc),
history_block(sc, 70),
health_block(),
service_block(sc),
settings_block(),
"[bold]Disclosures[/bold]\n" + "\n".join(" · " + d for d in disclosure_lines()),
"[dim]═" * 70 + "[/dim]",
])
self.w("vC-doc").update(doc)
# variant visibility + switcher
for i, vid in enumerate(("vA", "vB", "vC")):
self.query_one(f"#{vid}").styles.display = "block" if i == self.variant_idx else "none"
self.w("sb-variant").update(f" ← {key} · {name} → ")
self.w("sb-state").update(f" state [{self.scenario_idx + 1}/{len(self.scenarios)}]: {sc['name']} (s to cycle) ")
self.w("sb-actions").update(" " + actions_legend(paused))
# ---- actions
def action_prev_variant(self) -> None:
self.variant_idx = (self.variant_idx - 1) % len(VARIANTS)
self.render_all()
def action_next_variant(self) -> None:
self.variant_idx = (self.variant_idx + 1) % len(VARIANTS)
self.render_all()
def action_cycle_state(self) -> None:
self.scenario_idx = (self.scenario_idx + 1) % len(self.scenarios)
self.render_all()
def action_page(self, page: str) -> None:
if VARIANTS[self.variant_idx][0] != "B":
return
self.b_page = page
self.render_all()
def action_scroll_down(self) -> None:
if VARIANTS[self.variant_idx][0] == "C":
self.query_one("#vC").scroll_down(animated=False)
def action_scroll_up(self) -> None:
if VARIANTS[self.variant_idx][0] == "C":
self.query_one("#vC").scroll_up(animated=False)
def action_disclose(self) -> None:
self.push_screen(Disclosures())
def action_pause(self) -> None:
self.push_screen(ConfirmPause(), callback=self._pause_confirmed)
def _pause_confirmed(self, confirmed: bool) -> None:
if not confirmed:
self.tty_log.append("pause: cancelled at confirmation")
self.render_all()
return
result = self._run_tty_stub("pause (disable --now)")
if "OK" in result:
self.scenarios[self.scenario_idx]["service"] = {
"enabled": False, "timer": False,
"outcome": "ok · period closed by pause",
"freshness": "paused · no collection while disabled",
"period": "closed just now · end cause: deliberate disable",
}
self.tty_log.append(result)
self.render_all()
def action_resume(self) -> None:
result = self._run_tty_stub("resume (enable --now)") # no confirmation (ADR 0003 §8)
if "OK" in result or "skipped" in result:
self.scenarios[self.scenario_idx]["service"] = {
"enabled": True, "timer": True,
"outcome": "ok · resumed just now",
"freshness": "fresh · collection resuming",
"period": "open just now",
}
self.tty_log.append(result)
self.render_all()
def action_collect(self) -> None:
result = self._run_tty_stub("collect now", blocking=True) # synchronous outcome (ADR 0003 §7)
self.tty_log.append(result)
self.render_all()
# ---- tty passthrough validation (the point of the stub)
def _run_tty_stub(self, op: str, blocking: bool = False) -> str:
if os.environ.get("FENRIS_PROTOTYPE_NO_TTY") or not sys.stdin.isatty():
return f"{op}: tty stub SKIPPED (headless run — mechanism not exercised)"
try:
with self.suspend():
proc = subprocess.run([sys.executable, str(STUB), op])
verdict = "OK — suspend + terminal passthrough works" if proc.returncode == 0 else f"FAILED (exit {proc.returncode})"
return f"{op}: {verdict}"
except Exception as exc: # SuspendNotSupported and friends
return f"{op}: suspend failed: {type(exc).__name__} — this terminal may not support passthrough"
if __name__ == "__main__":
app = FenrisPrototypeApp()
print("[fenris prototype] throwaway UI for ticket #3 — variants switch with ←/→, states with s\n")
app.run()