Compare commits

..
Author SHA1 Message Date
xavierk a0e4690703 prototype(tui): throwaway TUI information-architecture prototype (ticket #3)
Three structurally different variants (Panes / Pages / Ledger), switchable
live, plus a six-state scenario rotator (steady, warming up, habit changed,
stale, no baseline, paused) driving the ADR-0002 section-13 contract and the
four separate ADR-0003 section-8 service facts. Pause/resume/collect suspend
the TUI and run a polkit stand-in on the real terminal to validate tty
passthrough. Headless smoke test + SVG screenshots included.
2026-08-31 16:30:28 +05:30
xavierk 26a6703152 docs(adr): 0003 service lifecycle — timer-driven collector, sanctioned control helper; glossary terms for collection run, deliberate disable 2026-08-31 16:14:52 +05:30
xavierk 43467f8957 docs(adr): 0002 projection model — sustained regime, categorical confidence; glossary terms for regime, habit change, scenario range, coverage 2026-08-31 15:38:41 +05:30
xavierk a63da74e44 docs(adr): 0001 observation store in SQLite; glossary terms for store entities 2026-08-31 14:48:14 +05:30
28 changed files with 4009 additions and 259 deletions
+69
View File
@@ -0,0 +1,69 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Endurance baseline**:
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
+42
View File
@@ -0,0 +1,42 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -0,0 +1,49 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -0,0 +1,31 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger and monitoring-period bookkeeping; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
@@ -1,259 +0,0 @@
# Fenris: systemd lifecycle and privilege constraints
**Ticket:** “Verify systemd lifecycle and privilege constraints”
**Status:** Research and planning only; no product implementation is included
**Research date:** 2026-08-31
## Executive recommendation
Run collection as a **system timer plus a short-lived system service**, not as a user service and not as the TUI's child process. Enable `fenris-collect.timer` at installation so PID 1 schedules a collection shortly after every boot and thereafter at the configured interval. Keep the interactive TUI an ordinary, on-demand, unprivileged process.
The collection service should invoke an absolute, administrator-owned `smartctl` binary directly—never `sudo`—and should have no listener or TUI code. It should read root-owned configuration from `/etc/fenris/`, use `/var/lib/fenris/` for durable history, use `/run/fenris/` only for ephemeral status/locking, and log to the journal. `StateDirectory=` and `RuntimeDirectory=` create and lifecycle-manage those standard locations and add the mount dependencies needed to reach them; state directories persist after service stop, while runtime directories normally do not. [systemd.exec(5), directory options](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=)
The TUI should:
1. inspect a deliberately small, non-secret status surface without elevation;
2. display **runtime state**, **boot enablement**, **last sample outcome**, and **freshness** separately;
3. offer only explicit “Enable collection at boot” and “Disable collection at boot” actions, with confirmation; and
4. ask systemd to make that change, allowing the platform's normal polkit authentication to occur.
Do **not** install a permissive polkit rule granting `org.freedesktop.systemd1.manage-unit-files` to a Fenris group. systemd uses that action for enable/disable/mask/preset and related unit-file operations generally, and its current unit-file authorization check supplies no unit detail with which a rule could safely limit authorization to Fenris. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security); [systemd `dbus-util.c`, pinned source: `manage-unit-files` has `details = NULL`](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L209-L223). If passwordless delegated startup control becomes a requirement, add a purpose-built, root-owned helper exposing only the two fixed Fenris operations and authorize that helper with a Fenris-specific polkit action; do not grant the generic systemd action.
## Product context observed in this repository
The current prototype combines sampling, persistence, HTTP serving, process daemonization, PID-file management, status, and control in [`fenris.py`](../../fenris.py). It starts a detached Python process itself, stores history/PID/log files under the checkout's `data/`, binds the dashboard to `0.0.0.0`, and runs `sudo -n smartctl -a -j DEVICE`. [`fenris.sh`](../../fenris.sh) starts/stops that process and currently recommends a passwordless sudoers entry for `/usr/sbin/smartctl` without argument constraints. The README describes the intended five-minute continuous collection and on-demand menu/dashboard.
That prototype shape is unsuitable for a boot-persistent privileged installation: a privileged process would also contain the HTTP server and large dashboard surface, checkout-relative state has no system ownership boundary, PID files duplicate service-manager state, and the broad sudoers example permits more than Fenris's read-only query. The recommendation below separates those concerns rather than wrapping the existing `start` command in a unit.
## Verified constraints
| Area | Verified constraint | Design consequence |
|---|---|---|
| System vs. user manager | A non-root user service cannot switch to another identity with `User=`; system services default to root and may select another user. [systemd.exec(5), `User=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#User=) | A normal user service is not a reliable privilege boundary for SMART access. Collection belongs in the system manager. |
| User-service persistence | User lingering causes that user's manager to be spawned at boot and kept after logout; without this extra policy, a user manager is session-oriented. [loginctl(1), `enable-linger`](https://www.freedesktop.org/software/systemd/man/latest/loginctl.html#enable-linger%20USER%E2%80%A6) | A user unit either fails the reboot/no-login requirement or requires lingering while still not solving device privilege. Reject it for the collector. |
| Boot enablement | `[Install]` directives do not execute at runtime; `enable` materializes them as symlinks. `WantedBy=` creates a `.wants/` link, and `timers.target` is the recommended boot target for application timers. [systemd.unit(5), `[Install]`](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#%5BInstall%5D%20Section%20Options); [systemd.special(7), `timers.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#timers.target) | Install with `WantedBy=timers.target`, then explicitly enable the timer. Merely shipping the files does not enable startup. |
| Enabled vs. running | `systemctl enable` does not start a unit, and `disable` does not stop it; `--now` couples those otherwise separate changes. [systemctl(1), `enable`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#enable%20UNIT%E2%80%A6); [systemctl(1), `disable`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#disable%20UNIT%E2%80%A6); [systemctl(1), `--now`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#--now) | The TUI must label startup and immediate runtime effects separately. Default startup actions should not silently use `--now`. |
| Timer recurrence | Combining `OnBootSec=` with a relative trigger provides post-boot and recurring activation. Timer expiry is coalesced within `AccuracySec=` (one minute by default). [systemd.timer(5), monotonic timers](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#OnActiveSec=); [systemd.timer(5), `AccuracySec=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#AccuracySec=) | Use `OnBootSec=` plus `OnUnitInactiveSec=` (or `OnUnitActiveSec=` after overlap behavior is tested), and choose accuracy deliberately. Do not promise exact-second polling. |
| Missed runs | `Persistent=` records the last trigger and can fire when an inactive timer is reactivated; its documented clock behavior must be considered with monotonic timers. [systemd.timer(5), `Persistent=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#Persistent=) | Fenris reads lifetime counters, so replaying every missed five-minute sample is neither possible nor useful. One prompt post-boot sample is sufficient; decide whether suspend catch-up warrants `Persistent=yes`. |
| Boot ordering | `timers.target` exists to activate timers after boot. `StateDirectory=`/`RuntimeDirectory=` automatically add `Requires=` and `After=` dependencies for mounts needed by their paths. `network-online.target` is for consumers that strictly require configured networking. [systemd.special(7), `timers.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#timers.target); [systemd.exec(5), implicit dependencies](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#Implicit%20Dependencies); [systemd.special(7), `network-online.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#network-online.target) | Do not order local SMART collection after the network. Let managed data directories establish filesystem ordering. Treat a not-yet-present NVMe device as a failed sample retried at the next interval unless a fixed device-unit dependency is proven necessary. |
| `smartctl` Linux path | smartmontools opens a Linux NVMe device read-only and sends `NVME_IOCTL_ADMIN_CMD`; failure is returned from the ioctl. [smartmontools `os_linux.cpp`, pinned source](https://github.com/smartmontools/smartmontools/blob/618fcaede4478bc7d17fa2a8db5fd18af3744e20/lib/os_linux.cpp#L2817-L2861) | Ordinary file read permission alone does not prove the admin ioctl will be authorized. Validate the shipped unit on every supported kernel/device transport. Root execution is the robust initial compatibility choice. |
| Capability substitution | `CAP_SYS_ADMIN` is intentionally overloaded and is described as plausibly “the new root”; Linux man-pages explicitly advise avoiding it where possible. [capabilities(7), `CAP_SYS_ADMIN`](https://man7.org/linux/man-pages/man7/capabilities.7.html#CAP_SYS_ADMIN); [capabilities(7), developer notes](https://man7.org/linux/man-pages/man7/capabilities.7.html#NOTES) | Do not move `CAP_SYS_ADMIN` into the long-lived TUI or combined web process merely to avoid UID 0. A short-lived, sandboxed root collector has a smaller practical exposure. A minimal capability set remains a test item, not an assumption. |
| Device sandboxing | `PrivateDevices=yes` supplies only pseudo-devices and excludes physical devices. [systemd.exec(5), `PrivateDevices=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#PrivateDevices=) | The collector must not set `PrivateDevices=yes`. Prefer a device cgroup allow-list for the configured node where supported and tested. The TUI/dashboard can use `PrivateDevices=yes`. |
| sudo non-interactivity | `sudo -n` never prompts and fails when authentication is required. [sudo(8), pinned upstream source](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudo.man.in#L629-L635) | `sudo -n` prevents a daemon hang but does not create authorization. It is unnecessary inside a root system service and gives poor interactive UX in the TUI. |
| sudoers command scope | If a sudoers command omits arguments, the user may supply any arguments; argument wildcards require care. `NOPASSWD` removes authentication for matching entries. [sudoers(5), pinned upstream source](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L1107-L1129); [sudoers(5), `PASSWD`/`NOPASSWD`](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L2063-L2089) | The README's path-only `NOPASSWD: /usr/sbin/smartctl` rule is too broad. Do not retain it as the installed architecture. If sudo is retained as a fallback, use a root-owned fixed-argument helper, not user-controlled smartctl arguments. |
| systemd authorization | Read access to systemd's D-Bus objects is generally available; state-changing unit operations require `manage-units`, while enablement operations require `manage-unit-files`. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security) | Status needs no blanket elevation. Starting/stopping and enabling/disabling are separate privileged action classes. |
| polkit rules | polkit loads JavaScript rules from `/etc/polkit-1/rules.d` and `/usr/share/polkit-1/rules.d`; a rule may return `AUTH_ADMIN`, while `AUTH_ADMIN_KEEP` caches authorization briefly for the same action/subject. [polkit(8), authorization rules](https://polkit.pages.freedesktop.org/polkit/polkit.8.html#AUTHORIZATION-RULES) | Prefer normal admin authentication and avoid cached authorization for a sensitive toggle. A custom delegated helper needs its own narrow action and rule. |
| Status parsing | `systemctl status` is human-readable and may include recent journal lines; `systemctl show` is the computer-parsable interface and supports selecting properties. [systemctl(1), `status`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#status%20PATTERN%E2%80%A6); [systemctl(1), `show`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#show%20PATTERN%E2%80%A6) | Never scrape `status` in the TUI. Query an explicit property allow-list and do not echo arbitrary logs in the default status screen. |
| Durable and ephemeral paths | systemd maps `StateDirectory=` to `/var/lib/` and `RuntimeDirectory=` to `/run/`; state/config/log directories remain after stop, while runtime directories are normally removed. [systemd.exec(5), directory table and lifecycle](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=) | Durable samples belong in `/var/lib/fenris`; locks/sockets belong in `/run/fenris`. A checkout-relative `data/` directory and PID file should not survive migration. |
## Proposed lifecycle and control boundary
```text
boot
└─ system systemd
└─ enabled fenris-collect.timer (unprivileged users may inspect)
└─ periodically activates fenris-collect.service
├─ short-lived privileged collector only
├─ reads /etc/fenris/fenris.conf
├─ runs /usr/sbin/smartctl with fixed read-only query arguments
├─ validates and sanitizes JSON
└─ atomically updates /var/lib/fenris/{history,latest,status}
interactive login (independent of collection)
└─ fenris TUI, ordinary invoking user
├─ reads safe status and permitted data
├─ queries selected systemd read-only properties
└─ on explicit confirmed request only:
└─ systemctl enable|disable fenris-collect.timer
└─ system bus → polkit admin authentication → PID 1
```
### Boundary rules
- **Collector:** may access the configured block/controller device and write only Fenris state. It must not bind a network socket, render a UI, edit its configuration, alter unit enablement, or run caller-supplied commands.
- **TUI/dashboard:** may read sanitized state. It must not inherit collector privilege. If a dashboard remains, run it as a separate unprivileged process bound to loopback by default; the current `0.0.0.0` binding must not become part of the privileged unit.
- **systemd/PID 1:** owns process lifecycle and boot enablement. Remove application daemonization, kill-by-PID, and checkout PID files. A service process should remain in the foreground; systemd's service model assumes the started process remains until termination unless a forking type is deliberately used. [systemd.service(5), examples](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Examples)
- **Administrator:** owns installation, `/etc/fenris`, unit files, authorization policy, membership in any read group, and startup-state changes.
A timer/oneshot split is preferable to a continuously privileged daemon because Fenris samples cumulative device counters and needs privilege only during a sample. If later requirements demand a live HTTP server, keep it in a distinct unprivileged service rather than extending collector lifetime.
## Unit sketch (planning only)
The following is intentionally illustrative. Paths, device cgroup syntax, hardening compatibility, and interval behavior must be validated on supported distributions before shipping.
```ini
# /usr/lib/systemd/system/fenris-collect.service
[Unit]
Description=Fenris NVMe SMART sample collector
Documentation=man:smartctl(8)
[Service]
Type=oneshot
ExecStart=/usr/libexec/fenris/fenris-collect --config /etc/fenris/fenris.conf
# Packaging creates the non-privileged read group; the process remains UID 0.
Group=fenris-readers
StateDirectory=fenris
StateDirectoryMode=0750
RuntimeDirectory=fenris
RuntimeDirectoryMode=0750
UMask=0027
# Hardening candidates; validate with smartctl on every supported transport.
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=no
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
LockPersonality=yes
RestrictAddressFamilies=AF_UNIX
# DevicePolicy=closed
# DeviceAllow=/dev/nvme0 r
# CapabilityBoundingSet=... # unresolved; do not guess CAP_SYS_ADMIN-only portability
```
`ProtectSystem=strict` makes the hierarchy read-only except API filesystems, while managed state/log directories are excluded so they remain writable; systemd recommends the protection for long-running services, and it is still useful defense-in-depth for this short-lived one. [systemd.exec(5), `ProtectSystem=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#ProtectSystem=) `NoNewPrivileges=yes` prevents this process and its descendants from gaining new privilege through `execve` mechanisms such as set-user-ID bits or file capabilities. [systemd.exec(5), `NoNewPrivileges=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#NoNewPrivileges=)
```ini
# /usr/lib/systemd/system/fenris-collect.timer
[Unit]
Description=Periodically collect Fenris NVMe SMART samples
[Timer]
OnBootSec=2min
OnUnitInactiveSec=5min
AccuracySec=30s
Unit=fenris-collect.service
[Install]
WantedBy=timers.target
```
`OnUnitInactiveSec=` measures from deactivation, which avoids overlapping a slow oneshot at the cost of interval drift; `OnUnitActiveSec=` measures from activation. Those distinct bases are defined by `systemd.timer(5)`. [systemd.timer(5), monotonic timer table](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#OnActiveSec=) Choose between them after measuring collection duration and confirming the desired interval semantics.
Do not add `After=network-online.target`; collection is local. Do not add `Requires=/dev/...` as pseudo-syntax. If strict device binding is needed, use the escaped `.device` unit generated for the configured node only after testing replacement/hotplug behavior; otherwise record a bounded failure and let the next timer activation retry.
## Authorization and TUI control sketch (planning only)
### Default: systemctl plus normal polkit authentication
Read path:
```text
systemctl show fenris-collect.service \
--property=LoadState,ActiveState,SubState,Result,ExecMainStatus
systemctl show fenris-collect.timer \
--property=LoadState,ActiveState,SubState,UnitFileState,NextElapseUSecRealtime,NextElapseUSecMonotonic
```
Mutation path, only after a confirmation screen naming the exact effect:
```text
Enable at future boots: systemctl enable fenris-collect.timer
Disable at future boots: systemctl disable fenris-collect.timer
```
Do not silently append `--now`. “Disable at boot” does not mean “cancel a currently executing sample,” and systemctl explicitly keeps enablement separate from start/stop unless `--now` is requested. [systemctl(1), `--now`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#--now)
The TUI must pass a fixed unit name and fixed verb without a shell. On cancellation, authentication failure, timeout, or non-zero exit, it should report no successful change and re-read `UnitFileState`; it must not infer success from the requested action.
### Why not a broad polkit group rule
`manage-units` calls can include `unit` and `verb` details in current systemd source, but the `manage-unit-files` check currently has no details. [systemd `dbus-util.c`, pinned source](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L160-L223) Therefore a rule such as “members of `fenris` may perform `org.freedesktop.systemd1.manage-unit-files`” would delegate generic unit enablement/masking operations, not only Fenris startup state. That is outside the required boundary.
### If delegated passwordless control is later required
Use a small root-owned D-Bus/helper mechanism with a Fenris-specific action, for example `com.bongbetic.fenris.manage-startup`. Its API should accept only an enum `{enable, disable}`, internally target the constant `fenris-collect.timer`, reject options/paths/extra units, perform the unit-file operation, and return the observed resulting state. A polkit rule may then grant that custom action to a designated local group, optionally requiring `subject.local && subject.active`; polkit exposes subject locality/activity and action details to JavaScript rules. [polkit(8), authorization rules and `Subject`](https://polkit.pages.freedesktop.org/polkit/polkit.8.html#AUTHORIZATION-RULES)
A sudo fallback should follow the same fixed helper design. Do not authorize `/usr/bin/systemctl` or `/usr/sbin/smartctl` without exact argument control: sudoers explicitly permits arbitrary arguments when none are specified. [sudoers(5), command arguments](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L1107-L1129)
## Ownership and path choices
| Path | Proposed owner/mode | Purpose and rationale |
|---|---|---|
| `/usr/libexec/fenris/fenris-collect` (or distribution-equivalent) | `root:root`, `0755`, package-managed | Privileged entry point must not be writable by TUI users or the service's read group. Use an absolute `ExecStart`; systemd does not provide shell syntax by default. [systemd.service(5), command lines](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Command%20lines) |
| `/usr/lib/systemd/system/fenris-collect.{service,timer}` | `root:root`, `0644`, package-managed | Vendor unit definitions. Local administrator overrides belong under `/etc/systemd/system/`, which has higher load-path precedence. [systemd.unit(5), unit load path](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#Unit%20File%20Load%20Path) |
| `/etc/fenris/fenris.conf` | `root:root`, `0644` if strictly non-secret; otherwise `0640` | Persistent host configuration. The collector reads but cannot write it under `ProtectSystem=strict`. The TUI should receive only safe fields through status rather than requiring config write access. |
| `/var/lib/fenris/` | created by `StateDirectory=fenris`, `0750`; `root:fenris-readers` if direct group reads are retained | Durable history, aggregates, latest sample, and machine-readable sample status. systemd maps state directories here and leaves them after stop. [systemd.exec(5), directory table/lifecycle](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=) |
| `/run/fenris/` | created by `RuntimeDirectory=fenris`, `0750` | Ephemeral lock/socket only. Do not use a PID file as authority; systemd already tracks the process/unit. Runtime directories are removed on stop by default. [systemd.exec(5), `RuntimeDirectoryPreserve=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectoryPreserve=) |
| Journal | journal ACL/policy | Operational diagnostics. Avoid a separate root-owned `/var/log/fenris.log` unless retention/export requirements demand it. Do not expose arbitrary journal messages through safe status. |
These locations also match FHS semantics: `/etc` holds host-specific configuration, `/run` is run-time variable data cleared at boot, and `/var/lib` holds application state that persists across restarts. [FHS 3.0, `/etc`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch03s07.html); [FHS 3.0, `/run`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch03s15.html); [FHS 3.0, `/var/lib`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch05s08.html)
Prefer a narrow read-only IPC/status endpoint over group-readable raw files if multi-user confidentiality matters. The sketch instead uses a dedicated `fenris-readers` primary group for the UID-0 collector so managed directories become `root:fenris-readers`; packaging must create that group, and only approved users should join it. Do not reuse the collector's privileged identity as a user-facing authorization group. Use atomic replace for `latest.json`/`status.json`, append safely for history, set a restrictive umask, and omit serial numbers, command lines, environment, and raw stderr from the shared surface.
## Safe status design
The default TUI status should be an allow-listed composition, not a dump of `systemctl status`, journal output, raw smartctl JSON, configuration, or process command lines.
Recommended fields:
```text
Installation: loaded | not-installed | error
Startup: enabled | disabled | static | masked | unknown
Scheduler runtime: active | inactive | failed
Collection runtime: active | inactive | failed
Last attempt: RFC3339 timestamp
Last success: RFC3339 timestamp
Freshness: fresh | stale | never (show threshold and age)
Last result: success | device-unavailable | permission-denied |
timeout | invalid-output | storage-error | internal-error
Next scheduled: timestamp if systemd reports one
Samples/history: count and range, if readable
Device: configured stable identifier or node; no serial by default
```
Design requirements:
- Treat `UnitFileState` (startup policy), `ActiveState`/`SubState` (runtime), and last sample result as independent axes. `is-enabled` documents multiple states—including enabled, disabled, static, indirect, generated, transient, and masked—so a Boolean loses actionable information. [systemctl(1), `is-enabled` state table](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#is-enabled%20UNIT%E2%80%A6)
- Parse only `systemctl show --property=...` or the equivalent D-Bus properties. `status` is explicitly human-oriented and includes journal data. [systemctl(1), `show`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#show%20PATTERN%E2%80%A6)
- Compute freshness from `last_success`, not merely timer activity. An active timer can coexist with repeated collection failures.
- Publish bounded error categories and a short administrator hint; keep raw smartctl stderr and tracebacks in the journal. This prevents device identifiers, paths, malformed device output, or command details from crossing the read boundary.
- Never report “enabled” immediately after a requested mutation without re-querying the authoritative state. Never equate “disabled” with “stopped.”
- If status data is unreadable, say `permission-denied`/`unknown`; do not elevate merely to render the screen.
## Alternatives and trade-offs
### Long-running system service
A foreground `fenris-collector.service` with `Restart=on-failure` and `WantedBy=multi-user.target` also satisfies reboot persistence; `Restart=` controls automatic restart after process failure, while `WantedBy=multi-user.target` is the standard installation relationship for a multi-user service. [systemd.service(5), `Restart=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Restart=); [systemd.special(7), `multi-user.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#multi-user.target)
Trade-off: it preserves today's loop model and precise in-process scheduling, but leaves a privileged Python process resident continuously and requires restart/backoff handling. Use it only if future collection needs persistent in-memory state that cannot be reconstructed from durable samples.
### Unprivileged daemon plus privileged smartctl helper
This can reduce privileged code if the helper accepts no uncontrolled path or options, validates a configured device allow-list, emits bounded sanitized output, and exits. It adds an IPC/protocol and another authorization surface. Giving the whole Python process `CAP_SYS_ADMIN` is not an equivalent reduction because that capability is exceptionally broad. [capabilities(7), notes](https://man7.org/linux/man-pages/man7/capabilities.7.html#NOTES)
### User service with lingering
Lingering can run a user manager from boot through logout, but the user manager cannot switch an ordinary user's unit to root and some system-service sandboxing features are unavailable in user services. [loginctl(1), lingering](https://www.freedesktop.org/software/systemd/man/latest/loginctl.html#enable-linger%20USER%E2%80%A6); [systemd.exec(5), user-service sandboxing limitations](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#Sandboxing) This adds account coupling and still needs a separate privilege mechanism. Reject for collection; it remains acceptable for an optional per-user dashboard client.
### Root daemon calling `sudo -n smartctl`
This is redundant: root already crosses the privilege boundary, while sudo introduces policy/path/configuration failure modes. For an unprivileged daemon, `sudo -n` avoids blocking but only works with pre-authorized sudoers policy and produces no authentication opportunity. [sudo(8), `--non-interactive`](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudo.man.in#L629-L635) Reject as the installed systemd architecture.
### Direct broad polkit delegation
Convenient but unsafe for startup-state delegation: `manage-unit-files` covers generic enable/disable/mask/preset operations and currently lacks unit-scoping details in systemd's authorization call. Use normal administrator authentication or a custom narrow action instead. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security); [pinned systemd authorization source](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L209-L223)
## Unresolved questions and required validation
1. **Supported distributions/systemd floor:** What is the minimum systemd version? Confirm every selected hardening and directory directive exists there; unknown directives can weaken the intended sandbox.
2. **Device identity:** Is configuration a mutable `/dev/nvmeN` node, namespace node, controller, `/dev/disk/by-id` link, or discovered set? smartctl documents Linux NVMe controller and namespace forms, but the stable product identity and hotplug behavior remain a product decision. [smartctl(8), pinned source](https://github.com/smartmontools/smartmontools/blob/618fcaede4478bc7d17fa2a8db5fd18af3744e20/src/smartctl.8.in#L72-L86)
3. **Privilege matrix:** On each supported kernel, packaging of smartmontools, NVMe/SATA/USB bridge, and device permission setup, record which open/ioctl fails as an unprivileged service and which exact capability/device allow-list is sufficient. Do not generalize an NVMe result to every smartmontools transport.
4. **Root vs. reduced-capability collector:** After that matrix exists, decide whether a non-root static service user plus a minimal capability/device set works portably. Reject any result that requires putting broad capability into the TUI/dashboard.
5. **Timer semantics:** Should five minutes be measured from sample start or completion? What maximum runtime and timeout are acceptable? Should resume from suspend trigger immediately? Decide `OnUnitActiveSec` vs. `OnUnitInactiveSec`, `Persistent=`, and `AccuracySec` from those answers.
6. **Startup toggle semantics:** Does “disable startup” leave the timer active until reboot, or should the TUI offer a separate, clearly labeled “disable and stop now”? The systemctl semantics intentionally separate these effects. [systemctl(1), enable/disable](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#enable%20UNIT%E2%80%A6)
7. **Who may read health history:** All local users, a `fenris-readers` group, or only an authenticated local client? This determines state modes and whether a read-only Unix socket is preferable to files.
8. **Dashboard scope:** Is the HTTP dashboard retained, and if so must it be local-only or remotely accessible? Remote access needs a separate threat model, authentication, transport security, and unprivileged service; it must not enlarge the collector boundary.
9. **Packaging paths:** Confirm `/usr/libexec` and `/usr/lib/systemd/system` equivalents per target distribution, the absolute smartctl path, and whether administrator overrides use environment files or a dedicated validated config format.
10. **Data durability:** Define atomicity, fsync policy, retention, corruption recovery, migration from checkout-relative `data/`, and behavior on read-only/full filesystems.
11. **Authorization UX:** Is ordinary admin authentication acceptable? If not, specify exactly which local principals may toggle startup and commission the custom helper/action rather than broad `manage-unit-files` delegation.
12. **Status confidentiality:** Decide whether model name, device path, capacity, wear, temperatures, and error counters are safe for every local reader; serial numbers and raw smartctl output should remain excluded by default.
## Decision summary
The verified boundary is: **system systemd owns scheduling and boot state; a short-lived privileged collector owns only device interrogation and state writes; an unprivileged on-demand TUI owns presentation; polkit/admin authentication mediates explicit startup changes.** This satisfies collection across reboot without tying it to login, avoids embedding sudo in unattended code, keeps generic service control out of the TUI's ambient privilege, and makes startup state observable and alterable without conflating it with current execution.
+2
View File
@@ -0,0 +1,2 @@
.venv/
__pycache__/
+46
View File
@@ -0,0 +1,46 @@
# Fenris TUI information-architecture PROTOTYPE (throwaway)
**This is throwaway code answering [ticket #3](https://git.bongbetic.com/xavierk/Fenris/issues/3).** It is not the redesign, reads nothing real, and never ships. Branch: `prototype/tui-information-architecture`.
## Question
What screen hierarchy, navigation, and action model makes Fenris's projection contract (ADR 0002 §13), the four separate service facts (ADR 0003 §8), warming-up, unexplained gaps, and changing habits understandable in a keyboard-first terminal?
## Run (one command)
```sh
./run
```
(creates `.venv` and installs `textual` on first use)
## What to flip through
**Variants (← / →)** — three structurally different answers, not restylings:
| Key | Variant | Idea |
|-----|---------|------|
| A | **Panes** | everything on one dense screen, btop-style; no navigation, panes are zones |
| B | **Pages** | persistent three-fact header (lifespan · confidence · freshness) + pages 1–5 |
| C | **Ledger** | one scrolling document in reading order, headline sentence first |
**States (s)** — same variants, six shapes of the contract:
1. steady · Supported (with one unexplained 3-hour gap)
2. warming up · Limited (11 of 14 days)
3. habit changed · Limited (regime 6 days old, scenario spread visible)
4. stale · Supported→Limited (last collect FAILED, 61 h old)
5. no baseline · Unavailable (wear too coarse to imply endurance)
6. paused · Limited (period closed by deliberate disable)
**Actions** — `p` pause (asks confirmation) · `r` resume (doesn't) · `c` collect now (synchronous outcome). Each suspends the TUI and runs `polkit_stub.py` on the real terminal: this validates the tty passthrough ADR 0003 requires for the polkit prompt. Results land in the tty log (variant B service page; every variant's log is the same list).
## What to react to
- Which variant's hierarchy matches how you think about the drive? (Mixing — "header from B, density of A" — is a valid answer and the point.)
- Are the four service facts separable at a glance?
- Do confidence states + contributing facts read as evidence, not as a percentage?
- Is the pause confirmation the right amount of friction?
- Did the polkit tty stub actually prompt in your terminal? (That's the mechanism check.)
`screenshots/` holds headless captures (`smoke_test.py`) of each variant at 80×24 and 140×40.
+29
View File
@@ -0,0 +1,29 @@
#!/usr/bin/env python3
"""STUB polkit-agent stand-in for the Fenris TUI prototype (throwaway).
Runs attached to the real terminal while the Textual app is suspended, exactly
where the platform polkit agent would prompt for `com.bongbetic.fenris.monitor`.
Accepts any password; the point is validating tty passthrough, not auth.
"""
import getpass
import sys
import time
op = sys.argv[1] if len(sys.argv) > 1 else "unknown"
print("=" * 56)
print(" polkit STUB · com.bongbetic.fenris.monitor")
print(f" operation: {op}")
print(" Authentication required to manage Fenris monitoring")
print("=" * 56)
try:
getpass.getpass(" password (anything works): ")
except (EOFError, KeyboardInterrupt):
print("\n(cancelled — operation not performed)")
sys.exit(1)
time.sleep(0.6) # pretend systemctl + monitoring-period bookkeeping
print(f" fenris-monitor {op}: done")
try:
input(" [press Enter to return to the TUI] ")
except EOFError:
pass
sys.exit(0)
+1
View File
@@ -0,0 +1 @@
textual>=0.60
+14
View File
@@ -0,0 +1,14 @@
#!/usr/bin/env bash
# PROTOTYPE runner — throwaway, see README.md
set -euo pipefail
cd "$(dirname "$0")"
if [ ! -x .venv/bin/python ]; then
if command -v uv >/dev/null 2>&1; then
uv venv -q .venv
uv pip install -q --python .venv/bin/python -r requirements.txt
else
python3 -m venv .venv
.venv/bin/pip -q install -r requirements.txt
fi
fi
exec .venv/bin/python tui_prototype.py
File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 61 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 40 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 56 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 56 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 43 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 42 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 57 KiB

File diff suppressed because one or more lines are too long

After

Width:  |  Height:  |  Size: 42 KiB

+67
View File
@@ -0,0 +1,67 @@
# Headless smoke test for the prototype: drives every variant × state through
# Textual's test pilot, exports SVG screenshots, and asserts contract strings render.
import asyncio, inspect, os, sys
os.environ["FENRIS_PROTOTYPE_NO_TTY"] = "1"
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from tui_prototype import FenrisPrototypeApp, VARIANTS
OUT = os.path.join(os.path.dirname(os.path.abspath(__file__)), "screenshots")
os.makedirs(OUT, exist_ok=True)
async def snap(app, name):
r = app.export_screenshot()
if inspect.isawaitable(r):
r = await r
with open(os.path.join(OUT, name + ".svg"), "w") as f:
f.write(r)
async def main():
checks = []
for size in [(80, 24), (140, 40)]:
app = FenrisPrototypeApp()
async with app.run_test(size=size) as pilot:
await pilot.pause()
for i, (key, _name) in enumerate(VARIANTS):
if i:
await pilot.press("right"); await pilot.pause()
tag = f"v{key}_{size[0]}x{size[1]}"
await snap(app, tag + "_steady")
# rotate states on variant B
await pilot.press("left"); await pilot.pause() # back to A
await pilot.press("right"); await pilot.pause() # B
for st in ["warming", "changed", "stale", "nobaseline", "paused", "steady"]:
await pilot.press("s"); await pilot.pause()
await snap(app, f"vB_{size[0]}x{size[1]}_{st}")
# pages on B
for k in ["1", "2", "3", "4", "5"]:
await pilot.press(k); await pilot.pause()
body = app.query_one("#vB-service").content
checks.append(("service facts", "boot enablement" in str(body) and "freshness" in str(body)))
# pause flow: confirm modal, y, stub skipped headless -> state becomes paused? (skip branch)
await pilot.press("4"); await pilot.pause()
await pilot.press("p"); await pilot.pause()
await pilot.press("n"); await pilot.pause() # cancel, no state change
head = str(app.query_one("#vB-overview").content)
checks.append(("confidence renders", "Projection confidence" in head))
# disclosure modal
await pilot.press("d"); await pilot.pause()
await pilot.press("escape"); await pilot.pause()
steady = FenrisPrototypeApp()
async with steady.run_test(size=(140, 40)) as pilot:
await pilot.pause()
doc = str(steady.query_one("#vC-doc").content)
checks += [
("headline", "Usage-adjusted theoretical lifespan" in doc),
("scenario range", "Scenario range" in doc),
("wear line", "Vendor wear" in doc),
("gap marker", "unexplained gap" in doc),
]
failed = [n for n, ok in checks if not ok]
print("CHECKS:", "all ok" if not failed else f"FAILED: {failed}")
for n, ok in checks:
print(f" {'ok ' if ok else 'FAIL'} {n}")
if failed:
sys.exit(1)
asyncio.run(main())
+573
View File
@@ -0,0 +1,573 @@
# PROTOTYPE (THROWAWAY) — Fenris TUI information-architecture prototype.
# Question (ticket #3): what screen hierarchy, navigation, and action model makes
# the projection contract (ADR 0002 §13) and the service facts (ADR 0003 §8)
# understandable in a keyboard-first terminal?
# Plan: three structurally different variants (A Panes / B Pages / C Ledger),
# switchable live with ←/→, plus a scenario rotator (s) that drives the same
# variants through warming-up / habit-change / stale / paused / no-baseline states.
# Data is synthetic but modeled on the real drive (Micron 2400 512GB, ~101 TB
# written, 50 % used). Nothing here reads or writes the observation store.
from __future__ import annotations
import os
import random
import subprocess
import sys
from pathlib import Path
from textual.app import App, ComposeResult
from textual.binding import Binding
from textual.containers import Horizontal, Vertical, VerticalScroll
from textual.screen import ModalScreen
from textual.widgets import Static
STUB = Path(__file__).with_name("polkit_stub.py")
# ---------------------------------------------------------------- fake data
DRIVE = {
"model": "Micron_2400_MTFDKBA512QFM",
"capacity": "512 GB",
"percentage_used": 50,
"written_tb": 101.1,
"temp": 38,
"spare": 100,
"media_errors": 0,
"power_on_hours": 11126,
"power_cycles": 5188,
"unsafe_shutdowns": 142,
}
BASELINE = {"tb": 220.0, "source": "Micron 2400 datasheet (PROTOTYPE placeholder provenance)"}
def _years(gb_per_day: float) -> str:
days = (BASELINE["tb"] - DRIVE["written_tb"]) * 1000.0 / gb_per_day
if days >= 365.25:
return f"~{days / 365.25:.1f} years"
return f"~{days:.0f} days"
def _days_hist(n: int, rate: float, seed: int) -> list[float]:
rng = random.Random(seed)
return [max(2.0, rate + rng.gauss(0, rate * 0.18)) for _ in range(n)]
def _sparkline(vals: list[float], width: int = 40) -> str:
if not vals:
return ""
mx = max(vals) or 1.0
blocks = " ▁▂▃▄▅▆▇█"
step = max(1, len(vals) // width or 1)
picked = vals[-width * step:][::step][-width:]
return "".join(blocks[min(len(blocks) - 1, int(v / mx * (len(blocks) - 1)) + (1 if v > 0 else 0))] for v in picked)
def _habit_bar(active: float, idle: float, off: float, unknown: float, width: int = 44) -> str:
total = active + idle + off + unknown or 1.0
segs = [("active", active, "green"), ("idle", idle, "yellow"), ("off", off, "cyan"), ("?", unknown, "magenta")]
out = []
for label, v, color in segs:
n = max(1 if v else 0, round(v / total * width))
out.append((f"[{color}]{label[0] * n}[/{color}]"))
legend = f" active {active / total:.0%} · idle {idle / total:.0%} · powered-off {off / total:.0%} · unknown {unknown / total:.0%}"
return "".join(out) + "\n " + legend
def build_scenarios() -> list[dict]:
"""Six states of the ADR-0002 §13 contract + ADR-0003 service facts."""
hist_steady = _days_hist(34, 55, seed=7)
hist_warm = _days_hist(11, 62, seed=11)
hist_changed = _days_hist(35, 70, seed=3) + _days_hist(6, 130, seed=4)
hist_stale = _days_hist(34, 55, seed=7)
hist_nobase = _days_hist(26, 48, seed=5)
hist_paused = _days_hist(28, 51, seed=9)
def svc(enabled, timer, outcome, freshness, period):
return {"enabled": enabled, "timer": timer, "outcome": outcome, "freshness": freshness, "period": period}
return [
{
"key": "steady",
"name": "steady · Supported",
"days": 34,
"history": hist_steady,
"rate": 55,
"headline": _years(55),
"confidence": "Supported",
"facts": [
"34 qualifying days (≥ 14), coverage 92 %",
"7- and 28-day rates within a factor of 2",
"no single day ≥ 50 % of trailing 28-day writes",
"1 day with 3 unknown hours — unexplained gap inside the period",
],
"horizons": [("last 7 days", _years(48)), ("last 28 days", _years(57))],
"habit": (0.34, 0.52, 0.10, 0.04),
"gap_days": {-9},
"habit_change": None,
"service": svc(True, True, "ok · 3 min ago (5 min cadence)", "fresh · newest sample 3 min old", "open since Aug 3 · deliberate disables: 0"),
},
{
"key": "warming",
"name": "warming up · Limited",
"days": 11,
"history": hist_warm,
"rate": 62,
"headline": _years(62),
"confidence": "Limited",
"facts": [
"warming up: 11 of 14 qualifying days",
"coverage 84 %",
],
"horizons": [("last 7 days", _years(66))],
"habit": (0.38, 0.46, 0.12, 0.04),
"gap_days": set(),
"habit_change": None,
"service": svc(True, True, "ok · 2 min ago", "fresh · newest sample 2 min old", "open since Aug 20"),
},
{
"key": "changed",
"name": "habit changed · Limited",
"days": 41,
"history": hist_changed,
"rate": 130,
"headline": _years(130),
"confidence": "Limited",
"facts": [
"usage habit changed 6 days ago — new regime adopted",
"regime 6 days old (young — Limited evidence)",
"coverage 88 %",
"vendor wear line disagrees ×2.1 with observed write rate",
],
"horizons": [("last 7 days", _years(128)), ("last 28 days", _years(71))],
"habit": (0.47, 0.41, 0.08, 0.04),
"gap_days": set(),
"habit_change": 6,
"service": svc(True, True, "ok · 4 min ago", "fresh · newest sample 4 min old", "open since Jul 15 · habit change noted Aug 26"),
},
{
"key": "stale",
"name": "stale · Supported→Limited",
"days": 34,
"history": hist_stale,
"rate": 55,
"headline": _years(55),
"confidence": "Limited",
"facts": [
"newest evidence 61 h old (missed — older than 48 h)",
"34 qualifying days, coverage 92 %",
],
"horizons": [("last 7 days", _years(48)), ("last 28 days", _years(57))],
"habit": (0.34, 0.52, 0.10, 0.04),
"gap_days": {-9},
"habit_change": None,
"service": svc(True, True, "FAILED · exit 1 · 61 h ago (device busy)", "missed · newest sample 61 h old", "open since Aug 3 · gap is unknown time inside the period"),
},
{
"key": "nobaseline",
"name": "no baseline · Unavailable",
"days": 26,
"history": hist_nobase,
"rate": 48,
"headline": None,
"confidence": "Unavailable",
"facts": [
"no verified rated-TBW override on record",
"vendor wear estimate too coarse to imply endurance (1 of ≥ 2 Percentage Used increments)",
],
"horizons": [("last 7 days", "48 GB/day (no baseline to project)"), ("last 28 days", "44 GB/day (no baseline to project)")],
"habit": (0.31, 0.55, 0.10, 0.04),
"gap_days": set(),
"habit_change": None,
"service": svc(True, True, "ok · 3 min ago", "fresh · newest sample 3 min old", "open since Aug 8"),
},
{
"key": "paused",
"name": "paused · Limited",
"days": 28,
"history": hist_paused,
"rate": 51,
"headline": _years(51),
"confidence": "Limited",
"facts": [
"no open monitoring period — paused 2 days ago",
"paused time is excluded from the usage habit by your choice",
],
"horizons": [("last 7 days (pre-pause)", _years(53))],
"habit": (0.33, 0.51, 0.12, 0.04),
"gap_days": set(),
"habit_change": None,
"service": svc(False, False, "ok · 2 d ago (period closed by pause)", "stale · monitoring paused 2 days ago", "closed 2 days ago · end cause: deliberate disable"),
},
]
# ------------------------------------------------------------- renderers
def headline_block(sc: dict) -> str:
if sc["headline"]:
return (
f"[bold]Usage-adjusted theoretical lifespan: [white]{sc['headline']}[/white][/bold]\n"
f" if current habits continue · sustained regime: {sc['days'] if not sc['habit_change'] else sc['habit_change']} days at {sc['rate']} GB/day"
)
return "[bold]Usage-adjusted theoretical lifespan: [red]no projection from this history yet[/red][/bold]\n " + "\n ".join(sc["facts"][:2])
def confidence_block(sc: dict) -> str:
color = {"Supported": "green", "Limited": "yellow", "Unavailable": "red"}[sc["confidence"]]
lines = [f"[bold]Projection confidence: [{color}]{sc['confidence']}[/{color}][/bold]"]
lines += [f" · {f}" for f in sc["facts"]]
return "\n".join(lines)
def horizon_block(sc: dict) -> str:
rows = [f" {label:<28} → [cyan]{value}[/cyan]" for label, value in sc["horizons"]]
return "[bold]Scenario range[/bold] (same endurance, other horizons)\n" + "\n".join(rows)
def wear_line(sc: dict) -> str:
return f"Vendor wear: {DRIVE['percentage_used']} % used · {DRIVE['written_tb']} TB of {BASELINE['tb']:.0f} TB rated (context, not a second projection)"
def disclosure_lines() -> list[str]:
return [
"Rated endurance is a vendor guarantee boundary, not a predicted failure date.",
"Powered-off time counts toward the projection while monitoring is enabled; deliberately paused time does not.",
"The scenario range is a spread of horizons, not a statistical interval.",
f"Baseline provenance: {BASELINE['source']}.",
]
def history_block(sc: dict, width: int = 60) -> str:
vals = sc["history"]
spark = _sparkline(vals, width)
marks = [" "] * len(spark)
if sc["habit_change"]:
idx = len(spark) - max(1, round(sc["habit_change"] / max(1, len(vals) // width or 1)))
if 0 <= idx < len(marks):
marks[idx] = "▲"
for g in sc["gap_days"]:
idx = len(spark) + g - 1
if 0 <= idx < len(marks) and marks[idx] == " ":
marks[idx] = "?"
head = f"[bold]Usage history[/bold] · {sc['days']} days · {min(vals):.0f}–{max(vals):.0f} GB/day"
bar = f"[green]{spark}[/green]"
markline = "".join(marks)
a, i, o, u = sc["habit"]
return head + "\n " + bar + "\n " + markline + " ▲ habit change · ? unexplained gap\n " + _habit_bar(a, i, o, u)
def health_block() -> str:
d = DRIVE
rows = [
f"[bold]Drive health[/bold] · {d['model']}",
f" temperature {d['temp']} °C · spare {d['spare']} %",
f" media errors {d['media_errors']} · unsafe shutdowns {d['unsafe_shutdowns']}",
f" power-on {d['power_on_hours']:,} h · {d['power_cycles']:,} cycles · {d['capacity']}",
]
return "\n".join(rows)
def service_block(sc: dict) -> str:
s = sc["service"]
en = "[green]enabled[/green]" if s["enabled"] else "[red]disabled[/red]"
tm = "[green]timer active[/green]" if s["timer"] else "[red]timer inactive[/red]"
return "\n".join([
"[bold]Service[/bold] (four separate facts)",
f" boot enablement: {en}",
f" runtime activity: {tm}",
f" last collect outcome: {s['outcome']}",
f" freshness: {s['freshness']}",
f" monitoring period: {s['period']}",
])
def settings_block() -> str:
return "\n".join([
"[bold]Settings[/bold] (read view · edit via CLI / drop-ins)",
" device: /dev/disk/by-id/nvme-Micron_2400_MTFDKBA512QFM_2341ABCD",
" cadence: every 5 min (systemd drop-in to change) · raw retention 14 d",
f" endurance baseline: {BASELINE['tb']:.0f} TB rated — {BASELINE['source']}",
])
def actions_legend(paused: bool) -> str:
resume = "[bold green]r resume[/bold green]" if paused else "r resume"
pause = "[bold yellow]p pause[/bold yellow]" if not paused else "p pause"
return f"{pause} (asks) · {resume} · c collect now · s state · ←/→ variant · d disclosures · q quit"
VARIANTS = [
("A", "Panes — one dense screen"),
("B", "Pages — persistent header + tabbed body"),
("C", "Ledger — scrolling narrative document"),
]
# ----------------------------------------------------------------- screens
class ConfirmPause(ModalScreen[bool]):
"""Pause asks for confirmation (ADR 0003 §8)."""
BINDINGS = [
Binding("y", "yes", "Pause"),
Binding("n", "no", "Cancel"),
Binding("escape", "no", "Cancel", show=False),
]
def compose(self) -> ComposeResult:
yield Static(
"[bold]Pause monitoring?[/bold]\n\n"
"This closes the current monitoring period.\n"
"Paused time is [bold]excluded[/bold] from your usage habit\n"
"(powered-off time would still count).\n\n"
"[dim]y pause · n cancel[/dim]",
id="confirm-text",
)
def action_yes(self) -> None:
self.dismiss(True)
def action_no(self) -> None:
self.dismiss(False)
class Disclosures(ModalScreen):
BINDINGS = [Binding("escape", "close", "Close"), Binding("d", "close", "Close")]
def compose(self) -> ComposeResult:
yield VerticalScroll(Static("\n".join(["[bold]Disclosures[/bold]"] + [f" · {d}" for d in disclosure_lines()]) + "\n\n[dim]esc to close[/dim]", id="disc-text"), id="disc-wrap")
def action_close(self) -> None:
self.dismiss()
# -------------------------------------------------------------------- app
class FenrisPrototypeApp(App):
TITLE = "Fenris — PROTOTYPE (throwaway)"
SUB_TITLE = "TUI information architecture · ticket #3"
BINDINGS = [
Binding("left", "prev_variant", "‹ variant", show=False),
Binding("right", "next_variant", "variant ›", show=False),
Binding("s", "cycle_state", "state", show=False),
Binding("p", "pause", "pause", show=False),
Binding("r", "resume", "resume", show=False),
Binding("c", "collect", "collect now", show=False),
Binding("d", "disclose", "disclosures", show=False),
Binding("q", "quit", "quit", show=False),
Binding("1", "page('overview')", "overview", show=False),
Binding("2", "page('history')", "history", show=False),
Binding("3", "page('drive')", "drive", show=False),
Binding("4", "page('service')", "service", show=False),
Binding("5", "page('settings')", "settings", show=False),
Binding("j", "scroll_down", "scroll down", show=False),
Binding("k", "scroll_up", "scroll up", show=False),
]
CSS = """
#vhost { height: 1fr; }
#vA { layout: grid; grid-size: 2 3; grid-columns: 3fr 2fr; grid-rows: 8 1fr 7; height: 1fr; }
#vA-head, #vA-bottom { column-span: 2; }
#vA-history { overflow-y: auto; }
.pane { border: round #555555; padding: 0 1; }
#vB-head { height: 6; border-bottom: thick #555555; padding: 0 1; }
#vB-body { height: 1fr; padding: 0 1; }
#vB-legend { height: 3; }
.page { height: 1fr; padding: 0 1; }
#vC { height: 1fr; padding: 0 2; }
#switchbar { height: 3; dock: bottom; background: $boost; }
#sb-variant { width: 1fr; text-style: reverse; }
#sb-state { width: 1fr; }
#sb-actions { width: 2fr; }
#confirm-text { padding: 1 2; }
#disc-wrap { padding: 1 2; height: auto; max-height: 80%; }
"""
def __init__(self) -> None:
super().__init__()
self.scenarios = build_scenarios()
self.scenario_idx = 0
self.variant_idx = 0
self.b_page = "overview"
self.tty_log: list[str] = []
# ---- composition
def compose(self) -> ComposeResult:
with Vertical(id="vhost"):
with Vertical(id="vA"):
yield Static("", id="vA-head", classes="pane")
yield Static("", id="vA-history", classes="pane")
yield Static("", id="vA-health", classes="pane")
yield Static("", id="vA-bottom", classes="pane")
with Vertical(id="vB"):
yield Static("", id="vB-head")
with Vertical(id="vB-body"):
yield Static("", id="vB-overview", classes="page")
yield Static("", id="vB-history", classes="page")
yield Static("", id="vB-drive", classes="page")
yield Static("", id="vB-service", classes="page")
yield Static("", id="vB-settings", classes="page")
yield Static("", id="vB-legend")
with VerticalScroll(id="vC"):
yield Static("", id="vC-doc")
with Horizontal(id="switchbar"):
yield Static("", id="sb-variant")
yield Static("", id="sb-state")
yield Static("", id="sb-actions")
def on_mount(self) -> None:
for vid in ("vA-head", "vA-history", "vA-health", "vA-bottom"):
w = self.query_one(f"#{vid}", Static)
w.border_title = {"vA-head": "headline", "vA-history": "usage history", "vA-health": "drive", "vA-bottom": "service + actions"}[vid]
self.render_all()
# ---- helpers
@property
def sc(self) -> dict:
return self.scenarios[self.scenario_idx]
def w(self, vid: str) -> Static:
return self.query_one(f"#{vid}", Static)
def render_all(self) -> None:
sc = self.sc
paused = not sc["service"]["enabled"]
key, name = VARIANTS[self.variant_idx]
# variant A: everything on one dense screen
self.w("vA-head").update(headline_block(sc) + "\n" + confidence_block(sc))
self.w("vA-history").update(history_block(sc) + "\n" + horizon_block(sc))
self.w("vA-health").update(health_block() + "\n\n" + settings_block())
self.w("vA-bottom").update(service_block(sc) + "\n " + actions_legend(paused))
# variant B: persistent header, tabbed pages
head = "\n".join([
headline_block(sc),
confidence_block(sc).split("\n")[0] + f" · {sc['confidence']}",
f"freshness: {sc['service']['freshness']} · wear: {DRIVE['percentage_used']} %",
])
self.w("vB-head").update(head)
self.w("vB-overview").update(confidence_block(sc) + "\n\n" + horizon_block(sc) + "\n\n" + wear_line(sc))
self.w("vB-history").update(history_block(sc, 70))
self.w("vB-drive").update(health_block() + "\n\n" + wear_line(sc))
self.w("vB-service").update(service_block(sc) + "\n\n " + actions_legend(paused) + "\n\n tty log:\n" + ("\n".join(self.tty_log) if self.tty_log else " (no privileged action taken yet)"))
self.w("vB-settings").update(settings_block() + "\n\n" + "\n".join(" · " + d for d in disclosure_lines()))
self.w("vB-legend").update(f"pages: 1 overview · 2 history · 3 drive · 4 service · 5 settings [now: {self.b_page}]")
for p in ("overview", "history", "drive", "service", "settings"):
self.w(f"vB-{p}").styles.display = "block" if p == self.b_page else "none"
# variant C: one scrolling document in reading order
doc = "\n\n".join([
"[dim]═" * 70 + "[/dim]",
headline_block(sc),
confidence_block(sc),
horizon_block(sc),
wear_line(sc),
history_block(sc, 70),
health_block(),
service_block(sc),
settings_block(),
"[bold]Disclosures[/bold]\n" + "\n".join(" · " + d for d in disclosure_lines()),
"[dim]═" * 70 + "[/dim]",
])
self.w("vC-doc").update(doc)
# variant visibility + switcher
for i, vid in enumerate(("vA", "vB", "vC")):
self.query_one(f"#{vid}").styles.display = "block" if i == self.variant_idx else "none"
self.w("sb-variant").update(f" ← {key} · {name} → ")
self.w("sb-state").update(f" state [{self.scenario_idx + 1}/{len(self.scenarios)}]: {sc['name']} (s to cycle) ")
self.w("sb-actions").update(" " + actions_legend(paused))
# ---- actions
def action_prev_variant(self) -> None:
self.variant_idx = (self.variant_idx - 1) % len(VARIANTS)
self.render_all()
def action_next_variant(self) -> None:
self.variant_idx = (self.variant_idx + 1) % len(VARIANTS)
self.render_all()
def action_cycle_state(self) -> None:
self.scenario_idx = (self.scenario_idx + 1) % len(self.scenarios)
self.render_all()
def action_page(self, page: str) -> None:
if VARIANTS[self.variant_idx][0] != "B":
return
self.b_page = page
self.render_all()
def action_scroll_down(self) -> None:
if VARIANTS[self.variant_idx][0] == "C":
self.query_one("#vC").scroll_down(animated=False)
def action_scroll_up(self) -> None:
if VARIANTS[self.variant_idx][0] == "C":
self.query_one("#vC").scroll_up(animated=False)
def action_disclose(self) -> None:
self.push_screen(Disclosures())
def action_pause(self) -> None:
self.push_screen(ConfirmPause(), callback=self._pause_confirmed)
def _pause_confirmed(self, confirmed: bool) -> None:
if not confirmed:
self.tty_log.append("pause: cancelled at confirmation")
self.render_all()
return
result = self._run_tty_stub("pause (disable --now)")
if "OK" in result:
self.scenarios[self.scenario_idx]["service"] = {
"enabled": False, "timer": False,
"outcome": "ok · period closed by pause",
"freshness": "paused · no collection while disabled",
"period": "closed just now · end cause: deliberate disable",
}
self.tty_log.append(result)
self.render_all()
def action_resume(self) -> None:
result = self._run_tty_stub("resume (enable --now)") # no confirmation (ADR 0003 §8)
if "OK" in result or "skipped" in result:
self.scenarios[self.scenario_idx]["service"] = {
"enabled": True, "timer": True,
"outcome": "ok · resumed just now",
"freshness": "fresh · collection resuming",
"period": "open just now",
}
self.tty_log.append(result)
self.render_all()
def action_collect(self) -> None:
result = self._run_tty_stub("collect now", blocking=True) # synchronous outcome (ADR 0003 §7)
self.tty_log.append(result)
self.render_all()
# ---- tty passthrough validation (the point of the stub)
def _run_tty_stub(self, op: str, blocking: bool = False) -> str:
if os.environ.get("FENRIS_PROTOTYPE_NO_TTY") or not sys.stdin.isatty():
return f"{op}: tty stub SKIPPED (headless run — mechanism not exercised)"
try:
with self.suspend():
proc = subprocess.run([sys.executable, str(STUB), op])
verdict = "OK — suspend + terminal passthrough works" if proc.returncode == 0 else f"FAILED (exit {proc.returncode})"
return f"{op}: {verdict}"
except Exception as exc: # SuspendNotSupported and friends
return f"{op}: suspend failed: {type(exc).__name__} — this terminal may not support passthrough"
if __name__ == "__main__":
app = FenrisPrototypeApp()
print("[fenris prototype] throwaway UI for ticket #3 — variants switch with ←/→, states with s\n")
app.run()