Compare commits

..
Author SHA1 Message Date
xavierk b2bfbfea4f docs: research NVMe endurance signals 2026-08-31 13:50:07 +05:30
2 changed files with 127 additions and 259 deletions
+127
View File
@@ -0,0 +1,127 @@
# NVMe endurance signals and projection constraints
Research for [Verify NVMe endurance signals and projection constraints](https://git.bongbetic.com/xavierk/Fenris/issues/5).
## Decision
Fenris can defensibly project **when a host-write endurance baseline would be consumed if the observed usage habit continues**. It cannot predict SSD failure.
Use lifetime **Data Units Written (DUW)** as the write counter, a model-and-capacity-specific **rated-TBW override** as the preferred baseline, and wall-clock observation history as the rate denominator. Keep **Percentage Used** as a separate manufacturer wear signal; only use it for a coarse, explicitly labeled implied baseline when no rated baseline exists. Treat **Power On Hours** as context, not elapsed calendar time. Express projection confidence as categorical evidence backed by visible facts, never as an accuracy percentage.
This decision removes two unsupported assumptions in the current calculation: inferring endurance as `capacity × 600` when Percentage Used is zero, and presenting `DUW / Percentage Used` as non-estimated endurance ([current Fenris calculation](https://git.bongbetic.com/xavierk/Fenris/src/commit/91db519148a6359f6968b758a76ce30b480aaae0/fenris.py#L261-L270)).
## What the NVMe signals support
| Signal | Standard semantics and precision | Defensible Fenris use |
|---|---|---|
| **Percentage Used** | An unsigned one-byte, vendor-specific estimate based on actual use and the manufacturer's prediction of NVM life. `100` means estimated endurance consumed but may not mean subsystem failure; values may exceed 100, and values above 254 are represented as 255. It is updated once per power-on hour while the controller is not asleep ([NVM Express Base Specification 2.0e, SMART / Health Information](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); [official libnvme field documentation](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)). | Preserve the raw integer. Do not clamp at 100; render 255 as `≥255%`, not an exact value. Do not call 100 a failure point. Zero is too coarse to establish zero wear or infer a baseline. |
| **Data Units Written** | A 128-bit cumulative count of host-written 512-byte data units, excluding metadata, reported in thousands and rounded upward. For the NVM command set, Write logical blocks count; Write Uncorrectable and Write Zeroes do not. Zero means the counter is not reported ([official libnvme structure and semantics](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L15-L23), [field definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)). | Store the raw integer and derive `reported_host_bytes = DUW × 512,000`. Call it **reported host writes**, not physical NAND writes or exact bytes. Treat zero as unsupported/ambiguous unless later positive samples prove support. |
| **Power On Hours** | Integer power-on hours; the controller may omit time powered in a non-operational power state ([official libnvme field definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L159-L162)). | Display as drive context and use changes as a diagnostic. Do not use it as exact active time, idle time, powered-off time, or the denominator of a calendar-life projection. |
The SMART / Health log describes controller-level lifetime information; Fenris should therefore bind a history segment to a stable controller identity and avoid implying filesystem-level or physical-NAND-write precision ([NVM Express Base Specification 2.0e](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
### DUW quantization
Let `q = 512,000 bytes`, raw counter `U_t`, and reported cumulative host writes `W_t = qU_t`. Since each cumulative endpoint is rounded upward, the reported interval delta is:
```text
ΔW = q(U_b - U_a)
```
Its error from endpoint quantization alone is less than one quantum: `|ΔW - actual interval writes| < 512,000 bytes`. Thus a zero hourly delta does not prove no writes below that resolution, and “exact bytes written” is not defensible. This bound follows directly from the standard's upward-rounded cumulative representation ([official libnvme DUW definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)).
## Endurance baseline precedence
Use these sources in order:
1. **Verified rated-TBW override** for the exact manufacturer, model, and capacity, with source URL and document revision.
2. **Unverified manual override**, visibly labeled as user-supplied.
3. **Implied endurance from Percentage Used**, visibly labeled as a coarse heuristic.
4. Otherwise, **projection unavailable**. Never synthesize TBW from capacity alone.
Manufacturer TBW values are model- and capacity-specific: Samsung, for example, rates the 1 TB 990 PRO at 600 TBW and the 2 TB model at 1,200 TBW, and states that its warranty is limited by the stated period or TBW, whichever comes first ([Samsung 990 PRO data sheet, pp. 3–4](https://download.semiconductor.samsung.com/resources/data-sheet/Samsung_NVMe_SSD_990_PRO_Datasheet_Rev.1.0.pdf)). Samsung's warranty treats crossing TBW as a warranty-limit condition, not as a predicted failure event ([Samsung SSD Limited Warranty, sections A–B](https://download.semiconductor.samsung.com/resources/warranty/SAMSUNG_SSD_Limited_Warranty_English_US.pdf)). Fenris must therefore call rated TBW an endurance/warranty baseline rather than physical end of life.
Store an override as bytes plus provenance. If the input is labeled TBW, define it explicitly as decimal terabytes:
```text
E_rated = entered_TBW × 10^12 bytes
R_rated = max(E_rated - W_t, 0)
```
Keep rated-budget consumption and manufacturer Percentage Used separate; disagreement is useful evidence, not a reason to blend them into a fabricated wear percentage.
### Implied endurance constraints
Only for `1 ≤ p ≤ 254`:
```text
E_implied = 100W_t / p
R_implied = max(E_implied - W_t, 0)
```
This assumes the vendor's Percentage Used estimate is proportional to host writes, which NVMe does **not** require: the field is explicitly vendor-specific and based on the manufacturer's life prediction ([NVM Express Base Specification 2.0e](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); [libnvme documentation](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)). Do not compute it for 0 or saturated 255. Label it **implied from vendor wear estimate**, show few significant digits, and do not promote it until multiple wear increments make the estimate less dominated by one-percentage-point quantization. At low values it is intrinsically unstable: changing `p` from 1 to 2 halves the result.
## Usage-adjusted projection
For a selected valid wall-clock history interval from `a` to `b`:
```text
rate = q(U_b - U_a) / elapsed_wall_clock_seconds
projected_seconds = remaining_baseline_bytes / rate (rate > 0)
```
If the rate is zero, report **no finite projection from this history**, not infinity. Required wording should be equivalent to:
> Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
Use wall-clock elapsed time because powered-off and idle periods are part of the observed usage habit, whereas Power On Hours may exclude non-operational powered states ([libnvme Power On Hours definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L159-L162)). Deliberately disabled monitoring must be excluded or marked unknown by lifecycle records; a SMART counter pair can recover aggregate writes across a collector gap but cannot reveal when within that gap the writes occurred.
## History and changing habits
Persist interval observations rather than pretending every delta belongs to a clock-hour bucket:
- Compute deltas only within one controller-identity segment and only when DUW is monotonic. A decrease is a segment boundary or data fault, never a delta to clamp to zero.
- Preserve both endpoints, elapsed wall time, counter delta, and gap/monitoring state. Writes across a gap cannot be assigned exactly to individual hours; any proportional allocation must be labeled estimated.
- Use daily aggregates for habit evidence and retain hourly intervals for display/diagnostics. Hourly samples are time-dependent, so raw sample count is not independent evidence; NIST warns that autocorrelation can invalidate standard `s/√N` uncertainty calculations and other statistical conclusions ([NIST Autocorrelation Plot](https://www.itl.nist.gov/div898/handbook/eda/section3/eda331.htm)).
- Compare descriptive recent, medium, and longer horizons (for example 7, 28, and 90 days) and expose their projection spread as a **scenario range**, not a statistical confidence interval.
- Flag a changing habit when recent and earlier daily-rate windows diverge materially for a sustained period. Prefer the recent sustained regime for the headline projection while retaining the older regime as comparison. This is necessary because a stationary time series has stable mean, variance, and autocorrelation structure; trend, changing variance, and seasonality violate that assumption ([NIST Stationarity](https://www.itl.nist.gov/div898/handbook/pmc/section4/pmc442.htm)).
- Before treating days as interchangeable, account for weekly or other periodic patterns; NIST describes seasonality as regular periodic behavior that must be addressed in a time-series model ([NIST Seasonality](https://www.itl.nist.gov/div898/handbook/pmc/section4/pmc443.htm)).
Exact horizon lengths and change thresholds are product guardrails to validate later, not statistically guaranteed constants.
## Projection confidence
Use four categorical states:
- **Unavailable** — no applicable baseline, unsupported DUW, identity/counter discontinuity, or no positive usable rate.
- **Warming up** — too little history to represent ordinary usage cycles.
- **Limited evidence** — implied or unverified baseline, short/incomplete history, substantial horizon spread, stale observations, or a recent habit change.
- **Supported evidence** — verified baseline, multiple representative usage cycles, good interval coverage, current observations, and stable rates across relevant horizons.
Always show the contributing facts, for example:
> Supported evidence · verified manufacturer TBW · 42 calendar days · 96% interval coverage · 6 weekly cycles · recent and 28-day rates agree
Do not display “82% confidence” or “95% accurate.” NIST defines confidence level through the long-run coverage of an interval procedure, not as the probability that this particular estimate is correct ([NIST Confidence Limits](https://www.itl.nist.gov/div898/handbook/eda/section3/eda352.htm)). A future statistical rate interval would cover rate-estimation uncertainty only; it would not validate the endurance baseline or guarantee that habits remain unchanged.
## Required disclosures
1. This is an endurance projection, not a predicted hardware-failure date.
2. Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated ([NVMe definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)).
3. Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold ([Samsung warranty](https://download.semiconductor.samsung.com/resources/warranty/SAMSUNG_SSD_Limited_Warranty_English_US.pdf)).
4. DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes ([NVMe definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)).
5. Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
6. Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
## Newly surfaced questions
Carry these to the next Wayfinder session rather than expanding this ticket:
- Should rated-budget and manufacturer Percentage Used projections appear side by side when they disagree?
- Which provenance fields are mandatory for a TBW override: URL, revision, model, capacity, region, and entry date?
- What exact warming-up, coverage, horizon, and changing-habit thresholds should the projection-model specification adopt?
- How should powered-off periods, deliberately disabled monitoring, and unexplained gaps be represented separately?
- Which stable controller identity prevents observation history from crossing a drive replacement?
- Should a detected recent regime automatically replace the long-term rate or require acknowledgement?
- Should the TUI expose multi-horizon scenarios only, or also a model-based statistical rate interval?
- How should unsupported DUW, saturated Percentage Used, and counter discontinuities appear in the TUI?
@@ -1,259 +0,0 @@
# Fenris: systemd lifecycle and privilege constraints
**Ticket:** “Verify systemd lifecycle and privilege constraints”
**Status:** Research and planning only; no product implementation is included
**Research date:** 2026-08-31
## Executive recommendation
Run collection as a **system timer plus a short-lived system service**, not as a user service and not as the TUI's child process. Enable `fenris-collect.timer` at installation so PID 1 schedules a collection shortly after every boot and thereafter at the configured interval. Keep the interactive TUI an ordinary, on-demand, unprivileged process.
The collection service should invoke an absolute, administrator-owned `smartctl` binary directly—never `sudo`—and should have no listener or TUI code. It should read root-owned configuration from `/etc/fenris/`, use `/var/lib/fenris/` for durable history, use `/run/fenris/` only for ephemeral status/locking, and log to the journal. `StateDirectory=` and `RuntimeDirectory=` create and lifecycle-manage those standard locations and add the mount dependencies needed to reach them; state directories persist after service stop, while runtime directories normally do not. [systemd.exec(5), directory options](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=)
The TUI should:
1. inspect a deliberately small, non-secret status surface without elevation;
2. display **runtime state**, **boot enablement**, **last sample outcome**, and **freshness** separately;
3. offer only explicit “Enable collection at boot” and “Disable collection at boot” actions, with confirmation; and
4. ask systemd to make that change, allowing the platform's normal polkit authentication to occur.
Do **not** install a permissive polkit rule granting `org.freedesktop.systemd1.manage-unit-files` to a Fenris group. systemd uses that action for enable/disable/mask/preset and related unit-file operations generally, and its current unit-file authorization check supplies no unit detail with which a rule could safely limit authorization to Fenris. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security); [systemd `dbus-util.c`, pinned source: `manage-unit-files` has `details = NULL`](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L209-L223). If passwordless delegated startup control becomes a requirement, add a purpose-built, root-owned helper exposing only the two fixed Fenris operations and authorize that helper with a Fenris-specific polkit action; do not grant the generic systemd action.
## Product context observed in this repository
The current prototype combines sampling, persistence, HTTP serving, process daemonization, PID-file management, status, and control in [`fenris.py`](../../fenris.py). It starts a detached Python process itself, stores history/PID/log files under the checkout's `data/`, binds the dashboard to `0.0.0.0`, and runs `sudo -n smartctl -a -j DEVICE`. [`fenris.sh`](../../fenris.sh) starts/stops that process and currently recommends a passwordless sudoers entry for `/usr/sbin/smartctl` without argument constraints. The README describes the intended five-minute continuous collection and on-demand menu/dashboard.
That prototype shape is unsuitable for a boot-persistent privileged installation: a privileged process would also contain the HTTP server and large dashboard surface, checkout-relative state has no system ownership boundary, PID files duplicate service-manager state, and the broad sudoers example permits more than Fenris's read-only query. The recommendation below separates those concerns rather than wrapping the existing `start` command in a unit.
## Verified constraints
| Area | Verified constraint | Design consequence |
|---|---|---|
| System vs. user manager | A non-root user service cannot switch to another identity with `User=`; system services default to root and may select another user. [systemd.exec(5), `User=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#User=) | A normal user service is not a reliable privilege boundary for SMART access. Collection belongs in the system manager. |
| User-service persistence | User lingering causes that user's manager to be spawned at boot and kept after logout; without this extra policy, a user manager is session-oriented. [loginctl(1), `enable-linger`](https://www.freedesktop.org/software/systemd/man/latest/loginctl.html#enable-linger%20USER%E2%80%A6) | A user unit either fails the reboot/no-login requirement or requires lingering while still not solving device privilege. Reject it for the collector. |
| Boot enablement | `[Install]` directives do not execute at runtime; `enable` materializes them as symlinks. `WantedBy=` creates a `.wants/` link, and `timers.target` is the recommended boot target for application timers. [systemd.unit(5), `[Install]`](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#%5BInstall%5D%20Section%20Options); [systemd.special(7), `timers.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#timers.target) | Install with `WantedBy=timers.target`, then explicitly enable the timer. Merely shipping the files does not enable startup. |
| Enabled vs. running | `systemctl enable` does not start a unit, and `disable` does not stop it; `--now` couples those otherwise separate changes. [systemctl(1), `enable`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#enable%20UNIT%E2%80%A6); [systemctl(1), `disable`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#disable%20UNIT%E2%80%A6); [systemctl(1), `--now`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#--now) | The TUI must label startup and immediate runtime effects separately. Default startup actions should not silently use `--now`. |
| Timer recurrence | Combining `OnBootSec=` with a relative trigger provides post-boot and recurring activation. Timer expiry is coalesced within `AccuracySec=` (one minute by default). [systemd.timer(5), monotonic timers](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#OnActiveSec=); [systemd.timer(5), `AccuracySec=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#AccuracySec=) | Use `OnBootSec=` plus `OnUnitInactiveSec=` (or `OnUnitActiveSec=` after overlap behavior is tested), and choose accuracy deliberately. Do not promise exact-second polling. |
| Missed runs | `Persistent=` records the last trigger and can fire when an inactive timer is reactivated; its documented clock behavior must be considered with monotonic timers. [systemd.timer(5), `Persistent=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#Persistent=) | Fenris reads lifetime counters, so replaying every missed five-minute sample is neither possible nor useful. One prompt post-boot sample is sufficient; decide whether suspend catch-up warrants `Persistent=yes`. |
| Boot ordering | `timers.target` exists to activate timers after boot. `StateDirectory=`/`RuntimeDirectory=` automatically add `Requires=` and `After=` dependencies for mounts needed by their paths. `network-online.target` is for consumers that strictly require configured networking. [systemd.special(7), `timers.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#timers.target); [systemd.exec(5), implicit dependencies](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#Implicit%20Dependencies); [systemd.special(7), `network-online.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#network-online.target) | Do not order local SMART collection after the network. Let managed data directories establish filesystem ordering. Treat a not-yet-present NVMe device as a failed sample retried at the next interval unless a fixed device-unit dependency is proven necessary. |
| `smartctl` Linux path | smartmontools opens a Linux NVMe device read-only and sends `NVME_IOCTL_ADMIN_CMD`; failure is returned from the ioctl. [smartmontools `os_linux.cpp`, pinned source](https://github.com/smartmontools/smartmontools/blob/618fcaede4478bc7d17fa2a8db5fd18af3744e20/lib/os_linux.cpp#L2817-L2861) | Ordinary file read permission alone does not prove the admin ioctl will be authorized. Validate the shipped unit on every supported kernel/device transport. Root execution is the robust initial compatibility choice. |
| Capability substitution | `CAP_SYS_ADMIN` is intentionally overloaded and is described as plausibly “the new root”; Linux man-pages explicitly advise avoiding it where possible. [capabilities(7), `CAP_SYS_ADMIN`](https://man7.org/linux/man-pages/man7/capabilities.7.html#CAP_SYS_ADMIN); [capabilities(7), developer notes](https://man7.org/linux/man-pages/man7/capabilities.7.html#NOTES) | Do not move `CAP_SYS_ADMIN` into the long-lived TUI or combined web process merely to avoid UID 0. A short-lived, sandboxed root collector has a smaller practical exposure. A minimal capability set remains a test item, not an assumption. |
| Device sandboxing | `PrivateDevices=yes` supplies only pseudo-devices and excludes physical devices. [systemd.exec(5), `PrivateDevices=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#PrivateDevices=) | The collector must not set `PrivateDevices=yes`. Prefer a device cgroup allow-list for the configured node where supported and tested. The TUI/dashboard can use `PrivateDevices=yes`. |
| sudo non-interactivity | `sudo -n` never prompts and fails when authentication is required. [sudo(8), pinned upstream source](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudo.man.in#L629-L635) | `sudo -n` prevents a daemon hang but does not create authorization. It is unnecessary inside a root system service and gives poor interactive UX in the TUI. |
| sudoers command scope | If a sudoers command omits arguments, the user may supply any arguments; argument wildcards require care. `NOPASSWD` removes authentication for matching entries. [sudoers(5), pinned upstream source](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L1107-L1129); [sudoers(5), `PASSWD`/`NOPASSWD`](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L2063-L2089) | The README's path-only `NOPASSWD: /usr/sbin/smartctl` rule is too broad. Do not retain it as the installed architecture. If sudo is retained as a fallback, use a root-owned fixed-argument helper, not user-controlled smartctl arguments. |
| systemd authorization | Read access to systemd's D-Bus objects is generally available; state-changing unit operations require `manage-units`, while enablement operations require `manage-unit-files`. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security) | Status needs no blanket elevation. Starting/stopping and enabling/disabling are separate privileged action classes. |
| polkit rules | polkit loads JavaScript rules from `/etc/polkit-1/rules.d` and `/usr/share/polkit-1/rules.d`; a rule may return `AUTH_ADMIN`, while `AUTH_ADMIN_KEEP` caches authorization briefly for the same action/subject. [polkit(8), authorization rules](https://polkit.pages.freedesktop.org/polkit/polkit.8.html#AUTHORIZATION-RULES) | Prefer normal admin authentication and avoid cached authorization for a sensitive toggle. A custom delegated helper needs its own narrow action and rule. |
| Status parsing | `systemctl status` is human-readable and may include recent journal lines; `systemctl show` is the computer-parsable interface and supports selecting properties. [systemctl(1), `status`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#status%20PATTERN%E2%80%A6); [systemctl(1), `show`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#show%20PATTERN%E2%80%A6) | Never scrape `status` in the TUI. Query an explicit property allow-list and do not echo arbitrary logs in the default status screen. |
| Durable and ephemeral paths | systemd maps `StateDirectory=` to `/var/lib/` and `RuntimeDirectory=` to `/run/`; state/config/log directories remain after stop, while runtime directories are normally removed. [systemd.exec(5), directory table and lifecycle](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=) | Durable samples belong in `/var/lib/fenris`; locks/sockets belong in `/run/fenris`. A checkout-relative `data/` directory and PID file should not survive migration. |
## Proposed lifecycle and control boundary
```text
boot
└─ system systemd
└─ enabled fenris-collect.timer (unprivileged users may inspect)
└─ periodically activates fenris-collect.service
├─ short-lived privileged collector only
├─ reads /etc/fenris/fenris.conf
├─ runs /usr/sbin/smartctl with fixed read-only query arguments
├─ validates and sanitizes JSON
└─ atomically updates /var/lib/fenris/{history,latest,status}
interactive login (independent of collection)
└─ fenris TUI, ordinary invoking user
├─ reads safe status and permitted data
├─ queries selected systemd read-only properties
└─ on explicit confirmed request only:
└─ systemctl enable|disable fenris-collect.timer
└─ system bus → polkit admin authentication → PID 1
```
### Boundary rules
- **Collector:** may access the configured block/controller device and write only Fenris state. It must not bind a network socket, render a UI, edit its configuration, alter unit enablement, or run caller-supplied commands.
- **TUI/dashboard:** may read sanitized state. It must not inherit collector privilege. If a dashboard remains, run it as a separate unprivileged process bound to loopback by default; the current `0.0.0.0` binding must not become part of the privileged unit.
- **systemd/PID 1:** owns process lifecycle and boot enablement. Remove application daemonization, kill-by-PID, and checkout PID files. A service process should remain in the foreground; systemd's service model assumes the started process remains until termination unless a forking type is deliberately used. [systemd.service(5), examples](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Examples)
- **Administrator:** owns installation, `/etc/fenris`, unit files, authorization policy, membership in any read group, and startup-state changes.
A timer/oneshot split is preferable to a continuously privileged daemon because Fenris samples cumulative device counters and needs privilege only during a sample. If later requirements demand a live HTTP server, keep it in a distinct unprivileged service rather than extending collector lifetime.
## Unit sketch (planning only)
The following is intentionally illustrative. Paths, device cgroup syntax, hardening compatibility, and interval behavior must be validated on supported distributions before shipping.
```ini
# /usr/lib/systemd/system/fenris-collect.service
[Unit]
Description=Fenris NVMe SMART sample collector
Documentation=man:smartctl(8)
[Service]
Type=oneshot
ExecStart=/usr/libexec/fenris/fenris-collect --config /etc/fenris/fenris.conf
# Packaging creates the non-privileged read group; the process remains UID 0.
Group=fenris-readers
StateDirectory=fenris
StateDirectoryMode=0750
RuntimeDirectory=fenris
RuntimeDirectoryMode=0750
UMask=0027
# Hardening candidates; validate with smartctl on every supported transport.
NoNewPrivileges=yes
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=no
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectControlGroups=yes
RestrictSUIDSGID=yes
LockPersonality=yes
RestrictAddressFamilies=AF_UNIX
# DevicePolicy=closed
# DeviceAllow=/dev/nvme0 r
# CapabilityBoundingSet=... # unresolved; do not guess CAP_SYS_ADMIN-only portability
```
`ProtectSystem=strict` makes the hierarchy read-only except API filesystems, while managed state/log directories are excluded so they remain writable; systemd recommends the protection for long-running services, and it is still useful defense-in-depth for this short-lived one. [systemd.exec(5), `ProtectSystem=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#ProtectSystem=) `NoNewPrivileges=yes` prevents this process and its descendants from gaining new privilege through `execve` mechanisms such as set-user-ID bits or file capabilities. [systemd.exec(5), `NoNewPrivileges=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#NoNewPrivileges=)
```ini
# /usr/lib/systemd/system/fenris-collect.timer
[Unit]
Description=Periodically collect Fenris NVMe SMART samples
[Timer]
OnBootSec=2min
OnUnitInactiveSec=5min
AccuracySec=30s
Unit=fenris-collect.service
[Install]
WantedBy=timers.target
```
`OnUnitInactiveSec=` measures from deactivation, which avoids overlapping a slow oneshot at the cost of interval drift; `OnUnitActiveSec=` measures from activation. Those distinct bases are defined by `systemd.timer(5)`. [systemd.timer(5), monotonic timer table](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#OnActiveSec=) Choose between them after measuring collection duration and confirming the desired interval semantics.
Do not add `After=network-online.target`; collection is local. Do not add `Requires=/dev/...` as pseudo-syntax. If strict device binding is needed, use the escaped `.device` unit generated for the configured node only after testing replacement/hotplug behavior; otherwise record a bounded failure and let the next timer activation retry.
## Authorization and TUI control sketch (planning only)
### Default: systemctl plus normal polkit authentication
Read path:
```text
systemctl show fenris-collect.service \
--property=LoadState,ActiveState,SubState,Result,ExecMainStatus
systemctl show fenris-collect.timer \
--property=LoadState,ActiveState,SubState,UnitFileState,NextElapseUSecRealtime,NextElapseUSecMonotonic
```
Mutation path, only after a confirmation screen naming the exact effect:
```text
Enable at future boots: systemctl enable fenris-collect.timer
Disable at future boots: systemctl disable fenris-collect.timer
```
Do not silently append `--now`. “Disable at boot” does not mean “cancel a currently executing sample,” and systemctl explicitly keeps enablement separate from start/stop unless `--now` is requested. [systemctl(1), `--now`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#--now)
The TUI must pass a fixed unit name and fixed verb without a shell. On cancellation, authentication failure, timeout, or non-zero exit, it should report no successful change and re-read `UnitFileState`; it must not infer success from the requested action.
### Why not a broad polkit group rule
`manage-units` calls can include `unit` and `verb` details in current systemd source, but the `manage-unit-files` check currently has no details. [systemd `dbus-util.c`, pinned source](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L160-L223) Therefore a rule such as “members of `fenris` may perform `org.freedesktop.systemd1.manage-unit-files`” would delegate generic unit enablement/masking operations, not only Fenris startup state. That is outside the required boundary.
### If delegated passwordless control is later required
Use a small root-owned D-Bus/helper mechanism with a Fenris-specific action, for example `com.bongbetic.fenris.manage-startup`. Its API should accept only an enum `{enable, disable}`, internally target the constant `fenris-collect.timer`, reject options/paths/extra units, perform the unit-file operation, and return the observed resulting state. A polkit rule may then grant that custom action to a designated local group, optionally requiring `subject.local && subject.active`; polkit exposes subject locality/activity and action details to JavaScript rules. [polkit(8), authorization rules and `Subject`](https://polkit.pages.freedesktop.org/polkit/polkit.8.html#AUTHORIZATION-RULES)
A sudo fallback should follow the same fixed helper design. Do not authorize `/usr/bin/systemctl` or `/usr/sbin/smartctl` without exact argument control: sudoers explicitly permits arbitrary arguments when none are specified. [sudoers(5), command arguments](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L1107-L1129)
## Ownership and path choices
| Path | Proposed owner/mode | Purpose and rationale |
|---|---|---|
| `/usr/libexec/fenris/fenris-collect` (or distribution-equivalent) | `root:root`, `0755`, package-managed | Privileged entry point must not be writable by TUI users or the service's read group. Use an absolute `ExecStart`; systemd does not provide shell syntax by default. [systemd.service(5), command lines](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Command%20lines) |
| `/usr/lib/systemd/system/fenris-collect.{service,timer}` | `root:root`, `0644`, package-managed | Vendor unit definitions. Local administrator overrides belong under `/etc/systemd/system/`, which has higher load-path precedence. [systemd.unit(5), unit load path](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#Unit%20File%20Load%20Path) |
| `/etc/fenris/fenris.conf` | `root:root`, `0644` if strictly non-secret; otherwise `0640` | Persistent host configuration. The collector reads but cannot write it under `ProtectSystem=strict`. The TUI should receive only safe fields through status rather than requiring config write access. |
| `/var/lib/fenris/` | created by `StateDirectory=fenris`, `0750`; `root:fenris-readers` if direct group reads are retained | Durable history, aggregates, latest sample, and machine-readable sample status. systemd maps state directories here and leaves them after stop. [systemd.exec(5), directory table/lifecycle](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=) |
| `/run/fenris/` | created by `RuntimeDirectory=fenris`, `0750` | Ephemeral lock/socket only. Do not use a PID file as authority; systemd already tracks the process/unit. Runtime directories are removed on stop by default. [systemd.exec(5), `RuntimeDirectoryPreserve=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectoryPreserve=) |
| Journal | journal ACL/policy | Operational diagnostics. Avoid a separate root-owned `/var/log/fenris.log` unless retention/export requirements demand it. Do not expose arbitrary journal messages through safe status. |
These locations also match FHS semantics: `/etc` holds host-specific configuration, `/run` is run-time variable data cleared at boot, and `/var/lib` holds application state that persists across restarts. [FHS 3.0, `/etc`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch03s07.html); [FHS 3.0, `/run`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch03s15.html); [FHS 3.0, `/var/lib`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch05s08.html)
Prefer a narrow read-only IPC/status endpoint over group-readable raw files if multi-user confidentiality matters. The sketch instead uses a dedicated `fenris-readers` primary group for the UID-0 collector so managed directories become `root:fenris-readers`; packaging must create that group, and only approved users should join it. Do not reuse the collector's privileged identity as a user-facing authorization group. Use atomic replace for `latest.json`/`status.json`, append safely for history, set a restrictive umask, and omit serial numbers, command lines, environment, and raw stderr from the shared surface.
## Safe status design
The default TUI status should be an allow-listed composition, not a dump of `systemctl status`, journal output, raw smartctl JSON, configuration, or process command lines.
Recommended fields:
```text
Installation: loaded | not-installed | error
Startup: enabled | disabled | static | masked | unknown
Scheduler runtime: active | inactive | failed
Collection runtime: active | inactive | failed
Last attempt: RFC3339 timestamp
Last success: RFC3339 timestamp
Freshness: fresh | stale | never (show threshold and age)
Last result: success | device-unavailable | permission-denied |
timeout | invalid-output | storage-error | internal-error
Next scheduled: timestamp if systemd reports one
Samples/history: count and range, if readable
Device: configured stable identifier or node; no serial by default
```
Design requirements:
- Treat `UnitFileState` (startup policy), `ActiveState`/`SubState` (runtime), and last sample result as independent axes. `is-enabled` documents multiple states—including enabled, disabled, static, indirect, generated, transient, and masked—so a Boolean loses actionable information. [systemctl(1), `is-enabled` state table](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#is-enabled%20UNIT%E2%80%A6)
- Parse only `systemctl show --property=...` or the equivalent D-Bus properties. `status` is explicitly human-oriented and includes journal data. [systemctl(1), `show`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#show%20PATTERN%E2%80%A6)
- Compute freshness from `last_success`, not merely timer activity. An active timer can coexist with repeated collection failures.
- Publish bounded error categories and a short administrator hint; keep raw smartctl stderr and tracebacks in the journal. This prevents device identifiers, paths, malformed device output, or command details from crossing the read boundary.
- Never report “enabled” immediately after a requested mutation without re-querying the authoritative state. Never equate “disabled” with “stopped.”
- If status data is unreadable, say `permission-denied`/`unknown`; do not elevate merely to render the screen.
## Alternatives and trade-offs
### Long-running system service
A foreground `fenris-collector.service` with `Restart=on-failure` and `WantedBy=multi-user.target` also satisfies reboot persistence; `Restart=` controls automatic restart after process failure, while `WantedBy=multi-user.target` is the standard installation relationship for a multi-user service. [systemd.service(5), `Restart=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Restart=); [systemd.special(7), `multi-user.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#multi-user.target)
Trade-off: it preserves today's loop model and precise in-process scheduling, but leaves a privileged Python process resident continuously and requires restart/backoff handling. Use it only if future collection needs persistent in-memory state that cannot be reconstructed from durable samples.
### Unprivileged daemon plus privileged smartctl helper
This can reduce privileged code if the helper accepts no uncontrolled path or options, validates a configured device allow-list, emits bounded sanitized output, and exits. It adds an IPC/protocol and another authorization surface. Giving the whole Python process `CAP_SYS_ADMIN` is not an equivalent reduction because that capability is exceptionally broad. [capabilities(7), notes](https://man7.org/linux/man-pages/man7/capabilities.7.html#NOTES)
### User service with lingering
Lingering can run a user manager from boot through logout, but the user manager cannot switch an ordinary user's unit to root and some system-service sandboxing features are unavailable in user services. [loginctl(1), lingering](https://www.freedesktop.org/software/systemd/man/latest/loginctl.html#enable-linger%20USER%E2%80%A6); [systemd.exec(5), user-service sandboxing limitations](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#Sandboxing) This adds account coupling and still needs a separate privilege mechanism. Reject for collection; it remains acceptable for an optional per-user dashboard client.
### Root daemon calling `sudo -n smartctl`
This is redundant: root already crosses the privilege boundary, while sudo introduces policy/path/configuration failure modes. For an unprivileged daemon, `sudo -n` avoids blocking but only works with pre-authorized sudoers policy and produces no authentication opportunity. [sudo(8), `--non-interactive`](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudo.man.in#L629-L635) Reject as the installed systemd architecture.
### Direct broad polkit delegation
Convenient but unsafe for startup-state delegation: `manage-unit-files` covers generic enable/disable/mask/preset operations and currently lacks unit-scoping details in systemd's authorization call. Use normal administrator authentication or a custom narrow action instead. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security); [pinned systemd authorization source](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L209-L223)
## Unresolved questions and required validation
1. **Supported distributions/systemd floor:** What is the minimum systemd version? Confirm every selected hardening and directory directive exists there; unknown directives can weaken the intended sandbox.
2. **Device identity:** Is configuration a mutable `/dev/nvmeN` node, namespace node, controller, `/dev/disk/by-id` link, or discovered set? smartctl documents Linux NVMe controller and namespace forms, but the stable product identity and hotplug behavior remain a product decision. [smartctl(8), pinned source](https://github.com/smartmontools/smartmontools/blob/618fcaede4478bc7d17fa2a8db5fd18af3744e20/src/smartctl.8.in#L72-L86)
3. **Privilege matrix:** On each supported kernel, packaging of smartmontools, NVMe/SATA/USB bridge, and device permission setup, record which open/ioctl fails as an unprivileged service and which exact capability/device allow-list is sufficient. Do not generalize an NVMe result to every smartmontools transport.
4. **Root vs. reduced-capability collector:** After that matrix exists, decide whether a non-root static service user plus a minimal capability/device set works portably. Reject any result that requires putting broad capability into the TUI/dashboard.
5. **Timer semantics:** Should five minutes be measured from sample start or completion? What maximum runtime and timeout are acceptable? Should resume from suspend trigger immediately? Decide `OnUnitActiveSec` vs. `OnUnitInactiveSec`, `Persistent=`, and `AccuracySec` from those answers.
6. **Startup toggle semantics:** Does “disable startup” leave the timer active until reboot, or should the TUI offer a separate, clearly labeled “disable and stop now”? The systemctl semantics intentionally separate these effects. [systemctl(1), enable/disable](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#enable%20UNIT%E2%80%A6)
7. **Who may read health history:** All local users, a `fenris-readers` group, or only an authenticated local client? This determines state modes and whether a read-only Unix socket is preferable to files.
8. **Dashboard scope:** Is the HTTP dashboard retained, and if so must it be local-only or remotely accessible? Remote access needs a separate threat model, authentication, transport security, and unprivileged service; it must not enlarge the collector boundary.
9. **Packaging paths:** Confirm `/usr/libexec` and `/usr/lib/systemd/system` equivalents per target distribution, the absolute smartctl path, and whether administrator overrides use environment files or a dedicated validated config format.
10. **Data durability:** Define atomicity, fsync policy, retention, corruption recovery, migration from checkout-relative `data/`, and behavior on read-only/full filesystems.
11. **Authorization UX:** Is ordinary admin authentication acceptable? If not, specify exactly which local principals may toggle startup and commission the custom helper/action rather than broad `manage-unit-files` delegation.
12. **Status confidentiality:** Decide whether model name, device path, capacity, wear, temperatures, and error counters are safe for every local reader; serial numbers and raw smartctl output should remain excluded by default.
## Decision summary
The verified boundary is: **system systemd owns scheduling and boot state; a short-lived privileged collector owns only device interrogation and state writes; an unprivileged on-demand TUI owns presentation; polkit/admin authentication mediates explicit startup changes.** This satisfies collection across reboot without tying it to login, avoids embedding sudo in unattended code, keeps generic service control out of the TUI's ambient privilege, and makes startup state observable and alterable without conflating it with current execution.