Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
4394d0c892 |
@@ -0,0 +1,259 @@
|
|||||||
|
# Fenris: systemd lifecycle and privilege constraints
|
||||||
|
|
||||||
|
**Ticket:** “Verify systemd lifecycle and privilege constraints”
|
||||||
|
**Status:** Research and planning only; no product implementation is included
|
||||||
|
**Research date:** 2026-08-31
|
||||||
|
|
||||||
|
## Executive recommendation
|
||||||
|
|
||||||
|
Run collection as a **system timer plus a short-lived system service**, not as a user service and not as the TUI's child process. Enable `fenris-collect.timer` at installation so PID 1 schedules a collection shortly after every boot and thereafter at the configured interval. Keep the interactive TUI an ordinary, on-demand, unprivileged process.
|
||||||
|
|
||||||
|
The collection service should invoke an absolute, administrator-owned `smartctl` binary directly—never `sudo`—and should have no listener or TUI code. It should read root-owned configuration from `/etc/fenris/`, use `/var/lib/fenris/` for durable history, use `/run/fenris/` only for ephemeral status/locking, and log to the journal. `StateDirectory=` and `RuntimeDirectory=` create and lifecycle-manage those standard locations and add the mount dependencies needed to reach them; state directories persist after service stop, while runtime directories normally do not. [systemd.exec(5), directory options](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=)
|
||||||
|
|
||||||
|
The TUI should:
|
||||||
|
|
||||||
|
1. inspect a deliberately small, non-secret status surface without elevation;
|
||||||
|
2. display **runtime state**, **boot enablement**, **last sample outcome**, and **freshness** separately;
|
||||||
|
3. offer only explicit “Enable collection at boot” and “Disable collection at boot” actions, with confirmation; and
|
||||||
|
4. ask systemd to make that change, allowing the platform's normal polkit authentication to occur.
|
||||||
|
|
||||||
|
Do **not** install a permissive polkit rule granting `org.freedesktop.systemd1.manage-unit-files` to a Fenris group. systemd uses that action for enable/disable/mask/preset and related unit-file operations generally, and its current unit-file authorization check supplies no unit detail with which a rule could safely limit authorization to Fenris. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security); [systemd `dbus-util.c`, pinned source: `manage-unit-files` has `details = NULL`](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L209-L223). If passwordless delegated startup control becomes a requirement, add a purpose-built, root-owned helper exposing only the two fixed Fenris operations and authorize that helper with a Fenris-specific polkit action; do not grant the generic systemd action.
|
||||||
|
|
||||||
|
## Product context observed in this repository
|
||||||
|
|
||||||
|
The current prototype combines sampling, persistence, HTTP serving, process daemonization, PID-file management, status, and control in [`fenris.py`](../../fenris.py). It starts a detached Python process itself, stores history/PID/log files under the checkout's `data/`, binds the dashboard to `0.0.0.0`, and runs `sudo -n smartctl -a -j DEVICE`. [`fenris.sh`](../../fenris.sh) starts/stops that process and currently recommends a passwordless sudoers entry for `/usr/sbin/smartctl` without argument constraints. The README describes the intended five-minute continuous collection and on-demand menu/dashboard.
|
||||||
|
|
||||||
|
That prototype shape is unsuitable for a boot-persistent privileged installation: a privileged process would also contain the HTTP server and large dashboard surface, checkout-relative state has no system ownership boundary, PID files duplicate service-manager state, and the broad sudoers example permits more than Fenris's read-only query. The recommendation below separates those concerns rather than wrapping the existing `start` command in a unit.
|
||||||
|
|
||||||
|
## Verified constraints
|
||||||
|
|
||||||
|
| Area | Verified constraint | Design consequence |
|
||||||
|
|---|---|---|
|
||||||
|
| System vs. user manager | A non-root user service cannot switch to another identity with `User=`; system services default to root and may select another user. [systemd.exec(5), `User=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#User=) | A normal user service is not a reliable privilege boundary for SMART access. Collection belongs in the system manager. |
|
||||||
|
| User-service persistence | User lingering causes that user's manager to be spawned at boot and kept after logout; without this extra policy, a user manager is session-oriented. [loginctl(1), `enable-linger`](https://www.freedesktop.org/software/systemd/man/latest/loginctl.html#enable-linger%20USER%E2%80%A6) | A user unit either fails the reboot/no-login requirement or requires lingering while still not solving device privilege. Reject it for the collector. |
|
||||||
|
| Boot enablement | `[Install]` directives do not execute at runtime; `enable` materializes them as symlinks. `WantedBy=` creates a `.wants/` link, and `timers.target` is the recommended boot target for application timers. [systemd.unit(5), `[Install]`](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#%5BInstall%5D%20Section%20Options); [systemd.special(7), `timers.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#timers.target) | Install with `WantedBy=timers.target`, then explicitly enable the timer. Merely shipping the files does not enable startup. |
|
||||||
|
| Enabled vs. running | `systemctl enable` does not start a unit, and `disable` does not stop it; `--now` couples those otherwise separate changes. [systemctl(1), `enable`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#enable%20UNIT%E2%80%A6); [systemctl(1), `disable`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#disable%20UNIT%E2%80%A6); [systemctl(1), `--now`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#--now) | The TUI must label startup and immediate runtime effects separately. Default startup actions should not silently use `--now`. |
|
||||||
|
| Timer recurrence | Combining `OnBootSec=` with a relative trigger provides post-boot and recurring activation. Timer expiry is coalesced within `AccuracySec=` (one minute by default). [systemd.timer(5), monotonic timers](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#OnActiveSec=); [systemd.timer(5), `AccuracySec=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#AccuracySec=) | Use `OnBootSec=` plus `OnUnitInactiveSec=` (or `OnUnitActiveSec=` after overlap behavior is tested), and choose accuracy deliberately. Do not promise exact-second polling. |
|
||||||
|
| Missed runs | `Persistent=` records the last trigger and can fire when an inactive timer is reactivated; its documented clock behavior must be considered with monotonic timers. [systemd.timer(5), `Persistent=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#Persistent=) | Fenris reads lifetime counters, so replaying every missed five-minute sample is neither possible nor useful. One prompt post-boot sample is sufficient; decide whether suspend catch-up warrants `Persistent=yes`. |
|
||||||
|
| Boot ordering | `timers.target` exists to activate timers after boot. `StateDirectory=`/`RuntimeDirectory=` automatically add `Requires=` and `After=` dependencies for mounts needed by their paths. `network-online.target` is for consumers that strictly require configured networking. [systemd.special(7), `timers.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#timers.target); [systemd.exec(5), implicit dependencies](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#Implicit%20Dependencies); [systemd.special(7), `network-online.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#network-online.target) | Do not order local SMART collection after the network. Let managed data directories establish filesystem ordering. Treat a not-yet-present NVMe device as a failed sample retried at the next interval unless a fixed device-unit dependency is proven necessary. |
|
||||||
|
| `smartctl` Linux path | smartmontools opens a Linux NVMe device read-only and sends `NVME_IOCTL_ADMIN_CMD`; failure is returned from the ioctl. [smartmontools `os_linux.cpp`, pinned source](https://github.com/smartmontools/smartmontools/blob/618fcaede4478bc7d17fa2a8db5fd18af3744e20/lib/os_linux.cpp#L2817-L2861) | Ordinary file read permission alone does not prove the admin ioctl will be authorized. Validate the shipped unit on every supported kernel/device transport. Root execution is the robust initial compatibility choice. |
|
||||||
|
| Capability substitution | `CAP_SYS_ADMIN` is intentionally overloaded and is described as plausibly “the new root”; Linux man-pages explicitly advise avoiding it where possible. [capabilities(7), `CAP_SYS_ADMIN`](https://man7.org/linux/man-pages/man7/capabilities.7.html#CAP_SYS_ADMIN); [capabilities(7), developer notes](https://man7.org/linux/man-pages/man7/capabilities.7.html#NOTES) | Do not move `CAP_SYS_ADMIN` into the long-lived TUI or combined web process merely to avoid UID 0. A short-lived, sandboxed root collector has a smaller practical exposure. A minimal capability set remains a test item, not an assumption. |
|
||||||
|
| Device sandboxing | `PrivateDevices=yes` supplies only pseudo-devices and excludes physical devices. [systemd.exec(5), `PrivateDevices=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#PrivateDevices=) | The collector must not set `PrivateDevices=yes`. Prefer a device cgroup allow-list for the configured node where supported and tested. The TUI/dashboard can use `PrivateDevices=yes`. |
|
||||||
|
| sudo non-interactivity | `sudo -n` never prompts and fails when authentication is required. [sudo(8), pinned upstream source](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudo.man.in#L629-L635) | `sudo -n` prevents a daemon hang but does not create authorization. It is unnecessary inside a root system service and gives poor interactive UX in the TUI. |
|
||||||
|
| sudoers command scope | If a sudoers command omits arguments, the user may supply any arguments; argument wildcards require care. `NOPASSWD` removes authentication for matching entries. [sudoers(5), pinned upstream source](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L1107-L1129); [sudoers(5), `PASSWD`/`NOPASSWD`](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L2063-L2089) | The README's path-only `NOPASSWD: /usr/sbin/smartctl` rule is too broad. Do not retain it as the installed architecture. If sudo is retained as a fallback, use a root-owned fixed-argument helper, not user-controlled smartctl arguments. |
|
||||||
|
| systemd authorization | Read access to systemd's D-Bus objects is generally available; state-changing unit operations require `manage-units`, while enablement operations require `manage-unit-files`. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security) | Status needs no blanket elevation. Starting/stopping and enabling/disabling are separate privileged action classes. |
|
||||||
|
| polkit rules | polkit loads JavaScript rules from `/etc/polkit-1/rules.d` and `/usr/share/polkit-1/rules.d`; a rule may return `AUTH_ADMIN`, while `AUTH_ADMIN_KEEP` caches authorization briefly for the same action/subject. [polkit(8), authorization rules](https://polkit.pages.freedesktop.org/polkit/polkit.8.html#AUTHORIZATION-RULES) | Prefer normal admin authentication and avoid cached authorization for a sensitive toggle. A custom delegated helper needs its own narrow action and rule. |
|
||||||
|
| Status parsing | `systemctl status` is human-readable and may include recent journal lines; `systemctl show` is the computer-parsable interface and supports selecting properties. [systemctl(1), `status`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#status%20PATTERN%E2%80%A6); [systemctl(1), `show`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#show%20PATTERN%E2%80%A6) | Never scrape `status` in the TUI. Query an explicit property allow-list and do not echo arbitrary logs in the default status screen. |
|
||||||
|
| Durable and ephemeral paths | systemd maps `StateDirectory=` to `/var/lib/` and `RuntimeDirectory=` to `/run/`; state/config/log directories remain after stop, while runtime directories are normally removed. [systemd.exec(5), directory table and lifecycle](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=) | Durable samples belong in `/var/lib/fenris`; locks/sockets belong in `/run/fenris`. A checkout-relative `data/` directory and PID file should not survive migration. |
|
||||||
|
|
||||||
|
## Proposed lifecycle and control boundary
|
||||||
|
|
||||||
|
```text
|
||||||
|
boot
|
||||||
|
└─ system systemd
|
||||||
|
└─ enabled fenris-collect.timer (unprivileged users may inspect)
|
||||||
|
└─ periodically activates fenris-collect.service
|
||||||
|
├─ short-lived privileged collector only
|
||||||
|
├─ reads /etc/fenris/fenris.conf
|
||||||
|
├─ runs /usr/sbin/smartctl with fixed read-only query arguments
|
||||||
|
├─ validates and sanitizes JSON
|
||||||
|
└─ atomically updates /var/lib/fenris/{history,latest,status}
|
||||||
|
|
||||||
|
interactive login (independent of collection)
|
||||||
|
└─ fenris TUI, ordinary invoking user
|
||||||
|
├─ reads safe status and permitted data
|
||||||
|
├─ queries selected systemd read-only properties
|
||||||
|
└─ on explicit confirmed request only:
|
||||||
|
└─ systemctl enable|disable fenris-collect.timer
|
||||||
|
└─ system bus → polkit admin authentication → PID 1
|
||||||
|
```
|
||||||
|
|
||||||
|
### Boundary rules
|
||||||
|
|
||||||
|
- **Collector:** may access the configured block/controller device and write only Fenris state. It must not bind a network socket, render a UI, edit its configuration, alter unit enablement, or run caller-supplied commands.
|
||||||
|
- **TUI/dashboard:** may read sanitized state. It must not inherit collector privilege. If a dashboard remains, run it as a separate unprivileged process bound to loopback by default; the current `0.0.0.0` binding must not become part of the privileged unit.
|
||||||
|
- **systemd/PID 1:** owns process lifecycle and boot enablement. Remove application daemonization, kill-by-PID, and checkout PID files. A service process should remain in the foreground; systemd's service model assumes the started process remains until termination unless a forking type is deliberately used. [systemd.service(5), examples](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Examples)
|
||||||
|
- **Administrator:** owns installation, `/etc/fenris`, unit files, authorization policy, membership in any read group, and startup-state changes.
|
||||||
|
|
||||||
|
A timer/oneshot split is preferable to a continuously privileged daemon because Fenris samples cumulative device counters and needs privilege only during a sample. If later requirements demand a live HTTP server, keep it in a distinct unprivileged service rather than extending collector lifetime.
|
||||||
|
|
||||||
|
## Unit sketch (planning only)
|
||||||
|
|
||||||
|
The following is intentionally illustrative. Paths, device cgroup syntax, hardening compatibility, and interval behavior must be validated on supported distributions before shipping.
|
||||||
|
|
||||||
|
```ini
|
||||||
|
# /usr/lib/systemd/system/fenris-collect.service
|
||||||
|
[Unit]
|
||||||
|
Description=Fenris NVMe SMART sample collector
|
||||||
|
Documentation=man:smartctl(8)
|
||||||
|
|
||||||
|
[Service]
|
||||||
|
Type=oneshot
|
||||||
|
ExecStart=/usr/libexec/fenris/fenris-collect --config /etc/fenris/fenris.conf
|
||||||
|
# Packaging creates the non-privileged read group; the process remains UID 0.
|
||||||
|
Group=fenris-readers
|
||||||
|
StateDirectory=fenris
|
||||||
|
StateDirectoryMode=0750
|
||||||
|
RuntimeDirectory=fenris
|
||||||
|
RuntimeDirectoryMode=0750
|
||||||
|
UMask=0027
|
||||||
|
|
||||||
|
# Hardening candidates; validate with smartctl on every supported transport.
|
||||||
|
NoNewPrivileges=yes
|
||||||
|
ProtectSystem=strict
|
||||||
|
ProtectHome=yes
|
||||||
|
PrivateTmp=yes
|
||||||
|
PrivateDevices=no
|
||||||
|
ProtectKernelTunables=yes
|
||||||
|
ProtectKernelModules=yes
|
||||||
|
ProtectControlGroups=yes
|
||||||
|
RestrictSUIDSGID=yes
|
||||||
|
LockPersonality=yes
|
||||||
|
RestrictAddressFamilies=AF_UNIX
|
||||||
|
# DevicePolicy=closed
|
||||||
|
# DeviceAllow=/dev/nvme0 r
|
||||||
|
# CapabilityBoundingSet=... # unresolved; do not guess CAP_SYS_ADMIN-only portability
|
||||||
|
```
|
||||||
|
|
||||||
|
`ProtectSystem=strict` makes the hierarchy read-only except API filesystems, while managed state/log directories are excluded so they remain writable; systemd recommends the protection for long-running services, and it is still useful defense-in-depth for this short-lived one. [systemd.exec(5), `ProtectSystem=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#ProtectSystem=) `NoNewPrivileges=yes` prevents this process and its descendants from gaining new privilege through `execve` mechanisms such as set-user-ID bits or file capabilities. [systemd.exec(5), `NoNewPrivileges=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#NoNewPrivileges=)
|
||||||
|
|
||||||
|
```ini
|
||||||
|
# /usr/lib/systemd/system/fenris-collect.timer
|
||||||
|
[Unit]
|
||||||
|
Description=Periodically collect Fenris NVMe SMART samples
|
||||||
|
|
||||||
|
[Timer]
|
||||||
|
OnBootSec=2min
|
||||||
|
OnUnitInactiveSec=5min
|
||||||
|
AccuracySec=30s
|
||||||
|
Unit=fenris-collect.service
|
||||||
|
|
||||||
|
[Install]
|
||||||
|
WantedBy=timers.target
|
||||||
|
```
|
||||||
|
|
||||||
|
`OnUnitInactiveSec=` measures from deactivation, which avoids overlapping a slow oneshot at the cost of interval drift; `OnUnitActiveSec=` measures from activation. Those distinct bases are defined by `systemd.timer(5)`. [systemd.timer(5), monotonic timer table](https://www.freedesktop.org/software/systemd/man/latest/systemd.timer.html#OnActiveSec=) Choose between them after measuring collection duration and confirming the desired interval semantics.
|
||||||
|
|
||||||
|
Do not add `After=network-online.target`; collection is local. Do not add `Requires=/dev/...` as pseudo-syntax. If strict device binding is needed, use the escaped `.device` unit generated for the configured node only after testing replacement/hotplug behavior; otherwise record a bounded failure and let the next timer activation retry.
|
||||||
|
|
||||||
|
## Authorization and TUI control sketch (planning only)
|
||||||
|
|
||||||
|
### Default: systemctl plus normal polkit authentication
|
||||||
|
|
||||||
|
Read path:
|
||||||
|
|
||||||
|
```text
|
||||||
|
systemctl show fenris-collect.service \
|
||||||
|
--property=LoadState,ActiveState,SubState,Result,ExecMainStatus
|
||||||
|
systemctl show fenris-collect.timer \
|
||||||
|
--property=LoadState,ActiveState,SubState,UnitFileState,NextElapseUSecRealtime,NextElapseUSecMonotonic
|
||||||
|
```
|
||||||
|
|
||||||
|
Mutation path, only after a confirmation screen naming the exact effect:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Enable at future boots: systemctl enable fenris-collect.timer
|
||||||
|
Disable at future boots: systemctl disable fenris-collect.timer
|
||||||
|
```
|
||||||
|
|
||||||
|
Do not silently append `--now`. “Disable at boot” does not mean “cancel a currently executing sample,” and systemctl explicitly keeps enablement separate from start/stop unless `--now` is requested. [systemctl(1), `--now`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#--now)
|
||||||
|
|
||||||
|
The TUI must pass a fixed unit name and fixed verb without a shell. On cancellation, authentication failure, timeout, or non-zero exit, it should report no successful change and re-read `UnitFileState`; it must not infer success from the requested action.
|
||||||
|
|
||||||
|
### Why not a broad polkit group rule
|
||||||
|
|
||||||
|
`manage-units` calls can include `unit` and `verb` details in current systemd source, but the `manage-unit-files` check currently has no details. [systemd `dbus-util.c`, pinned source](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L160-L223) Therefore a rule such as “members of `fenris` may perform `org.freedesktop.systemd1.manage-unit-files`” would delegate generic unit enablement/masking operations, not only Fenris startup state. That is outside the required boundary.
|
||||||
|
|
||||||
|
### If delegated passwordless control is later required
|
||||||
|
|
||||||
|
Use a small root-owned D-Bus/helper mechanism with a Fenris-specific action, for example `com.bongbetic.fenris.manage-startup`. Its API should accept only an enum `{enable, disable}`, internally target the constant `fenris-collect.timer`, reject options/paths/extra units, perform the unit-file operation, and return the observed resulting state. A polkit rule may then grant that custom action to a designated local group, optionally requiring `subject.local && subject.active`; polkit exposes subject locality/activity and action details to JavaScript rules. [polkit(8), authorization rules and `Subject`](https://polkit.pages.freedesktop.org/polkit/polkit.8.html#AUTHORIZATION-RULES)
|
||||||
|
|
||||||
|
A sudo fallback should follow the same fixed helper design. Do not authorize `/usr/bin/systemctl` or `/usr/sbin/smartctl` without exact argument control: sudoers explicitly permits arbitrary arguments when none are specified. [sudoers(5), command arguments](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudoers.man.in#L1107-L1129)
|
||||||
|
|
||||||
|
## Ownership and path choices
|
||||||
|
|
||||||
|
| Path | Proposed owner/mode | Purpose and rationale |
|
||||||
|
|---|---|---|
|
||||||
|
| `/usr/libexec/fenris/fenris-collect` (or distribution-equivalent) | `root:root`, `0755`, package-managed | Privileged entry point must not be writable by TUI users or the service's read group. Use an absolute `ExecStart`; systemd does not provide shell syntax by default. [systemd.service(5), command lines](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Command%20lines) |
|
||||||
|
| `/usr/lib/systemd/system/fenris-collect.{service,timer}` | `root:root`, `0644`, package-managed | Vendor unit definitions. Local administrator overrides belong under `/etc/systemd/system/`, which has higher load-path precedence. [systemd.unit(5), unit load path](https://www.freedesktop.org/software/systemd/man/latest/systemd.unit.html#Unit%20File%20Load%20Path) |
|
||||||
|
| `/etc/fenris/fenris.conf` | `root:root`, `0644` if strictly non-secret; otherwise `0640` | Persistent host configuration. The collector reads but cannot write it under `ProtectSystem=strict`. The TUI should receive only safe fields through status rather than requiring config write access. |
|
||||||
|
| `/var/lib/fenris/` | created by `StateDirectory=fenris`, `0750`; `root:fenris-readers` if direct group reads are retained | Durable history, aggregates, latest sample, and machine-readable sample status. systemd maps state directories here and leaves them after stop. [systemd.exec(5), directory table/lifecycle](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectory=) |
|
||||||
|
| `/run/fenris/` | created by `RuntimeDirectory=fenris`, `0750` | Ephemeral lock/socket only. Do not use a PID file as authority; systemd already tracks the process/unit. Runtime directories are removed on stop by default. [systemd.exec(5), `RuntimeDirectoryPreserve=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#RuntimeDirectoryPreserve=) |
|
||||||
|
| Journal | journal ACL/policy | Operational diagnostics. Avoid a separate root-owned `/var/log/fenris.log` unless retention/export requirements demand it. Do not expose arbitrary journal messages through safe status. |
|
||||||
|
|
||||||
|
These locations also match FHS semantics: `/etc` holds host-specific configuration, `/run` is run-time variable data cleared at boot, and `/var/lib` holds application state that persists across restarts. [FHS 3.0, `/etc`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch03s07.html); [FHS 3.0, `/run`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch03s15.html); [FHS 3.0, `/var/lib`](https://refspecs.linuxfoundation.org/FHS_3.0/fhs/ch05s08.html)
|
||||||
|
|
||||||
|
Prefer a narrow read-only IPC/status endpoint over group-readable raw files if multi-user confidentiality matters. The sketch instead uses a dedicated `fenris-readers` primary group for the UID-0 collector so managed directories become `root:fenris-readers`; packaging must create that group, and only approved users should join it. Do not reuse the collector's privileged identity as a user-facing authorization group. Use atomic replace for `latest.json`/`status.json`, append safely for history, set a restrictive umask, and omit serial numbers, command lines, environment, and raw stderr from the shared surface.
|
||||||
|
|
||||||
|
## Safe status design
|
||||||
|
|
||||||
|
The default TUI status should be an allow-listed composition, not a dump of `systemctl status`, journal output, raw smartctl JSON, configuration, or process command lines.
|
||||||
|
|
||||||
|
Recommended fields:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Installation: loaded | not-installed | error
|
||||||
|
Startup: enabled | disabled | static | masked | unknown
|
||||||
|
Scheduler runtime: active | inactive | failed
|
||||||
|
Collection runtime: active | inactive | failed
|
||||||
|
Last attempt: RFC3339 timestamp
|
||||||
|
Last success: RFC3339 timestamp
|
||||||
|
Freshness: fresh | stale | never (show threshold and age)
|
||||||
|
Last result: success | device-unavailable | permission-denied |
|
||||||
|
timeout | invalid-output | storage-error | internal-error
|
||||||
|
Next scheduled: timestamp if systemd reports one
|
||||||
|
Samples/history: count and range, if readable
|
||||||
|
Device: configured stable identifier or node; no serial by default
|
||||||
|
```
|
||||||
|
|
||||||
|
Design requirements:
|
||||||
|
|
||||||
|
- Treat `UnitFileState` (startup policy), `ActiveState`/`SubState` (runtime), and last sample result as independent axes. `is-enabled` documents multiple states—including enabled, disabled, static, indirect, generated, transient, and masked—so a Boolean loses actionable information. [systemctl(1), `is-enabled` state table](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#is-enabled%20UNIT%E2%80%A6)
|
||||||
|
- Parse only `systemctl show --property=...` or the equivalent D-Bus properties. `status` is explicitly human-oriented and includes journal data. [systemctl(1), `show`](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#show%20PATTERN%E2%80%A6)
|
||||||
|
- Compute freshness from `last_success`, not merely timer activity. An active timer can coexist with repeated collection failures.
|
||||||
|
- Publish bounded error categories and a short administrator hint; keep raw smartctl stderr and tracebacks in the journal. This prevents device identifiers, paths, malformed device output, or command details from crossing the read boundary.
|
||||||
|
- Never report “enabled” immediately after a requested mutation without re-querying the authoritative state. Never equate “disabled” with “stopped.”
|
||||||
|
- If status data is unreadable, say `permission-denied`/`unknown`; do not elevate merely to render the screen.
|
||||||
|
|
||||||
|
## Alternatives and trade-offs
|
||||||
|
|
||||||
|
### Long-running system service
|
||||||
|
|
||||||
|
A foreground `fenris-collector.service` with `Restart=on-failure` and `WantedBy=multi-user.target` also satisfies reboot persistence; `Restart=` controls automatic restart after process failure, while `WantedBy=multi-user.target` is the standard installation relationship for a multi-user service. [systemd.service(5), `Restart=`](https://www.freedesktop.org/software/systemd/man/latest/systemd.service.html#Restart=); [systemd.special(7), `multi-user.target`](https://www.freedesktop.org/software/systemd/man/latest/systemd.special.html#multi-user.target)
|
||||||
|
|
||||||
|
Trade-off: it preserves today's loop model and precise in-process scheduling, but leaves a privileged Python process resident continuously and requires restart/backoff handling. Use it only if future collection needs persistent in-memory state that cannot be reconstructed from durable samples.
|
||||||
|
|
||||||
|
### Unprivileged daemon plus privileged smartctl helper
|
||||||
|
|
||||||
|
This can reduce privileged code if the helper accepts no uncontrolled path or options, validates a configured device allow-list, emits bounded sanitized output, and exits. It adds an IPC/protocol and another authorization surface. Giving the whole Python process `CAP_SYS_ADMIN` is not an equivalent reduction because that capability is exceptionally broad. [capabilities(7), notes](https://man7.org/linux/man-pages/man7/capabilities.7.html#NOTES)
|
||||||
|
|
||||||
|
### User service with lingering
|
||||||
|
|
||||||
|
Lingering can run a user manager from boot through logout, but the user manager cannot switch an ordinary user's unit to root and some system-service sandboxing features are unavailable in user services. [loginctl(1), lingering](https://www.freedesktop.org/software/systemd/man/latest/loginctl.html#enable-linger%20USER%E2%80%A6); [systemd.exec(5), user-service sandboxing limitations](https://www.freedesktop.org/software/systemd/man/latest/systemd.exec.html#Sandboxing) This adds account coupling and still needs a separate privilege mechanism. Reject for collection; it remains acceptable for an optional per-user dashboard client.
|
||||||
|
|
||||||
|
### Root daemon calling `sudo -n smartctl`
|
||||||
|
|
||||||
|
This is redundant: root already crosses the privilege boundary, while sudo introduces policy/path/configuration failure modes. For an unprivileged daemon, `sudo -n` avoids blocking but only works with pre-authorized sudoers policy and produces no authentication opportunity. [sudo(8), `--non-interactive`](https://github.com/sudo-project/sudo/blob/194042b55c54055c7337fbbd93a518e8da69866f/docs/sudo.man.in#L629-L635) Reject as the installed systemd architecture.
|
||||||
|
|
||||||
|
### Direct broad polkit delegation
|
||||||
|
|
||||||
|
Convenient but unsafe for startup-state delegation: `manage-unit-files` covers generic enable/disable/mask/preset operations and currently lacks unit-scoping details in systemd's authorization call. Use normal administrator authentication or a custom narrow action instead. [systemd D-Bus API, Security](https://www.freedesktop.org/software/systemd/man/latest/org.freedesktop.systemd1.html#Security); [pinned systemd authorization source](https://github.com/systemd/systemd/blob/a6a831d0d9ce304619b8937b27b1286109b5e625/src/core/dbus-util.c#L209-L223)
|
||||||
|
|
||||||
|
## Unresolved questions and required validation
|
||||||
|
|
||||||
|
1. **Supported distributions/systemd floor:** What is the minimum systemd version? Confirm every selected hardening and directory directive exists there; unknown directives can weaken the intended sandbox.
|
||||||
|
2. **Device identity:** Is configuration a mutable `/dev/nvmeN` node, namespace node, controller, `/dev/disk/by-id` link, or discovered set? smartctl documents Linux NVMe controller and namespace forms, but the stable product identity and hotplug behavior remain a product decision. [smartctl(8), pinned source](https://github.com/smartmontools/smartmontools/blob/618fcaede4478bc7d17fa2a8db5fd18af3744e20/src/smartctl.8.in#L72-L86)
|
||||||
|
3. **Privilege matrix:** On each supported kernel, packaging of smartmontools, NVMe/SATA/USB bridge, and device permission setup, record which open/ioctl fails as an unprivileged service and which exact capability/device allow-list is sufficient. Do not generalize an NVMe result to every smartmontools transport.
|
||||||
|
4. **Root vs. reduced-capability collector:** After that matrix exists, decide whether a non-root static service user plus a minimal capability/device set works portably. Reject any result that requires putting broad capability into the TUI/dashboard.
|
||||||
|
5. **Timer semantics:** Should five minutes be measured from sample start or completion? What maximum runtime and timeout are acceptable? Should resume from suspend trigger immediately? Decide `OnUnitActiveSec` vs. `OnUnitInactiveSec`, `Persistent=`, and `AccuracySec` from those answers.
|
||||||
|
6. **Startup toggle semantics:** Does “disable startup” leave the timer active until reboot, or should the TUI offer a separate, clearly labeled “disable and stop now”? The systemctl semantics intentionally separate these effects. [systemctl(1), enable/disable](https://www.freedesktop.org/software/systemd/man/latest/systemctl.html#enable%20UNIT%E2%80%A6)
|
||||||
|
7. **Who may read health history:** All local users, a `fenris-readers` group, or only an authenticated local client? This determines state modes and whether a read-only Unix socket is preferable to files.
|
||||||
|
8. **Dashboard scope:** Is the HTTP dashboard retained, and if so must it be local-only or remotely accessible? Remote access needs a separate threat model, authentication, transport security, and unprivileged service; it must not enlarge the collector boundary.
|
||||||
|
9. **Packaging paths:** Confirm `/usr/libexec` and `/usr/lib/systemd/system` equivalents per target distribution, the absolute smartctl path, and whether administrator overrides use environment files or a dedicated validated config format.
|
||||||
|
10. **Data durability:** Define atomicity, fsync policy, retention, corruption recovery, migration from checkout-relative `data/`, and behavior on read-only/full filesystems.
|
||||||
|
11. **Authorization UX:** Is ordinary admin authentication acceptable? If not, specify exactly which local principals may toggle startup and commission the custom helper/action rather than broad `manage-unit-files` delegation.
|
||||||
|
12. **Status confidentiality:** Decide whether model name, device path, capacity, wear, temperatures, and error counters are safe for every local reader; serial numbers and raw smartctl output should remain excluded by default.
|
||||||
|
|
||||||
|
## Decision summary
|
||||||
|
|
||||||
|
The verified boundary is: **system systemd owns scheduling and boot state; a short-lived privileged collector owns only device interrogation and state writes; an unprivileged on-demand TUI owns presentation; polkit/admin authentication mediates explicit startup changes.** This satisfies collection across reboot without tying it to login, avoids embedding sudo in unattended code, keeps generic service control out of the TUI's ambient privilege, and makes startup state observable and alterable without conflating it with current execution.
|
||||||
Reference in New Issue
Block a user