feat: Fenris persistent TUI monitoring redesign
Implement the complete redesign per fenris-redesign spec: - Observation store: SQLite WAL mode, six entities, schema versioning - Collector: smartctl acquisition, sysfs identity, normalization - Projection: sustained regime rate, habit change, confidence states - Panes TUI: Textual keyboard-first layout with four normative regions - Status CLI: read-only composition with four service facts - Monitor helper: polkit-guarded toggle, collect, baseline ops - Legacy migration: idempotent single-transaction import - Hour classification, day aggregates, monitoring periods - Pruning, segmentation, drive health facts Cross-cutting acceptance sweep (CI-1 through CI-4): - 59 tests covering state matrix, TUI/CLI parity, prohibition set, required wording and six disclosures - Full suite: 289 tests, all green Issues #20, #32 closed.
This commit is contained in:
@@ -1,157 +1,126 @@
|
||||
<p align="center">
|
||||
<picture>
|
||||
<source srcset="assets/bongbetic-brand/wordmark-light.png" media="(prefers-color-scheme: dark)">
|
||||
<img src="assets/bongbetic-brand/wordmark-dark.png" alt="Bongbetic" width="260">
|
||||
</picture>
|
||||
<br>
|
||||
<sub>crafted with stubborn curiosity by <a href="https://bongbetic.com">Bongbetic</a></sub>
|
||||
</p>
|
||||
# Fenris 🐺
|
||||
|
||||
<p align="center">
|
||||
<img src="assets/bongbetic-brand/b_glyph.svg" width="48" alt="Fenris glyph">
|
||||
</p>
|
||||
*Observes an NVMe drive's real-world use and translates that history into an understandable endurance outlook.*
|
||||
|
||||
<h1 align="center">Fenris 🐺 — Your SSD's Tell-All Diary</h1>
|
||||
|
||||
<p align="center">
|
||||
<em>Your NVMe drive has been keeping secrets. Fenris makes it confess — in real time.</em>
|
||||
<br>
|
||||
<em>How much did you write today? How long until it taps out? No fairy dust — just your actual bytes.</em>
|
||||
</p>
|
||||
Fenris is a persistent TUI monitor backed by a short-lived privileged collector on a systemd timer. It reads SMART data every few minutes, stores compact observation history in SQLite, and recomputes a usage-adjusted theoretical lifespan on every screen render — no fairy dust, just your actual bytes.
|
||||
|
||||
---
|
||||
|
||||
Fenris is a tiny, stubborn daemon that eavesdrops on your NVMe drive's SMART gossip, writes it down every few minutes, and serves you a live dashboard that actually means something. Not "vibes" — **real GB written in the last 24 hours, real GB/hour, and a real countdown in hours, days, and years until your drive's endurance runs out**.
|
||||
## Requirements
|
||||
|
||||
> Think of it as a Fitbit for your SSD. Except it doesn't nag you to drink water.
|
||||
- **Python ≥ 3.9** (verified at install time)
|
||||
- **smartmontools** (`smartctl` — verified at install time)
|
||||
- **systemd** with a polkit agent (the collector runs as root oneshot; elevation is exclusively polkit)
|
||||
|
||||
## What it actually does (no hand-waving)
|
||||
No other OS packages or Python dependencies beyond [Textual](https://textual.textualize.io/) (pinned in the lockfile).
|
||||
|
||||
- **Listens** — polls `smartctl -j` on your NVMe device (default every 5 minutes, you pick).
|
||||
- **Remembers** — appends every sample to `data/history.jsonl` and rolls up per-hour totals into `data/hourly.jsonl` (survives restarts, rebuilds itself if you yank the power).
|
||||
- **Calculates** — rolling 24-hour window: *exact* bytes written in the last 24h, GB/h, GB/day, implied total TBW from `percentage_used`, remaining TB, and a projected life-remaining breakdown. Warming-up badge until it has 24h of coverage — no fake confidence.
|
||||
- **Shows off** — dense, live dashboard with wear-over-time + trailing-24h per-hour bars, sticky header, live countdown, and stale warnings if the daemon dozes off.
|
||||
|
||||
## You need
|
||||
|
||||
- **Python 3.7+**
|
||||
- **smartmontools** (`smartctl`)
|
||||
- Root-ish access to read NVMe SMART (passwordless `smartctl` or just run with `sudo` — your call)
|
||||
|
||||
### The sudo dance (one time)
|
||||
|
||||
Fenris runs `sudo -n smartctl ...` so it doesn't get stuck asking for a password mid-nap:
|
||||
## Install
|
||||
|
||||
```bash
|
||||
sudo visudo
|
||||
# add this line (swap in your username):
|
||||
youruser ALL=(root) NOPASSWD: /usr/sbin/smartctl
|
||||
sudo make install
|
||||
```
|
||||
|
||||
No sudo? Run the whole thing with `sudo` and it'll still behave.
|
||||
What it does:
|
||||
1. Builds a wheel from the checkout and installs it — with pinned dependencies — into the dedicated venv at `/opt/fenris`.
|
||||
2. Places the `fenris` wrapper in `/usr/local/bin`, helpers in `/usr/libexec/fenris`, systemd units in `/etc/systemd/system`, and the polkit policy in `/usr/share/polkit-1/actions/`.
|
||||
3. Creates `/var/lib/fenris` (root-written, group-readable) — the observation store is created lazily by the first collection run.
|
||||
4. Records every placed file in a manifest consumed by upgrade and uninstall.
|
||||
5. Detects `./data/history.jsonl` beside the source checkout and runs the idempotent legacy import if present.
|
||||
|
||||
## Get it running — 30 seconds
|
||||
|
||||
### The cozy way
|
||||
**A fresh install is fully dormant.** Units are present but disabled; nothing runs. The only opt-in is the sanctioned toggle:
|
||||
|
||||
```bash
|
||||
./fenris.sh
|
||||
# pick 1) Start monitoring → choose device / interval / port → done
|
||||
fenris monitor resume # enable timer + open first monitoring period
|
||||
fenris monitor pause # close the period, disable timer
|
||||
```
|
||||
|
||||
### The no-nonsense way
|
||||
## Upgrade
|
||||
|
||||
```bash
|
||||
python3 fenris.py start # defaults: /dev/nvme0, every 300s, port 8420
|
||||
python3 fenris.py start --interval 60 --port 9000 # if you're impatient
|
||||
python3 fenris.py status # "are we live? how's the drive?"
|
||||
python3 fenris.py sample # one sneaky sample right now
|
||||
python3 fenris.py stop # tuck it back in
|
||||
sudo make upgrade
|
||||
```
|
||||
|
||||
Dashboard lives at **http://localhost:8420** (or whatever port you chose).
|
||||
What it does:
|
||||
1. Snapshots `observations.db` to a one-generation backup (`.bak`).
|
||||
2. Installs the new wheel into the same venv with pinned dependencies.
|
||||
3. Syncs units and polkit against the manifest; runs `daemon-reload`.
|
||||
4. Restarts the timer **only** if unit contents changed **and** it is active — a running collection run finishes on its mapped interpreter; the next run uses the new code.
|
||||
5. Applies forward-only schema migrations (the store directory is never rebuilt; automatic downgrade does not exist).
|
||||
|
||||
## The menu, demystified
|
||||
Rollback: reinstall the previous version and restore `observations.db.bak`.
|
||||
|
||||
Run `./fenris.sh` and you'll get:
|
||||
## Uninstall and purge
|
||||
|
||||
```
|
||||
1) Start monitoring (background daemon + dashboard)
|
||||
2) Stop monitoring
|
||||
3) Status / current wear stats
|
||||
4) Take one sample right now
|
||||
5) Open dashboard URL
|
||||
---
|
||||
h) Help / how this works
|
||||
q) Exit (go touch grass)
|
||||
```bash
|
||||
make uninstall # removes artifacts, preserves config and observation history
|
||||
make purge # also removes /etc/fenris and /var/lib/fenris
|
||||
```
|
||||
|
||||
## What Fenris jots down
|
||||
Uninstall performs the sanctioned disable first (`fenris-monitor disable --now`) — an open period closes `user_disabled` — then removes the venv, helpers, units, polkit policy, and wrapper while keeping `/etc/fenris` and the observation store. Reinstalling resumes from the preserved store.
|
||||
|
||||
| Field | What's the gossip? |
|
||||
|-------|---------------------|
|
||||
| `percentage_used` | The drive's own wear-o-meter (0–100%) |
|
||||
| `bytes_written` / `bytes_read` | Lifetime totals — the receipts |
|
||||
| `available_spare` | Spare blocks left (%) |
|
||||
| `media_errors` | Uncorrectable boo-boos |
|
||||
| `power_on_hours` | How long it's been awake |
|
||||
| `temperature_c` | Is it sweating? |
|
||||
| `critical_warning` | NVMe's panic flags |
|
||||
## Cadence drop-ins
|
||||
|
||||
Hourly rollups also stash `bytes_written` per hour, `pct_start`/`pct_end`, and temp peaks — so the 24h math stays honest.
|
||||
The default collection cadence is **5 minutes** (`OnUnitInactiveSec=5min` in the timer unit). To change it, place a systemd drop-in:
|
||||
|
||||
## The dashboard — what's on screen
|
||||
```bash
|
||||
sudo systemctl edit fenris-collect.timer
|
||||
# Add:
|
||||
# [Timer]
|
||||
# OnUnitInactiveSec=10min
|
||||
```
|
||||
|
||||
- **Hero card: Projected life remaining** — big, friendly `361 d 2 h` (plus `≈ 361 days · ≈ 8666 hours · ≈ 0.99 years`), backed by `~280 GB/day` and `~101 TB left of ~202 TB total` on the test box.
|
||||
- **Data written (24h)** — exact GB in the rolling window + coverage (`10.4h of 24h` until warmed up).
|
||||
- **Write rate** — GB/h and GB/day, live.
|
||||
- **Wear, spare, temp, errors, power-on** — the usual suspects, with progress bars and polite color-coding.
|
||||
- **Two charts, side by side:** wear over time + trailing-24h hourly write bars (with a cheeky "now" bar for the current partial hour).
|
||||
- **Live plumbing:** polling synced to your interval, ETag-cached, countdown to next sample, warming-up + stale banners, pauses when you hide the tab (saves your battery, you're welcome).
|
||||
No interval key exists in `/etc/fenris/fenris.conf`. Cadence is a systemd concern, not a Fenris configuration key.
|
||||
|
||||
**API for the tinkerers:** `GET /api/data` · `/api/hourly` · `/api/summary` · `/api/config` · `/api/status` — all JSON, all friendly.
|
||||
## CLI reference
|
||||
|
||||
| Command | Behavior |
|
||||
|---|---|
|
||||
| `fenris` | Opens the TUI (no arguments). |
|
||||
| `fenris status` | Projection facts, enabled/active state, last collect outcome, journal hint on failure or staleness. Never auto-samples. |
|
||||
| `fenris sample` | On-demand collection via the privileged helper. Blocks until the run completes. |
|
||||
| `fenris monitor pause` | Sanctioned disable — asks for confirmation, then disables the timer and closes the monitoring period. |
|
||||
| `fenris monitor resume` | Sanctioned enable — enables the timer and opens a monitoring period. No confirmation. |
|
||||
| `fenris baseline set <json>` | CLI-side validation, then polkit-guarded persistence. |
|
||||
| `fenris baseline clear` | Remove the endurance baseline. |
|
||||
| `fenris import <path>` | Idempotent single-transaction legacy import. |
|
||||
| `fenris start` / `stop` / `run` | Rejected with a one-line migration pointer — never aliased. |
|
||||
| `fenris --device` | Rejected with a pointer to the configuration file. |
|
||||
|
||||
## Retired menu options
|
||||
|
||||
The legacy `fenris.sh` menu script and the `fenris.py` monolith have been removed. Here's where the old options went:
|
||||
|
||||
| Legacy option | Successor |
|
||||
|---|---|
|
||||
| 1) Start monitoring | `fenris monitor resume` |
|
||||
| 2) Stop monitoring | `fenris monitor pause` |
|
||||
| 3) Status / current wear stats | `fenris status` |
|
||||
| 4) Take one sample right now | `fenris sample` |
|
||||
| 5) Open dashboard URL | Removed — the HTML dashboard and HTTP server are gone; the TUI is the primary interface. |
|
||||
|
||||
## Configuration
|
||||
|
||||
`/etc/fenris/fenris.conf` holds exactly one key — the device selector:
|
||||
|
||||
```
|
||||
device = /dev/disk/by-id/nvme-Samsung_SSD_980_PRO_2TB_S6BENS0Txxxxx
|
||||
```
|
||||
|
||||
Use a stable `/dev/disk/by-id/` path. Raw `/dev/nvmeX` paths are warned against. The file is re-read every collection run.
|
||||
|
||||
## Where's my stuff?
|
||||
|
||||
```
|
||||
fenris/
|
||||
├── fenris.py # the whole show — daemon + server + math
|
||||
├── fenris.sh # the cozy menu
|
||||
├── README.md # hi — you're here
|
||||
├── assets/bongbetic-brand/ # Bongbetic wordmarks & glyphs (for Gitea + dashboard)
|
||||
└── data/
|
||||
├── history.jsonl # raw samples (JSONL, append-only)
|
||||
├── hourly.jsonl # per-hour rollups (auto-rebuilt on restart)
|
||||
├── fenris.pid # daemon PID
|
||||
└── fenris.log # daemon chatter
|
||||
```
|
||||
|
||||
## CLI cheat sheet
|
||||
|
||||
```bash
|
||||
python3 fenris.py start [--device /dev/nvme0] [--interval 300] [--port 8420]
|
||||
python3 fenris.py stop
|
||||
python3 fenris.py status
|
||||
python3 fenris.py sample [--device /dev/nvme0]
|
||||
python3 fenris.py run # foreground mode — what `start` spawns internally
|
||||
```
|
||||
|
||||
## Oops — troubleshooting without the tears
|
||||
|
||||
**"smartctl not found"**
|
||||
```bash
|
||||
sudo apt install smartmontools # Debian/Ubuntu
|
||||
sudo pacman -S smartmontools # Arch — you already knew
|
||||
```
|
||||
|
||||
**"needs root" / permission denied**
|
||||
Set up the passwordless line above, or just `sudo ./fenris.sh`.
|
||||
|
||||
**Dashboard says "stale"**
|
||||
Daemon napped or crashed. `python3 fenris.py status` will tell you. Kick it again with `start`.
|
||||
|
||||
**Only 10 hours of data and it says "preliminary"?**
|
||||
That's honesty, not a bug. It needs 24h of real writes to give a tight estimate. Let it simmer — the number gets sharper every hour.
|
||||
| Artifact | Location |
|
||||
|---|---|
|
||||
| Wrapper | `/usr/local/bin/fenris` |
|
||||
| Helpers | `/usr/libexec/fenris/fenris-collect`, `fenris-monitor` |
|
||||
| Units | `/etc/systemd/system/fenris-collect.{timer,service}` |
|
||||
| Polkit policy | `/usr/share/polkit-1/actions/com.bongbetic.fenris.monitor.policy` |
|
||||
| Configuration | `/etc/fenris/fenris.conf` |
|
||||
| Observation store | `/var/lib/fenris/observations.db` |
|
||||
| Venv | `/opt/fenris` |
|
||||
| Manifest | `/var/lib/fenris/manifest.txt` |
|
||||
| Legacy history | `./data/history.jsonl` (auto-imported on install if present) |
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user