Initial commit: Fenris NVMe wear monitor with rolling-24h dashboard

This commit is contained in:
xavierk
2026-08-23 18:47:12 +05:30
commit 53cc6c81b9
5 changed files with 1534 additions and 0 deletions
+7
View File
@@ -0,0 +1,7 @@
__pycache__/
*.pyc
.commandcode/
data/fenris.pid
data/fenris.log
data/history.jsonl
data/hourly.jsonl
+145
View File
@@ -0,0 +1,145 @@
# Fenris — NVMe Wear Monitor & Live Dashboard
Created by Bongbetic.
Periodically reads your NVMe drive's SMART health data, logs it over time, and serves a self-contained HTML dashboard estimating SSD lifespan from your actual daily usage trend.
## Requirements
- **Python 3.7+**
- **smartmontools** (`smartctl`) installed
- Root access to read NVMe SMART logs
### Setting up passwordless smartctl
Fenris runs `sudo -n smartctl ...` (no-prompt sudo). Either run with sudo or allow passwordless access:
```bash
sudo visudo
# Add this line (replace youruser with your username):
youruser ALL=(root) NOPASSWD: /usr/sbin/smartctl
```
## Quick Start
### Interactive Menu
```bash
./fenris.sh
```
### CLI
```bash
# Start daemon + dashboard in background
python3 fenris.py start
# Check status and latest wear stats
python3 fenris.py status
# Take one sample now
python3 fenris.py sample
# Stop the daemon
python3 fenris.py stop
```
Dashboard available at: `http://localhost:8420`
## CLI Reference
### fenris.py start
Start monitoring in background (daemon + dashboard).
```bash
python3 fenris.py start [OPTIONS]
Options:
--device PATH NVMe device (default: auto-detect, e.g. /dev/nvme0)
--interval SEC Seconds between samples (default: 300)
--port PORT Dashboard HTTP port (default: 8420)
```
### fenris.py stop
Stop background monitoring and clean up PID file.
### fenris.py status
Show daemon status and latest wear statistics.
### fenris.py sample
Take one sample immediately and print it.
```bash
python3 fenris.py sample [--device /dev/nvme0]
```
### fenris.py run
Run in foreground (used internally by `start`). Not intended for direct use.
## Interactive Menu
Run `./fenris.sh` for a guided interface:
```
1) Start monitoring (background daemon + dashboard)
2) Stop monitoring
3) Status / current wear stats
4) Take one sample right now
5) Open dashboard URL
---
h) Help / how this works
q) Exit
```
## Data Collected
| Field | Description |
|-------|-------------|
| `percentage_used` | SSD's own wear indicator (0-100%) |
| `bytes_written` / `bytes_read` | Total data written/read |
| `available_spare` | Remaining spare capacity (%) |
| `media_errors` | Number of uncorrectable errors |
| `power_on_hours` | Total power-on hours |
| `temperature_c` | Current temperature |
| `critical_warning` | NVMe critical warning flags |
## Dashboard Features
- Real-time wear level + projected life remaining in **hours / days / years** from the actual **rolling-24h write rate** and implied TBW endurance
- Exact **GB written in the last 24 hours** + GB/h and GB/day rate (updates every poll interval)
- Wear-over-time chart + **trailing-24h per-hour write bars**
- Dense layout with sticky header, ETag-cached polling synced to the daemon interval, countdown and live badge, stale/preliminary banners
- API: `GET /api/data`, `/api/hourly`, `/api/summary`, `/api/config`, `/api/status`
## File Structure
```
fenris/
├── fenris.py # Main Python script
├── fenris.sh # Interactive menu wrapper
├── README.md
└── data/
├── history.jsonl # Raw sample log (JSONL)
├── hourly.jsonl # Per-hour aggregates (rebuilt from history on restart)
├── fenris.pid # Daemon PID file
└── fenris.log # Daemon log output
```
## Troubleshooting
**"smartctl not found"**
```bash
sudo apt install smartmontools # Debian/Ubuntu
sudo pacman -S smartmontools # Arch
```
**"needs root" / permission denied**
Set up passwordless sudo (see Requirements) or run with sudo.
**Dashboard shows "stale"**
Daemon not running. Check with `python3 fenris.py status` and restart if needed.
Executable
+1035
View File
File diff suppressed because it is too large Load Diff
Executable
+128
View File
@@ -0,0 +1,128 @@
#!/usr/bin/env bash
# Fenris — interactive menu for the NVMe wear monitor & dashboard.
# Created by Bongbetic.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
PY="$SCRIPT_DIR/fenris.py"
PORT_DEFAULT=8420
INTERVAL_DEFAULT=300
banner() {
cat <<'EOF'
_____ _
| __|___ ___ _| |___
| __| -_| | . | _|
|__| |___|_|_|_|___|_|
NVMe wear monitor & live dashboard
Created by Bongbetic
EOF
}
pause() { read -rp "Press Enter to continue..." _; }
detect_device() {
# Query Python's auto-detect for the default device.
python3 -c "import sys; sys.path.insert(0,'$SCRIPT_DIR'); from fenris import detect_device; print(detect_device())" 2>/dev/null || echo /dev/nvme0
}
menu() {
clear
banner
echo
echo " 1) Start monitoring (background daemon + dashboard)"
echo " 2) Stop monitoring"
echo " 3) Status / current wear stats"
echo " 4) Take one sample right now"
echo " 5) Open dashboard URL"
echo " ---"
echo " h) Help / how this works"
echo " q) Exit"
echo
read -rp "Choose an option: " choice
echo
case "$choice" in
1) start_flow ;;
2) python3 "$PY" stop; pause ;;
3) python3 "$PY" status; pause ;;
4) read -rp "Device [default: auto-detect]: " dev
if [ -z "$dev" ]; then python3 "$PY" sample; else python3 "$PY" sample --device "$dev"; fi
pause ;;
5) show_url; pause ;;
h|H) help_text; pause ;;
q|Q) echo "Bye. — Fenris, by Bongbetic"; exit 0 ;;
*) echo "Invalid choice."; pause ;;
esac
}
start_flow() {
read -rp "NVMe device [Enter = auto-detect]: " dev
read -rp "Sample interval in seconds [Enter = ${INTERVAL_DEFAULT}]: " interval
read -rp "Dashboard port [Enter = ${PORT_DEFAULT}]: " port
interval="${interval:-$INTERVAL_DEFAULT}"
port="${port:-$PORT_DEFAULT}"
args=(start --interval "$interval" --port "$port")
if [ -n "${dev:-}" ]; then args+=(--device "$dev"); fi
echo
echo "Note: reading NVMe SMART data needs root."
echo "Fenris runs 'sudo -n smartctl ...' (no-prompt sudo). If this fails,"
echo "either run this menu with sudo, or allow passwordless smartctl via:"
echo " sudo visudo -> youruser ALL=(root) NOPASSWD: /usr/sbin/smartctl"
echo
python3 "$PY" "${args[@]}"
pause
}
show_url() {
if [ -f "$SCRIPT_DIR/data/fenris.pid" ]; then
# Try to read actual port from running process cmdline, else guess default.
local pid port
pid=$(<"$SCRIPT_DIR/data/fenris.pid")
port=$(tr '\0' '\n' < /proc/"$pid"/cmdline 2>/dev/null | grep -A1 -- '--port' | tail -1 || true)
port="${port:-$PORT_DEFAULT}"
echo "Dashboard: http://localhost:${port}"
else
echo "Fenris is not currently running. Start it first (option 1)."
fi
}
help_text() {
cat <<EOF
What Fenris does:
- Periodically reads your NVMe drive's SMART health data (via smartctl),
including "percentage_used" (the drive's own wear indicator), total
bytes written/read, temperature, spare capacity, and error counts.
- Logs every sample to: $SCRIPT_DIR/data/history.jsonl
and per-hour aggregates to: $SCRIPT_DIR/data/hourly.jsonl (GB/hour,
rebuilt from history on restart). Used for the trailing-24h bar chart
and exact rolling-24h write volume.
- Serves a live HTML dashboard (dense layout, interval-synced polling)
with wear-over-time and trailing-24h hourly-write charts, plus a
projected life-remaining estimate in hours/days/years derived from
your actual rolling-24h write rate and implied TBW endurance.
Requirements:
- smartmontools (smartctl) installed.
- Root access to read NVMe SMART logs — either run Fenris via sudo,
or set up passwordless sudo for smartctl (see option 1).
CLI usage (equivalent to this menu):
python3 fenris.py start [--device /dev/nvme0] [--interval 300] [--port 8420]
python3 fenris.py stop
python3 fenris.py status
python3 fenris.py sample [--device /dev/nvme0]
Leave it running in the background (option 1) and check back after a
few days/weeks of normal use — more samples = a more accurate lifespan
estimate.
Fenris — created by Bongbetic.
EOF
}
# Loop instead of recurse to avoid stack overflow.
while true; do menu; done
+219
View File
@@ -0,0 +1,219 @@
# Plan: Fenris Live Dashboard — Dynamic Polling + Bongbetic Brand + shadcn/ui
> Status: PLAN ONLY — no code touched. **Option A approved.** Data-management removed; cumulative write chart X = hours-in-day (0–24); dense layout; hourly diagnostics + real-time forecast model.
## 1. Context & Goals
**User request (consolidated):**
- Dashboard live/dynamic, updating at each poll interval without refresh — **Option A (zero-build).**
- Bongbetic branding + logo everywhere (`~/Documents/bongbetic/Logo`).
- shadcn/ui styling (CSS parity, no React build).
- **Remove data-management feature** (backup/restore/backups/repair/purge).
- **Cumulative data-written graph X-axis = hours in a day (0–24).**
- **Graphs smaller, side-by-side, flexible.** **Cards tighter/dense.**
- **Estimated life remaining = real-time, based on live hourly GB usage.** After 24h of run, app must keep recording hourly, store diagnostics, log GB/hour, and forecast remaining days from current hourly rate.
**Current state (verified 2026-08-19):**
- `fenris.py` single `DASHBOARD_HTML` via `ThreadingHTTPServer`. Hand-rolled dark CSS, 2× `<canvas>` stacked vertically (`height=220`, full-width), `setInterval(render,30000)` hardcoded, full `innerHTML` replace.
- `GET /api/data` → full `history.jsonl`; no `/api/config`; interval not exposed.
- Forecast: `linearForecast()` on `percentage_used` trend (simple linear regression, whole history). No hourly bucket, no GB/hour model, no diagnostics persistence beyond raw history.
- Cards: `minmax(220px,1fr)` gap 1rem, padding `1rem 1.2rem`, value `1.7rem` — loose.
- No asset vendoring, no Tailwind/shadcn.
**Goals:**
1. Interval-synced live refresh (no reload), ETag/diff, pulse.
2. Bongbetic header/logo/footer + shadcn tokens.
3. Dense cards + two graphs side-by-side, responsive, reduced height.
4. After 24h: hourly diagnostics log (`GB/hour`), real-time days-remaining forecast from live hourly write rate (not just wear %).
## 2. Non-Goals & Explicit Removals
- No auth/multi-tenant, no SSE/WebSocket v1, no Vite/React (Option B rejected).
- **Data-management removed:** delete `BACKUP_DIR`, 5 commands, parsers, `data/backups/`, `fenris.sh` items 6–10. Data = append-only `data/history.jsonl` (+ new `data/hourly.jsonl` per §5.2); orphan `backups/` logs “safe to delete”.
## 3. Logo / Brand Audit
```
bongbetic-logo-dark.svg (751×220, #F7F1E7 currentColor) — dark bg
bongbetic-symbol-dark.svg (192×192) — favicon/header mark
bongbetic-brand/favicon-32.png, icon-dark-512.png, site.webmanifest, b_glyph.svg
```
Vendor into `assets/` and serve via `Handler` static branch; inline SVG so `currentColor` follows `text-foreground`.
## 4. Target Architecture — Option A (Approved)
Zero-build: Tailwind CDN + shadcn CSS variables (Slate/Zinc dark). Semantic HTML + Tailwind mimics `Card/Badge/Progress/Alert/Skeleton`. Canvas charts wrapped in `Card`; wear + write side-by-side flex grid (see §6.1). Python exposes `/api/config`, `/api/status`, `/api/hourly`, static `/assets/*`.
## 5. Live / Dynamic Behavior
### 5.0 Polling (unchanged from prior plan)
- `GET /api/config → {interval, device, port, version}` → `POLL_MS = interval*1000` (clamp 5s–3600s, 30s fallback).
- Diff: hash `last_ts+length`; skip render if unchanged; else patch cards (textContent morph + `ring-2` pulse), append chart point via `requestAnimationFrame`.
- `visibilitychange` pause/resume, `navigator.onLine` backoff 1/2/4…60s, header badge “Live • every 5m • next in 03:42”.
- `GET /api/data` gains `ETag` (mtime+length) + 304.
### 5.1 Cumulative Write Chart — Hours-in-Day (Requirement)
- **Intraday 0–24h view:** `hours = h + m/60 + s/3600` (0–23.99). `todayRows = rows.filter(same DateString)` → `[hours, (bytes_written - midnightBaseline)/1e9]`.
- Axis: `minX=0, maxX=24`, ticks `00:00/06:00/12:00/18:00/24:00` via `opts.xIsHoursInDay`. Title “Cumulative written today (GB) — hours of day”. Empty → “No samples today yet”. Midnight → reset to 0 GB, caption “Resets at midnight — today only”.
- Wear chart stays absolute-time (weeks trend). Both canvases slimmed (see §6.1).
### 5.2 Real-Time Hourly Diagnostics & Forecast Model (New — Core Requirement)
**Problem with current forecast:** `linearForecast(points)` regresses `percentage_used` over full history; single slope, insensitive to bursty hourly writes, no GB/hour visibility.
**Required model:**
- Real-time on every poll, not daily batch.
- After 24h wall time since first sample, switch to **hourly GB/hour regime**; before 24h, show warming-up estimate.
- Persist per-hour diagnostics and GB/hour log.
**Design:**
1. **Raw source stays `data/history.jsonl`** (poll interval samples, e.g., every 300s).
2. **New derived log `data/hourly.jsonl`** (append-only, one record per wall hour):
```json
{"hour":"2026-08-19T14:00:00Z","samples":12,"gb_written":4.21,"gb_read":1.03,
"pct_start":1.2,"pct_end":1.21,"pct_delta":0.01,
"temp_avg":42.1,"temp_max":48,"spare":98,"media_errors":0}
```
Fields: hour bucket start (UTC), count, deltas from first/last sample in hour, averages. Written by `collector_loop` helper `flush_hourly()`.
3. **Hourly rollup logic (Python):**
- In-memory `current_hour_bucket`; on each `sample()`, accumulate `bytes_written` delta vs bucket start.
- On hour boundary (or every 60min since daemon start if clock not trusted), `append_hourly(rec);` also `append_sample(raw)`.
- On daemon (re)start, rebuild missing hours by scanning `history.jsonl` and aggregating by `hour = ts truncated to hour` (idempotent — dedupe by hour string).
- `/api/hourly` returns parsed `hourly.jsonl` (array, sorted). Add ETag similarly.
4. **Forecast — two complementary signals, UI shows primary “days remaining (hourly write model)” + secondary wear model:**
- **GB/hour → TBW model (primary, per user ask “present hourly usage → days”):**
```
hourly_avg = mean(last 24 hourly gb_written) // rolling 24, or EWMA α=0.3 if <24h
// derive endurance from vendor wear if available:
if pct_used>0: endurance_TB = bytes_written / (pct_used/100) // total TBW implied
else: endurance_TB = capacity_bytes * 600 // fallback: ~600× capacity (conservative), or mark unknown
remaining_TB = max(endurance_TB - bytes_written/1e12, 0)
days_remaining = remaining_TB / (hourly_avg *24) // hourly_avg in TB
```
If endurance derivable, show; else show wear-based only and badge “TBW unknown — using wear rate”.
- **Wear/hour model (secondary, cross-check):**
```
wear_per_hour = mean(last 24 pct_delta per hour)
hours_to_100 = (100 - pct_now) / wear_per_hour
days_wear = hours_to_100/24
```
Shown as tooltip / small “also ~X days at current wear rate”.
- **<24h warming up:** `hourly_avg` over available hours (n<24); badge “Warming up — Xh to confident forecast (now ~Y days, n=Nh)”. `days_remaining` still computed but flagged `preliminary`.
- **Real-time update:** every poll, frontend refetches `/api/data` + `/api/hourly`, recomputes hourly_avg client-side too (so UI reflects instantly even before next hourly flush); backend hourly file ensures persistence across restarts.
5. **Storage & retention:** `hourly.jsonl` append 24 records/day → ~9k/year, trivial. Keep forever; same manual-truncate philosophy (no purge command). Document `jq` one-liner to trim.
6. **UI integration:** new card “Est. days remaining (live hourly)” with large `N days` + sub `3.2 GB/hour avg (24h) • 1.1 TB remaining • updates each 5m`; secondary line wear model. New mini sparkline/bar inside card showing last 24h `gb_written` per hour (or when `<24h`, show available). Wear chart tooltip cross-links.
**API additions:**
```python
GET /api/config → {interval, device, port, version}
GET /api/status → {alive, samples, last_ts, pid, uptime_hours, hourly_samples}
GET /api/hourly → [hourlyRec, ...] # ETag + no-store
GET /api/data → unchanged + ETag
# no backup/restore
```
## 6. Layout & shadcn Mapping (Dense + Side-by-Side)
### 6.1 Dense Cards + Flexible Graphs
- **Cards:** tighter — Tailwind: `grid gap-3` (was 1rem), `grid-cols-2 md:grid-cols-3 xl:grid-cols-5`, card `p-3` (was 1rem 1.2rem), `rounded-lg` (was 10px), label `text-[0.70rem] tracking-wide`, value `text-xl font-semibold` (was 1.7rem), sub `text-xs`. Row height uniform via `min-h-[96px]`.
- **Graphs:** `section.grid.grid-cols-1.lg:grid-cols-2.gap-4` — two `Card`s side-by-side on ≥1024px, stacked below. Each Card: `p-4`, header `CardTitle` 0.9rem, canvas `h-[200px] lg:h-[220px] w-full` (down from 220 full-width), `flex-1 min-w-0` so canvases shrink. `drawLine` canvas `width=clientWidth, height=200` → responsive. No fixed page width; container `max-w-[1400px] mx-auto px-4`.
- **Flex guarantees:** `canvas { width:100%; height:100%; display:block }`, `chart-wrap { flex:1 min-w-0 }`, `canvas` DPR scaled but CSS size flexes. ResizeObserver re-draw on container resize.
- **Overall page:** less vertical scroll — header sticky + dense cards + two graphs in one row + footer.
| Current | shadcn | Tight classes |
|---|---|---|
| `.grid .card` | `Card` | `rounded-lg border bg-card p-3 shadow-sm gap-3` |
| value | `CardTitle` numeric | `text-xl font-semibold tabular-nums` |
| wear % | `Progress` | `h-1.5 bg-primary` |
| Spare/errors | `Badge` + `Alert` | `text-xs px-1.5 py-0` |
| charts | `Card` flex | `h-[200px] lg:h-[220px] p-4` |
| footer | muted | `text-xs text-muted-foreground` |
## 7. Branding Spec
Header sticky `[symbol 28px | wordmark] Fenris — NVMe Wear Monitor` left, `[Live ● next 03:42] [every 5m]` right, `bg-background/80 backdrop-blur`. Favicon `favicon-32.png`, `icon-dark-512`. Footer “© Bongbetic — Fenris · interval 300s · v0.1” + `b_glyph.svg` 16px + link to diagnostics count. No `data/backups`.
## 8. File / Code Changes (Option A)
**`fenris.py`:**
- Delete data-mgmt: `BACKUP_DIR`, 5 cmds, parsers; keep `start/stop/status/run/sample`.
- Add globals `_interval`, `_device`, `_port`; vendored `assets/` static branch.
- New `data/hourly.jsonl` + helpers `append_hourly()`, `load_hourly()`, `flush_hourly()`, `rebuild_hourly_from_history()`.
- `collector_loop`: on each sample also `update_hour_bucket`; hourly flush.
- Endpoints: `/api/config`, `/api/status`, `/api/hourly`, ETag on `/api/data`+`/api/hourly`, static `/assets/*`.
- Replace `DASHBOARD_HTML`: Tailwind CDN + shadcn vars, sticky header logos, dense grid, side-by-side Cards (200px canvases, flex), hours-in-day write chart (§5.1), live hourly forecast card (§5.2) with `hourly_avg` + `days_remaining` + 24h bar sparkline, midnight caption, footer glyph.
**`fenris.sh`:** remove items 6–10 branches, renumber Help/Exit, update `help_text`.
**`README.md`:** drop Data Management; add “Hourly diagnostics `data/hourly.jsonl` (GB/hour) + live days-remaining forecast after 24h” + `jq` truncate note.
**Assets:**
```bash
mkdir -p assets
cp ~/Documents/bongbetic/Logo/bongbetic-logo-dark.svg assets/
cp ~/Documents/bongbetic/Logo/bongbetic-symbol-dark.svg assets/
cp ~/Documents/bongbetic/Logo/bongbetic-brand/favicon-32.png assets/favicon.ico
cp ~/Documents/bongbetic/Logo/bongbetic-brand/icon-dark-512.png assets/
```
Keep `collector_loop` otherwise unchanged; no Vite.
## 9. Implementation Phases (Option A)
**Phase 0 — Approved** (this plan): Option A + removals + hours-in-day + dense layout + hourly model.
**Phase 1 — Deletion + API + hourly store:** delete data-mgmt, add config/status/hourly endpoints, hourly rollup helpers + rebuild, verify `curl /api/hourly | jq`.
**Phase 2 — Shell/layout:** dense cards, side-by-side flex graphs (200px), Tailwind CDN + shadcn vars, header/footer logos.
**Phase 3 — Live JS + charts:** interval-synced poll, diff patch, writeChart 0–24h (§5.1), wear chart, visibility/backoff, ETag 304.
**Phase 4 — Hourly forecast live:** `hourly_avg` rolling 24h, TBW-derived `days_remaining` + wear cross-check, warming-up badge <24h, per-hour GB bar in forecast card, midnight reset; wire `GET /api/hourly`.
**Phase 5 — Polish:** `Progress/Badge` variants, `Skeleton`, temp alert, responsive, `prefers-color-scheme`, ResizeObserver.
## 10. Risks
| Risk | Mitigation |
|---|---|
| Tailwind CDN offline | Vendor `assets/tailwind.css` offline fallback. |
| Logo currentColor | Inline SVG, test Chrome/Firefox. |
| Interval desync | Python global single source. |
| Flicker | `requestAnimationFrame` + keep rows. |
| History large | `hourly.jsonl` is compact; keep `/api/data` full v1, add `?since=` v2. |
| Data-mgmt removal confusion | `--help` no backup/restore; log orphan `backups/`. |
| Midnight reset | Caption “Resets at midnight — today only”. |
| <24h forecast noisy | Flag “preliminary (n=Xh)”, EWMA; show both GB and wear models. |
| TBW derivation unstable when pct=0 | Fallback to wear model only; badge “TBW unknown”. |
| Side-by-side overflow | `min-w-0 flex-1` + `grid-cols-1 lg:grid-cols-2` + 200px height; test 1024/1440. |
## 11. Verification
- [ ] `python3 -m py_compile fenris.py` / `bash -n fenris.sh` pass
- [ ] `python3 fenris.py backup` → unknown command; menu 1–5 only
- [ ] `python3 fenris.py start --interval 10` → badge `every 10s`
- [ ] `/api/config`, `/api/hourly`, `/api/data` 200; 304 when unchanged; `curl /assets/bongbetic-logo-dark.svg` svg+xml
- [ ] Cards tight: 5/col xl, p-3, no overflow; graphs side-by-side lg, stacked sm, each ~200px, flex resize no clipping
- [ ] writeChart X `00:00 06:00 12:00 18:00 24:00` today-only; midnight resets to 0
- [ ] Kill/restart daemon → `/api/hourly` rebuilt from history, no dupe hours
- [ ] `<24h`: forecast badge “Warming up (n=Xh) ~Y days (preliminary)”
- [ ] `≥24h`: let run 24h (or fake hourly file with 24 records) → card shows `hourly_avg` over 24, `days_remaining` updates each poll (change write load → forecast moves within one interval); bar of 24h gb/hour visible
- [ ] Tab hidden → pause, visible → immediate fetch; stale banner > interval*2
- [ ] `prefers-color-scheme` legible
## 12. Out of Scope
- Vite/React build, SSE, auth, purge — removed intentionally.
---
*Option A — dense side-by-side + hourly live forecast — ready to implement.*