feat: Fenris persistent TUI monitoring redesign

Implement the complete redesign per fenris-redesign spec:

- Observation store: SQLite WAL mode, six entities, schema versioning
- Collector: smartctl acquisition, sysfs identity, normalization
- Projection: sustained regime rate, habit change, confidence states
- Panes TUI: Textual keyboard-first layout with four normative regions
- Status CLI: read-only composition with four service facts
- Monitor helper: polkit-guarded toggle, collect, baseline ops
- Legacy migration: idempotent single-transaction import
- Hour classification, day aggregates, monitoring periods
- Pruning, segmentation, drive health facts

Cross-cutting acceptance sweep (CI-1 through CI-4):
- 59 tests covering state matrix, TUI/CLI parity, prohibition set,
  required wording and six disclosures
- Full suite: 289 tests, all green

Issues #20, #32 closed.
This commit is contained in:
xavierk
2026-09-02 11:48:08 +05:30
parent 1b141206b1
commit 2b05267690
7 changed files with 841 additions and 1317 deletions
+16 -5
View File
@@ -122,7 +122,7 @@ upgrade: dist/fenris-*.whl
@echo "=== Installing new wheel with pinned dependencies ==="
@sudo $(VENV_DIR)/bin/pip install dist/fenris-*.whl --quiet
@echo "=== Syncing units ==="
@echo "=== Syncing units against manifest ==="
@sudo install -m 0644 units/fenris-collect.timer $(UNIT_DIR)/
@sudo install -m 0644 units/fenris-collect.service $(UNIT_DIR)/
@sudo install -m 0644 polkit/com.bongbetic.fenris.monitor.policy $(POLKIT_DIR)/
@@ -146,10 +146,21 @@ upgrade: dist/fenris-*.whl
@echo "$(CONF_DIR)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(MANIFEST)" | sudo tee -a $(MANIFEST) > /dev/null
@if systemctl is-active --quiet fenris-collect.timer; then \
echo "=== Restarting timer (active) ==="; \
sudo systemctl restart fenris-collect.timer; \
fi
@echo "=== Restarting timer only if contents changed and active (IN-5) ==="
@for unit in fenris-collect.timer fenris-collect.service; do \
TMPFILE=$$(mktemp); \
sudo systemctl cat $$unit > $$TMPFILE 2>/dev/null || true; \
if ! diff -q $$TMPFILE $(UNIT_DIR)/$$unit > /dev/null 2>&1; then \
if systemctl is-active --quiet $$unit; then \
echo " $$unit changed and active — restarting"; \
sudo systemctl restart $$unit; \
fi; \
fi; \
rm -f $$TMPFILE; \
done
@echo "=== Applying forward-only schema migrations (IN-5, IN-6) ==="
@sudo $(VENV_DIR)/bin/python3 -c "from fenris.store import migrate_to_latest; from pathlib import Path; n = migrate_to_latest(Path('$(DATA_DIR)/observations.db')); print(f' Migration steps applied: {n}') if n else print(' Schema already current')"
@echo "=== Upgrade complete ==="
+89 -120
View File
@@ -1,157 +1,126 @@
<p align="center">
<picture>
<source srcset="assets/bongbetic-brand/wordmark-light.png" media="(prefers-color-scheme: dark)">
<img src="assets/bongbetic-brand/wordmark-dark.png" alt="Bongbetic" width="260">
</picture>
<br>
<sub>crafted with stubborn curiosity by <a href="https://bongbetic.com">Bongbetic</a></sub>
</p>
# Fenris 🐺
<p align="center">
<img src="assets/bongbetic-brand/b_glyph.svg" width="48" alt="Fenris glyph">
</p>
*Observes an NVMe drive's real-world use and translates that history into an understandable endurance outlook.*
<h1 align="center">Fenris 🐺 — Your SSD's Tell-All Diary</h1>
<p align="center">
<em>Your NVMe drive has been keeping secrets. Fenris makes it confess — in real time.</em>
<br>
<em>How much did you write today? How long until it taps out? No fairy dust — just your actual bytes.</em>
</p>
Fenris is a persistent TUI monitor backed by a short-lived privileged collector on a systemd timer. It reads SMART data every few minutes, stores compact observation history in SQLite, and recomputes a usage-adjusted theoretical lifespan on every screen render — no fairy dust, just your actual bytes.
---
Fenris is a tiny, stubborn daemon that eavesdrops on your NVMe drive's SMART gossip, writes it down every few minutes, and serves you a live dashboard that actually means something. Not "vibes" — **real GB written in the last 24 hours, real GB/hour, and a real countdown in hours, days, and years until your drive's endurance runs out**.
## Requirements
> Think of it as a Fitbit for your SSD. Except it doesn't nag you to drink water.
- **Python ≥ 3.9** (verified at install time)
- **smartmontools** (`smartctl` — verified at install time)
- **systemd** with a polkit agent (the collector runs as root oneshot; elevation is exclusively polkit)
## What it actually does (no hand-waving)
No other OS packages or Python dependencies beyond [Textual](https://textual.textualize.io/) (pinned in the lockfile).
- **Listens** — polls `smartctl -j` on your NVMe device (default every 5 minutes, you pick).
- **Remembers** — appends every sample to `data/history.jsonl` and rolls up per-hour totals into `data/hourly.jsonl` (survives restarts, rebuilds itself if you yank the power).
- **Calculates** — rolling 24-hour window: *exact* bytes written in the last 24h, GB/h, GB/day, implied total TBW from `percentage_used`, remaining TB, and a projected life-remaining breakdown. Warming-up badge until it has 24h of coverage — no fake confidence.
- **Shows off** — dense, live dashboard with wear-over-time + trailing-24h per-hour bars, sticky header, live countdown, and stale warnings if the daemon dozes off.
## You need
- **Python 3.7+**
- **smartmontools** (`smartctl`)
- Root-ish access to read NVMe SMART (passwordless `smartctl` or just run with `sudo` — your call)
### The sudo dance (one time)
Fenris runs `sudo -n smartctl ...` so it doesn't get stuck asking for a password mid-nap:
## Install
```bash
sudo visudo
# add this line (swap in your username):
youruser ALL=(root) NOPASSWD: /usr/sbin/smartctl
sudo make install
```
No sudo? Run the whole thing with `sudo` and it'll still behave.
What it does:
1. Builds a wheel from the checkout and installs it — with pinned dependencies — into the dedicated venv at `/opt/fenris`.
2. Places the `fenris` wrapper in `/usr/local/bin`, helpers in `/usr/libexec/fenris`, systemd units in `/etc/systemd/system`, and the polkit policy in `/usr/share/polkit-1/actions/`.
3. Creates `/var/lib/fenris` (root-written, group-readable) — the observation store is created lazily by the first collection run.
4. Records every placed file in a manifest consumed by upgrade and uninstall.
5. Detects `./data/history.jsonl` beside the source checkout and runs the idempotent legacy import if present.
## Get it running — 30 seconds
### The cozy way
**A fresh install is fully dormant.** Units are present but disabled; nothing runs. The only opt-in is the sanctioned toggle:
```bash
./fenris.sh
# pick 1) Start monitoring → choose device / interval / port → done
fenris monitor resume # enable timer + open first monitoring period
fenris monitor pause # close the period, disable timer
```
### The no-nonsense way
## Upgrade
```bash
python3 fenris.py start # defaults: /dev/nvme0, every 300s, port 8420
python3 fenris.py start --interval 60 --port 9000 # if you're impatient
python3 fenris.py status # "are we live? how's the drive?"
python3 fenris.py sample # one sneaky sample right now
python3 fenris.py stop # tuck it back in
sudo make upgrade
```
Dashboard lives at **http://localhost:8420** (or whatever port you chose).
What it does:
1. Snapshots `observations.db` to a one-generation backup (`.bak`).
2. Installs the new wheel into the same venv with pinned dependencies.
3. Syncs units and polkit against the manifest; runs `daemon-reload`.
4. Restarts the timer **only** if unit contents changed **and** it is active — a running collection run finishes on its mapped interpreter; the next run uses the new code.
5. Applies forward-only schema migrations (the store directory is never rebuilt; automatic downgrade does not exist).
## The menu, demystified
Rollback: reinstall the previous version and restore `observations.db.bak`.
Run `./fenris.sh` and you'll get:
## Uninstall and purge
```
1) Start monitoring (background daemon + dashboard)
2) Stop monitoring
3) Status / current wear stats
4) Take one sample right now
5) Open dashboard URL
---
h) Help / how this works
q) Exit (go touch grass)
```bash
make uninstall # removes artifacts, preserves config and observation history
make purge # also removes /etc/fenris and /var/lib/fenris
```
## What Fenris jots down
Uninstall performs the sanctioned disable first (`fenris-monitor disable --now`) — an open period closes `user_disabled` — then removes the venv, helpers, units, polkit policy, and wrapper while keeping `/etc/fenris` and the observation store. Reinstalling resumes from the preserved store.
| Field | What's the gossip? |
|-------|---------------------|
| `percentage_used` | The drive's own wear-o-meter (0–100%) |
| `bytes_written` / `bytes_read` | Lifetime totals — the receipts |
| `available_spare` | Spare blocks left (%) |
| `media_errors` | Uncorrectable boo-boos |
| `power_on_hours` | How long it's been awake |
| `temperature_c` | Is it sweating? |
| `critical_warning` | NVMe's panic flags |
## Cadence drop-ins
Hourly rollups also stash `bytes_written` per hour, `pct_start`/`pct_end`, and temp peaks — so the 24h math stays honest.
The default collection cadence is **5 minutes** (`OnUnitInactiveSec=5min` in the timer unit). To change it, place a systemd drop-in:
## The dashboard — what's on screen
```bash
sudo systemctl edit fenris-collect.timer
# Add:
# [Timer]
# OnUnitInactiveSec=10min
```
- **Hero card: Projected life remaining** — big, friendly `361 d 2 h` (plus `≈ 361 days · ≈ 8666 hours · ≈ 0.99 years`), backed by `~280 GB/day` and `~101 TB left of ~202 TB total` on the test box.
- **Data written (24h)** — exact GB in the rolling window + coverage (`10.4h of 24h` until warmed up).
- **Write rate** — GB/h and GB/day, live.
- **Wear, spare, temp, errors, power-on** — the usual suspects, with progress bars and polite color-coding.
- **Two charts, side by side:** wear over time + trailing-24h hourly write bars (with a cheeky "now" bar for the current partial hour).
- **Live plumbing:** polling synced to your interval, ETag-cached, countdown to next sample, warming-up + stale banners, pauses when you hide the tab (saves your battery, you're welcome).
No interval key exists in `/etc/fenris/fenris.conf`. Cadence is a systemd concern, not a Fenris configuration key.
**API for the tinkerers:** `GET /api/data` · `/api/hourly` · `/api/summary` · `/api/config` · `/api/status` — all JSON, all friendly.
## CLI reference
| Command | Behavior |
|---|---|
| `fenris` | Opens the TUI (no arguments). |
| `fenris status` | Projection facts, enabled/active state, last collect outcome, journal hint on failure or staleness. Never auto-samples. |
| `fenris sample` | On-demand collection via the privileged helper. Blocks until the run completes. |
| `fenris monitor pause` | Sanctioned disable — asks for confirmation, then disables the timer and closes the monitoring period. |
| `fenris monitor resume` | Sanctioned enable — enables the timer and opens a monitoring period. No confirmation. |
| `fenris baseline set <json>` | CLI-side validation, then polkit-guarded persistence. |
| `fenris baseline clear` | Remove the endurance baseline. |
| `fenris import <path>` | Idempotent single-transaction legacy import. |
| `fenris start` / `stop` / `run` | Rejected with a one-line migration pointer — never aliased. |
| `fenris --device` | Rejected with a pointer to the configuration file. |
## Retired menu options
The legacy `fenris.sh` menu script and the `fenris.py` monolith have been removed. Here's where the old options went:
| Legacy option | Successor |
|---|---|
| 1) Start monitoring | `fenris monitor resume` |
| 2) Stop monitoring | `fenris monitor pause` |
| 3) Status / current wear stats | `fenris status` |
| 4) Take one sample right now | `fenris sample` |
| 5) Open dashboard URL | Removed — the HTML dashboard and HTTP server are gone; the TUI is the primary interface. |
## Configuration
`/etc/fenris/fenris.conf` holds exactly one key — the device selector:
```
device = /dev/disk/by-id/nvme-Samsung_SSD_980_PRO_2TB_S6BENS0Txxxxx
```
Use a stable `/dev/disk/by-id/` path. Raw `/dev/nvmeX` paths are warned against. The file is re-read every collection run.
## Where's my stuff?
```
fenris/
├── fenris.py # the whole show — daemon + server + math
├── fenris.sh # the cozy menu
├── README.md # hi — you're here
├── assets/bongbetic-brand/ # Bongbetic wordmarks & glyphs (for Gitea + dashboard)
└── data/
├── history.jsonl # raw samples (JSONL, append-only)
├── hourly.jsonl # per-hour rollups (auto-rebuilt on restart)
├── fenris.pid # daemon PID
└── fenris.log # daemon chatter
```
## CLI cheat sheet
```bash
python3 fenris.py start [--device /dev/nvme0] [--interval 300] [--port 8420]
python3 fenris.py stop
python3 fenris.py status
python3 fenris.py sample [--device /dev/nvme0]
python3 fenris.py run # foreground mode — what `start` spawns internally
```
## Oops — troubleshooting without the tears
**"smartctl not found"**
```bash
sudo apt install smartmontools # Debian/Ubuntu
sudo pacman -S smartmontools # Arch — you already knew
```
**"needs root" / permission denied**
Set up the passwordless line above, or just `sudo ./fenris.sh`.
**Dashboard says "stale"**
Daemon napped or crashed. `python3 fenris.py status` will tell you. Kick it again with `start`.
**Only 10 hours of data and it says "preliminary"?**
That's honesty, not a bug. It needs 24h of real writes to give a tight estimate. Let it simmer — the number gets sharper every hour.
| Artifact | Location |
|---|---|
| Wrapper | `/usr/local/bin/fenris` |
| Helpers | `/usr/libexec/fenris/fenris-collect`, `fenris-monitor` |
| Units | `/etc/systemd/system/fenris-collect.{timer,service}` |
| Polkit policy | `/usr/share/polkit-1/actions/com.bongbetic.fenris.monitor.policy` |
| Configuration | `/etc/fenris/fenris.conf` |
| Observation store | `/var/lib/fenris/observations.db` |
| Venv | `/opt/fenris` |
| Manifest | `/var/lib/fenris/manifest.txt` |
| Legacy history | `./data/history.jsonl` (auto-imported on install if present) |
---
-1061
View File
File diff suppressed because it is too large Load Diff
-128
View File
@@ -1,128 +0,0 @@
#!/usr/bin/env bash
# Fenris — interactive menu for the NVMe wear monitor & dashboard.
# Created by Bongbetic.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
PY="$SCRIPT_DIR/fenris.py"
PORT_DEFAULT=8420
INTERVAL_DEFAULT=300
banner() {
cat <<'EOF'
_____ _
| __|___ ___ _| |___
| __| -_| | . | _|
|__| |___|_|_|_|___|_|
NVMe wear monitor & live dashboard
Created by Bongbetic
EOF
}
pause() { read -rp "Press Enter to continue..." _; }
detect_device() {
# Query Python's auto-detect for the default device.
python3 -c "import sys; sys.path.insert(0,'$SCRIPT_DIR'); from fenris import detect_device; print(detect_device())" 2>/dev/null || echo /dev/nvme0
}
menu() {
clear
banner
echo
echo " 1) Start monitoring (background daemon + dashboard)"
echo " 2) Stop monitoring"
echo " 3) Status / current wear stats"
echo " 4) Take one sample right now"
echo " 5) Open dashboard URL"
echo " ---"
echo " h) Help / how this works"
echo " q) Exit"
echo
read -rp "Choose an option: " choice
echo
case "$choice" in
1) start_flow ;;
2) python3 "$PY" stop; pause ;;
3) python3 "$PY" status; pause ;;
4) read -rp "Device [default: auto-detect]: " dev
if [ -z "$dev" ]; then python3 "$PY" sample; else python3 "$PY" sample --device "$dev"; fi
pause ;;
5) show_url; pause ;;
h|H) help_text; pause ;;
q|Q) echo "Bye. — Fenris, by Bongbetic"; exit 0 ;;
*) echo "Invalid choice."; pause ;;
esac
}
start_flow() {
read -rp "NVMe device [Enter = auto-detect]: " dev
read -rp "Sample interval in seconds [Enter = ${INTERVAL_DEFAULT}]: " interval
read -rp "Dashboard port [Enter = ${PORT_DEFAULT}]: " port
interval="${interval:-$INTERVAL_DEFAULT}"
port="${port:-$PORT_DEFAULT}"
args=(start --interval "$interval" --port "$port")
if [ -n "${dev:-}" ]; then args+=(--device "$dev"); fi
echo
echo "Note: reading NVMe SMART data needs root."
echo "Fenris runs 'sudo -n smartctl ...' (no-prompt sudo). If this fails,"
echo "either run this menu with sudo, or allow passwordless smartctl via:"
echo " sudo visudo -> youruser ALL=(root) NOPASSWD: /usr/sbin/smartctl"
echo
python3 "$PY" "${args[@]}"
pause
}
show_url() {
if [ -f "$SCRIPT_DIR/data/fenris.pid" ]; then
# Try to read actual port from running process cmdline, else guess default.
local pid port
pid=$(<"$SCRIPT_DIR/data/fenris.pid")
port=$(tr '\0' '\n' < /proc/"$pid"/cmdline 2>/dev/null | grep -A1 -- '--port' | tail -1 || true)
port="${port:-$PORT_DEFAULT}"
echo "Dashboard: http://localhost:${port}"
else
echo "Fenris is not currently running. Start it first (option 1)."
fi
}
help_text() {
cat <<EOF
What Fenris does:
- Periodically reads your NVMe drive's SMART health data (via smartctl),
including "percentage_used" (the drive's own wear indicator), total
bytes written/read, temperature, spare capacity, and error counts.
- Logs every sample to: $SCRIPT_DIR/data/history.jsonl
and per-hour aggregates to: $SCRIPT_DIR/data/hourly.jsonl (GB/hour,
rebuilt from history on restart). Used for the trailing-24h bar chart
and exact rolling-24h write volume.
- Serves a live HTML dashboard (dense layout, interval-synced polling)
with wear-over-time and trailing-24h hourly-write charts, plus a
projected life-remaining estimate in hours/days/years derived from
your actual rolling-24h write rate and implied TBW endurance.
Requirements:
- smartmontools (smartctl) installed.
- Root access to read NVMe SMART logs — either run Fenris via sudo,
or set up passwordless sudo for smartctl (see option 1).
CLI usage (equivalent to this menu):
python3 fenris.py start [--device /dev/nvme0] [--interval 300] [--port 8420]
python3 fenris.py stop
python3 fenris.py status
python3 fenris.py sample [--device /dev/nvme0]
Leave it running in the background (option 1) and check back after a
few days/weeks of normal use — more samples = a more accurate lifespan
estimate.
Fenris — created by Bongbetic.
EOF
}
# Loop instead of recurse to avoid stack overflow.
while true; do menu; done
+26
View File
@@ -94,6 +94,26 @@ def cmd_import(args: argparse.Namespace) -> None:
print("Legacy import: use fenris-import directly")
def cmd_migrate(args: argparse.Namespace) -> None:
"""Apply forward-only schema migrations (IN-5, IN-6).
Called by 'sudo make upgrade'. Raises on newer-schema store.
"""
from fenris.store import migrate_to_latest
from pathlib import Path
store_path = Path("/var/lib/fenris/observations.db")
if not store_path.exists():
print("No observation store found — nothing to migrate.")
return
steps = migrate_to_latest(store_path)
if steps:
print(f"Migration complete: {steps} step(s) applied.")
else:
print("Schema already current.")
def main() -> None:
parser = argparse.ArgumentParser(
prog="fenris",
@@ -145,6 +165,10 @@ def main() -> None:
import_parser.add_argument("path", help="Path to history.jsonl")
import_parser.set_defaults(func=cmd_import)
# Migrate (IN-5, IN-6) — called by upgrade, not for human use
migrate_parser = subparsers.add_parser("migrate", help=argparse.SUPPRESS)
migrate_parser.set_defaults(func=cmd_migrate)
# Rejected commands
for cmd in ["start", "stop", "run"]:
reject_parser = subparsers.add_parser(cmd, help=argparse.SUPPRESS)
@@ -172,6 +196,8 @@ def main() -> None:
args.func(args)
elif args.command == "import":
cmd_import(args)
elif args.command == "migrate":
cmd_migrate(args)
if __name__ == "__main__":
+44 -3
View File
@@ -176,12 +176,53 @@ def _create_schema(conn: sqlite3.Connection):
def _apply_migrations(conn: sqlite3.Connection, current_version: int):
"""Apply forward-only migrations from current_version to SCHEMA_VERSION."""
# Future migrations will go here
# For now, just upgrade to current version
"""Apply forward-only migrations from current_version to SCHEMA_VERSION.
Each migration step is a transactional block. Add new steps as sequential
elif branches when SCHEMA_VERSION increases.
Spec: §3.6, §10.2
"""
# Migration 1→2: example placeholder
# if current_version < 2:
# conn.execute("ALTER TABLE ...")
# current_version = 2
pass
def migrate_to_latest(store_path: Path) -> int:
"""Apply forward-only migrations to bring the store to SCHEMA_VERSION.
Called by the upgrade target (§10.2). Returns the number of migration
steps applied. Raises ValueError on newer-schema store (§3.6, §9.5).
Spec: §3.6, §10.2, §10.3
"""
conn = sqlite3.connect(str(store_path))
conn.execute("PRAGMA journal_mode=WAL")
cursor = conn.execute("PRAGMA user_version")
current_version = cursor.fetchone()[0]
if current_version > SCHEMA_VERSION:
conn.close()
raise ValueError(
f"Observation store written by a newer Fenris (version {current_version}) "
f"— upgrade Fenris"
)
if current_version == SCHEMA_VERSION:
conn.close()
return 0 # Already up to date
steps = SCHEMA_VERSION - current_version
_apply_migrations(conn, current_version)
conn.execute(f"PRAGMA user_version={SCHEMA_VERSION}")
conn.commit()
conn.close()
return steps
def is_store_faulty(store_path: Path) -> bool:
"""Check if the store is present but cannot be read or trusted."""
if not store_path.exists():
+666
View File
@@ -0,0 +1,666 @@
"""Cross-cutting acceptance sweep (issue #32).
Systematic verification of every acceptance criterion that spans multiple
subsystems. Grouped by criterion ID; each test cites its clause.
CI-1 Exhaustive state matrix: confidence × freshness × baseline tier
CI-2 TUI/CLI parity: identical outcomes and wording
CI-3 Prohibition set: automated structural checks
CI-4 Required wording and six disclosures in both views
"""
import re
import sqlite3
from datetime import datetime, timedelta, timezone
from pathlib import Path
from unittest.mock import patch
import pytest
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store, SCHEMA_VERSION
from fenris.monitoring_periods import ensure_period_open
from fenris.projection import (
compute_projection,
ConfidenceState,
BaselineTier,
DISCLOSURES,
STALENESS_HOURS,
WARMING_MIN_DAYS,
YOUNG_REGIME_DAYS,
)
from fenris.status import (
grade_freshness,
get_status,
render_status,
format_disclosures,
FRESH_THRESHOLD_S,
STALENESS_THRESHOLD_S,
CADENCE_DEFAULT_S,
ACCURACY_SEC,
)
from fenris.tui import (
FenrisTuiApp,
_format_remaining,
)
SRC_DIR = Path(__file__).parent.parent / "src"
FENRIS_PKG = SRC_DIR / "fenris"
def _clock(year=2026, month=9, day=30, hour=12):
return datetime(year, month, day, hour, 0, 0, tzinfo=timezone.utc)
def _insert_baseline(conn, tbw_tb=1.0, verified=True,
model="Samsung SSD 970 EVO Plus 1TB",
source_url="https://example.com/spec",
doc_rev="v1.0", entry_date="2026-01-01",
nominal_cap=1024000000000):
conn.execute(
"INSERT INTO endurance_baseline "
"(tbw_terabytes, source_url, document_revision, entry_date, model_string, "
" nominal_capacity_bytes, validated_by, verified, created_at, updated_at) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(tbw_tb, source_url, doc_rev, entry_date, model, nominal_cap,
"machine_match" if verified else None, verified,
"2026-01-01T00:00:00+00:00", "2026-01-01T00:00:00+00:00"),
)
conn.commit()
def _insert_segment(conn, opened_at="2026-09-01T00:00:00+00:00",
identity_key="nqn.test", degraded=False,
mn="Samsung SSD 970 EVO Plus 1TB"):
conn.execute(
"INSERT INTO controller_segments "
"(opened_at, identity_key, identity_degraded, subnqn, sn, mn, fr, vid, ssvid, transport) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(opened_at, identity_key, degraded, "nqn.test", "SN123", mn, "FW1",
"0x144d", "0x144d", "pcie"),
)
conn.commit()
def _insert_day(conn, day, bw=1024*1024*100, coverage=0.95, samples=24):
conn.execute(
"INSERT INTO day_aggregates (day, active_seconds, idle_seconds, "
"powered_off_seconds, unknown_seconds, bytes_written_delta, "
"bytes_read_delta, sample_count, coverage) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
(day, 3600, 0, 0, 0, bw, 0, samples, coverage),
)
conn.commit()
def _insert_sample(conn, ts, pu=5):
conn.execute(
"INSERT INTO samples (ts, device, data_units_written, data_units_read, "
"percentage_used, bytes_written, bytes_read, power_on_hours) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?)",
(ts, "/dev/nvme0n1", 1000000, 500000, pu, 512000000000, 256000000000, 8765),
)
conn.commit()
def _open_period(conn, start="2026-09-01T00:00:00+00:00"):
ensure_period_open(conn, datetime.fromisoformat(start))
def _setup_full_store(conn, *, baseline=True, segment=True, days=30,
bw=1024*1024*100, coverage=0.95, samples_per_day=24,
sample_ts="2026-09-30T10:00:00+00:00",
period_start="2026-09-01T00:00:00+00:00",
segment_opened="2026-09-01T00:00:00+00:00",
baseline_kw=None, segment_kw=None):
if baseline:
_insert_baseline(conn, **(baseline_kw or {}))
if segment:
_insert_segment(conn, opened_at=segment_opened, **(segment_kw or {}))
_open_period(conn, start=period_start)
for i in range(days):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=bw, coverage=coverage, samples=samples_per_day)
if sample_ts:
_insert_sample(conn, sample_ts)
# ===================================================================
# CI-1: Exhaustive state matrix
# ===================================================================
class TestCI1StateMatrix:
"""Systematic walk of confidence x freshness x baseline tier combinations."""
def test_no_baseline_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert proj.headline_remaining_seconds is None
assert proj.baseline_tier == BaselineTier.NONE
conn.close()
def test_verified_baseline_possible_supported(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.SUPPORTED
assert proj.baseline_tier == BaselineTier.VERIFIED
assert proj.headline_remaining_seconds is not None
conn.close()
def test_unverified_baseline_possible_limited(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(
tbw_tb=10.0, verified=False, source_url=None))
proj = compute_projection(conn, _clock())
assert proj.baseline_tier == BaselineTier.UNVERIFIED
assert proj.confidence_state != ConfidenceState.SUPPORTED
conn.close()
def test_model_mismatch_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(model="Different Model"))
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert proj.baseline_tier == BaselineTier.NONE
conn.close()
def test_fresh_sample_grades_fresh(self, tmp_path):
now = _clock()
ts = (now - timedelta(seconds=FRESH_THRESHOLD_S - 10)).isoformat()
assert grade_freshness(ts, now) == "fresh"
def test_missed_sample_grades_missed(self, tmp_path):
now = _clock()
ts = (now - timedelta(hours=2)).isoformat()
assert grade_freshness(ts, now) == "missed"
def test_stale_sample_grades_stale(self, tmp_path):
now = _clock()
ts = (now - timedelta(hours=49)).isoformat()
assert grade_freshness(ts, now) == "stale"
def test_empty_store_grades_empty(self, tmp_path):
now = _clock()
assert grade_freshness(None, now) == "empty"
def test_unsupported_fresh(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100)
fresh_ts = (_clock() - timedelta(seconds=60)).isoformat()
_insert_sample(conn, fresh_ts)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert grade_freshness(fresh_ts, _clock()) == "fresh"
conn.close()
def test_limited_young_regime(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(5):
d = (datetime(2026, 9, 25) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
conn.close()
def test_limited_warming(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(10):
d = (datetime(2026, 9, 20) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
assert proj.warming_fact is not None
conn.close()
def test_limited_stale_data(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn, opened_at="2026-08-01T00:00:00+00:00")
_open_period(conn, start="2026-08-01T00:00:00+00:00")
for i in range(30):
d = (datetime(2026, 8, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
stale_ts = (_clock() - timedelta(days=5)).isoformat()
_insert_sample(conn, stale_ts)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
assert proj.staleness_fact is not None
conn.close()
def test_limited_degraded_identity(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn, identity_key=None, degraded=True)
_open_period(conn)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
assert proj.degraded_identity_fact is not None
conn.close()
def test_unsupported_zero_rate(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=0)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert proj.zero_rate_fact is not None
conn.close()
def test_headline_present_when_projection_exists(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
proj = compute_projection(conn, _clock())
assert proj.headline_remaining_seconds is not None
assert proj.headline_remaining_seconds > 0
conn.close()
def test_headline_absent_when_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
proj = compute_projection(conn, _clock())
assert proj.headline_remaining_seconds is None
conn.close()
def test_headline_absent_when_zero_rate(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=1.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=0)
proj = compute_projection(conn, _clock())
assert proj.headline_remaining_seconds is None
conn.close()
def test_facts_always_list(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
proj = compute_projection(conn, _clock())
assert isinstance(proj.contributing_facts, list)
conn.close()
def test_facts_never_empty_for_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
proj = compute_projection(conn, _clock())
assert len(proj.contributing_facts) > 0
conn.close()
def test_confidence_never_percentage(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
proj = compute_projection(conn, _clock())
assert proj.confidence_state in (
ConfidenceState.UNSUPPORTED, ConfidenceState.LIMITED, ConfidenceState.SUPPORTED)
for f in proj.contributing_facts:
if re.match(r"^\\d+%$", f.strip()):
pytest.fail("Bare percentage in facts: %r" % f)
conn.close()
def test_status_renders_same_state_as_projection(self, tmp_path):
db = tmp_path / "observations.db"
conn = init_store(db)
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
conn.close()
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": True, "timer_active": True,
"last_collect_ok": True, "last_collect_age_s": 60,
"last_collect_reason": None,
}):
status = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "Supported" in status or "supported" in status.lower()
assert "remaining" in status.lower()
# ===================================================================
# CI-2: TUI/CLI parity
# ===================================================================
class TestCI2Parity:
"""Verify TUI and CLI share the same constants, formatting, and wording."""
def test_freshness_constants_shared(self):
from fenris import tui as tui_mod
from fenris import status as status_mod
assert tui_mod.FRESH_THRESHOLD_S == status_mod.FRESH_THRESHOLD_S
assert tui_mod.STALENESS_THRESHOLD_S == status_mod.STALENESS_THRESHOLD_S
def test_grade_freshness_shared(self):
from fenris.tui import grade_freshness as tui_gf
from fenris.status import grade_freshness as status_gf
assert tui_gf is status_gf
def test_disclosures_shared(self):
from fenris.projection import DISCLOSURES as proj_disc
from fenris.status import format_disclosures
output = format_disclosures()
for d in proj_disc:
assert d in output
def test_status_four_facts_match_tui_strip(self, tmp_path):
db = tmp_path / "observations.db"
conn = init_store(db)
_insert_segment(conn)
_open_period(conn)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
conn.close()
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": True, "timer_active": True,
"last_collect_ok": True, "last_collect_age_s": 120,
"last_collect_reason": None,
}):
status = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "boot:" in status
assert "timer:" in status
assert "last collect:" in status
assert "freshness:" in status
def test_pause_resume_action_names(self):
tui_keys = {b.key for b in FenrisTuiApp.BINDINGS}
assert "p" in tui_keys
assert "r" in tui_keys
assert "c" in tui_keys
assert "q" in tui_keys
def test_empty_store_greeting_both_views(self, tmp_path):
db = tmp_path / "observations.db"
init_store(db)
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
status = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "no observations yet" in status.lower()
def test_store_fault_phrase_both_views(self, tmp_path):
status_src = (FENRIS_PKG / "status.py").read_text()
tui_src = (FENRIS_PKG / "tui.py").read_text()
phrase = "observation store unreadable"
assert phrase in status_src
assert phrase.lower() in tui_src.lower()
def test_newer_schema_phrase_both_views(self):
status_src = (FENRIS_PKG / "status.py").read_text()
phrase = "observation store written by a newer Fenris"
assert phrase in status_src
def test_status_never_prompts(self):
status_src = (FENRIS_PKG / "status.py").read_text()
assert "input(" not in status_src
# ===================================================================
# CI-3: Prohibition set
# ===================================================================
class TestCI3ProhibitionSet:
"""Structural codebase checks for every prohibition clause."""
def _read_all_sources(self):
files = {}
for py in FENRIS_PKG.glob("*.py"):
files[py.name] = py.read_text()
return files
def test_single_acquisition_path(self):
"""Only fenris-collect may interrogate the device. [2.1, 8.7]
collector.py contains the acquisition functions; collect.py is the
fenris-collect entry point that invokes them. No other module may
reference smartctl.
"""
sources = self._read_all_sources()
allowed = {"collector.py", "collect.py"}
for name, text in sources.items():
if name in allowed:
continue
assert "smartctl" not in text, (
"%s must not contain smartctl" % name
)
def test_no_run_surface(self):
"""No /run/fenris coordination surface. [1.2, 3]"""
sources = self._read_all_sources()
for name, text in sources.items():
assert "/run/fenris" not in text, (
"%s references /run/fenris" % name
)
def test_single_config_key(self):
"""Config holds exactly one key: device. [8.3]"""
status_src = (FENRIS_PKG / "status.py").read_text()
in_read_config = False
config_keys = []
for line in status_src.split("\n"):
if "def read_config" in line:
in_read_config = True
elif in_read_config and line.strip().startswith("def "):
break
elif in_read_config and "key ==" in line:
match = re.search(r'key\s*==\s*["\']([^"\']+)["\']', line)
if match:
config_keys.append(match.group(1))
assert "device" in config_keys
assert len(config_keys) == 1, "Found keys: %s" % config_keys
def test_no_alerting_machinery(self):
"""No alerting, notification, or escalation. [9.6]"""
sources = self._read_all_sources()
alert_keywords = ["send_email", "smtp", "webhook", "push_notification"]
for name, text in sources.items():
for kw in alert_keywords:
for line in text.split("\n"):
stripped = line.strip()
if kw in stripped and not stripped.startswith("#"):
pytest.fail(
"%s contains alerting keyword '%s': %s" % (name, kw, stripped)
)
def test_no_synthetic_baselines(self):
"""No synthetic or capacity-derived baseline. [6.1]"""
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert "synthetic" not in proj_src.lower()
def test_no_stored_projections(self):
"""Projections never stored; recomputed on read. [3.7, 6.10]"""
store_src = (FENRIS_PKG / "store.py").read_text()
create_tables = re.findall(r"CREATE TABLE.*?(?=\n\n|$)", store_src, re.DOTALL)
table_names = []
for ct in create_tables:
m = re.search(r"IF NOT EXISTS\s+(\w+)", ct)
if m:
table_names.append(m.group(1))
assert "projection" not in [t.lower() for t in table_names]
def test_no_partial_newer_schema_interpretation(self):
"""Readers refuse newer-schema stores. [3.6, 9.5]"""
status_src = (FENRIS_PKG / "status.py").read_text()
assert "NewerSchema" in status_src
assert "upgrade Fenris" in status_src
def test_polkit_authorizes_one_binary(self):
"""Polkit authorizes exactly one binary: fenris-monitor. [8.5]"""
monitor_src = (FENRIS_PKG / "monitor.py").read_text()
assert "fenris-monitor" in monitor_src or "fenris_monitor" in monitor_src
collect_src = (FENRIS_PKG / "collect.py").read_text()
assert "polkit" not in collect_src.lower()
def test_no_hour_interpolation(self):
"""No absent hour is interpolated or fabricated. [5.3]"""
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert "interpolat" not in proj_src.lower()
assert "fabricat" not in proj_src.lower()
def test_fenris_sh_not_shipped(self):
"""fenris.sh is not shipped. [8.8]"""
repo_root = Path(__file__).parent.parent
assert not (repo_root / "fenris.sh").exists()
# ===================================================================
# CI-4: Required wording and six disclosures
# ===================================================================
class TestCI4WordingAndDisclosures:
"""Verify exact fixed phrases and disclosures in both views."""
def test_exactly_six_disclosures(self):
assert len(DISCLOSURES) == 6
def test_disclosure_1_endurance_not_failure(self):
assert "endurance projection" in DISCLOSURES[0].lower()
assert "hardware-failure" in DISCLOSURES[0].lower() or "failure date" in DISCLOSURES[0].lower()
def test_disclosure_2_vendor_specific(self):
assert "vendor-specific" in DISCLOSURES[1]
assert "255 is saturated" in DISCLOSURES[1]
def test_disclosure_3_warranty_not_failure(self):
assert "warranty" in DISCLOSURES[2].lower() or "endurance threshold" in DISCLOSURES[2].lower()
assert "failure threshold" in DISCLOSURES[2].lower()
def test_disclosure_4_duw_rounding(self):
assert "DUW" in DISCLOSURES[3]
assert "upward-rounded" in DISCLOSURES[3]
assert "NAND" in DISCLOSURES[3]
def test_disclosure_5_quality_depends(self):
assert "baseline provenance" in DISCLOSURES[4]
assert "future workload" in DISCLOSURES[4]
def test_disclosure_6_gaps_and_disabled(self):
assert "Gaps" in DISCLOSURES[5]
assert "deliberately disabled" in DISCLOSURES[5]
def test_disclosures_render_in_status(self):
output = format_disclosures()
assert output.startswith("Disclosures")
for i in range(1, 7):
assert "%d." % i in output
for d in DISCLOSURES:
assert d in output
def test_disclosures_render_in_tui(self):
tui_src = (FENRIS_PKG / "tui.py").read_text()
assert "format_disclosures" in tui_src
def test_zero_rate_phrase(self):
phrase = "no finite projection from this history"
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert phrase in proj_src
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
def test_unavailable_no_baseline_phrase(self):
phrase = "no applicable endurance baseline"
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert phrase in proj_src
def test_store_fault_phrase(self):
phrase = "observation store unreadable"
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
tui_src = (FENRIS_PKG / "tui.py").read_text()
assert phrase.lower() in tui_src.lower()
def test_newer_schema_phrase(self):
phrase = "observation store written by a newer Fenris"
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
def test_no_observations_phrase(self):
phrase = "no observations yet"
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
tui_src = (FENRIS_PKG / "tui.py").read_text()
assert phrase in tui_src.lower()
def test_config_error_phrase(self):
phrase = "configuration error:"
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
def test_degraded_identity_phrase(self):
phrase = "controller identity unavailable"
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert phrase in proj_src
phrase2 = "replacement detection relies on write-counter continuity only"
assert phrase2 in proj_src
def test_scenario_range_only_spread(self):
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert "confidence interval" not in proj_src.lower()
def test_no_percentage_in_confidence_rendering(self):
for name in ["projection.py", "tui.py", "status.py"]:
src = (FENRIS_PKG / name).read_text()
assert not re.search(r"\\d+%\\s*confidence", src, re.IGNORECASE), (
"Found XX%% confidence in %s" % name
)
def test_status_disclosures_accessible(self):
db = Path("/tmp/_ci4_test.db")
conn = init_store(db)
conn.close()
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = render_status(store_path=db, clock_now=now,
query_services=True, query_journal=False,
show_disclosures=True)
assert "Disclosures" in result
assert "1." in result
assert "6." in result
db.unlink(missing_ok=True)