Compare commits

..
Author SHA1 Message Date
xavierk 2eb4c48fb6 research: evaluate Python TUI frameworks 2026-08-31 13:45:19 +05:30
12 changed files with 69 additions and 473 deletions
-1
View File
@@ -1,5 +1,4 @@
__pycache__/ __pycache__/
.pi/
*.pyc *.pyc
.commandcode/ .commandcode/
data/fenris.pid data/fenris.pid
-9
View File
@@ -1,9 +0,0 @@
## Agent skills
### Issue tracker
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
### Domain docs
This is a single-context repository. See `docs/agents/domain.md`.
-73
View File
@@ -1,73 +0,0 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Store fault**:
The condition where the observation store is present but cannot be read or trusted — unreadable, corrupt, or written by a newer Fenris — degrading every view that depends on it rather than crashing or guessing.
_Avoid_: Database error, corruption, broken data
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Endurance baseline**:
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
-42
View File
@@ -1,42 +0,0 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -1,49 +0,0 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -1,31 +0,0 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger and monitoring-period bookkeeping; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
@@ -1,31 +0,0 @@
# 4. Installation lifecycle: Makefile-delivered venv, dormant install, sanctioned teardown
## Status
Accepted — resolves [Define installation, upgrade, and removal behavior](https://git.bongbetic.com/xavierk/Fenris/issues/9) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris runs today from the source checkout (`fenris.py`, `fenris.sh`, `data/`): code, state, and control all live relative to wherever the checkout sits. The redesign fixes system artifacts — helpers in `/usr/libexec/fenris` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), the observation store at `/var/lib/fenris/observations.db` ([ADR 0001](0001-observation-store-sqlite.md)), configuration at `/etc/fenris/fenris.conf` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), Textual on Python 3.9+ ([framework decision](https://git.bongbetic.com/xavierk/Fenris/issues/6)) — but nothing says how those artifacts are delivered, upgraded, or removed, or what happens when the checkout moves or disappears.
## Decision
1. **Delivery.** `sudo make install` builds a wheel from the checkout and installs it, with pinned dependencies, into a dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only: after install, nothing references it.
2. **Layout and manifest.** Units in `/etc/systemd/system` (`fenris-collect.{timer,service}`, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)); helpers in `/usr/libexec/fenris`; polkit policy in `/usr/share/polkit-1/actions/`; configuration and observation store in their ADR-fixed locations. The installer records every file it places in an explicit manifest consumed by upgrade and uninstall.
3. **Privilege.** One root installer (`sudo make install`); at runtime, elevation is exclusively polkit (`auth_admin`, `fenris-monitor` only, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)). The installer never enables or starts units.
4. **Dormant install.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle (`fenris monitor resume [--now]`, or the first-run TUI prompt), which enables the timer and opens the first period in one step.
5. **Legacy import.** The installer detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs [ADR 0001](0001-observation-store-sqlite.md)'s idempotent single-transaction import, and reports imported counts — existing observations never depend on checkout survival. `fenris import <path>` remains available for later finds.
6. **Upgrade.** `sudo make upgrade` builds and installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed and it is active — safe with `Persistent=no`), leaves timer state untouched, and never kills an in-flight collection run: a running oneshot finishes on its mapped interpreter, so at worst one old-code run completes to the store and the next run uses the new code. It then applies forward-only observation-store schema migrations governed by a `schema_version` table. `/var/lib/fenris` is never rebuilt.
7. **Rollback.** Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade is explicitly unsupported.
8. **Removal.** `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. `make purge` additionally removes configuration and store. Journal entries age out naturally.
9. **Dependencies.** Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing.
10. **Scaffolding and floor.** The installer creates `/var/lib/fenris` with [ADR 0001](0001-observation-store-sqlite.md)'s root-written group-read permissions and verifies `python3 ≥ 3.9`, failing cleanly otherwise — the Textual contingency becomes an install-time gate rather than a runtime crash. The database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet.
## Consequences
- Installed Fenris survives checkout deletion; the checkout is only where builds happen.
- Teardown preserves monitoring-period semantics: deliberate removal excludes the uninstalled span from the usage habit instead of leaving it as unknown-inside.
- Installs are reproducible; dependency drift cannot ride in on an upgrade.
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
- Rollback support is exactly one generation deep, no further.
- The README documents install, upgrade, uninstall/purge, and legacy import alongside [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s menu-successor mapping.
@@ -1,27 +0,0 @@
# 5. Failure and recovery: visible degradation, never fabrication
## Status
Accepted — resolves [Define failure and recovery behavior](https://git.bongbetic.com/xavierk/Fenris/issues/10) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The observation store ([ADR 0001](0001-observation-store-sqlite.md)) and the service lifecycle ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) settled single-writer transactions, bounded single failed runs, freshness grading (fresh / missed / stale), and absent-hour classification through power-on-hours evidence. Left open by the [failure ticket](https://git.bongbetic.com/xavierk/Fenris/issues/10): behavior per failure class — malformed observations inside the store, missed observations, store faults (unreadable, corrupt, or newer-schema database), and repeated collector failures — and how the habit record re-anchors after the store itself is lost.
## Decision
1. **Malformed observations — refuse at the write boundary.** The collector validates every row it would write against the store's domain invariants (hour seconds sum to 3600, non-negative DUW delta within a controller segment, coverage consistent with sample count). A violating run writes nothing for that run, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact, but under a single trusted writer they should never see one. Store invariant: everything persisted is well-formed.
2. **Missed observations — never backfill.** Fenris never interpolates, estimates, or fabricates an hour. Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s power-on-hours classification is the only inference admitted.
3. **Store faults — degrade, never recreate over.** An unreadable or corrupt database is a store fault: readers surface a "observation store unreadable" fact with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside, the next run starts a fresh store, and if the legacy import never completed, the still-present `history.jsonl` is re-imported. No built-in destructive command exists.
4. **Newer schema — readers refuse symmetrically.** The TUI and `status` detect a `user_version` newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal in [ADR 0001](0001-observation-store-sqlite.md) and the forward-only upgrade rule of [ADR 0004](0004-install-upgrade-removal-lifecycle.md).
5. **Repeated collector failures — flat cadence, no escalation.** The timer's retry is the recovery path; the settled freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff and no notification machinery; a persistent failure reads as stale exactly like any other gap.
6. **Drive-reported anomalies — facts, not alerts.** `critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status`; no alerting or notification surface exists. Fenris observes and projects; it does not alarm. The projection is unaffected: endurance math consumes writes, not warnings.
7. **Orphaned samples — the collector re-anchors observed fact.** When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent: only the sanctioned path of [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) records a `user_disabled` close. Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
## Consequences
- Validation lives at one boundary — the collector — so the store's contract is "everything in it is well-formed" and readers only defend against the impossible.
- No synthetic data can ever enter the habit record; confidence categories can be trusted to reflect real evidence.
- Store-fault recovery can lose history; the mitigation is the one-file backup story of [ADR 0001](0001-observation-store-sqlite.md), kept human-sanctioned so loss is never silent.
- Period bookkeeping splits by epistemics: the helper records intent, the collector records observed fact.
- Fenris stays fully local and silent: no notification, escalation, or alerting machinery anywhere.
-26
View File
@@ -1,26 +0,0 @@
# Domain Docs
How engineering skills should consume this repository’s domain documentation.
## Layout
This is a single-context repository:
```text
/
├── CONTEXT.md
├── docs/adr/
└── ...
```
## Before exploring
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
## Use the glossary’s vocabulary
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
## Flag ADR conflicts
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
-85
View File
@@ -1,85 +0,0 @@
# Issue tracker: Gitea
Issues for this repository live in Gitea at:
https://git.bongbetic.com/xavierk/Fenris/issues
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
## General operations
- List: `tea issues list`
- Read: `tea issues <index> --comments`
- Create: `tea issues create --title "<title>" --description "<body>"`
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
- Assign: `tea issues edit <index> --add-assignees "<username>"`
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
- Comment: `tea comments add <index> --description "<comment>"`
- Close: `tea issues close <index>`
- Reopen: `tea issues reopen <index>`
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
## When a skill says “publish to the issue tracker”
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
## When a skill says “fetch the relevant ticket”
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
## Wayfinding operations
Wayfinder maps and decision tickets are Gitea issues.
### Map and ticket grouping
- A map has the label `wayfinder:map`.
- Create one milestone named `Wayfinder: <map title>` for the effort.
- Assign the map and all its tickets to that milestone.
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
### Blocking
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
```bash
tea api -X POST \
repos/{owner}/{repo}/issues/<blocked>/dependencies \
-F index=<blocker> \
-f owner=xavierk \
-f repo=Fenris
```
List blockers:
```bash
tea api repos/{owner}/{repo}/issues/<index>/dependencies
```
Remove the relationship with the same payload and `-X DELETE`.
### Frontier
List open issues in the map’s milestone. Exclude:
- the issue labelled `wayfinder:map`
- assigned tickets, because assignment is the claim
- tickets whose dependency query returns any open issue
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
### Claim
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
### Resolve
1. Add the answer as a resolution comment.
2. Close the ticket.
3. Re-fetch the map immediately before editing it.
4. Append a linked one-line context pointer to `Decisions so far`.
5. Create newly visible tickets, then wire dependencies in a second pass.
-99
View File
@@ -1,99 +0,0 @@
# Controller identity for observation-history segmentation
Research for [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11).
## Decision
Record the **subsystem NQN as exposed by the Linux kernel** — normalized, trailing-space stripped — as the controller-segment identity key:
```text
identity_key = strip(subnqn) # /sys/class/nvme-subsystem/…/subsysnqn,
# identical to /sys/class/nvme/nvmeX/subsysnqn
fallback: "nqn.2014.08.org.nvmexpress:" + hex4(vid) + hex4(ssvid)
+ raw20(sn) + raw40(mn) # byte-for-byte the kernel's synthesized NQN
last resort: strip(mn) + "|" + strip(sn) # when only smartctl-style fields exist
```
Store the raw `subnqn`, `sn`, `mn` strings (normalized) plus `fr` (firmware revision) as **segment metadata**, never `fr` inside the key: firmware revision is the one mandatory field that legitimately changes on the same drive ([Base Spec 2.0e §5.17.2.1: FR is the *currently active* firmware revision](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Namespace identifiers (NGUID, EUI-64, UUID) are excluded from the key: they are namespace-scoped while the SMART counters being segmented are controller-scoped, and each may be absent or reused. When the key is blank, record an empty key with a degraded marker and rely on the DUW-monotonic rule; a same-model same-serial replacement is unobservable by any identifier and is caught — as a boundary, not an identity change — by the DUW-decrease rule of ADR 0001.
This key satisfies the ticket's asymmetry: a drive replacement changes `subnqn` (real NQNs are unique per subsystem; kernel-generated ones embed serial+model), while firmware quirks and counter resets on the same drive leave it unchanged — counter discontinuities are already ADR 0001's second segmentation axis.
## What each candidate is
| Candidate | Spec definition | Scope | Mandatory | Verdict for the key |
|---|---|---|---|---|
| **SN + MN** | ASCII strings assigned by the vendor in Identify Controller, bytes 23:04 and 63:24; §4.3 shows them left-justified and space-padded ([2.0e §5.17.2.1, Identify Controller data structure](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf), [§4.3 Identifier Format and Layout](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | NVM subsystem | Mandatory for I/O and Admin controllers | Core of the fallback; uniqueness explicitly not guaranteed by the spec |
| **FR** | Currently active firmware revision, ASCII, bytes 71:64 ([2.0e §5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Domain (subsystem) | Mandatory | Never in the key — it is *meant* to change on the same drive |
| **SUBNQN** | NVM Subsystem NQN, UTF-8 null-terminated, bytes 1023:768; mandatory if the controller is ≥ 1.2.1, otherwise may be all zero ([2.0e §5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | NVM subsystem (shared by all its controllers) | Mandatory ≥ 1.2.1, optional below | **Chosen key**; the spec says hosts *should* use it as the subsystem's unique identifier |
| **NGUID / EUI-64** | IEEE-based identifiers in Identify Namespace, bytes 119:104 / 127:120 ([§4.3.4, §4.3.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Namespace | Optional; may both be zero | Rejected — wrong scope, may be missing, may be reused |
| **CNTLID** | Controller ID, unique only *within* a subsystem ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Controller | Mandatory | Rejected — not unique across subsystems (this host's drive reports 0) |
A note on the PDF: figure numbers in the table of contents of the 2.0e revision are offset from the body captions (e.g. the SN/MN figure is "Figure 128" in the TOC but "Figure 130" in the body), so this document cites section numbers, which are stable.
## Scope: the counters being segmented are controller-scoped
SMART / Health Information (LID 02h) is scope **Controller** (mandatory) with an optional namespace view ([2.0e §5.16.1, Get Log Page – Log Page Identifiers](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Section 5.16.1.3: "The information provided is over the life of the controller and is retained across power cycles"; hosts request the controller log page with NSID `FFFFFFFFh`/`0h`, the per-namespace view is optional (LPA bit 0), and "the controller log page and namespaces specific log page contain identical information" in 2.0e ([§5.16.1.3](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Data Units Written therefore accumulates per controller/subsystem, not per namespace — the identity key must be subsystem-scoped, and SN/MN/SUBNQN are all defined as NVM-subsystem fields ([§5.17.2.1: SN/MN are "the serial number/model number for the NVM subsystem"](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); the Persistent Event Log repeats that its SN/MN/SUBNQN copies are the same subsystem values).
NGUID and EUI-64 live in Identify **Namespace**, not the controller structure ([§4.3.4, §4.3.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); libnvme documents them on `struct nvme_id_ns`, while `sn`/`mn`/`fr` live on [`struct nvme_id_ctrl`](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L1467-L1477) and `subnqn` on the same structure ([types.h](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L1576))). Segmentation keyed on a namespace identifier would split or merge history whenever namespaces are attached, detached, formatted, or recreated, while the DUW counter — the thing being differenced — sails on unchanged. The spec's own namespace-identity guidance (§3.2.1.6: NSIDs "may change across power off conditions"; to detect the same namespace use UUID, NGUID, or EUI-64) addresses a different problem from ours.
## Stability verdicts (reboots, firmware, replacement)
1. **Reboots, same drive** — SN, MN, SUBNQN are stable: they are vendor-assigned subsystem fields, and an NQN "is permanent for the lifetime of the host or NVM subsystem" ([§4.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). SMART data "is retained across power cycles" ([§5.16.1.3](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
2. **Firmware update, same drive** — FR changes by definition (it reports the active revision). The spec guarantees persistence of SMART data across *power cycles*, and says nothing about firmware commits resetting counters — vendor behavior, which is precisely why ADR 0001 keeps the DUW-monotonic rule as an independent boundary. Identity (SN/MN/SUBNQN) is not specified to change with firmware; keeping FR out of the key means a firmware update never quarantines history as a "new drive", and a firmware-induced counter reset is caught by the DUW rule instead.
3. **Drive replacement, different model** — every candidate changes.
4. **Drive replacement, identical model** — MN unchanged; SN changes *if* vendor serials are unique; SUBNQN changes because both real NQNs (empirically this host's Micron embeds the serial: `nqn.2016-08.com.micron:nvme:nvm-subsystem-sn-233542F44436`) and kernel-generated NQNs (which concatenate SN and MN, [drivers/nvme/host/core.c nvme_init_subnqn](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3164-L3194)) derive from the serial.
5. **NGUID/EUI-64 under namespace churn** — not stable in the needed sense: if the UIDREUSE bit is 0 "a controller **may reuse** a non-zero NGUID/EUI64 value for a new namespace after the original namespace using the value has been deleted" ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)); libnvme's own field docs say the values hold only "throughout the life of the namespace", "preserved across namespace and controller operations" ([types.h nguid/eui64](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L2648-L2656)).
## Availability and known pathologies
- **SN/MN**: mandatory for I/O and Admin controllers ([§5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) and exposed by the kernel since 4.5 ([sysfs-nvme: /sys/class/nvme/nvmeX/{model,serial,firmware_rev}](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L1-L10)). But uniqueness is disclaimed: "The mechanism used by the vendor to assign Serial Number and Model Number values to ensure uniqueness is outside the scope of this specification" ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Duplicate serials across units are therefore spec-legal.
- **SUBNQN**: zero on pre-1.2.1 subsystems ([§5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); the Persistent Event Log likewise defines the not-supported case as all bytes cleared to 0h), and some real devices report garbage: the kernel carries `NVME_QUIRK_IGNORE_DEV_SUBNQN` for, among others, Intel P4500/P4600, Intel 760p/Pro 7600p, and a Silicon Motion device ([pci.c quirk table](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/pci.c#L4121-L4144), [nvme.h flag definition](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/nvme.h#L103-L106)). This is why Fenris should read the **kernel-exposed** value rather than the raw Identify bytes: when the device NQN is missing, invalid, or quirk-ignored, `nvme_init_subnqn` synthesizes `nqn.2014.08.org.nvmexpress:{vid}{ssvid}{sn}{mn}` — mirroring the spec's own construction for pre-1.2.1 subsystems ([§4.5.1, "NQN Construction for Older NVM Subsystems"](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf), which composes the NQN starting string, VID, SSVID, SN, MN) — so `/sys/.../subsysnqn` is populated on every kernel ≥ 4.8 for every controller ([sysfs-nvme subsysnqn entry, added 4.8](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L48-L60)). The kernel also uses the NQN as *the* subsystem key when building multipath heads ([core.c: subsystems are matched by `subsys->subnqn`](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3345-L3350)).
- **Virtual controllers / blank serials**: the kernel's own NVMe target (nvmet, the `loop` transport) sets model number to the literal `"Linux"` and generates a **random** serial per subsystem "as our controllers are ephemeral" ([target/core.c nvmet_subsys_alloc](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/core.c#L1840-L1856), [nvmet.h `NVMET_DEFAULT_CTRL_MODEL`](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/nvmet.h#L31)); Identify then reports those values verbatim ([target/admin-cmd.c](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/admin-cmd.c#L670-L677)). For such devices `model|serial` is unstable across target re-creation, while the NQN is the configured subsystem name. No identifier can make an ephemeral virtual drive look like stable hardware; the degraded marker covers it.
- **NGUID/EUI-64**: optional — the kernel sysfs attributes are documented as "Hidden if all zeros" ([sysfs-nvme nguid/eui entries](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L259-L293)), and the spec requires only that *at least one* of EUI64/NGUID/UUID be valid at namespace creation ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
## The key, normalization, and failure handling
1. **Primary key**: `strip(subnqn)` read from the kernel path (via libnvme; see next section). Values are ASCII/UTF-8 with code values 0x20–0x7E, left-justified and space-padded per the spec's string rules ([§1.4.2 ASCII/UTF-8 string conventions](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)); normalize by stripping trailing (and leading) spaces only. **No case folding**: NQNs are compared "as binary strings without any text processing (e.g., case folding)" ([§4.5, NQN processing rules](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)), and SN/MN have no canonical case either.
2. **Fallback ladder** (defensive; on Linux ≥ 4.8 the kernel fallback already fires before Fenris ever sees an empty value): (a) device-provided SUBNQN; (b) the kernel's composite `nqn.2014.08.org.nvmexpress:{vid}{ssvid}{sn}{mn}` built from Identify — the kernel spells the date with dots and concatenates the raw fixed-width SN and MN, "slightly different from the format specified" in §4.5.1's NQN construction "for historic reasons" ([core.c nvme_init_subnqn](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3164-L3194)); Fenris's fallback reproduces the kernel spelling so a fallback-built key equals the sysfs value byte-for-byte; (c) `strip(mn)|strip(sn)`. Every rung changes on drive replacement and survives firmware updates; the ladder exists so a missing rung never yields a blank key from a perfectly good physical drive.
3. **Blank key** (all components empty — e.g. a virtual controller reporting nothing): record `identity_key = ""` plus a `degraded` marker and segment by DUW monotonicity alone; surface as a fact, never fabricate uniqueness (matches ADR 0005's "facts, not alerts" posture).
4. **Duplicate keys across physical units** are undetectable by construction when both units report identical SN/MN/SUBNQN. The mitigation already exists in ADR 0001: a replacement drive almost certainly reports a **lower** DUW than the accumulated history, and any DUW decrease forces a segment boundary regardless of identity. Write deltas remain quarantined even though identity cannot distinguish the units.
5. **Recorded metadata per segment**: normalized `subnqn`, `sn`, `mn`, `fr`, plus `transport` — diagnostics for humans, not key components.
## How the collector obtains the key via libnvme
Three concrete paths, all first-party:
1. **libnvme Python bindings** (`from libnvme import nvme`, official SWIG bindings, ["python bindings for libnvme"](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/pyproject.toml#L1-L8)). `nvme.root()` scans sysfs ([nvme.i: nvme_root() calls nvme_scan](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L492-L501)); controller objects expose `model`, `serial`, `firmware`, `subsysnqn`, `name`, `sysfs_dir` as attributes ([nvme.i struct nvme_ctrl attributes](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L402-L424)); namespace objects expose `nsid`, `nguid`, `eui64`, `uuid` ([nvme.i struct nvme_ns](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L481-L490)). These getters read sysfs via `nvme_get_ctrl_attr` ([tree.c populates ctrl fields from sysfs attributes](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/tree.c#L2085-L2097)), and `__nvme_get_attr` **strips the trailing newline and trailing spaces and returns NULL when the result is empty** ([linux.c L526–552](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/linux.c#L526-L552)) — i.e., the bindings deliver exactly the normalization this decision requires, and a blank field arrives as `None`.
2. **Admin passthrough** for raw Identify: `nvme_ctrl_identify(c, &id)` fills `struct nvme_id_ctrl` ([man page: "Issues an 'identify controller' command"](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_ctrl_identify.2#L13-L17)), whose `sn[20]`/`mn[40]`/`fr[8]` and `subnqn[256]` members are documented in the [nvme_id_ctrl(2) man page](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_id_ctrl.2#L5-L15) ([subnqn member](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_id_ctrl.2#L616-L617)); namespace identifiers come from `nvme_ns_identify` filling `struct nvme_id_ns` ([tree.h nvme_ns_identify](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/tree.h#L823-L832)). Here the collector must strip trailing spaces itself.
3. **nvme-cli JSON** (built on libnvme): `nvme id-ctrl -o json` emits `sn`, `mn`, `fr`, `cntlid` with strings copied verbatim *including padding* ([nvme-print-json.c L409–L422](https://github.com/linux-nvme/nvme-cli/blob/c8ec7e41f3738b20828849228a549f4ed1d03fc0/src/nvme-print-json.c#L409-L422)) plus `subnqn`, which is included only when non-empty ([L522–L523](https://github.com/linux-nvme/nvme-cli/blob/c8ec7e41f3738b20828849228a549f4ed1d03fc0/src/nvme-print-json.c#L522-L523)); `nvme list -o json` emits per device `DevicePath`, `Firmware`, `ModelNumber`, `SerialNumber` ([v2.16 json_list_item_obj](https://github.com/linux-nvme/nvme-cli/blob/faf7326a2997dea91687fd3daa17fc405910a4c1/nvme-print-json.c#L4669-L4693)). So stripping trailing spaces is the collector's job in both nvme-cli paths.
What libnvme/nvme-cli offer beyond the current `smartctl -j` collector: `subnqn` (the chosen key), `cntlid`, `vid`/`ssvid`, and the namespace identifiers — smartmontools' NVMe JSON device section carries `model_name`, `serial_number`, `firmware_version` (from `id_ctrl.mn/sn/fr`, [nvmeprint.cpp print_drive_info](https://github.com/smartmontools/smartmontools/blob/9f83095a631ff71df44f8065c7a4a00134d3d404/smartmontools/nvmeprint.cpp#L108-L121)) and no NQN. smartmontools also trims the strings it copies ([utility.cpp format_char_array strips leading/trailing spaces](https://github.com/smartmontools/smartmontools/blob/9f83095a631ff71df44f8065c7a4a00134d3d404/smartmontools/utility.cpp#L692-L708)), so normalization is compatible across both collectors.
## Empirical check (one real drive, sysfs only)
```console
$ cat /sys/class/nvme/nvme0/{model,serial,firmware_rev,subsysnqn}
Micron_2400_MTFDKBA512QFM<spaces to 40> # sysfs preserves Identify padding
233542F44436<spaces to 20>
V3MA001<space> # FR padded to 8
nqn.2016-08.com.micron:nvme:nvm-subsystem-sn-233542F44436
$ cat /sys/block/nvme0n1/{nguid,eui,wwid}
00000000-0000-0001-00a0-752342f44436
00 a0 75 01 42 f4 44 36
eui.000000000000000100a0752342f44436
```
Confirms: raw sysfs keeps the spec's trailing-space padding (strip before storing); the vendor NQN embeds the serial; NGUID/EUI-64 are present here but are per-namespace; `/sys/class/nvme-subsystem/nvme-subsys0/{model,serial,firmware_rev,subsysnqn}` carries the same subsystem-level values ([sysfs-nvme nvme-subsystem entries, added 4.15](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L425-L434)). The kernel's own namespace `wwid` uses the priority ladder `uuid.{UUID}` → `eui.{NGUID}` → `eui.{EUI64}` → `nvme.{VID}-{SERIAL}-{MODEL}-{NSID}` ([sysfs-nvme wwid](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L277-L285)) — a first-party precedent for a serial+model fallback when better identifiers are absent.
## What ADR 0001's segment and migration rules must assume
1. **Legacy `history.jsonl` cannot carry a full identity.** The legacy collector records `device` (`/dev/nvme0`), `model` (smartctl `model_name`), `capacity_bytes`, and the counters — **no serial, no NQN, no firmware** ([fenris.py sample(): the identity-adjacent fields are `device` and `model` only](https://git.bongbetic.com/xavierk/Fenris/src/commit/d45d931a0deaaa4ff8c8ba0ff2454c0d2a211ef9/fenris.py#L73-L76)). Migration may therefore only import legacy samples under a **model-scoped legacy identity**, explicitly labeled incomplete. It cannot prove the legacy samples came from the current drive, cannot detect a same-model swap that happened before migration, and must not backfill serials retroactively.
2. **The first new-version collection run starts a new controller segment** for the monitored drive, recording full `identity_key` plus `sn`/`mn`/`fr`/`subnqn` metadata, because the legacy and new identities are not comparable — the identity change is an epistemic boundary, not a detected drive swap. Imported legacy rows keep their legacy segment; the DUW-monotonic rule already governs deltas within it.
3. **Two independent segmentation axes stay independent**: identity-key change ⇒ boundary (drive replacement); DUW decrease ⇒ boundary (counter reset, firmware quirk, or same-identity replacement). Neither implies the other; ADR 0001's `controller_segments` wording ("boundaries where controller identity changes or DUW decreases") already encodes this, and this research fixes "controller identity" to mean the normalized kernel-exposed subsystem NQN with the fallback ladder above.
4. **The key is a string with rules, not just a field choice**: normalization (space-stripping, no case folding) must be specified once and applied at write time, or the same drive could split its own history across a collector implementation change (smartctl, nvme-cli, and libnvme deliver differently-trimmed values, as shown above).
## Newly surfaced questions
- Should the store record `vid`/`ssvid`/`cntlid` alongside segment metadata now (cheap) for future composite needs, or keep segments minimal?
- Should a detected blank/degraded identity degrade the projection confidence category (it weakens the "identity/counter discontinuity" axis of the Unavailable state)?
- Does the collector pin to one acquisition path (libnvme bindings vs `nvme list -o json` vs continued `smartctl -j` plus a sysfs read for `subnqn`) — an operational choice this research does not settle?
+69
View File
@@ -0,0 +1,69 @@
# Python TUI frameworks for Fenris
_Research snapshot: 2026-08-31. Planning only; no product implementation._
## Decision
Use **Textual** for Fenris's later TUI prototype and, if the prototype checks below pass, for the persistent TUI.
Textual is the best fit for Fenris's keyboard-first Overview, Usage History, Drive Health, and Settings views because one framework supplies flexible grid/horizontal/vertical layouts, dashboard and form widgets, background workers, and a headless interaction driver ([layout](https://textual.textualize.io/guide/layout/), [widget gallery](https://textual.textualize.io/widget_gallery/), [workers](https://textual.textualize.io/guide/workers/), [testing](https://textual.textualize.io/guide/testing/)). Its principal costs are the largest direct dependency set in this shortlist and only a compact `Sparkline` as built-in charting; detailed historical plots would need a custom widget or the first-party `textual-plotext` integration ([PyPI metadata](https://pypi.org/pypi/textual/json), [Sparkline](https://textual.textualize.io/widgets/sparkline/), [`textual-plotext`](https://github.com/Textualize/textual-plotext)).
This recommendation is **conditional on raising Fenris's Python floor from its currently documented Python 3.7+ to Python 3.9+**: Textual 8.2.8 and Urwid 4.0.13 require Python 3.9+, while prompt_toolkit 3.0.53 requires Python 3.10+ ([Fenris README](../../README.md#you-need), [Textual metadata](https://pypi.org/pypi/textual/json), [Urwid metadata](https://pypi.org/pypi/urwid/json), [prompt_toolkit metadata](https://pypi.org/pypi/prompt-toolkit/json)). If retaining Python 3.7 is mandatory, none of the current versions evaluated here qualifies.
**Fallback:** choose **Urwid** if prototype evidence shows that explicit low-level terminal/display control or a smaller direct dependency set matters more than Textual's higher-level layouts, forms, workers, and test driver ([display modules](https://urwid.org/manual/displaymodules.html), [package metadata](https://pypi.org/pypi/urwid/json)).
## Scope and method
The shortlist covers maintained, full-screen-capable Python projects with first-party application/widget documentation: Textual, Urwid, and prompt_toolkit. The evaluation uses official documentation, repository release APIs, and PyPI package metadata. Versions, Python floors, dependencies, and release activity are a point-in-time snapshot and should be rechecked when dependencies are locked.
## Comparison
| Criterion | Textual 8.2.8 | Urwid 4.0.13 | prompt_toolkit 3.0.53 |
|---|---|---|---|
| **Terminal compatibility** | PyPI classifies Linux, macOS, and Windows 10/11 support; this exceeds Fenris's Linux-only scope ([metadata](https://pypi.org/pypi/textual/json)). | Offers pure-Python raw and OS curses displays. The manual compares UTF-8, color, mouse, and external-event-loop capabilities; curses is described as broadly terminal-compatible, while raw supports 88/256/24-bit color and external loops ([display modules](https://urwid.org/manual/displaymodules.html)). | Provides full-screen applications and platform-specific POSIX/Windows test input; official input docs describe cross-platform one-key-at-a-time input and Windows behavior ([full-screen apps](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/full_screen_apps.html), [input](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/asking_for_input.html), [testing](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/advanced_topics/unit_testing.html)). |
| **Responsive layout** | Vertical, horizontal, and grid layouts support fractional, percentage, and automatic sizing, nesting, spans, overflow, docking, and runtime layout changes ([layout](https://textual.textualize.io/guide/layout/)). | `Pile`, `Columns`, `GridFlow`, `Overlay`, `ListBox`, padding, and flow/box/fixed sizing provide procedural composition ([widgets](https://urwid.org/manual/widgets.html)). | `VSplit`, `HSplit`, `FloatContainer`, `ConditionalContainer`, and scrollable panes compose full-screen regions, but the reviewed guide does not document breakpoint-style responsiveness ([full-screen apps](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/full_screen_apps.html)). |
| **Dashboard, charts, history** | Built-ins include `DataTable`, `ProgressBar`, `Digits`, and `Sparkline`; Sparkline summarizes reactive numerical data into bars according to available widget width ([gallery](https://textual.textualize.io/widget_gallery/), [Sparkline](https://textual.textualize.io/widgets/sparkline/)). `textual-plotext` adds a `PlotextPlot` wrapper for richer plotting ([repository](https://github.com/Textualize/textual-plotext)). | Includes `BarGraph`, `GraphVScale`, and `ProgressBar`; richer history views still require composition or custom rendering ([graph widgets](https://urwid.org/reference/widget.html#graph-widgets)). | The documented reusable widget set supports text areas, buttons, frames, dialogs, and menus, but does not list a chart widget; Fenris charts would therefore be custom controls/rendering ([widget reference](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/reference.html#module-prompt_toolkit.widgets)). |
| **Forms and keyboard use** | Built-ins include `Input`, `MaskedInput`, `Select`, checkbox/radio/selection controls, switches, buttons, and text areas ([gallery](https://textual.textualize.io/widget_gallery/)). `Input` supports validation on change, blur, or submit and emits results with its messages ([Input](https://textual.textualize.io/widgets/input/)). | Supplies `Edit`, `Button`, `CheckBox`, `RadioButton`, focus handling, and keyboard/mouse event support; application code owns more form orchestration ([widgets](https://urwid.org/manual/widgets.html)). | Strong key-binding, editing, completion, history, suggestion, and validation facilities; full-screen reusable components include `TextArea` and `Button` ([input](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/asking_for_input.html), [full-screen apps](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/full_screen_apps.html)). |
| **Async/background updates** | Async workers run coroutines concurrently; thread workers cover blocking APIs. Workers support cancellation and exclusivity, while thread-originated UI updates use `call_from_thread()` or thread-safe messages ([workers](https://textual.textualize.io/guide/workers/)). | Supports Select, asyncio, Twisted, GLib, Tornado, and ZMQ event loops plus alarms and watched files; executor use varies by loop ([main loop](https://urwid.org/manual/mainloop.html)). | Uses asyncio natively and supports `Application.run_async()` inside an existing event loop ([asyncio](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/advanced_topics/asyncio.html)). |
| **Testability** | `App.run_test()` runs headlessly and returns a `Pilot` that can press keys, click, wait for queued messages, and run at specified terminal sizes ([testing](https://textual.textualize.io/guide/testing/)). | The public manuals reviewed do not expose a comparable application pilot; tests would be built around widgets, rendering, callbacks, and event-loop seams ([widgets](https://urwid.org/manual/widgets.html), [main loop](https://urwid.org/manual/mainloop.html)). | Official testing guidance supplies platform-specific pipe input, `DummyOutput`, and app sessions for asserting results or data changes, while warning against brittle stdout-byte assertions ([unit testing](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/advanced_topics/unit_testing.html)). |
| **Direct dependency cost** | Five required distributions: `markdown-it-py[linkify]`, `mdit-py-plugins`, `platformdirs`, `pygments`, and `rich`; syntax-tree packages are optional-extra dependencies ([metadata](https://pypi.org/pypi/textual/json)). | Two required distributions: `wcwidth` and `typing-extensions`; curses, GLib, Tornado, Trio, Twisted, ZMQ, serial, and LCD integrations are optional extras ([metadata](https://pypi.org/pypi/urwid/json)). | One required distribution: `wcwidth` ([metadata](https://pypi.org/pypi/prompt-toolkit/json)). |
| **Python floor** | `>=3.9,<4.0` ([metadata](https://pypi.org/pypi/textual/json)). | `>=3.9.0` ([metadata](https://pypi.org/pypi/urwid/json)). | `>=3.10` ([metadata](https://pypi.org/pypi/prompt-toolkit/json)). |
| **Project health signal** | Five non-draft releases from 8.2.4 through 8.2.8 were published between 2026-04-19 and 2026-06-30 ([release API](https://api.github.com/repos/Textualize/textual/releases?per_page=5)). | Five releases from 4.0.9 through 4.0.13 were published between 2026-08-14 and 2026-08-25 ([release API](https://api.github.com/repos/urwid/urwid/releases?per_page=5)). | Release 3.0.53 was published 2026-07-26, following 3.0.52 on 2025-08-27 ([release API](https://api.github.com/repos/prompt-toolkit/python-prompt-toolkit/releases?per_page=5)). |
All three show recent release activity, so maintenance does not eliminate a candidate. Textual wins on integrated application-level capabilities rather than on release recency alone.
## Fit to the four views
- **Overview:** use Textual layout containers with `Digits`, labels, progress indicators, and compact sparklines for monitoring status, usage-adjusted theoretical lifespan, projection confidence, and recent activity ([layout](https://textual.textualize.io/guide/layout/), [gallery](https://textual.textualize.io/widget_gallery/)).
- **Usage History:** prototype a `DataTable` plus `Sparkline` first. Add `textual-plotext` only if user testing establishes a need for axes, multiple series, or richer plots, keeping plotting out of the core dependency decision until then ([DataTable](https://textual.textualize.io/widgets/data_table/), [Sparkline](https://textual.textualize.io/widgets/sparkline/), [`textual-plotext`](https://github.com/Textualize/textual-plotext)).
- **Drive Health:** use a table and status/progress widgets for NVMe attributes; wording must continue to distinguish endurance projection from a hardware-failure prediction ([gallery](https://textual.textualize.io/widget_gallery/), [parent-map destination and terminology](https://git.bongbetic.com/xavierk/Fenris/issues/1)).
- **Settings:** Textual's selection, boolean, and validated text-entry controls cover a keyboard-only form without inventing controls ([gallery](https://textual.textualize.io/widget_gallery/), [Input validation](https://textual.textualize.io/widgets/input/)).
The TUI should consume snapshots from the independent collector rather than run NVMe collection in its UI loop; the ticket's parent map makes that independent-collector architecture part of the destination ([parent map](https://git.bongbetic.com/xavierk/Fenris/issues/1)). Textual workers remain useful for non-blocking database/IPC reads and cancellable refreshes ([workers](https://textual.textualize.io/guide/workers/)).
## Prototype acceptance checks
The later prototype should verify:
1. Complete keyboard-only operation and predictable focus order across all four views.
2. Reflow or deliberate scrolling at 80×24 and a representative wide terminal; Textual's test runner accepts explicit terminal sizes ([testing](https://textual.textualize.io/guide/testing/)).
3. Smooth refresh while database/IPC reads are delayed, plus clean cancellation and shutdown ([workers](https://textual.textualize.io/guide/workers/)).
4. Headless tests for view switching, settings validation, stale/error states, and collector disappearance ([testing](https://textual.textualize.io/guide/testing/), [Input](https://textual.textualize.io/widgets/input/)).
5. Readable behavior in the actual Linux terminal matrix, including SSH and tmux/screen if those are required; framework platform labels alone do not establish every terminal's behavior.
6. Usage History performance with realistic observation volume and gaps.
7. A table/text equivalent for every chart so observation gaps and projection confidence are not encoded by color or graphics alone.
8. Installation with Fenris's chosen packaging mechanism and the agreed Python floor.
## Newly surfaced questions
Do not open follow-up tickets during this research resolution; carry these into the next Wayfinder session:
- May the redesign raise Fenris's minimum Python version from 3.7 to 3.9?
- What is the minimum supported terminal contract: modern color terminals only, or also Linux virtual console, monochrome/16-color, SSH, tmux, and screen?
- Is a table plus sparkline sufficient for Usage History, or are axes, multiple series, zoom, and interval selection required?
- Should richer plotting be an optional install extra?
- What is the small-terminal policy: reflow, hide secondary content, scroll, or reject below a documented size?
- Which accessibility targets are required, including screen readers, reduced color, high contrast, and reduced animation?
- Does the TUI read observation storage directly or use an IPC snapshot API from the collector?
- Which settings are user-owned versus system/collector-owned, and which changes require privilege escalation?
- Are semantic interaction assertions sufficient, or are deterministic visual snapshots also required?