Compare commits

..
9 changed files with 312 additions and 127 deletions
+1
View File
@@ -1,4 +1,5 @@
__pycache__/ __pycache__/
.pi/
*.pyc *.pyc
.commandcode/ .commandcode/
data/fenris.pid data/fenris.pid
+9
View File
@@ -0,0 +1,9 @@
## Agent skills
### Issue tracker
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
### Domain docs
This is a single-context repository. See `docs/agents/domain.md`.
+69
View File
@@ -0,0 +1,69 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Endurance baseline**:
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
+42
View File
@@ -0,0 +1,42 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -0,0 +1,49 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -0,0 +1,31 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger and monitoring-period bookkeeping; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
+26
View File
@@ -0,0 +1,26 @@
# Domain Docs
How engineering skills should consume this repository’s domain documentation.
## Layout
This is a single-context repository:
```text
/
├── CONTEXT.md
├── docs/adr/
└── ...
```
## Before exploring
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
## Use the glossary’s vocabulary
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
## Flag ADR conflicts
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
+85
View File
@@ -0,0 +1,85 @@
# Issue tracker: Gitea
Issues for this repository live in Gitea at:
https://git.bongbetic.com/xavierk/Fenris/issues
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
## General operations
- List: `tea issues list`
- Read: `tea issues <index> --comments`
- Create: `tea issues create --title "<title>" --description "<body>"`
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
- Assign: `tea issues edit <index> --add-assignees "<username>"`
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
- Comment: `tea comments add <index> --description "<comment>"`
- Close: `tea issues close <index>`
- Reopen: `tea issues reopen <index>`
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
## When a skill says “publish to the issue tracker”
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
## When a skill says “fetch the relevant ticket”
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
## Wayfinding operations
Wayfinder maps and decision tickets are Gitea issues.
### Map and ticket grouping
- A map has the label `wayfinder:map`.
- Create one milestone named `Wayfinder: <map title>` for the effort.
- Assign the map and all its tickets to that milestone.
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
### Blocking
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
```bash
tea api -X POST \
repos/{owner}/{repo}/issues/<blocked>/dependencies \
-F index=<blocker> \
-f owner=xavierk \
-f repo=Fenris
```
List blockers:
```bash
tea api repos/{owner}/{repo}/issues/<index>/dependencies
```
Remove the relationship with the same payload and `-X DELETE`.
### Frontier
List open issues in the map’s milestone. Exclude:
- the issue labelled `wayfinder:map`
- assigned tickets, because assignment is the claim
- tickets whose dependency query returns any open issue
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
### Claim
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
### Resolve
1. Add the answer as a resolution comment.
2. Close the ticket.
3. Re-fetch the map immediately before editing it.
4. Append a linked one-line context pointer to `Decisions so far`.
5. Create newly visible tickets, then wire dependencies in a second pass.
-127
View File
@@ -1,127 +0,0 @@
# NVMe endurance signals and projection constraints
Research for [Verify NVMe endurance signals and projection constraints](https://git.bongbetic.com/xavierk/Fenris/issues/5).
## Decision
Fenris can defensibly project **when a host-write endurance baseline would be consumed if the observed usage habit continues**. It cannot predict SSD failure.
Use lifetime **Data Units Written (DUW)** as the write counter, a model-and-capacity-specific **rated-TBW override** as the preferred baseline, and wall-clock observation history as the rate denominator. Keep **Percentage Used** as a separate manufacturer wear signal; only use it for a coarse, explicitly labeled implied baseline when no rated baseline exists. Treat **Power On Hours** as context, not elapsed calendar time. Express projection confidence as categorical evidence backed by visible facts, never as an accuracy percentage.
This decision removes two unsupported assumptions in the current calculation: inferring endurance as `capacity × 600` when Percentage Used is zero, and presenting `DUW / Percentage Used` as non-estimated endurance ([current Fenris calculation](https://git.bongbetic.com/xavierk/Fenris/src/commit/91db519148a6359f6968b758a76ce30b480aaae0/fenris.py#L261-L270)).
## What the NVMe signals support
| Signal | Standard semantics and precision | Defensible Fenris use |
|---|---|---|
| **Percentage Used** | An unsigned one-byte, vendor-specific estimate based on actual use and the manufacturer's prediction of NVM life. `100` means estimated endurance consumed but may not mean subsystem failure; values may exceed 100, and values above 254 are represented as 255. It is updated once per power-on hour while the controller is not asleep ([NVM Express Base Specification 2.0e, SMART / Health Information](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); [official libnvme field documentation](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)). | Preserve the raw integer. Do not clamp at 100; render 255 as `≥255%`, not an exact value. Do not call 100 a failure point. Zero is too coarse to establish zero wear or infer a baseline. |
| **Data Units Written** | A 128-bit cumulative count of host-written 512-byte data units, excluding metadata, reported in thousands and rounded upward. For the NVM command set, Write logical blocks count; Write Uncorrectable and Write Zeroes do not. Zero means the counter is not reported ([official libnvme structure and semantics](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L15-L23), [field definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)). | Store the raw integer and derive `reported_host_bytes = DUW × 512,000`. Call it **reported host writes**, not physical NAND writes or exact bytes. Treat zero as unsupported/ambiguous unless later positive samples prove support. |
| **Power On Hours** | Integer power-on hours; the controller may omit time powered in a non-operational power state ([official libnvme field definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L159-L162)). | Display as drive context and use changes as a diagnostic. Do not use it as exact active time, idle time, powered-off time, or the denominator of a calendar-life projection. |
The SMART / Health log describes controller-level lifetime information; Fenris should therefore bind a history segment to a stable controller identity and avoid implying filesystem-level or physical-NAND-write precision ([NVM Express Base Specification 2.0e](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
### DUW quantization
Let `q = 512,000 bytes`, raw counter `U_t`, and reported cumulative host writes `W_t = qU_t`. Since each cumulative endpoint is rounded upward, the reported interval delta is:
```text
ΔW = q(U_b - U_a)
```
Its error from endpoint quantization alone is less than one quantum: `|ΔW - actual interval writes| < 512,000 bytes`. Thus a zero hourly delta does not prove no writes below that resolution, and “exact bytes written” is not defensible. This bound follows directly from the standard's upward-rounded cumulative representation ([official libnvme DUW definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)).
## Endurance baseline precedence
Use these sources in order:
1. **Verified rated-TBW override** for the exact manufacturer, model, and capacity, with source URL and document revision.
2. **Unverified manual override**, visibly labeled as user-supplied.
3. **Implied endurance from Percentage Used**, visibly labeled as a coarse heuristic.
4. Otherwise, **projection unavailable**. Never synthesize TBW from capacity alone.
Manufacturer TBW values are model- and capacity-specific: Samsung, for example, rates the 1 TB 990 PRO at 600 TBW and the 2 TB model at 1,200 TBW, and states that its warranty is limited by the stated period or TBW, whichever comes first ([Samsung 990 PRO data sheet, pp. 3–4](https://download.semiconductor.samsung.com/resources/data-sheet/Samsung_NVMe_SSD_990_PRO_Datasheet_Rev.1.0.pdf)). Samsung's warranty treats crossing TBW as a warranty-limit condition, not as a predicted failure event ([Samsung SSD Limited Warranty, sections A–B](https://download.semiconductor.samsung.com/resources/warranty/SAMSUNG_SSD_Limited_Warranty_English_US.pdf)). Fenris must therefore call rated TBW an endurance/warranty baseline rather than physical end of life.
Store an override as bytes plus provenance. If the input is labeled TBW, define it explicitly as decimal terabytes:
```text
E_rated = entered_TBW × 10^12 bytes
R_rated = max(E_rated - W_t, 0)
```
Keep rated-budget consumption and manufacturer Percentage Used separate; disagreement is useful evidence, not a reason to blend them into a fabricated wear percentage.
### Implied endurance constraints
Only for `1 ≤ p ≤ 254`:
```text
E_implied = 100W_t / p
R_implied = max(E_implied - W_t, 0)
```
This assumes the vendor's Percentage Used estimate is proportional to host writes, which NVMe does **not** require: the field is explicitly vendor-specific and based on the manufacturer's life prediction ([NVM Express Base Specification 2.0e](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); [libnvme documentation](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)). Do not compute it for 0 or saturated 255. Label it **implied from vendor wear estimate**, show few significant digits, and do not promote it until multiple wear increments make the estimate less dominated by one-percentage-point quantization. At low values it is intrinsically unstable: changing `p` from 1 to 2 halves the result.
## Usage-adjusted projection
For a selected valid wall-clock history interval from `a` to `b`:
```text
rate = q(U_b - U_a) / elapsed_wall_clock_seconds
projected_seconds = remaining_baseline_bytes / rate (rate > 0)
```
If the rate is zero, report **no finite projection from this history**, not infinity. Required wording should be equivalent to:
> Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
Use wall-clock elapsed time because powered-off and idle periods are part of the observed usage habit, whereas Power On Hours may exclude non-operational powered states ([libnvme Power On Hours definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L159-L162)). Deliberately disabled monitoring must be excluded or marked unknown by lifecycle records; a SMART counter pair can recover aggregate writes across a collector gap but cannot reveal when within that gap the writes occurred.
## History and changing habits
Persist interval observations rather than pretending every delta belongs to a clock-hour bucket:
- Compute deltas only within one controller-identity segment and only when DUW is monotonic. A decrease is a segment boundary or data fault, never a delta to clamp to zero.
- Preserve both endpoints, elapsed wall time, counter delta, and gap/monitoring state. Writes across a gap cannot be assigned exactly to individual hours; any proportional allocation must be labeled estimated.
- Use daily aggregates for habit evidence and retain hourly intervals for display/diagnostics. Hourly samples are time-dependent, so raw sample count is not independent evidence; NIST warns that autocorrelation can invalidate standard `s/√N` uncertainty calculations and other statistical conclusions ([NIST Autocorrelation Plot](https://www.itl.nist.gov/div898/handbook/eda/section3/eda331.htm)).
- Compare descriptive recent, medium, and longer horizons (for example 7, 28, and 90 days) and expose their projection spread as a **scenario range**, not a statistical confidence interval.
- Flag a changing habit when recent and earlier daily-rate windows diverge materially for a sustained period. Prefer the recent sustained regime for the headline projection while retaining the older regime as comparison. This is necessary because a stationary time series has stable mean, variance, and autocorrelation structure; trend, changing variance, and seasonality violate that assumption ([NIST Stationarity](https://www.itl.nist.gov/div898/handbook/pmc/section4/pmc442.htm)).
- Before treating days as interchangeable, account for weekly or other periodic patterns; NIST describes seasonality as regular periodic behavior that must be addressed in a time-series model ([NIST Seasonality](https://www.itl.nist.gov/div898/handbook/pmc/section4/pmc443.htm)).
Exact horizon lengths and change thresholds are product guardrails to validate later, not statistically guaranteed constants.
## Projection confidence
Use four categorical states:
- **Unavailable** — no applicable baseline, unsupported DUW, identity/counter discontinuity, or no positive usable rate.
- **Warming up** — too little history to represent ordinary usage cycles.
- **Limited evidence** — implied or unverified baseline, short/incomplete history, substantial horizon spread, stale observations, or a recent habit change.
- **Supported evidence** — verified baseline, multiple representative usage cycles, good interval coverage, current observations, and stable rates across relevant horizons.
Always show the contributing facts, for example:
> Supported evidence · verified manufacturer TBW · 42 calendar days · 96% interval coverage · 6 weekly cycles · recent and 28-day rates agree
Do not display “82% confidence” or “95% accurate.” NIST defines confidence level through the long-run coverage of an interval procedure, not as the probability that this particular estimate is correct ([NIST Confidence Limits](https://www.itl.nist.gov/div898/handbook/eda/section3/eda352.htm)). A future statistical rate interval would cover rate-estimation uncertainty only; it would not validate the endurance baseline or guarantee that habits remain unchanged.
## Required disclosures
1. This is an endurance projection, not a predicted hardware-failure date.
2. Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated ([NVMe definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L88-L101)).
3. Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold ([Samsung warranty](https://download.semiconductor.samsung.com/resources/warranty/SAMSUNG_SSD_Limited_Warranty_English_US.pdf)).
4. DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes ([NVMe definition](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_smart_log.2#L122-L138)).
5. Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
6. Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
## Newly surfaced questions
Carry these to the next Wayfinder session rather than expanding this ticket:
- Should rated-budget and manufacturer Percentage Used projections appear side by side when they disagree?
- Which provenance fields are mandatory for a TBW override: URL, revision, model, capacity, region, and entry date?
- What exact warming-up, coverage, horizon, and changing-habit thresholds should the projection-model specification adopt?
- How should powered-off periods, deliberately disabled monitoring, and unexplained gaps be represented separately?
- Which stable controller identity prevents observation history from crossing a drive replacement?
- Should a detected recent regime automatically replace the long-term rate or require acknowledgement?
- Should the TUI expose multi-horizon scenarios only, or also a model-based statistical rate interval?
- How should unsupported DUW, saturated Percentage Used, and counter discontinuities appear in the TUI?