diff --git a/CONTEXT.md b/CONTEXT.md index e997fdd..0df96e8 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -43,3 +43,19 @@ _Avoid_: Counter reset handling, drive swap detection **Endurance baseline**: The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such. _Avoid_: TBW value, failure threshold, max writes + +**Sustained regime**: +The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes. +_Avoid_: Current window, detection period + +**Habit change**: +A sustained divergence between recent and earlier daily write rates that starts a new sustained regime. +_Avoid_: Spike, anomaly + +**Scenario range**: +The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval. +_Avoid_: Confidence interval, error bar + +**Coverage**: +The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown. +_Avoid_: Uptime, sample count diff --git a/docs/adr/0002-projection-model-sustained-regime.md b/docs/adr/0002-projection-model-sustained-regime.md new file mode 100644 index 0000000..c1cb8ca --- /dev/null +++ b/docs/adr/0002-projection-model-sustained-regime.md @@ -0,0 +1,49 @@ +# 2. Projection model: sustained-regime rate with categorical confidence + +## Status + +Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). + +## Context + +Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes. + +## Decision + +1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped. +2. **Headline rate from the sustained regime.** + ```text + rate = regime DUW delta bytes / in-period wall-clock seconds + projected = max(E_baseline − W_t, 0) / rate (rate > 0) + E_rated = entered_TBW × 10¹² bytes + E_implied = 100 · W_t / p (1 ≤ p ≤ 254) + ``` + The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders). +3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence. +4. **Hour classification** (named constants, no configuration surface): + - **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span. + - **Active**: DUW delta ≥ 256 MiB in the hour. + - **Idle**: powered on, sampled, below the active threshold. + - **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters. + - Disabled time is not an hour state: it is wall-clock outside monitoring periods. +5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage. +6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number. +7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact. +8. **Confidence rule table.** + - **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change. + - **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old. + - **Limited**: every other case with a baseline and a positive rate; the failing facts are shown. + - Confidence always renders as state plus contributing facts, never a percentage. +9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive. +10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance". +11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero. +12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section. +13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored. + +## Consequences + +- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states. +- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history. +- Coverage becomes a first-class displayed fact rather than an internal heuristic. +- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface. +- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation. \ No newline at end of file