docs(adr): 0002 projection model — sustained regime, categorical confidence; glossary terms for regime, habit change, scenario range, coverage
This commit is contained in:
+16
@@ -43,3 +43,19 @@ _Avoid_: Counter reset handling, drive swap detection
|
|||||||
**Endurance baseline**:
|
**Endurance baseline**:
|
||||||
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
|
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
|
||||||
_Avoid_: TBW value, failure threshold, max writes
|
_Avoid_: TBW value, failure threshold, max writes
|
||||||
|
|
||||||
|
**Sustained regime**:
|
||||||
|
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
|
||||||
|
_Avoid_: Current window, detection period
|
||||||
|
|
||||||
|
**Habit change**:
|
||||||
|
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
|
||||||
|
_Avoid_: Spike, anomaly
|
||||||
|
|
||||||
|
**Scenario range**:
|
||||||
|
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
|
||||||
|
_Avoid_: Confidence interval, error bar
|
||||||
|
|
||||||
|
**Coverage**:
|
||||||
|
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
|
||||||
|
_Avoid_: Uptime, sample count
|
||||||
|
|||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# 2. Projection model: sustained-regime rate with categorical confidence
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
|
||||||
|
2. **Headline rate from the sustained regime.**
|
||||||
|
```text
|
||||||
|
rate = regime DUW delta bytes / in-period wall-clock seconds
|
||||||
|
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
|
||||||
|
E_rated = entered_TBW × 10¹² bytes
|
||||||
|
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
|
||||||
|
```
|
||||||
|
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
|
||||||
|
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
|
||||||
|
4. **Hour classification** (named constants, no configuration surface):
|
||||||
|
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
|
||||||
|
- **Active**: DUW delta ≥ 256 MiB in the hour.
|
||||||
|
- **Idle**: powered on, sampled, below the active threshold.
|
||||||
|
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
|
||||||
|
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
|
||||||
|
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
|
||||||
|
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
|
||||||
|
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
|
||||||
|
8. **Confidence rule table.**
|
||||||
|
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
|
||||||
|
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
|
||||||
|
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
|
||||||
|
- Confidence always renders as state plus contributing facts, never a percentage.
|
||||||
|
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
|
||||||
|
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
|
||||||
|
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
|
||||||
|
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
|
||||||
|
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
|
||||||
|
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
|
||||||
|
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
|
||||||
|
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
|
||||||
|
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
|
||||||
Reference in New Issue
Block a user