Files
Fenris/docs/adr/0002-projection-model-sustained-regime.md
T

6.3 KiB
Raw Blame History

2. Projection model: sustained-regime rate with categorical confidence

Status

Accepted — resolves Define the lifespan projection and confidence model on the Wayfinder map.

Context

Fenris's current compute_summary projects from a single trailing-24-hour write rate against endurance inferred as DUW / Percentage Used or synthesized as capacity × 600, alongside a second linear regression of Percentage Used toward 100. The endurance research established which signals can defensibly support a projection, and ADR 0001 fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.

Decision

  1. One projection. The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (wear_days) and the capacity × 600 synthesis are dropped.
  2. Headline rate from the sustained regime.
    rate          = regime DUW delta bytes / in-period wall-clock seconds
    projected     = max(E_baseline − W_t, 0) / rate      (rate > 0)
    E_rated       = entered_TBW × 10¹² bytes
    E_implied     = 100 · W_t / p                          (1 ≤ p ≤ 254)
    
    The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a scenario range; only horizons the history actually covers appear (no placeholders).
  3. Habit change. A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
  4. Hour classification (named constants, no configuration surface):
    • Powered-off: the hour's power-on-hours delta is below 90% of its wall-clock span.
    • Active: DUW delta ≥ 256 MiB in the hour.
    • Idle: powered on, sampled, below the active threshold.
    • Unknown: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
    • Disabled time is not an hour state: it is wall-clock outside monitoring periods.
  5. Denominator. Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
  6. Minimum evidence. Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
  7. Staleness. A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
  8. Confidence rule table.
    • Unavailable: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
    • Supported: verified baseline and ≥ 14 qualifying days and coverage ≥ 80% and fresh (< 48 h) and 7/28/90 rates within a factor of 2 across existing horizons and no single day ≥ 50% of trailing 28-day bytes and regime ≥ 7 days old.
    • Limited: every other case with a baseline and a positive rate; the failing facts are shown.
    • Confidence always renders as state plus contributing facts, never a percentage.
  9. Segment breaks. A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
  10. Implied-baseline eligibility. The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
  11. Uncertainty. The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
  12. Language. The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
  13. Contract. The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.

Consequences

  • The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
  • compute_summary's wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
  • Coverage becomes a first-class displayed fact rather than an internal heuristic.
  • All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
  • Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.