13 KiB
NVMe endurance signals and projection constraints
Research for Verify NVMe endurance signals and projection constraints.
Decision
Fenris can defensibly project when a host-write endurance baseline would be consumed if the observed usage habit continues. It cannot predict SSD failure.
Use lifetime Data Units Written (DUW) as the write counter, a model-and-capacity-specific rated-TBW override as the preferred baseline, and wall-clock observation history as the rate denominator. Keep Percentage Used as a separate manufacturer wear signal; only use it for a coarse, explicitly labeled implied baseline when no rated baseline exists. Treat Power On Hours as context, not elapsed calendar time. Express projection confidence as categorical evidence backed by visible facts, never as an accuracy percentage.
This decision removes two unsupported assumptions in the current calculation: inferring endurance as capacity × 600 when Percentage Used is zero, and presenting DUW / Percentage Used as non-estimated endurance (current Fenris calculation).
What the NVMe signals support
| Signal | Standard semantics and precision | Defensible Fenris use |
|---|---|---|
| Percentage Used | An unsigned one-byte, vendor-specific estimate based on actual use and the manufacturer's prediction of NVM life. 100 means estimated endurance consumed but may not mean subsystem failure; values may exceed 100, and values above 254 are represented as 255. It is updated once per power-on hour while the controller is not asleep (NVM Express Base Specification 2.0e, SMART / Health Information; official libnvme field documentation). |
Preserve the raw integer. Do not clamp at 100; render 255 as ≥255%, not an exact value. Do not call 100 a failure point. Zero is too coarse to establish zero wear or infer a baseline. |
| Data Units Written | A 128-bit cumulative count of host-written 512-byte data units, excluding metadata, reported in thousands and rounded upward. For the NVM command set, Write logical blocks count; Write Uncorrectable and Write Zeroes do not. Zero means the counter is not reported (official libnvme structure and semantics, field definition). | Store the raw integer and derive reported_host_bytes = DUW × 512,000. Call it reported host writes, not physical NAND writes or exact bytes. Treat zero as unsupported/ambiguous unless later positive samples prove support. |
| Power On Hours | Integer power-on hours; the controller may omit time powered in a non-operational power state (official libnvme field definition). | Display as drive context and use changes as a diagnostic. Do not use it as exact active time, idle time, powered-off time, or the denominator of a calendar-life projection. |
The SMART / Health log describes controller-level lifetime information; Fenris should therefore bind a history segment to a stable controller identity and avoid implying filesystem-level or physical-NAND-write precision (NVM Express Base Specification 2.0e).
DUW quantization
Let q = 512,000 bytes, raw counter U_t, and reported cumulative host writes W_t = qU_t. Since each cumulative endpoint is rounded upward, the reported interval delta is:
ΔW = q(U_b - U_a)
Its error from endpoint quantization alone is less than one quantum: |ΔW - actual interval writes| < 512,000 bytes. Thus a zero hourly delta does not prove no writes below that resolution, and “exact bytes written” is not defensible. This bound follows directly from the standard's upward-rounded cumulative representation (official libnvme DUW definition).
Endurance baseline precedence
Use these sources in order:
- Verified rated-TBW override for the exact manufacturer, model, and capacity, with source URL and document revision.
- Unverified manual override, visibly labeled as user-supplied.
- Implied endurance from Percentage Used, visibly labeled as a coarse heuristic.
- Otherwise, projection unavailable. Never synthesize TBW from capacity alone.
Manufacturer TBW values are model- and capacity-specific: Samsung, for example, rates the 1 TB 990 PRO at 600 TBW and the 2 TB model at 1,200 TBW, and states that its warranty is limited by the stated period or TBW, whichever comes first (Samsung 990 PRO data sheet, pp. 3–4). Samsung's warranty treats crossing TBW as a warranty-limit condition, not as a predicted failure event (Samsung SSD Limited Warranty, sections A–B). Fenris must therefore call rated TBW an endurance/warranty baseline rather than physical end of life.
Store an override as bytes plus provenance. If the input is labeled TBW, define it explicitly as decimal terabytes:
E_rated = entered_TBW × 10^12 bytes
R_rated = max(E_rated - W_t, 0)
Keep rated-budget consumption and manufacturer Percentage Used separate; disagreement is useful evidence, not a reason to blend them into a fabricated wear percentage.
Implied endurance constraints
Only for 1 ≤ p ≤ 254:
E_implied = 100W_t / p
R_implied = max(E_implied - W_t, 0)
This assumes the vendor's Percentage Used estimate is proportional to host writes, which NVMe does not require: the field is explicitly vendor-specific and based on the manufacturer's life prediction (NVM Express Base Specification 2.0e; libnvme documentation). Do not compute it for 0 or saturated 255. Label it implied from vendor wear estimate, show few significant digits, and do not promote it until multiple wear increments make the estimate less dominated by one-percentage-point quantization. At low values it is intrinsically unstable: changing p from 1 to 2 halves the result.
Usage-adjusted projection
For a selected valid wall-clock history interval from a to b:
rate = q(U_b - U_a) / elapsed_wall_clock_seconds
projected_seconds = remaining_baseline_bytes / rate (rate > 0)
If the rate is zero, report no finite projection from this history, not infinity. Required wording should be equivalent to:
Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
Use wall-clock elapsed time because powered-off and idle periods are part of the observed usage habit, whereas Power On Hours may exclude non-operational powered states (libnvme Power On Hours definition). Deliberately disabled monitoring must be excluded or marked unknown by lifecycle records; a SMART counter pair can recover aggregate writes across a collector gap but cannot reveal when within that gap the writes occurred.
History and changing habits
Persist interval observations rather than pretending every delta belongs to a clock-hour bucket:
- Compute deltas only within one controller-identity segment and only when DUW is monotonic. A decrease is a segment boundary or data fault, never a delta to clamp to zero.
- Preserve both endpoints, elapsed wall time, counter delta, and gap/monitoring state. Writes across a gap cannot be assigned exactly to individual hours; any proportional allocation must be labeled estimated.
- Use daily aggregates for habit evidence and retain hourly intervals for display/diagnostics. Hourly samples are time-dependent, so raw sample count is not independent evidence; NIST warns that autocorrelation can invalidate standard
s/√Nuncertainty calculations and other statistical conclusions (NIST Autocorrelation Plot). - Compare descriptive recent, medium, and longer horizons (for example 7, 28, and 90 days) and expose their projection spread as a scenario range, not a statistical confidence interval.
- Flag a changing habit when recent and earlier daily-rate windows diverge materially for a sustained period. Prefer the recent sustained regime for the headline projection while retaining the older regime as comparison. This is necessary because a stationary time series has stable mean, variance, and autocorrelation structure; trend, changing variance, and seasonality violate that assumption (NIST Stationarity).
- Before treating days as interchangeable, account for weekly or other periodic patterns; NIST describes seasonality as regular periodic behavior that must be addressed in a time-series model (NIST Seasonality).
Exact horizon lengths and change thresholds are product guardrails to validate later, not statistically guaranteed constants.
Projection confidence
Use four categorical states:
- Unavailable — no applicable baseline, unsupported DUW, identity/counter discontinuity, or no positive usable rate.
- Warming up — too little history to represent ordinary usage cycles.
- Limited evidence — implied or unverified baseline, short/incomplete history, substantial horizon spread, stale observations, or a recent habit change.
- Supported evidence — verified baseline, multiple representative usage cycles, good interval coverage, current observations, and stable rates across relevant horizons.
Always show the contributing facts, for example:
Supported evidence · verified manufacturer TBW · 42 calendar days · 96% interval coverage · 6 weekly cycles · recent and 28-day rates agree
Do not display “82% confidence” or “95% accurate.” NIST defines confidence level through the long-run coverage of an interval procedure, not as the probability that this particular estimate is correct (NIST Confidence Limits). A future statistical rate interval would cover rate-estimation uncertainty only; it would not validate the endurance baseline or guarantee that habits remain unchanged.
Required disclosures
- This is an endurance projection, not a predicted hardware-failure date.
- Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated (NVMe definition).
- Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold (Samsung warranty).
- DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes (NVMe definition).
- Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
- Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
Newly surfaced questions
Carry these to the next Wayfinder session rather than expanding this ticket:
- Should rated-budget and manufacturer Percentage Used projections appear side by side when they disagree?
- Which provenance fields are mandatory for a TBW override: URL, revision, model, capacity, region, and entry date?
- What exact warming-up, coverage, horizon, and changing-habit thresholds should the projection-model specification adopt?
- How should powered-off periods, deliberately disabled monitoring, and unexplained gaps be represented separately?
- Which stable controller identity prevents observation history from crossing a drive replacement?
- Should a detected recent regime automatically replace the long-term rate or require acknowledgement?
- Should the TUI expose multi-horizon scenarios only, or also a model-based statistical rate interval?
- How should unsupported DUW, saturated Percentage Used, and counter discontinuities appear in the TUI?