Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6e5af87a64 |
@@ -1,140 +0,0 @@
|
||||
# Defensible median scoring and comparison rules
|
||||
|
||||
Research for [Establish defensible median scoring and comparison rules](https://git.bongbetic.com/xavierk/odin/issues/7), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
|
||||
|
||||
**Accessed:** 25 September 2026. **Status:** decision evidence and recommendations; no final scoring formula, calibrated reference values, benchmark runs, or implementation. The user's requirement is a **median** score.
|
||||
|
||||
## What the median should mean
|
||||
|
||||
There are three separate choices:
|
||||
|
||||
| Level | Meaning | Main limitation |
|
||||
|---|---|---|
|
||||
| Median of repeated measurements | Typical result for one fixed workload under stated conditions | Hides occasional long stalls; does not combine CPU, storage and browser results |
|
||||
| Median across normalized workloads | Typical relative performance across a fixed test set | Domains with many tests gain more influence; changing references can change rankings |
|
||||
| Median across domain summaries | Typical relative performance across explicitly chosen domains | Domain definitions matter; poor performance in a minority of domains can disappear from the headline |
|
||||
|
||||
NIST defines the sample median as the middle observation, or the arithmetic average of the middle two for an even sample count. Its resistance to extremes is useful, but it is a measure of location, not completeness, reliability, or worst-case response. [NIST location][nist-location]
|
||||
|
||||
**Recommendation for discussion:** use medians to summarize repeated valid measurements, normalize against frozen references, form predefined domain summaries, then use a median across those domains for the requested headline. Preserve each stage and its raw inputs. This proposes an aggregation structure; the domain membership, weighting, reference values, repeat counts and numeric scale remain decisions.
|
||||
|
||||
Established suites demonstrate why the levels must stay explicit. SPEC CPU 2017 takes median execution times from three runs, or the slower of two, then uses a **geometric mean** across ratios. Speedometer uses inverse geometric means across test durations and arithmetic means across iterations. These are methodological precedents, not permission to substitute a geometric mean for Odin's requested median. Preserve a tool's native result under its original name; label Odin's further aggregation separately. [SPEC rules][spec-rules] [Speedometer methodology][speedometer]
|
||||
|
||||
## Normalization and category balance
|
||||
|
||||
Milliseconds, operations/second and GB/s cannot share a meaningful raw median. A candidate approach is a dimensionless ratio against the **same workload's** reference: observation/reference for positive higher-is-better measures, reference/observation for positive lower-is-better measures. SPEC uses the latter for elapsed-time ratios. The metric identity must include its units, workload, size, concurrency, timing boundaries and direction. Reject invalid/nonfinite inputs under declared validity rules; a zero measured duration must not become an infinite score. [SPEC overview][spec-overview]
|
||||
|
||||
**Fictional arithmetic examples throughout this report:** these numbers illustrate consequences only. They are not Odin calibration or measured hardware results. A displayed index of 100 at the reference is an arbitrary illustrative scale, not a recommendation.
|
||||
|
||||
| Fictional metric | Reference | Observed | Illustrative ratio / index |
|
||||
|---|---:|---:|---:|
|
||||
| Work throughput | 25 operations/s | 50 operations/s | 2.0 / 200 |
|
||||
| Memory bandwidth | 25 GB/s | 20 GB/s | 0.8 / 80 |
|
||||
| Request latency | 4 ms | 2 ms | 2.0 / 200 |
|
||||
|
||||
A median of these indices is 200, despite memory bandwidth being below reference. That is a consequence of the chosen statistic. Show domain detail and slow-tail measurements alongside it. “200 versus 100” describes this index; it does not establish that every application is twice as fast.
|
||||
|
||||
Fix the order of operations. For fictional repeated times `[1, 3]` ms and a 3 ms reference, normalizing the raw median gives `3 / 2 = 1.5`; taking the median of individual ratios `[3, 1]` gives 2. Even-count averaging and reciprocal normalization do not commute. The report format must specify which result it contains. [median definition][nist-location]
|
||||
|
||||
Balance domains before counting metrics. If ten CPU tests each score 160 and four other domains score 40, 80, 100 and 120, the flat fourteen-test median is 160. A median across five domain summaries is 100. Adding CPU subtests should not silently redefine the product's priorities. Similarly, adding Python/Rust/Java variants must not automatically multiply the language domain's influence. A median of domain medians is a deliberate hierarchical index, not the pooled median of all observations.
|
||||
|
||||
Exclude health counters, memory-test pass/fail, driver availability and installed RAM capacity from throughput arithmetic. Multiple correlated outputs from one workload—throughput, IOPS, average latency and several percentiles—also need an explicit selection rule before any becomes an independent scored contribution.
|
||||
|
||||
## Calibration is a substantive decision
|
||||
|
||||
SPEC establishes per-workload reference times on a named machine and publishes the calculation rules. Its documentation explains that reference changes preserve relative overall rankings for its geometric-mean calculation. **That invariance does not generally hold for a median across normalized metrics.** [SPEC reference explanation][spec-overview]
|
||||
|
||||
For three fictional higher-is-better workloads, let machine A produce `[1, 10, 10]` and B produce `[2, 2, 20]` in each workload's own units:
|
||||
|
||||
| Fictional reference vector | A's normalized results → median | B's normalized results → median | Ordering |
|
||||
|---|---|---|---|
|
||||
| `[1, 1, 1]` | `[1, 10, 10]` → 10 | `[2, 2, 20]` → 2 | A higher |
|
||||
| `[1, 10, 10]` | `[1, 1, 1]` → 1 | `[2, 0.2, 2]` → 2 | B higher |
|
||||
|
||||
The measurements did not change. Changing the reference altered the relative scales and which observations occupied the middle. Retain the requested median, make the reference rationale public, and version reference changes rather than treating them as cosmetic rescaling.
|
||||
|
||||
| Reference option | What it supports | Decision cost |
|
||||
|---|---|---|
|
||||
| Named reference configuration | Auditable, fixed anchor with documented per-test measurements | One machine's balance influences the median; configurations and repeatability need validation |
|
||||
| Frozen reference cohort | Per-test references from a documented collection of machines | Cohort selection, sampling bias and revision policy become part of the score |
|
||||
| User's own baseline | Local before/after comparisons | A score relative to oneself cannot rank different machines |
|
||||
|
||||
Recommend evaluating a frozen, locally distributable calibration manifest. Include reference measurements and provenance, reference conditions, workload and artifact digests, units/directions, aggregation order, required domains and calibration identity. Store it with results so calculation remains reproducible offline. Raw observations must survive changes; a recalculated score should identify its new calibration and preserve the original.
|
||||
|
||||
An arbitrary scale factor is acceptable if described as an index. A claim such as “100 is the median Linux machine,” a percentile rank, or a universal poor/good threshold requires representative population evidence that does not exist yet. Separately normalizing each architecture to its own average would also prevent interpreting those numbers as one common cross-architecture scale.
|
||||
|
||||
## Repetitions, warmup and uncertainty
|
||||
|
||||
Google Benchmark documents warmup, repetitions, median, standard deviation and coefficient of variation; it distinguishes user-visible wall time from CPU consumption. NIST recommends examining ordered observations for changing location/spread and says potential outliers should not simply be deleted when their cause is unknown. These support retaining all repeated observations and their conditions. [Google Benchmark][google-guide] [NIST run sequence][nist-runseq] [NIST outliers][nist-outliers]
|
||||
|
||||
Recommended measurement rules:
|
||||
|
||||
- Define warmup separately for each workload and retain its duration. Warm caches/JIT throughput, cold launch time and sustained thermal performance are different questions. Do not discard a slow first run from a declared cold-start test.
|
||||
- Fix repetition and stopping rules before observing scores. A quick run may estimate a median without enough evidence for a useful confidence interval; it should not claim the precision of the standard profile. Calibrate the minimum repeats against the 10–20 minute budget.
|
||||
- Retain run order, warmup, elapsed time, temperatures/power context, competing activity and invalidation reasons. A drifting sequence is not interchangeable independent noise. Repetitions within one process or thermal episode are not automatically independent runs.
|
||||
- Exclude observations only for declared validity failures such as incorrect output, changed workload, cancellation or protocol failure. Preserve them with reasons. A slow but valid run can represent the usability problem Odin is meant to reveal.
|
||||
- Show central spread such as MAD or IQR, plus tails where the workload supplies enough events. NIST defines MAD and IQR as distinct measures of spread; neither is itself a confidence interval. A median across repeated p99 values must not be labelled the p99 of all requests. [NIST scale][nist-scale]
|
||||
|
||||
NIST documents median confidence intervals based on order statistics/binomial probabilities, interpolated methods and bootstrap alternatives. Choose and validate a median-appropriate method; do not apply a mean's standard-error formula to a median. Confidence also depends on sample count and assumptions about the measurements. For an aggregate, account for shared run-level variation and state whether uncertainty in the calibration reference is included. The spread **between different domain scores** is not sampling uncertainty about the headline. [NIST median intervals][nist-median-ci]
|
||||
|
||||
Keep a graph of results in time order. Google documents CPU selection, boost, scheduler contention, SMT, caches and NUMA as variance sources. Its suggestions for controlled laboratory microbenchmarks include changing system settings; Odin's installed-system baseline should record existing conditions and label any tuned experiment separately. The existing [CPU/memory report][odin-cpu] and [portability/UI report][odin-portability] explain TUI interference and qualification needs. Stable repeated numbers alone do not prove that a workload represents real usability. [Google variance][google-variance]
|
||||
|
||||
## Comparability must be attached to every score
|
||||
|
||||
SPEC requires performance-relevant observation conditions and valid workload outputs; its CPU suite intentionally measures processor, memory subsystem **and compilers**. Even a fixed-toolchain comparison describes a defined software/hardware configuration. [SPEC rules][spec-rules] [SPEC overview][spec-overview]
|
||||
|
||||
Recommend two explicit comparison purposes:
|
||||
|
||||
- **Installed-system usability:** the chosen installed browser, shell, runtimes, drivers and kernel are part of what is measured. Version/configuration changes may explain a score change without any hardware change.
|
||||
- **Controlled reference workload:** fixed workload assets, runtime/compiler contracts, flags, input data and execution modes improve comparison across machines. Architecture-specific artifacts must implement equivalent declared work and validate outputs; different ISA policies need disclosure.
|
||||
|
||||
Neither mode needs to masquerade as a pure hardware measurement. Keep their result identities distinct. Kernel/libc/distro differences can be the subject of a comparison, but they must be visible and the workload contract must remain equivalent.
|
||||
|
||||
A comparison identity should include suite/scoring/calibration versions, workload set, run profile, tool/artifact versions, options and data digests, timing/aggregation rules, browser mode, hardware/virtualization context and validity/coverage. Require a documented equivalence decision before combining scores across changed tools or workloads. SPEC warns that scores across different suite generations generally cannot be converted. [SPEC overview][spec-overview]
|
||||
|
||||
Speedometer 3.1 instructs users to use a clean browser profile, close competing programs/tabs, keep its page focused, avoid device interaction, use AC power and allow cooling when needed. Its official UI computes a 95% interval around its **arithmetic mean**; that interval cannot be attached to Odin's median unchanged. Preserve native browser score/uncertainty and label any median of complete runs separately. Headed and headless measurements need separate identities until an equivalence study justifies any shared interpretation; background versus foreground execution is also material. [instructions][speedometer-instructions] [3.1 result code][speedometer-main]
|
||||
|
||||
VM results characterize the guest allocation and virtualization environment. Keep native, virtualized and emulated cohorts identifiable; VM compatibility success does not establish native performance. Storage cache mode, queue depth, engine, filesystem and durability policy similarly belong to the workload identity. A fallback such as buffered I/O cannot silently replace a direct-I/O measurement with the same scoring identity. [CPU/memory report][odin-cpu] [storage report][odin-storage] [portability report][odin-portability]
|
||||
|
||||
## Missing tests and eligibility
|
||||
|
||||
For fictional domain indices `[40, 80, 100, 120, 160]`, the complete median is 100. Omitting 40 produces 110; omitting both 40 and 80 produces 120. Available-only aggregation can reward absent or deliberately skipped weak components.
|
||||
|
||||
**Recommendation:** define a versioned required set for the full score. Permit a clearly named partial median and domain results when the full set is unavailable, with the exact subset identified. Compare partial scores only over the same compatible subset; a pairwise intersection comparison must recompute **both** results and label that narrower scope. An optional pack must not silently change the headline's membership.
|
||||
|
||||
Keep distinct outcomes: completed-valid, completed-with-limitations, unsupported, missing dependency, permission denied, unsafe to run, cancelled, timed out, and failed validation. The eventual validity contract decides whether a limited result remains score-eligible. Never impute missing results as zero, a reference score, or a healthy pass. A required workload failing verification makes the full score ineligible, while preserving completed measurements and the associated finding. Good numbers from other domains should not cancel that failure.
|
||||
|
||||
The user accepted reporting unavailable tests across Linux targets. That does not resolve which domains are mandatory, whether every machine should still display a partial number, or how partial results should look. Those are explicit product decisions.
|
||||
|
||||
## Presentation, recommendations and open decisions
|
||||
|
||||
Recommend a headline containing the median, score identity, full/partial state, and eligible-domain coverage. The next view should show domain values, raw units, repeat count/spread, tail latency, invalidations and reference details. Show health findings beside performance: a fast drive with serious SMART evidence still needs attention. Missing telemetry must stay unknown. The storage and CPU reports establish why speed cannot determine drive replacement or certify memory health. [storage][odin-storage] [CPU/memory][odin-cpu]
|
||||
|
||||
Optimization advice should cite the observation and matching rule: for example, measured foreground stalls plus pressure evidence can support investigating memory contention. A low normalized score alone does not identify its cause. Keep severity of health evidence, completeness of coverage, measurement uncertainty and performance position as separate concepts; avoid one synthetic “confidence/health” percentage that mixes them.
|
||||
|
||||
Before implementation, decide:
|
||||
|
||||
1. The headline's median level, domain membership and balancing rules; whether responsiveness contributes or remains an accompanying measurement.
|
||||
2. The reference configuration/cohort, scale and calibration-release policy; collect actual calibration data before inventing thresholds.
|
||||
3. Required versus optional coverage, partial-score display and exact comparison eligibility.
|
||||
4. Installed-system versus controlled-workload defaults, architecture/ISA policies, browser modes and VM cohorts.
|
||||
5. Repetition/warmup/stopping rules, outlier validity rules, median interval method and honest quick/standard/extended precision claims.
|
||||
6. Evidence requirements for optimization rules and separation of urgent health findings from the score.
|
||||
|
||||
**Evidence limits:** no calibration population, repeatability measurements, TUI-overhead budget or headed/headless equivalence study was produced. Examples are arithmetic demonstrations only. Context7 successfully resolved Google Benchmark; two BrowserBench/Speedometer searches returned unrelated packages, so its official repository and deployed 3.1 documentation were inspected directly. NIST and SPEC sources were inspected directly as statistical and benchmark-methodology references. This report neither adopts SPEC's workloads nor claims that their aggregation rules are Odin's final design.
|
||||
|
||||
[nist-location]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda351.htm
|
||||
[nist-scale]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda356.htm
|
||||
[nist-outliers]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda35h.htm
|
||||
[nist-runseq]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda33p.htm
|
||||
[nist-median-ci]: https://www.itl.nist.gov/div898/software/dataplot/refman1/auxillar/mediancl.htm
|
||||
[spec-rules]: https://www.spec.org/cpu2017/Docs/runrules.html
|
||||
[spec-overview]: https://www.spec.org/cpu2017/Docs/overview.html
|
||||
[google-guide]: https://github.com/google/benchmark/blob/main/docs/user_guide.md
|
||||
[google-variance]: https://github.com/google/benchmark/blob/main/docs/reducing_variance.md
|
||||
[speedometer]: https://github.com/WebKit/Speedometer/blob/main/README.md
|
||||
[speedometer-instructions]: https://browserbench.org/Speedometer3.1/instructions.html
|
||||
[speedometer-main]: https://browserbench.org/Speedometer3.1/resources/main.mjs
|
||||
[odin-cpu]: https://git.bongbetic.com/xavierk/odin/src/commit/90213d7f62cbd118f08ea8ff2f8042e94aa038a7/docs/research/cpu-memory.md
|
||||
[odin-storage]: https://git.bongbetic.com/xavierk/odin/src/commit/6e5af87a64faedd4a8ad31ba10d9be4b499e8349/docs/research/storage-health.md
|
||||
[odin-portability]: https://git.bongbetic.com/xavierk/odin/src/commit/20681cd4f184a9fc0164dd638a252de44ef230f5/docs/research/portability-ui.md
|
||||
@@ -0,0 +1,139 @@
|
||||
# Storage measurements and trustworthy health advice
|
||||
|
||||
Research for [Establish storage measurements and trustworthy health advice](https://git.bongbetic.com/xavierk/odin/issues/4), part of Odin's Wayfinder map. Access date for every source: **2026-09-25**.
|
||||
|
||||
This report establishes evidence and candidate policies. It does not select Odin's final workloads, thresholds, privileged execution design, or score. No benchmark, device query, self-test, installation, or hardware change was performed during this investigation.
|
||||
|
||||
## Findings that shape the decision
|
||||
|
||||
The strongest candidate is **fio for file-based performance measurements, smartmontools for cross-protocol health findings, and an optional nvme-cli adapter for additional NVMe evidence**. These tools cover different responsibilities. A fast benchmark cannot establish drive health; a passing SMART status cannot establish future reliability. Health findings should therefore remain visible independently of the performance score and capability coverage.
|
||||
|
||||
There is a material compatibility change already: nvme-cli **v3.1**, released September 18, 2026, documents `nvme log smart`; `nvme smart-log` is a deprecated compatibility alias. Its default output format version is now 2, with version 1 available. fio **3.43** was released September 23. A bleeding-edge development environment is compatible with reproducible measurements only if Odin records tool versions and keeps workload and parser versions explicit. Package-manager availability alone does not establish supported commands or JSON schemas. [S1][S3]
|
||||
|
||||
## Candidate tools and measurement scope
|
||||
|
||||
| Candidate | Useful responsibility | Limits and recommendation to consider |
|
||||
| --- | --- | --- |
|
||||
| fio | Sequential/random reads and writes, block sizes, queue depths, latency distributions, bounded I/O, optional verification | Best primary workload candidate. Select a small fixed workload vocabulary; do not accept arbitrary user-supplied job files into privileged execution. |
|
||||
| smartctl | ATA, SCSI and NVMe identity, health, existing error/self-test logs; JSON; many bridge/controller adapters | Best baseline health reader. Decode protocol-specific semantics and command status separately. Some transports are unsafe for automatic probing. |
|
||||
| nvme-cli | NVMe-specific identity, SMART and detailed logs | Useful optional supplement. Version 2/3 command and JSON differences need explicit compatibility handling. Avoid duplicating the same controller's health as several independent findings. |
|
||||
| Native `/proc` and `/sys` | I/O pressure, completed I/O, queue activity and available sensors | Low-dependency contextual evidence, not a workload or a substitute for SMART. |
|
||||
| GNU `dd` | Bounded sequential copying, optionally direct I/O and final synchronization | Possible explicitly labelled basic fallback. Its copy-oriented output does not supply fio's workload control or latency distributions; its result must not silently substitute into the same scored workload. |
|
||||
|
||||
Sources: fio HOWTO, smartctl manual, nvme-cli released documentation, kernel PSI/I/O documentation and GNU manual. [S1–S4][S8][S9][S14]
|
||||
|
||||
A compact candidate performance set is:
|
||||
|
||||
| Measurement | Candidate workload, still to be selected | What its result means |
|
||||
| --- | --- | --- |
|
||||
| Sequential read/write | Large blocks, for example 1 MiB, one job, depth 1 | Large-file throughput through the selected filesystem and storage path |
|
||||
| Random read/write | 4 KiB, one job, depth 1 | Small-request responsiveness; report latency and IOPS |
|
||||
| Queued random read | Same block size at a documented higher depth, such as 16 or 32 | Concurrency capability; a different workload from depth 1 |
|
||||
| Durable small writes | A separate small, bounded workload with defined sync frequency | Application-visible cost of requesting persistence |
|
||||
| Optional integrity check | Write and verify Odin-owned file blocks with fio checksums | Whether the tested data path returned those bytes correctly; not a full-surface drive or whole-RAM certification |
|
||||
|
||||
Record read and write throughput in explicit units, IOPS, completed bytes, operation count, errors, elapsed time, and p50/p95/p99 latency where sample counts support them. Distinguish fio completion latency from total latency: total includes submission latency. Record achieved queue-depth distribution; requesting depth greater than one does not make a synchronous engine asynchronous. fio `psync` is a useful depth-1 compatibility candidate; `io_uring` and `libaio` are queued-engine candidates where supported. Engine changes must be visible in results and comparability rules. [S1]
|
||||
|
||||
Measure one storage workload at a time during reference runs. Concurrent CPU or memory stress can instead be an explicitly identified contention experiment. Record filesystem, mount options, device topology, encryption/RAID/virtualization, kernel, selected engine, power/thermal state, background I/O, and pre/post free space. The measurement describes this path under these conditions, not the NVMe/HDD in isolation.
|
||||
|
||||
## Safe operation and comparability constraints
|
||||
|
||||
The following are proposed invariants, rather than finalized profile numbers:
|
||||
|
||||
1. **Own every writable byte.** Create a private run directory on a deliberately selected filesystem and exclusively create its regular files. Validate ownership, type and target identity; reject symlink redirection and raw block/character devices. `O_CREAT|O_EXCL` supplies exclusive creation semantics. Do not use arbitrary existing user files as write targets. Avoid selecting `/tmp` automatically: tmpfs stores files in virtual memory and may use swap. [S7][S11]
|
||||
2. **Budget storage space and cumulative writes separately.** fio `size` defines the working region, while `io_size` can independently bound I/O. `runtime` stops at the earlier of completion or time limit; `time_based` loops the workload. Thus a small file plus a timed loop can write many times its size. Prefer explicit byte and time bounds without `time_based` for ordinary write profiles. Count fixture preparation, repetitions and verification-related writes in a per-run host-write budget. Reserve free space, account for quotas and metadata, recheck during execution, and stop on ENOSPC or I/O errors. Space and byte thresholds remain product decisions. [S1]
|
||||
3. **Do not promise a physical NAND-write limit.** A workload's host bytes are measurable; filesystem/controller write amplification and unrelated host activity are additional. NVMe Data Units Written measures host data in units of 1,000 × 512 bytes, rounded up, excluding metadata; it is not a universal NAND-wear counter. Background activity also prevents attributing its entire delta to Odin. [S5]
|
||||
4. **Make cache and durability modes explicit.** fio `direct=1` normally requests `O_DIRECT`; support and alignment vary by filesystem and kernel, and misaligned requests can fail or fall back to buffered I/O. Direct I/O does not by itself provide `O_SYNC` persistence guarantees, bypass every device cache, or prove sustained media speed. `invalidate` is conditional on platform/file support. Avoid global `drop_caches`: kernel documentation warns of additional I/O and CPU costs. A buffered fallback must be labelled and excluded from direct-I/O comparisons. [S1][S7][S12]
|
||||
5. **Include preparation and flush costs honestly.** Read tests over newly created fixtures still require writes. A user choosing no writes can reuse an identified valid fixture or skip that workload; Odin should not create one silently. Do not measure unwritten sparse-file holes as disk reads. For writes, document `end_fsync` or other synchronization and report end-to-end time including the final flush separately from unsynchronized throughput. Control data compressibility/deduplication using a declared fio buffer policy; generating fresh data adds CPU cost. Preserve normal filesystem settings rather than silently disabling compression or copy-on-write. [S1][S7]
|
||||
6. **Treat cancellation and cleanup as part of the run.** Bound the entire job group; stop launching work on cancellation, retain partial status, reap workers, and remove only proven Odin-owned artifacts. A worker stuck in kernel I/O may not stop immediately. Crash recovery needs a manifest and ownership checks before deletion. Cleanup failure is a reported outcome, never a reason to recursively delete a user-selected directory.
|
||||
7. **Collect health before load and reduce work when evidence is serious.** A candidate policy is to skip storage stress when critical media/reliability findings or current unreadable data are already present. Pause on documented thermal alarms or loss of safety headroom. Display estimated host writes before a write run. Avoid automatic discard/TRIM, formatting, SMART feature changes, firmware updates, cache-policy changes or repair operations as benchmark preparation.
|
||||
|
||||
Short bounded tests cannot establish steady-state SSD performance after exhaustion of a large write cache, or scan every HDD sector. Making test data larger than all caches can conflict with a quick run and a conservative write budget. Report the actual duration and working set; do not extrapolate a short burst into an endurance or sustained-performance guarantee. The tradeoff between low impact and sustained measurements needs an explicit run-profile decision.
|
||||
|
||||
## Health evidence and field interpretation
|
||||
|
||||
Prefer structured output with the original tool version, schema identifier, command outcome, timestamp, device identity and transport. A missing field is unknown, not zero. Preserve large counters losslessly: smartctl JSON can emit string/byte-array companions for integers exceeding JavaScript's safe integer range; `--json=v` requests them consistently. This matters for an Ink/JavaScript consumer. [S2]
|
||||
|
||||
For smartctl NVMe output, the primary object is `nvme_smart_health_information_log`. The inspected source confirms the following keys and conversions. Raw NVMe temperature is Kelvin; smartctl's `temperature` here is already Celsius. Do not convert it twice. [S6]
|
||||
|
||||
| Evidence | Meaning | Candidate interpretation |
|
||||
| --- | --- | --- |
|
||||
| `critical_warning` bit 0; `available_spare` vs `available_spare_threshold` | Spare capacity below the controller's threshold | Urgent preservation/service finding; display the device-provided threshold |
|
||||
| Bit 1; `temperature`, warning/critical temperature time | Above an over-temperature or below an under-temperature threshold | Stop heat-producing tests; investigate cooling/environment. This alone is not proof that replacement is needed |
|
||||
| Bit 2 | NVM subsystem reliability degraded | Urgent backup and replacement/service assessment |
|
||||
| Bit 3 | Media placed read-only for a device reliability condition | Urgent preservation and replacement/service assessment; distinct from user namespace write protection |
|
||||
| Bits 4/5 | Volatile-memory backup failure; persistent-memory region read-only/unreliable | Urgent loss-of-protection/service finding when applicable; explain the specific subsystem |
|
||||
| `percentage_used` | Vendor estimate of endurance consumed | 100 means estimated endurance consumed, **not guaranteed failure**; values may exceed 100. Plan replacement according to manufacturer guidance and workload, without inventing days remaining |
|
||||
| `media_errors` | Unrecovered data-integrity errors, including ECC/CRC/tag errors | Investigate any nonzero history; escalating recent deltas plus failed I/O are much stronger urgent evidence than an isolated old count |
|
||||
| `num_err_log_entries` | Lifetime number of error-information entries | Inspect status/cause and recency. It is not interchangeable with media errors |
|
||||
| `unsafe_shutdowns` | Loss of power without shutdown notification | Investigate shutdown/power history and correlate with errors; not proof of failed media |
|
||||
| `data_units_written`, power-on hours, thermal counters | Usage/history with specified units and reporting limits | Useful trends and context; no universal lifespan formula |
|
||||
|
||||
NVMe warning bits are current state, not persistent event history; zero today does not erase yesterday's finding. Some temperature fields are optional, and zero can mean unsupported. Per-namespace SMART is optional; the global namespace identifier can describe a controller's aggregate. Preserve scope rather than assigning identical controller totals to every namespace. [S3][S5][S6]
|
||||
|
||||
**ATA needs a separate mapping.** Keep attribute ID, raw representation, normalized current/worst value, threshold, type and failure state. smartctl states that these meanings are vendor-specific; SSD meanings can differ and displayed names can be wrong for models absent from its drive database. The label `Pre-fail` by itself does not mean a drive is failing: the current normalized value must cross its threshold. [S2]
|
||||
|
||||
Common drive-database candidates include reallocated sectors (5), pending sectors (197), offline uncorrectable sectors (198), and interface CRC errors (199). Interpret them only with a matching model/firmware/database rule; do not apply a universal raw-count threshold or turn interface errors directly into a disk-replacement recommendation. Preserve lifetime history and recent deltas separately. SMART RETURN STATUS, failed applicable thresholds, existing self-test failures, and observed host I/O errors are stronger when they agree. SCSI health uses its own exception/sense reporting rather than ATA attribute assumptions. [S2][S15]
|
||||
|
||||
smartctl exit status is a bitmask. Bits 0–2 can describe invocation/access/command problems; bits 3–7 describe failing status, thresholds and historical error/self-test evidence. A nonzero exit must not discard usable JSON, and access failure must not become “bad drive.” Reading an existing self-test log is different from starting a test. The manual notes that running self-tests can degrade performance and normal I/O can extend their duration; any future self-test workflow needs a separate user decision and scheduling. [S2]
|
||||
|
||||
## Candidate advice rubric
|
||||
|
||||
This rubric is a proposed interpretation layer over the documented evidence, not a manufacturer's warranty or an adopted Odin policy.
|
||||
|
||||
| Finding class | Evidence sufficient to consider it | Appropriate wording/action |
|
||||
| --- | --- | --- |
|
||||
| **Replace/service now** | Credible ATA failing status/current applicable prefailure threshold; NVMe degraded reliability/read-only media; serious repeated data-integrity failures attributable to the device | “Preserve accessible data now; avoid further stress; arrange replacement or service.” Cite exact flags, device scope and timestamps. Hardware attribution may still need confirmation |
|
||||
| **Investigate urgently** | New media errors, pending/uncorrectable sectors, recent failed self-tests, resets/timeouts, thermal alarms, loss of power-loss protection | Identify the failing path; correlate controller, connection, power and filesystem evidence. Do not automatically blame the medium |
|
||||
| **Monitor / plan replacement** | Stable historical findings or vendor-estimated endurance consumed without current failure evidence | Retain trends, explain wear status and manufacturer limits, and plan according to importance/workload. No invented remaining-life percentage |
|
||||
| **No concerning evidence observed** | Successful supported collection with no relevant current finding | State what was checked and when; keep normal backup advice independent of a performance score |
|
||||
| **Unknown / limited coverage** | Missing permission/tool/field, sleeping drive, unsupported bridge/controller, virtual device, ambiguous identity | Explain the missing capability and a bounded next step. Never convert unavailable evidence into a healthy badge |
|
||||
|
||||
The smartctl manual recommends preserving data promptly when the drive reports failing health. Conversely, Google's primary HDD population study found that SMART-only models were unlikely to predict individual failures reliably. That older HDD result is not a calibrated modern-SSD failure model, but it reinforces the distinction between a useful warning and a guarantee of future health. NVMe's own endurance-field semantics explicitly reject equating 100% usage with failure. [S2][S5][S16]
|
||||
|
||||
## Compatibility, privilege and general health
|
||||
|
||||
USB, SAT and RAID support must follow known transport rules. The smartctl manual documents bridge-specific NVMe adapters and per-physical-disk MegaRAID addressing; a RAID logical volume is not automatically one physical drive. Particularly important: its **JMB39x/JMS56x transport uses READ/WRITE commands to a RAID-volume sector**. It warns that the wrong device can be overwritten and interruption can prevent restoration. Exclude these from routine automated probing; “try every device type” is not a safe compatibility strategy. Even standby-aware queries may wake a disk during autodetection, so record unsupported power-state handling. [S2]
|
||||
|
||||
VMs need an explicit virtual-device classification. QEMU's NVMe implementation constructs SMART data from its emulated controller and block-accounting state. A guest can therefore show valid-looking SMART without revealing the host drive's health. Guest tests establish guest-path performance; actual physical passthrough and device identity require separate verification. VM coverage cannot establish USB, physical RAID, real wear counters or thermal behavior. [S17]
|
||||
|
||||
Keep ordinary file workloads unprivileged. Device queries may require additional device permissions or kernel capabilities; NVMe's Linux passthrough code explicitly gates classes of commands. A future privileged mechanism should allow only validated read operations and selected devices, with no arbitrary shell or passthrough-command forwarding. Permission denial is an expected capability outcome, not an instruction to run the whole TUI as root. [S18]
|
||||
|
||||
For overall system health, useful complementary evidence is:
|
||||
|
||||
- `/proc/pressure/io`: `some` measures time with some stalled tasks, `full` time with all non-idle tasks stalled; use same-window deltas alongside workload latency. This detects pressure, not its sole cause. [S8]
|
||||
- `/proc/diskstats` or per-device sysfs statistics: completed I/O, time and queue context. Counters have concurrency/accounting caveats; busy percentage alone does not establish NVMe saturation. [S9]
|
||||
- Available hwmon readings, limits and alarm flags: retain sensor identity and units; chip-specific alarms and missing sensors preclude a universal hard-coded temperature cutoff. Standard hwmon ABI readings are intended to be readable by unprivileged applications. [S10]
|
||||
- Kernel errors and existing EDAC/RAS evidence: distinguish corrected errors from uncorrected/fatal errors and report available history. EDAC documentation explicitly says corrected errors may, but need not, predict later uncorrected errors. Missing reporting hardware/driver is unknown. Kernel log access can require `CAP_SYSLOG` when `dmesg_restrict=1`; do not assume systemd/journald on Void or other distributions. [S13][S19]
|
||||
|
||||
Kernel or mount optimizations should be suggestions tied to an observed limitation and a documented tradeoff, recorded for subsequent comparable runs. This research supports observing current settings and thermal/power/error evidence; it supplies no evidence for blanket scheduler, write-cache, governor, or filesystem changes.
|
||||
|
||||
## Remaining decisions and evidence gaps
|
||||
|
||||
The next human decisions are the ordinary run's write authorization/budget, minimum free-space reserve, workload lengths and repetitions, required versus optional queued/sync/verification tests, how reduced-capability results affect score eligibility, the supported transport list, the privilege interaction, and the exact advice wording. A sustained-media profile would need a separate impact budget.
|
||||
|
||||
Implementation work will need parser fixtures from supported smartctl/nvme-cli versions; success, partial and denied-permission results; real ATA/NVMe/USB/RAID samples; healthy and failing vendor examples; and proof of cancellation, space reservation, direct-I/O handling and cleanup across filesystems. No such hardware validation occurred here. No calibrated cross-device replacement thresholds or modern SSD remaining-life model were found or claimed.
|
||||
|
||||
Context7 library resolution succeeded for fio, smartmontools, nvme-cli, Linux kernel, GNU Coreutils and QEMU. Both allowed nvme-cli documentation fetches returned “Could not fetch documentation snippets”; its official released documents and source were inspected instead. The NVM Express specifications landing page returned HTTP 403, so NVMe field semantics here are grounded in maintained libnvme definitions and smartmontools implementation rather than a directly retrieved current specification PDF. Those are material evidence limits, not silently filled gaps.
|
||||
|
||||
## Sources inspected
|
||||
|
||||
- **S1:** [fio 3.43 HOWTO](https://github.com/axboe/fio/blob/fio-3.43/HOWTO.rst), relevant workload, size/runtime, buffering, engines, percentile, verification and error sections; [release](https://github.com/axboe/fio/releases/tag/fio-3.43).
|
||||
- **S2:** [smartctl manual source](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/smartctl.8.in), health, attributes, JSON, exit status, device transports, standby and self-tests.
|
||||
- **S3:** nvme-cli v3.1 [SMART log command](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/nvme-log-smart.txt), [global options](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/global-options.txt), [legacy alias](https://github.com/linux-nvme/nvme-cli/blob/master/Documentation/nvme-smart-log.txt), and [release](https://github.com/linux-nvme/nvme-cli/releases/tag/v3.1).
|
||||
- **S4:** [smartmontools NVMe support examples](https://www.smartmontools.org/wiki/NVMe_Support), inspected through Context7.
|
||||
- **S5:** [libnvme types](https://github.com/linux-nvme/libnvme/blob/master/src/nvme/types.h), `nvme_smart_log` and `nvme_smart_crit` documentation.
|
||||
- **S6:** [smartmontools NVMe JSON implementation](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/nvmeprint.cpp).
|
||||
- **S7:** [Linux man-pages open(2)](https://man7.org/linux/man-pages/man2/open.2.html), exclusive creation and direct/synchronized I/O.
|
||||
- **S8:** [Linux PSI documentation](https://www.kernel.org/doc/html/latest/accounting/psi.html).
|
||||
- **S9:** [Linux I/O statistics documentation](https://www.kernel.org/doc/html/latest/admin-guide/iostats.html).
|
||||
- **S10:** [Linux hwmon sysfs interface](https://www.kernel.org/doc/html/latest/hwmon/sysfs-interface.html).
|
||||
- **S11:** [Linux tmpfs documentation](https://docs.kernel.org/filesystems/tmpfs.html).
|
||||
- **S12:** [Linux VM sysctl documentation](https://docs.kernel.org/admin-guide/sysctl/vm.html), `drop_caches`.
|
||||
- **S13:** [Linux RAS documentation source](https://www.kernel.org/doc/html/latest/_sources/admin-guide/RAS/main.rst.txt), error categories and EDAC.
|
||||
- **S14:** [GNU Coreutils dd manual](https://www.gnu.org/software/coreutils/manual/html_node/dd-invocation.html).
|
||||
- **S15:** [smartmontools drive database](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/drivedb.h), default and model-dependent attribute mappings.
|
||||
- **S16:** [Google, Failure Trends in a Large Disk Drive Population](https://research.google/pubs/failure-trends-in-a-large-disk-drive-population/), primary publication abstract, 2007.
|
||||
- **S17:** [QEMU NVMe implementation](https://github.com/qemu/qemu/blob/master/hw/nvme/ctrl.c), `nvme_smart_info`; [NVMe device documentation](https://github.com/qemu/qemu/blob/master/docs/system/devices/nvme.rst), inspected through Context7.
|
||||
- **S18:** [Linux NVMe ioctl authorization](https://github.com/torvalds/linux/blob/master/drivers/nvme/host/ioctl.c).
|
||||
- **S19:** [Linux kernel sysctl documentation](https://www.kernel.org/doc/html/latest/admin-guide/sysctl/kernel.html), `dmesg_restrict`.
|
||||
Reference in New Issue
Block a user