Author SHA1 Message Date
Codex 6e5af87a64 Research storage measurements and health evidence 2026-09-25 23:30:34 +05:30
2 changed files with 139 additions and 139 deletions
-139
View File
@@ -1,139 +0,0 @@
# Safe CPU, memory, and kernel measurements for Odin
Research for [Establish safe CPU, memory, and kernel measurements](https://git.bongbetic.com/xavierk/odin/issues/2), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
**Accessed:** 25 September 2026. **Status:** recommendations for later decisions, not a selected workload suite. No benchmarks, stress tests, privilege changes, or host configuration changes were performed.
## Findings that shape the decision
Odin can reuse established workloads and Linux interfaces, but no single tool measures system usability, peak performance, and hardware health. Keep four outcomes distinct:
- **Performance measurement:** completed work per second or elapsed time under a specified workload.
- **Responsiveness:** foreground request or wakeup latency, including its tail, under a specified competing load.
- **Pressure observation:** time lost to CPU, memory, or I/O contention during that interval.
- **Health finding:** a detected verification failure or reported hardware error, with the observation's coverage.
This separation follows the tools' actual contracts: sysbench benchmarks operations; schbench measures artificial requests and scheduling delays; PSI measures stalls; memtester checks memory contents. Stress-ng explicitly says it was **not intended as a precise benchmark suite**. Successful stress completion establishes only that the chosen work completed without detected failures under those conditions. [sysbench][sysbench] [schbench][schbench] [PSI][psi] [memtester][memtester-man] [stress-ng][stress-readme]
## Candidate comparison
Maintenance observations describe inspected releases or repository activity, not a guarantee of future support. GPL notices and third-party dependencies need checking again for the exact distributable artifacts. Perf is distributed in the Linux source tree, whose COPYING specifies GPL-2.0-only with the syscall exception and notes that other licenses may also apply; preserve the selected tool and dependency notices. [Linux COPYING][kernel-license]
| Candidate | Useful measurement and limitations | Controls, output, portability, maintenance |
|---|---|---|
| **sysbench CPU / memory** | CPU events are a small prime-search workload, not a general application-performance model. The memory test defaults to a **1 KiB block**; its `100G` default is cumulative transfer size, not RAM allocation. A default run therefore cannot stand for DRAM bandwidth. | Thread, event, duration, warmup and percentile controls; human-readable built-in reports require a pinned parser. Upstream advertises x86_64 and aarch64 packages. GPL-2.0-or-later. Latest published release inspected: 1.0.20, April 2020; repository push activity March 2025. Void still packages 1.0.20. [README][sysbench] [CPU source][sysbench-cpu] [memory source][sysbench-memory] [release][sysbench-release] [metadata][sysbench-meta] [Void][void-sysbench] |
| **schbench** | Reports synthetic request latency, wakeup latency, and requests/second. Closer to responsiveness than peak throughput. Its server-inspired matrix workload deliberately penalizes preemption using per-CPU locks; this is not a desktop interaction model. | Runtime, worker/message counts, work size, request rate, affinity and JSON percentiles. Source has x86 and aarch64 paths; Linux pthread/futex interfaces, no normal root requirement found. Musl behavior remains unverified. GPLv2. Inspected upstream commit `6300b8f`, June 2025. [methodology][schbench] [source][schbench-source] |
| **stress-ng** | Appropriate candidate for controlled background load and selected verification diagnostics. Bogo operations are unsuitable as Odin's cross-test score foundation. | Explicit byte/worker/time caps, verification, YAML and differentiated exit codes. Upstream documents musl builds and testing on ARM64/x86-64. GPL-2.0-or-later; release 0.22.01, September 2026. Void's recipe has explicit musl handling. [README][stress-readme] [manual][stress-man] [release][stress-release] [Void][void-stress] |
| **STREAM** | Established sustainable memory-bandwidth kernels: Copy, Scale, Add, Triad. No memory-latency or comprehensive error-detection result. Each array must exceed cache requirements; a small cache-resident run is not a compliant STREAM result. | Portable C; multicore uses OpenMP and its runtime. Array size is a build parameter in reference 5.10. Text output and numerical validation. Reference source dates to 2013: stable method, not evidence of a modern portability test matrix. Custom license permits use/redistribution but imposes result-naming/run-rule conditions. [source and license][stream-source] [run rules][stream-rules] |
| **lmbench `lat_mem_rd`** | Pointer-chain latency over sizes/strides exposes cache, memory and TLB behavior. Its manual acknowledges vulnerability to stride-sensitive prefetchers; do not present this as an architecture-independent “true RAM latency.” | Warmup, repetitions, bounded size, text pairs. GPLv2 COPYING inspected; Intel's repository is active but contains old documentation. Portability and selected-file licensing need validation before adoption; not a recommended mandatory dependency yet. [manual][lmbench-memory] [README][lmbench-readme] [COPYING][lmbench-license] [metadata][lmbench-meta] |
| **cyclictest / perf** | Cyclictest measures timer wakeup latency; perf supplies diagnostic counters. Neither is a substitute for foreground application response. | Cyclictest offers duration, histogram and JSON; source is GPL-2.0-only, current rt-tests release 2.11. Its startup tests permission to enter SCHED_FIFO even when ordinary policy is selected. Perf depends on kernel/PMU access and can emit JSON. Treat both as optional diagnostic coverage. [cyclictest][cyclictest] [privilege check][rt-utils] [rt-tests release][rt-release] [perf][perf-stat] |
| **memtester** | Online checking of allocated memory, not all installed RAM. Failures can involve memory, CPU, temperature or power; the result does not identify a replaceable DIMM by itself. | Byte size and finite iteration count; default iterations are infinite. Text plus exit-bit mask. Record actual allocation and locking, not only exit status. GPL-2.0-only. Version 4.7.1's December 2024 fix addresses stricter C23/GCC 15 compilation; Void recipe inspected still selects 4.6.0. [manual][memtester-man] [source][memtester-source] [changelog][memtester-changelog] [Void][void-memtester] |
| **Memtest86+** | Offline, bootable diagnostics reach almost all memory without the resident OS. They cannot run as an ordinary in-terminal stage. | GPLv2. Stable v8.10, May 2026, lists x86, x86-64 and LoongArch64. Current main additionally lists AArch64 with UEFI boot. This is a **stable/development difference**, not certified aarch64 coverage. [stable README][memtest-stable] [development README][memtest-main] [release][memtest-release] |
| **EDAC / rasdaemon** | Hardware-error telemetry, complementary to active tests. Availability depends on hardware, firmware, drivers and exposed events. | Read EDAC counters where accessible; optionally consume an existing rasdaemon history. Rasdaemon monitors kernel trace events and has database backends; its repository shows September 2026 activity and GPLv2 metadata. Starting it is a separate privileged monitoring action, not necessary to read available counters. [EDAC ABI][edac-abi] [rasdaemon][rasdaemon] [metadata][ras-meta] |
## Measuring usability under contention
**Recommendation:** evaluate paired idle and loaded foreground-request measurements as a first-class candidate. Run the same bounded foreground work alone, with a fixed CPU background load, and with separately controlled memory pressure. Preserve normal scheduling policy, work size, thread count, placement, requested arrival rate, actual throughput, sample count, and latency distribution. Report absolute latency and degradation relative to idle; a fast idle result can coexist with poor responsiveness under contention.
Existing tools can supply the observations. Schbench records both request and wakeup latency, supports a fixed request rate, and can coexist with an independently bounded stress-ng background worker. Sysbench's rate-limited engine is another candidate: its source adds measured queue time to event duration and reports queue length. These remain synthetic proxies for foreground work; neither measures keyboard-to-pixel delay, terminal rendering, browser interaction, or application launch by itself. Those need their own workload definitions. [schbench methodology][schbench] [schbench source][schbench-source] [sysbench queue][sysbench-core] [sysbench timer][sysbench-timer]
Schbench needs qualification before selection. Its README says warmup defaults to five seconds, while the inspected source defaults to zero and bypasses its warmup reset in request-rate mode. It uses `gettimeofday`, so clock adjustments can contaminate timing. Set options explicitly, retain the revision, validate timing conditions, and decide whether its deliberate preemption penalty fits Odin's goal. Do not adapt work size independently on every system and then compare the resulting latency as equal work. [source][schbench-source]
Cyclictest is a useful separate scheduler diagnostic. Its default behavior can hold `/dev/cpu_dma_latency` at zero and suppress deep idle states; `--default-system` avoids that tuning. The source's unconditional real-time privilege check prevents assuming that an ordinary-policy configuration is universally unprivileged. Its timer latency, especially under SCHED_FIFO, is a different measurement from a normal foreground application's response. [manual][cyclictest] [source][cyclictest-source] [privileges][rt-utils]
## Kernel telemetry and compatibility
**PSI:** Read system and, where available, workload-cgroup `cpu`, `memory`, and `io` pressure. `some` is time when at least some tasks are stalled; `full` is time when all non-idle tasks are stalled together. The cumulative `total` counter permits interval deltas; rolling 10/60/300-second averages can smear a short benchmark across adjacent phases. **System-wide CPU `full` is undefined and exposed as zero for compatibility**—zero there cannot mean perfect responsiveness. PSI exists in the inspected Linux 4.20 source, but requires `CONFIG_PSI` and may be disabled by default pending `psi=1`. Probe actual files and readability, not only kernel version. [PSI][psi] [4.20 documentation][psi-420] [Kconfig][kconfig]
**Capacity and pressure:** Record usable RAM, `MemAvailable`, swap capacity/usage, process or cgroup memory, major faults, and swap/reclaim activity where exposed. `MemAvailable` is an estimate of memory available without swapping, not an allocation guarantee. Swap occupancy alone does not establish current pressure. `/proc/stat` supplies CPU time and steal time, but kernel documentation explicitly warns that `iowait` is unreliable. Correlate these observations with latency and PSI; do not derive a definitive bottleneck from CPU utilization or a single counter. [proc documentation][proc]
**Cgroups:** Observe effective CPU affinity/cpuset, CPU quota, memory/swap limits and relevant ancestors. Host RAM and online CPU counts may exceed what the benchmark is allowed to use. Cgroup v2 `cpu.stat` records throttling; `memory.events` separates high-limit reclaim, OOM conditions and kills. Parse by key: the kernel explicitly allows new `memory.stat` entries in the middle. These interfaces are independent of a particular init system, but writable delegation and enabled controllers are not guaranteed on Void, deb, rpm, containers, or user sessions. [cgroup v2][cgroup]
**Perf:** Treat hardware counters as enrichment. `CONFIG_PERF_EVENTS`, CPU PMU support, virtualization, `perf_event_paranoid`, capabilities and distribution policy can limit access. Kernel documentation recommends `CAP_PERFMON` over broad `CAP_SYS_ADMIN`; Odin should describe missing access rather than lowering system security settings. Record event identity and time-running percentage when multiplexing occurs. PMU-specific cache and pipeline events are not universal normalized scores. Perf's manual also warns of overhead at short sampling intervals, particularly below 100 ms. [security][perf-security] [Kconfig][kconfig] [perf stat][perf-stat]
**Architecture/libc:** Proc/sysfs/cgroup interfaces offer the strongest common layer across x86_64/aarch64 and glibc/musl. Workload binaries still need distinct, verified artifacts and recorded toolchain flags. Upstream stress-ng explicitly documents musl; sysbench's advertised architectures and Void recipes are useful evidence, but none of the inspected material certifies Odin's whole four-way architecture/libc matrix. STREAM additionally needs a compatible OpenMP runtime; rt-tests has library dependencies. A package recipe proves availability intent, not successful operation. Missing checks should carry reasons such as unsupported, permission denied, unavailable dependency, or insufficient safe resources. [stress-ng][stress-readme] [sysbench][sysbench] [Void recipes][void-sysbench] [STREAM][stream-source] [rt-tests Makefile][rt-makefile]
## Resource safety and cancellation
The following is a proposed execution contract, with exact caps left to the safety decision:
1. Calculate a conservative working budget from current `MemAvailable`, effective cgroup/ancestor headroom, expected tool/runtime overhead, and a retained reserve. Recheck while running. There is no sourced universal percentage that guarantees safety when other programs allocate concurrently.
2. Where delegated cgroup v2 control exists, put disposable workers in their own subtree and keep the supervisor outside that subtree. `memory.high` induces reclaim/throttling and **is not a hard cap**; `memory.max` bounds charged memory and can invoke OOM inside the cgroup. `memory.swap.max` separately controls swap. Use deliberate limits and record them because they change results. Caps reduce risk; they cannot guarantee that an unrelated system-wide shortage never kills a process. [cgroup v2][cgroup]
3. Reserve intentional pressure for explicitly selected, isolated work. Without reliable containment or enough reserve, recommend pressure **observation** and small bounded workloads, and report unavailable active-pressure coverage. Allocation success alone is insufficient: memtester's own manual warns about overcommit, swapping and OOM affecting other programs. [memtester][memtester-man]
4. Never expose memtester's physical-address/device modes in the normal benchmark path. They overwrite the mapped region and can crash the system when it belongs to another process or the kernel. For ordinary allocations, verify locked bytes and completed patterns/iterations. Linux permits unprivileged locking up to `RLIMIT_MEMLOCK`; larger locking requires suitable privilege, commonly `CAP_IPC_LOCK`. The tool's “run as root” advice should not force the entire TUI to run as root. [manual][memtester-man] [Linux mlock][mlock]
5. Use finite work/time limits and a supervisor deadline. For stress-ng, enable only reviewed stressors, verification where supported, and no OOM respawn (`--oomable`). Its `--oom-avoid` is a heuristic with measurement overhead, not containment. In 0.22.01, `--vm-bytes` describes a total across VM workers; other stressors have different allocation semantics, so retain the exact version and options. [stress-ng manual][stress-man]
6. Cancel the worker process group gracefully, then terminate remaining descendants after a defined grace period; use `cgroup.kill` when accessible. Preserve a cancelled/partial result. Stress-ng documents SIGINT cleanup, but its timeout can overrun during uninterruptible calls or cleanup. No userspace deadline guarantees immediate cancellation of an uninterruptible kernel task. [manual][stress-man] [cgroup kill][cgroup]
## Repetition, kernel settings, and TUI overhead
**Recommendation:** record warmup separately, repeat bounded measurements, retain all repetitions and dispersion, and report a median only at a clearly defined level. STREAM's official report takes the **best** iteration after discarding the first; a median of repeated STREAM run results is a different statistic. Do not silently relabel its internal minimum as a median, combine raw milliseconds with MB/s, or replace an unavailable result with zero. Overall score normalization belongs to the scoring decision. [STREAM source][stream-source]
Record kernel/build identity, visible preemption/scheduler settings, CPU topology and allowed CPUs, NUMA placement, THP policy, libc, workload/compiler version and flags, governor/driver, boost, power source and temperature observations. NUMA placement and THP policy affect what memory workload is actually measured. Kernel CPUFreq documentation explains that `scaling_cur_freq` can be a requested state rather than measured frequency, and boost depends on thermal/power conditions and package load. A governor name or falling frequency alone does not prove thermal throttling. Correlate sustained performance with available temperatures, thermal trip/cooling states, and power/frequency evidence; unavailable sensors remain unavailable. [CPUFreq][cpufreq] [NUMA][numa] [THP][thp] [thermal interfaces][thermal]
For “single core,” specify whether the worker is pinned and how the core is chosen on heterogeneous CPUs. For “multicore,” specify workers relative to allowed CPUs, SMT and quota; do not silently change those rules between systems. Preserve the machine's existing configuration for the baseline. Potential governor, scheduler, THP, affinity or kernel changes should be advice or separately labelled experiments, not automatic optimization before measuring.
The TUI competes for CPU time, memory bandwidth, cache and terminal I/O. Recommend throttled graph updates, buffered logs, no expensive animation during timed sections, and a quiet measurement mode that preserves cancellation. Avoid hiding this by reserving a core without recording it: that reduces tested capacity. Later validation should compare quiet versus normal rendering on the slowest supported machines and establish an overhead budget. This is a proposed qualification experiment, not evidence that a particular redraw rate is already safe. Lmbench explicitly warns about competing cache/CPU work; PSI's own Kconfig notes overhead can show up in synthetic scheduler stress tests. [lmbench][lmbench-readme] [Kconfig][kconfig]
## Memory errors, VM coverage, and remaining decisions
Online memtester cannot touch RAM occupied by the kernel or other processes. It may allocate less than requested and may continue unlocked; inspected 4.7.1 source can then still exit zero if its pattern checks succeed. Parse allocation/locking evidence alongside exit bits and report “no errors detected in the tested allocation,” with size, iterations and duration. A mismatch warrants investigation, not an automatic RAM-replacement diagnosis. [manual][memtester-man] [source][memtester-source]
EDAC counters reset at driver initialization or explicit reset; preserve counter baselines and `seconds_since_reset` without resetting them. Corrected errors merit attention, but uncorrected errors may cause a panic before a counter increments. DIMM labels can depend on board-specific userspace mapping. Therefore missing EDAC nodes, zero observed deltas, and an empty rasdaemon history cannot certify error-free RAM. Offer offline follow-up where supported; decide how development-only AArch64 Memtest86+ support should be presented. [EDAC ABI][edac-abi] [EDAC model][edac] [Memtest86+ stable][memtest-stable] [development][memtest-main]
VMs can validate packaging, libc/architecture execution, permissions, telemetry fallbacks, cgroup containment and result handling. Guest CPU/memory scores describe the guest allocation and host scheduling conditions; steal time is useful context. They do not certify the host's DIMMs, ECC pipeline, cooling, physical memory-channel bandwidth, or representative bare-metal scheduler tails. Rasdaemon's upstream QEMU tests intentionally inject virtual nonfatal events: useful for exercising decoding, not proving physical hardware health. [proc][proc] [rasdaemon CI description][rasdaemon]
**Recommended next decision:** shortlist sysbench for a narrow CPU baseline, STREAM for bandwidth, schbench for responsiveness qualification, selected stress-ng workers for bounded load, and memtester plus available EDAC/RAS for diagnostics. Keep perf/cyclictest optional; defer mandatory memory-latency scoring until a candidate is validated. Final selection remains open.
The human-facing decisions still needed are the foreground workload's meaning; fixed versus relative background load; inclusion of responsiveness in the median score; safe resource reserves and privileges; repetition/time allocation within the 10–20 minute standard run; treatment of heterogeneous cores and missing coverage; and whether offline/development-tool guidance belongs in the first release.
**Evidence limits:** no candidate was built or executed across the target matrix. Context7 resolved Linux kernel, sysbench, memtester and rt-tests documentation. Stress-ng, STREAM and schbench searches returned unrelated libraries, so no false library match was used; their owning sources were inspected directly. Memtester's upstream HTTPS site failed certificate validation; the report uses the original source/manpage/changelog preserved by Debian, cross-checked against the Void 4.6.0 source archive checksum, and identifies the version difference. Exact artifact compatibility, parser contracts, resource budgets, and score repeatability require later qualification.
[sysbench]: https://github.com/akopytov/sysbench/blob/master/README.md
[sysbench-cpu]: https://github.com/akopytov/sysbench/blob/master/src/tests/cpu/sb_cpu.c
[sysbench-memory]: https://github.com/akopytov/sysbench/blob/master/src/tests/memory/sb_memory.c
[sysbench-core]: https://github.com/akopytov/sysbench/blob/master/src/sysbench.c
[sysbench-timer]: https://github.com/akopytov/sysbench/blob/master/src/sb_timer.h
[sysbench-release]: https://github.com/akopytov/sysbench/releases/tag/1.0.20
[sysbench-meta]: https://api.github.com/repos/akopytov/sysbench
[void-sysbench]: https://github.com/void-linux/void-packages/blob/master/srcpkgs/sysbench/template
[schbench]: https://kernel.googlesource.com/pub/scm/linux/kernel/git/mason/schbench/+/6300b8f3a8922c61ea6bb2cdfa1901a42c0cc6fc/README.md
[schbench-source]: https://kernel.googlesource.com/pub/scm/linux/kernel/git/mason/schbench/+/6300b8f3a8922c61ea6bb2cdfa1901a42c0cc6fc/schbench.c
[stress-readme]: https://github.com/ColinIanKing/stress-ng/blob/V0.22.01/README.md
[stress-man]: https://github.com/ColinIanKing/stress-ng/blob/V0.22.01/stress-ng.1
[stress-release]: https://github.com/ColinIanKing/stress-ng/releases/tag/V0.22.01
[void-stress]: https://github.com/void-linux/void-packages/blob/master/srcpkgs/stress-ng/template
[stream-source]: https://www.cs.virginia.edu/stream/FTP/Code/stream.c
[stream-rules]: https://www.cs.virginia.edu/stream/ref.html
[lmbench-memory]: https://github.com/intel/lmbench/blob/master/doc/lat_mem_rd.8
[lmbench-readme]: https://github.com/intel/lmbench/blob/master/README
[lmbench-license]: https://github.com/intel/lmbench/blob/master/COPYING
[lmbench-meta]: https://api.github.com/repos/intel/lmbench
[cyclictest]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/src/cyclictest/cyclictest.8
[cyclictest-source]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/src/cyclictest/cyclictest.c
[rt-utils]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/src/lib/rt-utils.c
[rt-release]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6
[rt-makefile]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/Makefile
[perf-stat]: https://github.com/torvalds/linux/blob/master/tools/perf/Documentation/perf-stat.txt
[kernel-license]: https://github.com/torvalds/linux/blob/master/COPYING
[perf-security]: https://docs.kernel.org/admin-guide/perf-security.html
[memtester-man]: https://sources.debian.org/data/main/m/memtester/4.7.1-1/memtester.8
[memtester-source]: https://sources.debian.org/data/main/m/memtester/4.7.1-1/memtester.c
[memtester-changelog]: https://sources.debian.org/data/main/m/memtester/4.7.1-1/CHANGELOG
[void-memtester]: https://github.com/void-linux/void-packages/blob/master/srcpkgs/memtester/template
[mlock]: https://man7.org/linux/man-pages/man2/mlock.2.html
[memtest-stable]: https://github.com/memtest86plus/memtest86plus/blob/v8.10/README.md
[memtest-main]: https://github.com/memtest86plus/memtest86plus/blob/main/README.md
[memtest-release]: https://github.com/memtest86plus/memtest86plus/releases/tag/v8.10
[edac-abi]: https://github.com/torvalds/linux/blob/master/Documentation/ABI/testing/sysfs-devices-edac
[edac]: https://docs.kernel.org/driver-api/edac.html
[rasdaemon]: https://github.com/mchehab/rasdaemon/blob/master/README.rst
[ras-meta]: https://api.github.com/repos/mchehab/rasdaemon
[psi]: https://docs.kernel.org/accounting/psi.html
[psi-420]: https://github.com/torvalds/linux/blob/v4.20/Documentation/accounting/psi.txt
[kconfig]: https://github.com/torvalds/linux/blob/master/init/Kconfig
[proc]: https://docs.kernel.org/filesystems/proc.html
[cgroup]: https://docs.kernel.org/admin-guide/cgroup-v2.html
[cpufreq]: https://docs.kernel.org/admin-guide/pm/cpufreq.html
[numa]: https://docs.kernel.org/admin-guide/mm/numa_memory_policy.html
[thp]: https://docs.kernel.org/admin-guide/mm/transhuge.html
[thermal]: https://docs.kernel.org/driver-api/thermal/sysfs-api.html
+139
View File
@@ -0,0 +1,139 @@
# Storage measurements and trustworthy health advice
Research for [Establish storage measurements and trustworthy health advice](https://git.bongbetic.com/xavierk/odin/issues/4), part of Odin's Wayfinder map. Access date for every source: **2026-09-25**.
This report establishes evidence and candidate policies. It does not select Odin's final workloads, thresholds, privileged execution design, or score. No benchmark, device query, self-test, installation, or hardware change was performed during this investigation.
## Findings that shape the decision
The strongest candidate is **fio for file-based performance measurements, smartmontools for cross-protocol health findings, and an optional nvme-cli adapter for additional NVMe evidence**. These tools cover different responsibilities. A fast benchmark cannot establish drive health; a passing SMART status cannot establish future reliability. Health findings should therefore remain visible independently of the performance score and capability coverage.
There is a material compatibility change already: nvme-cli **v3.1**, released September 18, 2026, documents `nvme log smart`; `nvme smart-log` is a deprecated compatibility alias. Its default output format version is now 2, with version 1 available. fio **3.43** was released September 23. A bleeding-edge development environment is compatible with reproducible measurements only if Odin records tool versions and keeps workload and parser versions explicit. Package-manager availability alone does not establish supported commands or JSON schemas. [S1][S3]
## Candidate tools and measurement scope
| Candidate | Useful responsibility | Limits and recommendation to consider |
| --- | --- | --- |
| fio | Sequential/random reads and writes, block sizes, queue depths, latency distributions, bounded I/O, optional verification | Best primary workload candidate. Select a small fixed workload vocabulary; do not accept arbitrary user-supplied job files into privileged execution. |
| smartctl | ATA, SCSI and NVMe identity, health, existing error/self-test logs; JSON; many bridge/controller adapters | Best baseline health reader. Decode protocol-specific semantics and command status separately. Some transports are unsafe for automatic probing. |
| nvme-cli | NVMe-specific identity, SMART and detailed logs | Useful optional supplement. Version 2/3 command and JSON differences need explicit compatibility handling. Avoid duplicating the same controller's health as several independent findings. |
| Native `/proc` and `/sys` | I/O pressure, completed I/O, queue activity and available sensors | Low-dependency contextual evidence, not a workload or a substitute for SMART. |
| GNU `dd` | Bounded sequential copying, optionally direct I/O and final synchronization | Possible explicitly labelled basic fallback. Its copy-oriented output does not supply fio's workload control or latency distributions; its result must not silently substitute into the same scored workload. |
Sources: fio HOWTO, smartctl manual, nvme-cli released documentation, kernel PSI/I/O documentation and GNU manual. [S1–S4][S8][S9][S14]
A compact candidate performance set is:
| Measurement | Candidate workload, still to be selected | What its result means |
| --- | --- | --- |
| Sequential read/write | Large blocks, for example 1 MiB, one job, depth 1 | Large-file throughput through the selected filesystem and storage path |
| Random read/write | 4 KiB, one job, depth 1 | Small-request responsiveness; report latency and IOPS |
| Queued random read | Same block size at a documented higher depth, such as 16 or 32 | Concurrency capability; a different workload from depth 1 |
| Durable small writes | A separate small, bounded workload with defined sync frequency | Application-visible cost of requesting persistence |
| Optional integrity check | Write and verify Odin-owned file blocks with fio checksums | Whether the tested data path returned those bytes correctly; not a full-surface drive or whole-RAM certification |
Record read and write throughput in explicit units, IOPS, completed bytes, operation count, errors, elapsed time, and p50/p95/p99 latency where sample counts support them. Distinguish fio completion latency from total latency: total includes submission latency. Record achieved queue-depth distribution; requesting depth greater than one does not make a synchronous engine asynchronous. fio `psync` is a useful depth-1 compatibility candidate; `io_uring` and `libaio` are queued-engine candidates where supported. Engine changes must be visible in results and comparability rules. [S1]
Measure one storage workload at a time during reference runs. Concurrent CPU or memory stress can instead be an explicitly identified contention experiment. Record filesystem, mount options, device topology, encryption/RAID/virtualization, kernel, selected engine, power/thermal state, background I/O, and pre/post free space. The measurement describes this path under these conditions, not the NVMe/HDD in isolation.
## Safe operation and comparability constraints
The following are proposed invariants, rather than finalized profile numbers:
1. **Own every writable byte.** Create a private run directory on a deliberately selected filesystem and exclusively create its regular files. Validate ownership, type and target identity; reject symlink redirection and raw block/character devices. `O_CREAT|O_EXCL` supplies exclusive creation semantics. Do not use arbitrary existing user files as write targets. Avoid selecting `/tmp` automatically: tmpfs stores files in virtual memory and may use swap. [S7][S11]
2. **Budget storage space and cumulative writes separately.** fio `size` defines the working region, while `io_size` can independently bound I/O. `runtime` stops at the earlier of completion or time limit; `time_based` loops the workload. Thus a small file plus a timed loop can write many times its size. Prefer explicit byte and time bounds without `time_based` for ordinary write profiles. Count fixture preparation, repetitions and verification-related writes in a per-run host-write budget. Reserve free space, account for quotas and metadata, recheck during execution, and stop on ENOSPC or I/O errors. Space and byte thresholds remain product decisions. [S1]
3. **Do not promise a physical NAND-write limit.** A workload's host bytes are measurable; filesystem/controller write amplification and unrelated host activity are additional. NVMe Data Units Written measures host data in units of 1,000 × 512 bytes, rounded up, excluding metadata; it is not a universal NAND-wear counter. Background activity also prevents attributing its entire delta to Odin. [S5]
4. **Make cache and durability modes explicit.** fio `direct=1` normally requests `O_DIRECT`; support and alignment vary by filesystem and kernel, and misaligned requests can fail or fall back to buffered I/O. Direct I/O does not by itself provide `O_SYNC` persistence guarantees, bypass every device cache, or prove sustained media speed. `invalidate` is conditional on platform/file support. Avoid global `drop_caches`: kernel documentation warns of additional I/O and CPU costs. A buffered fallback must be labelled and excluded from direct-I/O comparisons. [S1][S7][S12]
5. **Include preparation and flush costs honestly.** Read tests over newly created fixtures still require writes. A user choosing no writes can reuse an identified valid fixture or skip that workload; Odin should not create one silently. Do not measure unwritten sparse-file holes as disk reads. For writes, document `end_fsync` or other synchronization and report end-to-end time including the final flush separately from unsynchronized throughput. Control data compressibility/deduplication using a declared fio buffer policy; generating fresh data adds CPU cost. Preserve normal filesystem settings rather than silently disabling compression or copy-on-write. [S1][S7]
6. **Treat cancellation and cleanup as part of the run.** Bound the entire job group; stop launching work on cancellation, retain partial status, reap workers, and remove only proven Odin-owned artifacts. A worker stuck in kernel I/O may not stop immediately. Crash recovery needs a manifest and ownership checks before deletion. Cleanup failure is a reported outcome, never a reason to recursively delete a user-selected directory.
7. **Collect health before load and reduce work when evidence is serious.** A candidate policy is to skip storage stress when critical media/reliability findings or current unreadable data are already present. Pause on documented thermal alarms or loss of safety headroom. Display estimated host writes before a write run. Avoid automatic discard/TRIM, formatting, SMART feature changes, firmware updates, cache-policy changes or repair operations as benchmark preparation.
Short bounded tests cannot establish steady-state SSD performance after exhaustion of a large write cache, or scan every HDD sector. Making test data larger than all caches can conflict with a quick run and a conservative write budget. Report the actual duration and working set; do not extrapolate a short burst into an endurance or sustained-performance guarantee. The tradeoff between low impact and sustained measurements needs an explicit run-profile decision.
## Health evidence and field interpretation
Prefer structured output with the original tool version, schema identifier, command outcome, timestamp, device identity and transport. A missing field is unknown, not zero. Preserve large counters losslessly: smartctl JSON can emit string/byte-array companions for integers exceeding JavaScript's safe integer range; `--json=v` requests them consistently. This matters for an Ink/JavaScript consumer. [S2]
For smartctl NVMe output, the primary object is `nvme_smart_health_information_log`. The inspected source confirms the following keys and conversions. Raw NVMe temperature is Kelvin; smartctl's `temperature` here is already Celsius. Do not convert it twice. [S6]
| Evidence | Meaning | Candidate interpretation |
| --- | --- | --- |
| `critical_warning` bit 0; `available_spare` vs `available_spare_threshold` | Spare capacity below the controller's threshold | Urgent preservation/service finding; display the device-provided threshold |
| Bit 1; `temperature`, warning/critical temperature time | Above an over-temperature or below an under-temperature threshold | Stop heat-producing tests; investigate cooling/environment. This alone is not proof that replacement is needed |
| Bit 2 | NVM subsystem reliability degraded | Urgent backup and replacement/service assessment |
| Bit 3 | Media placed read-only for a device reliability condition | Urgent preservation and replacement/service assessment; distinct from user namespace write protection |
| Bits 4/5 | Volatile-memory backup failure; persistent-memory region read-only/unreliable | Urgent loss-of-protection/service finding when applicable; explain the specific subsystem |
| `percentage_used` | Vendor estimate of endurance consumed | 100 means estimated endurance consumed, **not guaranteed failure**; values may exceed 100. Plan replacement according to manufacturer guidance and workload, without inventing days remaining |
| `media_errors` | Unrecovered data-integrity errors, including ECC/CRC/tag errors | Investigate any nonzero history; escalating recent deltas plus failed I/O are much stronger urgent evidence than an isolated old count |
| `num_err_log_entries` | Lifetime number of error-information entries | Inspect status/cause and recency. It is not interchangeable with media errors |
| `unsafe_shutdowns` | Loss of power without shutdown notification | Investigate shutdown/power history and correlate with errors; not proof of failed media |
| `data_units_written`, power-on hours, thermal counters | Usage/history with specified units and reporting limits | Useful trends and context; no universal lifespan formula |
NVMe warning bits are current state, not persistent event history; zero today does not erase yesterday's finding. Some temperature fields are optional, and zero can mean unsupported. Per-namespace SMART is optional; the global namespace identifier can describe a controller's aggregate. Preserve scope rather than assigning identical controller totals to every namespace. [S3][S5][S6]
**ATA needs a separate mapping.** Keep attribute ID, raw representation, normalized current/worst value, threshold, type and failure state. smartctl states that these meanings are vendor-specific; SSD meanings can differ and displayed names can be wrong for models absent from its drive database. The label `Pre-fail` by itself does not mean a drive is failing: the current normalized value must cross its threshold. [S2]
Common drive-database candidates include reallocated sectors (5), pending sectors (197), offline uncorrectable sectors (198), and interface CRC errors (199). Interpret them only with a matching model/firmware/database rule; do not apply a universal raw-count threshold or turn interface errors directly into a disk-replacement recommendation. Preserve lifetime history and recent deltas separately. SMART RETURN STATUS, failed applicable thresholds, existing self-test failures, and observed host I/O errors are stronger when they agree. SCSI health uses its own exception/sense reporting rather than ATA attribute assumptions. [S2][S15]
smartctl exit status is a bitmask. Bits 0–2 can describe invocation/access/command problems; bits 3–7 describe failing status, thresholds and historical error/self-test evidence. A nonzero exit must not discard usable JSON, and access failure must not become “bad drive.” Reading an existing self-test log is different from starting a test. The manual notes that running self-tests can degrade performance and normal I/O can extend their duration; any future self-test workflow needs a separate user decision and scheduling. [S2]
## Candidate advice rubric
This rubric is a proposed interpretation layer over the documented evidence, not a manufacturer's warranty or an adopted Odin policy.
| Finding class | Evidence sufficient to consider it | Appropriate wording/action |
| --- | --- | --- |
| **Replace/service now** | Credible ATA failing status/current applicable prefailure threshold; NVMe degraded reliability/read-only media; serious repeated data-integrity failures attributable to the device | “Preserve accessible data now; avoid further stress; arrange replacement or service.” Cite exact flags, device scope and timestamps. Hardware attribution may still need confirmation |
| **Investigate urgently** | New media errors, pending/uncorrectable sectors, recent failed self-tests, resets/timeouts, thermal alarms, loss of power-loss protection | Identify the failing path; correlate controller, connection, power and filesystem evidence. Do not automatically blame the medium |
| **Monitor / plan replacement** | Stable historical findings or vendor-estimated endurance consumed without current failure evidence | Retain trends, explain wear status and manufacturer limits, and plan according to importance/workload. No invented remaining-life percentage |
| **No concerning evidence observed** | Successful supported collection with no relevant current finding | State what was checked and when; keep normal backup advice independent of a performance score |
| **Unknown / limited coverage** | Missing permission/tool/field, sleeping drive, unsupported bridge/controller, virtual device, ambiguous identity | Explain the missing capability and a bounded next step. Never convert unavailable evidence into a healthy badge |
The smartctl manual recommends preserving data promptly when the drive reports failing health. Conversely, Google's primary HDD population study found that SMART-only models were unlikely to predict individual failures reliably. That older HDD result is not a calibrated modern-SSD failure model, but it reinforces the distinction between a useful warning and a guarantee of future health. NVMe's own endurance-field semantics explicitly reject equating 100% usage with failure. [S2][S5][S16]
## Compatibility, privilege and general health
USB, SAT and RAID support must follow known transport rules. The smartctl manual documents bridge-specific NVMe adapters and per-physical-disk MegaRAID addressing; a RAID logical volume is not automatically one physical drive. Particularly important: its **JMB39x/JMS56x transport uses READ/WRITE commands to a RAID-volume sector**. It warns that the wrong device can be overwritten and interruption can prevent restoration. Exclude these from routine automated probing; “try every device type” is not a safe compatibility strategy. Even standby-aware queries may wake a disk during autodetection, so record unsupported power-state handling. [S2]
VMs need an explicit virtual-device classification. QEMU's NVMe implementation constructs SMART data from its emulated controller and block-accounting state. A guest can therefore show valid-looking SMART without revealing the host drive's health. Guest tests establish guest-path performance; actual physical passthrough and device identity require separate verification. VM coverage cannot establish USB, physical RAID, real wear counters or thermal behavior. [S17]
Keep ordinary file workloads unprivileged. Device queries may require additional device permissions or kernel capabilities; NVMe's Linux passthrough code explicitly gates classes of commands. A future privileged mechanism should allow only validated read operations and selected devices, with no arbitrary shell or passthrough-command forwarding. Permission denial is an expected capability outcome, not an instruction to run the whole TUI as root. [S18]
For overall system health, useful complementary evidence is:
- `/proc/pressure/io`: `some` measures time with some stalled tasks, `full` time with all non-idle tasks stalled; use same-window deltas alongside workload latency. This detects pressure, not its sole cause. [S8]
- `/proc/diskstats` or per-device sysfs statistics: completed I/O, time and queue context. Counters have concurrency/accounting caveats; busy percentage alone does not establish NVMe saturation. [S9]
- Available hwmon readings, limits and alarm flags: retain sensor identity and units; chip-specific alarms and missing sensors preclude a universal hard-coded temperature cutoff. Standard hwmon ABI readings are intended to be readable by unprivileged applications. [S10]
- Kernel errors and existing EDAC/RAS evidence: distinguish corrected errors from uncorrected/fatal errors and report available history. EDAC documentation explicitly says corrected errors may, but need not, predict later uncorrected errors. Missing reporting hardware/driver is unknown. Kernel log access can require `CAP_SYSLOG` when `dmesg_restrict=1`; do not assume systemd/journald on Void or other distributions. [S13][S19]
Kernel or mount optimizations should be suggestions tied to an observed limitation and a documented tradeoff, recorded for subsequent comparable runs. This research supports observing current settings and thermal/power/error evidence; it supplies no evidence for blanket scheduler, write-cache, governor, or filesystem changes.
## Remaining decisions and evidence gaps
The next human decisions are the ordinary run's write authorization/budget, minimum free-space reserve, workload lengths and repetitions, required versus optional queued/sync/verification tests, how reduced-capability results affect score eligibility, the supported transport list, the privilege interaction, and the exact advice wording. A sustained-media profile would need a separate impact budget.
Implementation work will need parser fixtures from supported smartctl/nvme-cli versions; success, partial and denied-permission results; real ATA/NVMe/USB/RAID samples; healthy and failing vendor examples; and proof of cancellation, space reservation, direct-I/O handling and cleanup across filesystems. No such hardware validation occurred here. No calibrated cross-device replacement thresholds or modern SSD remaining-life model were found or claimed.
Context7 library resolution succeeded for fio, smartmontools, nvme-cli, Linux kernel, GNU Coreutils and QEMU. Both allowed nvme-cli documentation fetches returned “Could not fetch documentation snippets”; its official released documents and source were inspected instead. The NVM Express specifications landing page returned HTTP 403, so NVMe field semantics here are grounded in maintained libnvme definitions and smartmontools implementation rather than a directly retrieved current specification PDF. Those are material evidence limits, not silently filled gaps.
## Sources inspected
- **S1:** [fio 3.43 HOWTO](https://github.com/axboe/fio/blob/fio-3.43/HOWTO.rst), relevant workload, size/runtime, buffering, engines, percentile, verification and error sections; [release](https://github.com/axboe/fio/releases/tag/fio-3.43).
- **S2:** [smartctl manual source](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/smartctl.8.in), health, attributes, JSON, exit status, device transports, standby and self-tests.
- **S3:** nvme-cli v3.1 [SMART log command](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/nvme-log-smart.txt), [global options](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/global-options.txt), [legacy alias](https://github.com/linux-nvme/nvme-cli/blob/master/Documentation/nvme-smart-log.txt), and [release](https://github.com/linux-nvme/nvme-cli/releases/tag/v3.1).
- **S4:** [smartmontools NVMe support examples](https://www.smartmontools.org/wiki/NVMe_Support), inspected through Context7.
- **S5:** [libnvme types](https://github.com/linux-nvme/libnvme/blob/master/src/nvme/types.h), `nvme_smart_log` and `nvme_smart_crit` documentation.
- **S6:** [smartmontools NVMe JSON implementation](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/nvmeprint.cpp).
- **S7:** [Linux man-pages open(2)](https://man7.org/linux/man-pages/man2/open.2.html), exclusive creation and direct/synchronized I/O.
- **S8:** [Linux PSI documentation](https://www.kernel.org/doc/html/latest/accounting/psi.html).
- **S9:** [Linux I/O statistics documentation](https://www.kernel.org/doc/html/latest/admin-guide/iostats.html).
- **S10:** [Linux hwmon sysfs interface](https://www.kernel.org/doc/html/latest/hwmon/sysfs-interface.html).
- **S11:** [Linux tmpfs documentation](https://docs.kernel.org/filesystems/tmpfs.html).
- **S12:** [Linux VM sysctl documentation](https://docs.kernel.org/admin-guide/sysctl/vm.html), `drop_caches`.
- **S13:** [Linux RAS documentation source](https://www.kernel.org/doc/html/latest/_sources/admin-guide/RAS/main.rst.txt), error categories and EDAC.
- **S14:** [GNU Coreutils dd manual](https://www.gnu.org/software/coreutils/manual/html_node/dd-invocation.html).
- **S15:** [smartmontools drive database](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/drivedb.h), default and model-dependent attribute mappings.
- **S16:** [Google, Failure Trends in a Large Disk Drive Population](https://research.google/pubs/failure-trends-in-a-large-disk-drive-population/), primary publication abstract, 2007.
- **S17:** [QEMU NVMe implementation](https://github.com/qemu/qemu/blob/master/hw/nvme/ctrl.c), `nvme_smart_info`; [NVMe device documentation](https://github.com/qemu/qemu/blob/master/docs/system/devices/nvme.rst), inspected through Context7.
- **S18:** [Linux NVMe ioctl authorization](https://github.com/torvalds/linux/blob/master/drivers/nvme/host/ioctl.c).
- **S19:** [Linux kernel sysctl documentation](https://www.kernel.org/doc/html/latest/admin-guide/sysctl/kernel.html), `dmesg_restrict`.