Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6e5af87a64 |
@@ -1,135 +0,0 @@
|
||||
# Establish browser, shell, and language workload validity
|
||||
|
||||
Research for [Establish browser, shell, and language workload validity](https://git.bongbetic.com/xavierk/odin/issues/5), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
|
||||
|
||||
Access date: **2026-09-25**. This is a candidate analysis, not the final suite. No browsers, toolchains, or dependencies were installed; no benchmarks or builds were run.
|
||||
|
||||
## Recommendation for the selection decision
|
||||
|
||||
Preserve five distinct measurement purposes: browser application responsiveness, shell interaction, process startup, compilation, and warmed application execution. A single fast loop cannot represent all five. Retain upstream benchmark names, versions, metrics, and correctness rules; an Odin aggregate is a separate decision.
|
||||
|
||||
A practical initial option is **Speedometer 3.1**, a small Bash latency/throughput set, and optional language packs with pinned workload inputs. JetStream, MotionMark, representative compilation projects, and larger JVM applications fit extended investigation. Installed software answers “how usable is this machine as configured”; pinned reference software improves comparison across machines. Store these as distinct modes rather than treating them as interchangeable.
|
||||
|
||||
## 1. Browser candidates and official measurement rules
|
||||
|
||||
BrowserBench’s live index currently points to **Speedometer 3.1**, **JetStream 3.0**, and **MotionMark 1.3.2**. [1]
|
||||
|
||||
| Candidate | What it measures | Conditional role and interpretation |
|
||||
| --- | --- | --- |
|
||||
| Speedometer 3.1 | End-to-end simulated web-app interactions: TodoMVC frameworks, editors, charts, and news applications. | Strong default browser-responsiveness candidate. Preserve the full default suite and iteration protocol for an upstream-style result. It does not measure browser process startup or network loading. |
|
||||
| JetStream 3.0 | JavaScript/WebAssembly workloads, combining startup, average, and worst-case execution behavior. | Useful extended engine-compute view. “Startup” is workload first-iteration behavior, not opening the browser. Its 77 default workloads and upstream geometric aggregation should remain intact. |
|
||||
| MotionMark 1.3.2 | Browser graphics scenes adjusted toward a target frame rate, with confidence intervals. | Useful optional browser-rendering view where a suitable graphics/display path exists. It is not a standalone GPU benchmark or evidence that the intended hardware driver is active. |
|
||||
|
||||
Speedometer’s release/3.1 source currently resolves to commit `1386415be8fef2f6b6bbdbe1828872471c5d802a`. Its `package.json` still says `3.0.0-alpha`, demonstrating why package metadata alone is not a benchmark identity. The release source defaults to **10 iterations** and an **800×600 workload viewport**; URL parameters can change iterations, suites, warmup, and measurement behavior. The result includes individual measurements and JSON exports. Preserve these settings and raw results; do not silently shorten the suite and report an ordinary “Speedometer 3.1” score. [2]
|
||||
|
||||
Official Speedometer instructions recommend a latest stable browser, a separate clean profile, closed background applications/tabs, a focused benchmark page, AC power, no interaction, and cooling between runs where needed. A pinned historical browser is useful for longitudinal comparison but must be identified as a reference-browser experiment. Speedometer aggregates inverse geometric-mean durations, then averages scores across iterations and reports a 95% confidence interval around that mean. Retain that native statistic and uncertainty. Any Odin median across complete benchmark runs is a separate statistic; do not relabel either or average heterogeneous subtest times. [2][3]
|
||||
|
||||
JetStream explicitly says scores are not comparable across its versions. Its startup, worst-case, and average components are deliberately part of the metric; generic harness warmups must not erase those startup observations. MotionMark depends on drawing-area class, viewport, refresh behavior, and browser rendering. Its instructions recommend maximizing the window at the default resolution and 60 Hz for comparison across browsers; a changed display policy is a recorded experimental condition, not something Odin should silently impose. [4][5]
|
||||
|
||||
### Offline hosting and licensing
|
||||
|
||||
Speedometer documents a local static HTTP server, so a prepared source/assets snapshot can be served on loopback without fetching the public site during measurement. Its root license permits source/binary redistribution with the copyright, conditions, and disclaimer retained. That does not replace the license notices of bundled frameworks, fonts, datasets, or other third-party assets. JetStream and MotionMark’s deployed HTML also carries permissive two-condition notices, while their workload provenance points to multiple projects. Before distributing a pack, inventory the exact selected files and retain their individual notices. [2][4][5][6]
|
||||
|
||||
Recommended preparation contract: content-addressed assets, release/commit identity, dependency/license manifest, complete local resource closure, and a future offline-network audit. Bind the server to loopback and prepare everything before timing. A copied public page without its subresources is not a verified offline benchmark. Keep upstream workload code unchanged where feasible; record every adapter patch. The existence of localhost support does not prove the entire pack is self-contained until that artifact is checked.
|
||||
|
||||
### Automation and the Linux boundary
|
||||
|
||||
Two viable routes serve different purposes:
|
||||
|
||||
1. **Native installed browser plus matching WebDriver.** Speedometer itself documents Selenium testing with ChromeDriver, GeckoDriver, and other browser drivers. This route can use distro-supported native browsers on Void/musl or other systems outside Playwright’s matrix, provided the actual browser/driver pair exists and is qualified. Use explicit browser and driver paths and versions. Selenium Manager can download software by default; its documented offline mode disables network requests/downloads. Its documentation still describes limited Linux architecture support, so automatic management cannot be presumed to cover aarch64 or musl. Native automation remains a capability to prove, not a promise. [7]
|
||||
2. **Pinned browser supplied with Playwright on its supported platforms.** Current documentation lists Debian 12/13 and Ubuntu 22.04/24.04/26.04 on x86-64/arm64. Its bundled Firefox/WebKit builds require glibc; musl distributions are unsupported. Stock Firefox cannot simply be substituted because Playwright requires its patched Firefox. Branded Chrome/Edge support and custom executable paths also do not establish universal compatibility. [8]
|
||||
|
||||
**Odin can launch on a system even when its browser module is unavailable.** On musl, use a qualified native browser/driver pair or report an explicit capability skip. A glibc container/chroot or software-rendered browser is a separate environment and cannot transparently stand in for native desktop usability.
|
||||
|
||||
The user has approved **both headed and headless browser modes**, with terminal controls/results and distinct measurement labels. Chrome’s unified headless mode shares browser code with headed Chrome, but Playwright also provides a distinct headless-shell binary and documents behavioral differences. Sharing code does not establish equal scores, display scheduling, GPU access, or compositor behavior. Label a headless cohort explicitly and avoid claiming compliance with focused-window desktop run conditions. A headed cohort needs a graphical session and the official run conditions. Do not merge their reference distributions or invent a conversion factor. The default mode for each run profile remains a selection decision. [8][9]
|
||||
|
||||
Automation should start the pinned page, observe completion, and collect its full result without repeated polling/instrumentation inside timed work. Correctness failures, missing workloads, or timeouts invalidate a full-suite result. Shortened or filtered runs can be Odin browser probes with their own identities; they are not substitutes for official full-suite statistics.
|
||||
|
||||
Record browser binary/version/channel, driver and automation versions, full flags, profile policy, headed/headless/shell distinction, viewport/device scale, display resolution/refresh, X11/Wayland/virtual display, renderer and acceleration evidence when exposed, kernel/libc/architecture, power state, suite digest, and all query parameters. An unavailable GPU detail is unknown, not proof of hardware acceleration. Browser startup, if desired, should be a separate process-to-ready-page test with its cache/profile policy declared.
|
||||
|
||||
## 2. Meaningful Bash and shell measurements
|
||||
|
||||
Bash startup depends on invocation. Login shells read `/etc/profile` and the first readable user login profile; interactive non-login shells read `.bashrc`; noninteractive shells may execute the file named by `BASH_ENV`. `--noprofile` and `--norc` control different paths. Therefore “time bash” is underspecified. [10]
|
||||
|
||||
| Candidate | Measurement definition | Boundary to preserve |
|
||||
| --- | --- | --- |
|
||||
| Clean process startup | Directly launch a known Bash binary with controlled environment/startup-file policy and a trivial command; measure process launch through exit. | This includes loader/process/exit costs. It is not prompt readiness. Explicitly control `BASH_ENV`, inherited functions, locale, and working directory. |
|
||||
| Interactive readiness | Launch a clean interactive shell under a PTY; measure until a declared prompt-ready marker, then measure prompt return after a builtin command. | Account for PTY setup and marker instrumentation. `bash -i -c exit` does not observe a real prompt. Test login and non-login separately if both are offered. |
|
||||
| Builtin workload throughput | Fixed arithmetic, parameter expansion, arrays, and parsing over seeded in-memory inputs, returning a verified checksum. | Bounded output; no per-iteration `date`, `cat`, `grep`, or other subprocesses. Input size and Bash version are part of the workload identity. |
|
||||
| External-command/pipeline workload | Fixed commands, input corpus, and pipe structure, with exact executable versions. | Measures the combined shell, process-launch, tool, and I/O stack. Report it as such. |
|
||||
| User configuration readiness | Opt-in observation using the user’s real startup files and prompt configuration. | These files execute arbitrary user-configured commands, may contact networks or change state, and may not be repeatable. Report local usability separately from the clean reference result. |
|
||||
|
||||
**Harness trap:** Hyperfine defaults to an intermediate `/bin/sh` on Unix and calibrates/subtracts shell-spawn time. For Bash startup, use its direct-command mode or another direct process timer so the Bash process is the measured payload. Hyperfine also supports warmups and repetitions; warmed filesystem caches are a declared condition, not “cold boot.” Do not globally drop caches or alter system tuning just to manufacture a shell metric. [11]
|
||||
|
||||
Treat time-to-first-prompt and prompt-to-next-prompt as distributions, not a single best sample. Bound captured PTY output, support cancellation, and keep configured-shell timeout/error details. Store configuration fingerprints with care; benchmark records should not copy potentially secret startup-file contents.
|
||||
|
||||
## 3. Language workload candidates
|
||||
|
||||
Each result describes a **workload + implementation + toolchain + environment**, not an intrinsic language ranking. Numeric kernels are useful but should not be the only evidence. Libraries, parsing, allocation, and application behavior matter; Python’s C-backed JSON module and a Rust regex library do not represent the same implementation simply because both are called language tests.
|
||||
|
||||
| Language / candidate | Useful workload selection | Evidence, tradeoffs, and role |
|
||||
| --- | --- | --- |
|
||||
| Python: pyperformance + pyperf | `python_startup`, `python_startup_no_site`, `json_loads`/`json_dumps`, one regex workload, and optionally a pure-Python workload such as `richards`. | Maintained Python project favors real applications. Current PyPI metadata: pyperformance **1.14.0**, Python ≥3.10; pyperf **2.10.0**, Python ≥3.9. Suite docs contain older requirements, so prefer release metadata. MIT project license; inspect selected workload notices. Small subsets fit standard profiles. [12] |
|
||||
| Rust: rustc-perf subsets | Runtime groups include parsing, text search, compression, hash maps, and numeric/graphics kernels; compile candidates include pinned real crates such as `syn`. | Official compiler-performance project clearly separates compilation from generated-program execution. Runtime suite is explicitly **experimental**, with changing workloads; pin a commit and validate chosen outputs. MIT infrastructure, separate compile-benchmark licenses. Avoid adopting the full collector by default: it has profiling/environment dependencies. [13] |
|
||||
| C++: LLVM test-suite subset | A small C++ application/proxy workload such as the serial miniFE variant, plus a smaller verified workload; optionally time compiling the same pinned sources. | Suite compiles/runs whole programs, checks reference output, and records compile and execution times separately. CMake/lit and dependencies increase setup cost. Apache-2.0 with LLVM exceptions for project code; third-party directories have separate terms. Good standard/extended candidate after footprint qualification. [14] |
|
||||
| Java: selected workloads under JMH | A packaged, reviewed data-processing/allocation workload with validated results and configured warmup/forks. | JMH **1.37** is the current published Maven release inspected. It is a harness, not a representative suite by itself; samples explain pitfalls rather than defining an Odin score. GPLv2 with the Classpath exception on designated files. Suitable where a small, controlled JVM workload is desired. [15] |
|
||||
| Java: DaCapo application subset | `lusearch` for search, `h2` for database-style execution, or another justified application. | Latest inspected release **23.11-MR2-chopin**; application suite with nontrivial memory use and output validation. Release notes establish Java 11–21 compatibility, not arbitrary newer JDKs. Harness Apache-2.0; component programs retain their licenses. Better extended candidate than a compulsory quick test. [16] |
|
||||
|
||||
pyperformance explicitly says it is not tuned for PyPy. CPython, PyPy, free-threaded builds, optional JIT builds, and distro build choices require distinct metadata and qualification. Interpreter startup with and without `site` measures different initialization; preserve both names if offered. Some pyperformance dependencies have native extensions, so a Python package being source-available does not prove wheel availability on every musl/architecture combination. [12]
|
||||
|
||||
## 4. Fair measurement protocol
|
||||
|
||||
**Startup, compilation, and warmed execution need separate records.** For Rust/C++, compile once before execution measurements unless compilation is the workload. For Java, bytecode compilation with `javac` differs from runtime JIT work. For Python, process startup/imports differ from repeated execution in an initialized interpreter. If the workload intentionally includes setup, say so and apply the same rule on every system.
|
||||
|
||||
pyperf uses calibration, multiple processes, warmups, values, and metadata. Its fast mode explicitly trades accuracy for speed. JMH provides forks, warmup/measurement iterations, state setup, and result consumption to avoid dead-code elimination; its maintainers warn that a harness does not eliminate benchmarking mistakes. Retain cold/startup samples and warmed samples separately. Use predetermined warmup/measurement rules and report instability; do not keep warming until a favorable result appears. [15][17]
|
||||
|
||||
Proposed compilation policy for selection:
|
||||
|
||||
- Pin sources, datasets, dependency lockfiles, language standard/edition, compiler identity, target triple, optimization flags, linker, libc, and library versions.
|
||||
- Use a documented release optimization configuration. Cargo’s release defaults include `opt-level=3`, incremental off, and 16 codegen units; its bench profile inherits release. Record overrides and dependency profiles. C++ should use an explicitly chosen optimization level and standard; `-Ofast` changes standards-compliance assumptions and cannot silently replace a strict floating-point contract. [18]
|
||||
- Keep baseline-ISA builds and host-tuned builds distinct. `target-cpu=native`/`-march=native`, SIMD dispatch, LTO, PGO, allocator choice, and thread counts can materially change results. A single binary is not portable across x86_64 and aarch64; equal flags do not mean equal generated instructions.
|
||||
- Compilation tests distinguish clean builds, no-op rebuilds, and controlled incremental edits. Pre-fetch dependencies and hold build parallelism/cache policy fixed. Network resolution is not compilation speed; compiler caches must be controlled or explicitly measured.
|
||||
- Validate outputs/checksums and numerical tolerances. Prevent constant folding/dead work, but avoid timing validation if the workload definition excludes it. Do not compare implementations with different precision, algorithms, data sizes, or hidden thread counts under a shared metric name.
|
||||
|
||||
An **installed-toolchain mode** best reveals current developer experience. A **reference-toolchain mode** uses pinned, prepared artifacts for stronger cross-machine comparison, with acquisition size, libc/ISA support, licenses, and security updates owned explicitly. It must not install into or replace the user’s default toolchain during a run. Report compiler/runtime version changes as changes in execution conditions, not unexplained hardware improvements.
|
||||
|
||||
## 5. Profile options, budgets, and unavailable tools
|
||||
|
||||
The following are provisional **application-domain budgets**, not measured runtime promises or the whole Odin run duration:
|
||||
|
||||
| Profile | Candidate scope | Budget policy to evaluate |
|
||||
| --- | --- | --- |
|
||||
| Quick | Browser capability probe; clean Bash startup and one builtin workload; startup/small verified work on selected installed languages. | Approximately 30–60 seconds for this domain. No full browser score if the official suite does not fit. Short observations carry lower-confidence status. |
|
||||
| Standard | Full Speedometer 3.1 on one qualified browser; Bash set; selected language pack(s) with normal warmup/repetitions. | Reserve minutes rather than seconds; a pilot target is 2–5 minutes for browser and 1–3 minutes per selected language pack. Slow hosts may exceed it. |
|
||||
| Extended | JetStream, optional MotionMark, more process forks, compilation workloads, selected DaCapo applications, additional browsers/toolchains. | User-visible per-module time/memory/disk ceilings; an initial 15–30-minute application allocation needs calibration on low-end physical hardware. |
|
||||
|
||||
Preparation/download/compilation costs are shown separately unless compilation is the named workload. Check tools and resource requirements before beginning. Do not shorten an upstream test behind the user’s back to meet a deadline: mark it incomplete and preserve diagnostics.
|
||||
|
||||
Missing interpreter/compiler/browser/driver, unsupported ABI, unavailable display, insufficient resources, timeout, and failed output validation are different outcomes. None is a zero performance measurement. Skips reduce **capability coverage**; the score-design ticket must determine eligibility for an aggregate. Optional-tool absence should leave completed results usable. No hardware/browser/type of VM was executed here to substantiate duration estimates.
|
||||
|
||||
## Remaining decisions and evidence gaps
|
||||
|
||||
Both browser modes are already approved. Select the default mode for each profile; the native-browser/driver and reference-browser support matrix; exact upstream full-suite versus Odin-probe identities; language subsets; installed/reference toolchain policy; fixed compilation/warmup rules; and measured profile budgets. Browser automation overhead, all-assets-offline closure, per-asset redistribution inventory, selected-suite behavior on musl/aarch64 and newer JDKs, and correctness/variance qualification remain future proof work. An application score must not conceal these unresolved cohort boundaries.
|
||||
|
||||
## Sources and method
|
||||
|
||||
Primary sources inspected **2026-09-25**. Context7 resolution preceded queries for Speedometer/WebKit, Bash, Hyperfine, Playwright, pyperformance/pyperf, Rust, LLVM, JMH, GCC, and Selenium. Exact Speedometer and pyperformance lookups returned unrelated entries; those were rejected and their own repositories inspected. Queries stayed within the three-command limit per lookup; no quota error occurred.
|
||||
|
||||
1. [BrowserBench current index](https://browserbench.org/).
|
||||
2. Speedometer release source: [commit](https://github.com/WebKit/Speedometer/commit/1386415be8fef2f6b6bbdbe1828872471c5d802a), [parameters](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/params.mjs), [result client](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/main.mjs), [package metadata](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/package.json), [local hosting](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Development.md).
|
||||
3. [Speedometer 3.1 instructions](https://browserbench.org/Speedometer3.1/instructions.html), [workload descriptions](https://browserbench.org/Speedometer3.1/about.html), [measurement/scoring explanation](https://github.com/WebKit/Speedometer/blob/main/README.md).
|
||||
4. [JetStream 3.0 implementation page](https://browserbench.org/JetStream3.0/), [in-depth methodology and provenance](https://browserbench.org/JetStream3.0/in-depth.html).
|
||||
5. [MotionMark 1.3.2 implementation page](https://browserbench.org/MotionMark1.3.2/), [methodology, display requirements, and version history](https://browserbench.org/MotionMark1.3.2/about.html).
|
||||
6. [Speedometer release license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE).
|
||||
7. [Speedometer browser-driver testing documentation](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Testing.md), [Selenium Manager: explicit drivers, offline mode, and architecture limitations](https://www.selenium.dev/documentation/selenium_manager/).
|
||||
8. [Playwright system requirements](https://playwright.dev/docs/intro), [browser distributions and headless modes](https://github.com/microsoft/playwright/blob/main/docs/src/browsers.md), [musl limitations](https://github.com/microsoft/playwright/blob/main/docs/src/docker.md).
|
||||
9. [Chrome unified headless and headless-shell distinction](https://developer.chrome.com/docs/chromium/headless).
|
||||
10. [GNU Bash startup files](https://www.gnu.org/software/bash/manual/html_node/Bash-Startup-Files.html), [invocation options](https://www.gnu.org/software/bash/manual/html_node/Invoking-Bash.html).
|
||||
11. [Hyperfine README: intermediate-shell correction, direct mode, warmups, repetitions](https://github.com/sharkdp/hyperfine/blob/master/README.md).
|
||||
12. [pyperformance purpose and license](https://github.com/python/pyperformance/blob/main/README.rst), [PyPI release metadata](https://pypi.org/pypi/pyperformance/json), [pyperf release metadata](https://pypi.org/pypi/pyperf/json), [benchmark manifest](https://github.com/python/pyperformance/blob/main/pyperformance/data-files/benchmarks/MANIFEST), [workload descriptions](https://github.com/python/pyperformance/blob/main/doc/benchmarks.rst), [dependency/runtime guidance](https://github.com/python/pyperformance/blob/main/doc/usage.rst), [startup implementation](https://github.com/python/pyperformance/blob/main/pyperformance/data-files/benchmarks/bm_python_startup/run_benchmark.py).
|
||||
13. [rustc-perf purpose/licenses](https://github.com/rust-lang/rustc-perf/blob/master/README.md), [compile suite](https://github.com/rust-lang/rustc-perf/blob/master/collector/compile-benchmarks/README.md), [experimental runtime suite](https://github.com/rust-lang/rustc-perf/blob/master/collector/runtime-benchmarks/README.md), [runtime groups](https://github.com/rust-lang/rustc-perf/tree/master/collector/runtime-benchmarks), [collector requirements](https://github.com/rust-lang/rustc-perf/blob/master/collector/README.md).
|
||||
14. [LLVM test-suite guide](https://llvm.org/docs/TestSuiteGuide.html), [testing/output-validation model](https://llvm.org/docs/TestingGuide.html), [serial miniFE description](https://github.com/llvm/llvm-test-suite/blob/main/MultiSource/Benchmarks/DOE-ProxyApps-C++/miniFE/README), [license and third-party exceptions](https://github.com/llvm/llvm-test-suite/blob/main/LICENSE.TXT).
|
||||
15. [JMH purpose and cautions](https://github.com/openjdk/jmh/blob/master/README.md), [published version metadata](https://repo.maven.apache.org/maven2/org/openjdk/jmh/jmh-core/maven-metadata.xml), [parameter/warmup/fork example](https://github.com/openjdk/jmh/blob/master/jmh-samples/src/main/java/org/openjdk/jmh/samples/JMHSample_27_Params.java), [result consumption and profilers](https://github.com/openjdk/jmh/blob/master/jmh-samples/src/main/java/org/openjdk/jmh/samples/JMHSample_35_Profilers.java), [license](https://github.com/openjdk/jmh/blob/master/LICENSE).
|
||||
16. [DaCapo latest release](https://github.com/dacapobench/dacapobench/releases/tag/v23.11-MR2-chopin), [release compatibility and workload notes](https://github.com/dacapobench/dacapobench/blob/0db32562cf169730c163d88df2eeb28217ca7d03/benchmarks/RELEASE_NOTES.md), [purpose/reporting/license guidance](https://github.com/dacapobench/dacapobench/blob/master/README.md).
|
||||
17. [pyperf runner configuration](https://github.com/psf/pyperf/blob/main/doc/runner.rst), [process/calibration architecture](https://github.com/psf/pyperf/blob/main/doc/run_benchmark.rst).
|
||||
18. [Cargo profile defaults](https://doc.rust-lang.org/cargo/reference/profiles.html), [rustc code-generation options](https://doc.rust-lang.org/rustc/codegen-options/index.html), [GCC optimization semantics](https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html).
|
||||
@@ -0,0 +1,139 @@
|
||||
# Storage measurements and trustworthy health advice
|
||||
|
||||
Research for [Establish storage measurements and trustworthy health advice](https://git.bongbetic.com/xavierk/odin/issues/4), part of Odin's Wayfinder map. Access date for every source: **2026-09-25**.
|
||||
|
||||
This report establishes evidence and candidate policies. It does not select Odin's final workloads, thresholds, privileged execution design, or score. No benchmark, device query, self-test, installation, or hardware change was performed during this investigation.
|
||||
|
||||
## Findings that shape the decision
|
||||
|
||||
The strongest candidate is **fio for file-based performance measurements, smartmontools for cross-protocol health findings, and an optional nvme-cli adapter for additional NVMe evidence**. These tools cover different responsibilities. A fast benchmark cannot establish drive health; a passing SMART status cannot establish future reliability. Health findings should therefore remain visible independently of the performance score and capability coverage.
|
||||
|
||||
There is a material compatibility change already: nvme-cli **v3.1**, released September 18, 2026, documents `nvme log smart`; `nvme smart-log` is a deprecated compatibility alias. Its default output format version is now 2, with version 1 available. fio **3.43** was released September 23. A bleeding-edge development environment is compatible with reproducible measurements only if Odin records tool versions and keeps workload and parser versions explicit. Package-manager availability alone does not establish supported commands or JSON schemas. [S1][S3]
|
||||
|
||||
## Candidate tools and measurement scope
|
||||
|
||||
| Candidate | Useful responsibility | Limits and recommendation to consider |
|
||||
| --- | --- | --- |
|
||||
| fio | Sequential/random reads and writes, block sizes, queue depths, latency distributions, bounded I/O, optional verification | Best primary workload candidate. Select a small fixed workload vocabulary; do not accept arbitrary user-supplied job files into privileged execution. |
|
||||
| smartctl | ATA, SCSI and NVMe identity, health, existing error/self-test logs; JSON; many bridge/controller adapters | Best baseline health reader. Decode protocol-specific semantics and command status separately. Some transports are unsafe for automatic probing. |
|
||||
| nvme-cli | NVMe-specific identity, SMART and detailed logs | Useful optional supplement. Version 2/3 command and JSON differences need explicit compatibility handling. Avoid duplicating the same controller's health as several independent findings. |
|
||||
| Native `/proc` and `/sys` | I/O pressure, completed I/O, queue activity and available sensors | Low-dependency contextual evidence, not a workload or a substitute for SMART. |
|
||||
| GNU `dd` | Bounded sequential copying, optionally direct I/O and final synchronization | Possible explicitly labelled basic fallback. Its copy-oriented output does not supply fio's workload control or latency distributions; its result must not silently substitute into the same scored workload. |
|
||||
|
||||
Sources: fio HOWTO, smartctl manual, nvme-cli released documentation, kernel PSI/I/O documentation and GNU manual. [S1–S4][S8][S9][S14]
|
||||
|
||||
A compact candidate performance set is:
|
||||
|
||||
| Measurement | Candidate workload, still to be selected | What its result means |
|
||||
| --- | --- | --- |
|
||||
| Sequential read/write | Large blocks, for example 1 MiB, one job, depth 1 | Large-file throughput through the selected filesystem and storage path |
|
||||
| Random read/write | 4 KiB, one job, depth 1 | Small-request responsiveness; report latency and IOPS |
|
||||
| Queued random read | Same block size at a documented higher depth, such as 16 or 32 | Concurrency capability; a different workload from depth 1 |
|
||||
| Durable small writes | A separate small, bounded workload with defined sync frequency | Application-visible cost of requesting persistence |
|
||||
| Optional integrity check | Write and verify Odin-owned file blocks with fio checksums | Whether the tested data path returned those bytes correctly; not a full-surface drive or whole-RAM certification |
|
||||
|
||||
Record read and write throughput in explicit units, IOPS, completed bytes, operation count, errors, elapsed time, and p50/p95/p99 latency where sample counts support them. Distinguish fio completion latency from total latency: total includes submission latency. Record achieved queue-depth distribution; requesting depth greater than one does not make a synchronous engine asynchronous. fio `psync` is a useful depth-1 compatibility candidate; `io_uring` and `libaio` are queued-engine candidates where supported. Engine changes must be visible in results and comparability rules. [S1]
|
||||
|
||||
Measure one storage workload at a time during reference runs. Concurrent CPU or memory stress can instead be an explicitly identified contention experiment. Record filesystem, mount options, device topology, encryption/RAID/virtualization, kernel, selected engine, power/thermal state, background I/O, and pre/post free space. The measurement describes this path under these conditions, not the NVMe/HDD in isolation.
|
||||
|
||||
## Safe operation and comparability constraints
|
||||
|
||||
The following are proposed invariants, rather than finalized profile numbers:
|
||||
|
||||
1. **Own every writable byte.** Create a private run directory on a deliberately selected filesystem and exclusively create its regular files. Validate ownership, type and target identity; reject symlink redirection and raw block/character devices. `O_CREAT|O_EXCL` supplies exclusive creation semantics. Do not use arbitrary existing user files as write targets. Avoid selecting `/tmp` automatically: tmpfs stores files in virtual memory and may use swap. [S7][S11]
|
||||
2. **Budget storage space and cumulative writes separately.** fio `size` defines the working region, while `io_size` can independently bound I/O. `runtime` stops at the earlier of completion or time limit; `time_based` loops the workload. Thus a small file plus a timed loop can write many times its size. Prefer explicit byte and time bounds without `time_based` for ordinary write profiles. Count fixture preparation, repetitions and verification-related writes in a per-run host-write budget. Reserve free space, account for quotas and metadata, recheck during execution, and stop on ENOSPC or I/O errors. Space and byte thresholds remain product decisions. [S1]
|
||||
3. **Do not promise a physical NAND-write limit.** A workload's host bytes are measurable; filesystem/controller write amplification and unrelated host activity are additional. NVMe Data Units Written measures host data in units of 1,000 × 512 bytes, rounded up, excluding metadata; it is not a universal NAND-wear counter. Background activity also prevents attributing its entire delta to Odin. [S5]
|
||||
4. **Make cache and durability modes explicit.** fio `direct=1` normally requests `O_DIRECT`; support and alignment vary by filesystem and kernel, and misaligned requests can fail or fall back to buffered I/O. Direct I/O does not by itself provide `O_SYNC` persistence guarantees, bypass every device cache, or prove sustained media speed. `invalidate` is conditional on platform/file support. Avoid global `drop_caches`: kernel documentation warns of additional I/O and CPU costs. A buffered fallback must be labelled and excluded from direct-I/O comparisons. [S1][S7][S12]
|
||||
5. **Include preparation and flush costs honestly.** Read tests over newly created fixtures still require writes. A user choosing no writes can reuse an identified valid fixture or skip that workload; Odin should not create one silently. Do not measure unwritten sparse-file holes as disk reads. For writes, document `end_fsync` or other synchronization and report end-to-end time including the final flush separately from unsynchronized throughput. Control data compressibility/deduplication using a declared fio buffer policy; generating fresh data adds CPU cost. Preserve normal filesystem settings rather than silently disabling compression or copy-on-write. [S1][S7]
|
||||
6. **Treat cancellation and cleanup as part of the run.** Bound the entire job group; stop launching work on cancellation, retain partial status, reap workers, and remove only proven Odin-owned artifacts. A worker stuck in kernel I/O may not stop immediately. Crash recovery needs a manifest and ownership checks before deletion. Cleanup failure is a reported outcome, never a reason to recursively delete a user-selected directory.
|
||||
7. **Collect health before load and reduce work when evidence is serious.** A candidate policy is to skip storage stress when critical media/reliability findings or current unreadable data are already present. Pause on documented thermal alarms or loss of safety headroom. Display estimated host writes before a write run. Avoid automatic discard/TRIM, formatting, SMART feature changes, firmware updates, cache-policy changes or repair operations as benchmark preparation.
|
||||
|
||||
Short bounded tests cannot establish steady-state SSD performance after exhaustion of a large write cache, or scan every HDD sector. Making test data larger than all caches can conflict with a quick run and a conservative write budget. Report the actual duration and working set; do not extrapolate a short burst into an endurance or sustained-performance guarantee. The tradeoff between low impact and sustained measurements needs an explicit run-profile decision.
|
||||
|
||||
## Health evidence and field interpretation
|
||||
|
||||
Prefer structured output with the original tool version, schema identifier, command outcome, timestamp, device identity and transport. A missing field is unknown, not zero. Preserve large counters losslessly: smartctl JSON can emit string/byte-array companions for integers exceeding JavaScript's safe integer range; `--json=v` requests them consistently. This matters for an Ink/JavaScript consumer. [S2]
|
||||
|
||||
For smartctl NVMe output, the primary object is `nvme_smart_health_information_log`. The inspected source confirms the following keys and conversions. Raw NVMe temperature is Kelvin; smartctl's `temperature` here is already Celsius. Do not convert it twice. [S6]
|
||||
|
||||
| Evidence | Meaning | Candidate interpretation |
|
||||
| --- | --- | --- |
|
||||
| `critical_warning` bit 0; `available_spare` vs `available_spare_threshold` | Spare capacity below the controller's threshold | Urgent preservation/service finding; display the device-provided threshold |
|
||||
| Bit 1; `temperature`, warning/critical temperature time | Above an over-temperature or below an under-temperature threshold | Stop heat-producing tests; investigate cooling/environment. This alone is not proof that replacement is needed |
|
||||
| Bit 2 | NVM subsystem reliability degraded | Urgent backup and replacement/service assessment |
|
||||
| Bit 3 | Media placed read-only for a device reliability condition | Urgent preservation and replacement/service assessment; distinct from user namespace write protection |
|
||||
| Bits 4/5 | Volatile-memory backup failure; persistent-memory region read-only/unreliable | Urgent loss-of-protection/service finding when applicable; explain the specific subsystem |
|
||||
| `percentage_used` | Vendor estimate of endurance consumed | 100 means estimated endurance consumed, **not guaranteed failure**; values may exceed 100. Plan replacement according to manufacturer guidance and workload, without inventing days remaining |
|
||||
| `media_errors` | Unrecovered data-integrity errors, including ECC/CRC/tag errors | Investigate any nonzero history; escalating recent deltas plus failed I/O are much stronger urgent evidence than an isolated old count |
|
||||
| `num_err_log_entries` | Lifetime number of error-information entries | Inspect status/cause and recency. It is not interchangeable with media errors |
|
||||
| `unsafe_shutdowns` | Loss of power without shutdown notification | Investigate shutdown/power history and correlate with errors; not proof of failed media |
|
||||
| `data_units_written`, power-on hours, thermal counters | Usage/history with specified units and reporting limits | Useful trends and context; no universal lifespan formula |
|
||||
|
||||
NVMe warning bits are current state, not persistent event history; zero today does not erase yesterday's finding. Some temperature fields are optional, and zero can mean unsupported. Per-namespace SMART is optional; the global namespace identifier can describe a controller's aggregate. Preserve scope rather than assigning identical controller totals to every namespace. [S3][S5][S6]
|
||||
|
||||
**ATA needs a separate mapping.** Keep attribute ID, raw representation, normalized current/worst value, threshold, type and failure state. smartctl states that these meanings are vendor-specific; SSD meanings can differ and displayed names can be wrong for models absent from its drive database. The label `Pre-fail` by itself does not mean a drive is failing: the current normalized value must cross its threshold. [S2]
|
||||
|
||||
Common drive-database candidates include reallocated sectors (5), pending sectors (197), offline uncorrectable sectors (198), and interface CRC errors (199). Interpret them only with a matching model/firmware/database rule; do not apply a universal raw-count threshold or turn interface errors directly into a disk-replacement recommendation. Preserve lifetime history and recent deltas separately. SMART RETURN STATUS, failed applicable thresholds, existing self-test failures, and observed host I/O errors are stronger when they agree. SCSI health uses its own exception/sense reporting rather than ATA attribute assumptions. [S2][S15]
|
||||
|
||||
smartctl exit status is a bitmask. Bits 0–2 can describe invocation/access/command problems; bits 3–7 describe failing status, thresholds and historical error/self-test evidence. A nonzero exit must not discard usable JSON, and access failure must not become “bad drive.” Reading an existing self-test log is different from starting a test. The manual notes that running self-tests can degrade performance and normal I/O can extend their duration; any future self-test workflow needs a separate user decision and scheduling. [S2]
|
||||
|
||||
## Candidate advice rubric
|
||||
|
||||
This rubric is a proposed interpretation layer over the documented evidence, not a manufacturer's warranty or an adopted Odin policy.
|
||||
|
||||
| Finding class | Evidence sufficient to consider it | Appropriate wording/action |
|
||||
| --- | --- | --- |
|
||||
| **Replace/service now** | Credible ATA failing status/current applicable prefailure threshold; NVMe degraded reliability/read-only media; serious repeated data-integrity failures attributable to the device | “Preserve accessible data now; avoid further stress; arrange replacement or service.” Cite exact flags, device scope and timestamps. Hardware attribution may still need confirmation |
|
||||
| **Investigate urgently** | New media errors, pending/uncorrectable sectors, recent failed self-tests, resets/timeouts, thermal alarms, loss of power-loss protection | Identify the failing path; correlate controller, connection, power and filesystem evidence. Do not automatically blame the medium |
|
||||
| **Monitor / plan replacement** | Stable historical findings or vendor-estimated endurance consumed without current failure evidence | Retain trends, explain wear status and manufacturer limits, and plan according to importance/workload. No invented remaining-life percentage |
|
||||
| **No concerning evidence observed** | Successful supported collection with no relevant current finding | State what was checked and when; keep normal backup advice independent of a performance score |
|
||||
| **Unknown / limited coverage** | Missing permission/tool/field, sleeping drive, unsupported bridge/controller, virtual device, ambiguous identity | Explain the missing capability and a bounded next step. Never convert unavailable evidence into a healthy badge |
|
||||
|
||||
The smartctl manual recommends preserving data promptly when the drive reports failing health. Conversely, Google's primary HDD population study found that SMART-only models were unlikely to predict individual failures reliably. That older HDD result is not a calibrated modern-SSD failure model, but it reinforces the distinction between a useful warning and a guarantee of future health. NVMe's own endurance-field semantics explicitly reject equating 100% usage with failure. [S2][S5][S16]
|
||||
|
||||
## Compatibility, privilege and general health
|
||||
|
||||
USB, SAT and RAID support must follow known transport rules. The smartctl manual documents bridge-specific NVMe adapters and per-physical-disk MegaRAID addressing; a RAID logical volume is not automatically one physical drive. Particularly important: its **JMB39x/JMS56x transport uses READ/WRITE commands to a RAID-volume sector**. It warns that the wrong device can be overwritten and interruption can prevent restoration. Exclude these from routine automated probing; “try every device type” is not a safe compatibility strategy. Even standby-aware queries may wake a disk during autodetection, so record unsupported power-state handling. [S2]
|
||||
|
||||
VMs need an explicit virtual-device classification. QEMU's NVMe implementation constructs SMART data from its emulated controller and block-accounting state. A guest can therefore show valid-looking SMART without revealing the host drive's health. Guest tests establish guest-path performance; actual physical passthrough and device identity require separate verification. VM coverage cannot establish USB, physical RAID, real wear counters or thermal behavior. [S17]
|
||||
|
||||
Keep ordinary file workloads unprivileged. Device queries may require additional device permissions or kernel capabilities; NVMe's Linux passthrough code explicitly gates classes of commands. A future privileged mechanism should allow only validated read operations and selected devices, with no arbitrary shell or passthrough-command forwarding. Permission denial is an expected capability outcome, not an instruction to run the whole TUI as root. [S18]
|
||||
|
||||
For overall system health, useful complementary evidence is:
|
||||
|
||||
- `/proc/pressure/io`: `some` measures time with some stalled tasks, `full` time with all non-idle tasks stalled; use same-window deltas alongside workload latency. This detects pressure, not its sole cause. [S8]
|
||||
- `/proc/diskstats` or per-device sysfs statistics: completed I/O, time and queue context. Counters have concurrency/accounting caveats; busy percentage alone does not establish NVMe saturation. [S9]
|
||||
- Available hwmon readings, limits and alarm flags: retain sensor identity and units; chip-specific alarms and missing sensors preclude a universal hard-coded temperature cutoff. Standard hwmon ABI readings are intended to be readable by unprivileged applications. [S10]
|
||||
- Kernel errors and existing EDAC/RAS evidence: distinguish corrected errors from uncorrected/fatal errors and report available history. EDAC documentation explicitly says corrected errors may, but need not, predict later uncorrected errors. Missing reporting hardware/driver is unknown. Kernel log access can require `CAP_SYSLOG` when `dmesg_restrict=1`; do not assume systemd/journald on Void or other distributions. [S13][S19]
|
||||
|
||||
Kernel or mount optimizations should be suggestions tied to an observed limitation and a documented tradeoff, recorded for subsequent comparable runs. This research supports observing current settings and thermal/power/error evidence; it supplies no evidence for blanket scheduler, write-cache, governor, or filesystem changes.
|
||||
|
||||
## Remaining decisions and evidence gaps
|
||||
|
||||
The next human decisions are the ordinary run's write authorization/budget, minimum free-space reserve, workload lengths and repetitions, required versus optional queued/sync/verification tests, how reduced-capability results affect score eligibility, the supported transport list, the privilege interaction, and the exact advice wording. A sustained-media profile would need a separate impact budget.
|
||||
|
||||
Implementation work will need parser fixtures from supported smartctl/nvme-cli versions; success, partial and denied-permission results; real ATA/NVMe/USB/RAID samples; healthy and failing vendor examples; and proof of cancellation, space reservation, direct-I/O handling and cleanup across filesystems. No such hardware validation occurred here. No calibrated cross-device replacement thresholds or modern SSD remaining-life model were found or claimed.
|
||||
|
||||
Context7 library resolution succeeded for fio, smartmontools, nvme-cli, Linux kernel, GNU Coreutils and QEMU. Both allowed nvme-cli documentation fetches returned “Could not fetch documentation snippets”; its official released documents and source were inspected instead. The NVM Express specifications landing page returned HTTP 403, so NVMe field semantics here are grounded in maintained libnvme definitions and smartmontools implementation rather than a directly retrieved current specification PDF. Those are material evidence limits, not silently filled gaps.
|
||||
|
||||
## Sources inspected
|
||||
|
||||
- **S1:** [fio 3.43 HOWTO](https://github.com/axboe/fio/blob/fio-3.43/HOWTO.rst), relevant workload, size/runtime, buffering, engines, percentile, verification and error sections; [release](https://github.com/axboe/fio/releases/tag/fio-3.43).
|
||||
- **S2:** [smartctl manual source](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/smartctl.8.in), health, attributes, JSON, exit status, device transports, standby and self-tests.
|
||||
- **S3:** nvme-cli v3.1 [SMART log command](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/nvme-log-smart.txt), [global options](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/global-options.txt), [legacy alias](https://github.com/linux-nvme/nvme-cli/blob/master/Documentation/nvme-smart-log.txt), and [release](https://github.com/linux-nvme/nvme-cli/releases/tag/v3.1).
|
||||
- **S4:** [smartmontools NVMe support examples](https://www.smartmontools.org/wiki/NVMe_Support), inspected through Context7.
|
||||
- **S5:** [libnvme types](https://github.com/linux-nvme/libnvme/blob/master/src/nvme/types.h), `nvme_smart_log` and `nvme_smart_crit` documentation.
|
||||
- **S6:** [smartmontools NVMe JSON implementation](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/nvmeprint.cpp).
|
||||
- **S7:** [Linux man-pages open(2)](https://man7.org/linux/man-pages/man2/open.2.html), exclusive creation and direct/synchronized I/O.
|
||||
- **S8:** [Linux PSI documentation](https://www.kernel.org/doc/html/latest/accounting/psi.html).
|
||||
- **S9:** [Linux I/O statistics documentation](https://www.kernel.org/doc/html/latest/admin-guide/iostats.html).
|
||||
- **S10:** [Linux hwmon sysfs interface](https://www.kernel.org/doc/html/latest/hwmon/sysfs-interface.html).
|
||||
- **S11:** [Linux tmpfs documentation](https://docs.kernel.org/filesystems/tmpfs.html).
|
||||
- **S12:** [Linux VM sysctl documentation](https://docs.kernel.org/admin-guide/sysctl/vm.html), `drop_caches`.
|
||||
- **S13:** [Linux RAS documentation source](https://www.kernel.org/doc/html/latest/_sources/admin-guide/RAS/main.rst.txt), error categories and EDAC.
|
||||
- **S14:** [GNU Coreutils dd manual](https://www.gnu.org/software/coreutils/manual/html_node/dd-invocation.html).
|
||||
- **S15:** [smartmontools drive database](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/drivedb.h), default and model-dependent attribute mappings.
|
||||
- **S16:** [Google, Failure Trends in a Large Disk Drive Population](https://research.google/pubs/failure-trends-in-a-large-disk-drive-population/), primary publication abstract, 2007.
|
||||
- **S17:** [QEMU NVMe implementation](https://github.com/qemu/qemu/blob/master/hw/nvme/ctrl.c), `nvme_smart_info`; [NVMe device documentation](https://github.com/qemu/qemu/blob/master/docs/system/devices/nvme.rst), inspected through Context7.
|
||||
- **S18:** [Linux NVMe ioctl authorization](https://github.com/torvalds/linux/blob/master/drivers/nvme/host/ioctl.c).
|
||||
- **S19:** [Linux kernel sysctl documentation](https://www.kernel.org/doc/html/latest/admin-guide/sysctl/kernel.html), `dmesg_restrict`.
|
||||
Reference in New Issue
Block a user