Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7ff9d74e8f |
@@ -0,0 +1,115 @@
|
||||
# GPU driver evidence and performance workloads
|
||||
|
||||
Research for [Establish GPU driver evidence and performance workloads](https://git.bongbetic.com/xavierk/odin/issues/3). All sources were inspected on **2026-09-25**. These are conditional recommendations for the workload decision, not an adopted suite or support guarantee. No drivers were installed, settings changed, device queries executed, or benchmarks run.
|
||||
|
||||
## Decision summary
|
||||
|
||||
Odin can establish **which device and API worked for a specified operation under the current session**. It cannot certify one universally “correct” driver from a package name, loaded module, advertised API version or benchmark score. Keep discovery, successful execution, output validation, presentation and performance as distinct evidence.
|
||||
|
||||
The strongest initial options are a small **headless Vulkan compute profile** using a qualified clpeak build, plus **API-specific rendering workloads** where supported. vkmark and glmark2 offer useful scenes, but their backend requirements and relatively infrequent releases need qualification. No candidate alone measures compute throughput, rendering, compositor behavior, video acceleration and browser usability.
|
||||
|
||||
Visible GPU windows remain a decision for **Choose Odin’s workload suite and run profiles**. The user's approval of headed/headless browser modes does not settle GPU presentation. Terminal orchestration can support headless workloads without opening a window; it cannot thereby prove the desktop presentation path works.
|
||||
|
||||
## 1. Evidence to collect per device
|
||||
|
||||
| Stage | Evidence | Justified conclusion |
|
||||
| --- | --- | --- |
|
||||
| Hardware and kernel path | DRM/sysfs device, associated PCI or platform identity, bound kernel driver, accessible render node | Device and kernel path are present; userspace API operation is still unproven |
|
||||
| API discovery | Vulkan physical-device properties/features/queues; OpenGL vendor/renderer/version from the actual context; OpenCL platform/device type | This implementation advertises these capabilities in this environment |
|
||||
| Execution | Bounded allocation, command submission, completion and error outcome on the selected device | The tested operation completed; report device loss, allocation failure and timeout distinctly |
|
||||
| Correctness | A known-output compute or rendering check with a defined tolerance | The sampled operation returned an expected result; throughput alone does not supply this proof |
|
||||
| Presentation | Successful creation and presentation to a particular X11/Wayland surface, if this mode is selected | That surface/session path worked; headless success is a different finding |
|
||||
|
||||
Linux render nodes permit non-global rendering without DRM-master authentication, subject to ordinary filesystem permissions. They do not grant modesetting rights. Treat denied access as capability coverage, not “bad GPU,” and do not respond by elevating the entire TUI or changing device permissions. DRM discovery must include platform devices: ARM GPUs need not be PCI devices. [1]
|
||||
|
||||
Vulkan properties include device type, vendor/device identifiers, device UUID, and driver identification. `deviceType=CPU` identifies a typically host-processor implementation; the specification calls device type informational, so combine it with driver/renderer evidence. `driverVersion` is vendor-specified, not universal semantic versioning. `conformanceVersion` describes the implementer's prior conformance testing, not a test of this installation. Where supported, `VK_EXT_physical_device_drm` connects API devices to DRM node major/minor numbers. [2]
|
||||
|
||||
**vulkaninfo** is a useful discovery candidate. Its `--summary` covers enumerated devices; `--json=<index>` writes a Vulkan Profiles JSON file for one device. Plain `--json` defaults to the first device, so one successful invocation does not inventory every GPU. Use a private output directory and record the tool/schema version. The inspected SDK tag is `vulkan-sdk-1.4.357.0`, with Apache-2.0 project licensing. [3]
|
||||
|
||||
Mesa LLVMpipe/Softpipe are software renderers. Zink is an OpenGL implementation over Vulkan and can use a hardware Vulkan driver: the word “Mesa” or “Zink” is not evidence of software rendering. Mesa and the Vulkan loader also expose selection/override variables, including `LIBGL_ALWAYS_SOFTWARE`, `DRI_PRIME`, `MESA_VK_DEVICE_SELECT` and `VK_DRIVER_FILES`. Record relevant effective overrides and selected devices; do not silently change the user's stack during baseline measurement. Vulkan/OpenCL CPU devices and mock drivers must not contribute a hardware-GPU score. [4][5][6]
|
||||
|
||||
## 2. Driver and architecture scope
|
||||
|
||||
| Hardware family | Paths worth supporting conditionally | Boundary |
|
||||
| --- | --- | --- |
|
||||
| Intel | Appropriate Linux kernel driver plus Mesa OpenGL/ANV; an independently available compute runtime | Working OpenGL does not establish Vulkan or OpenCL support. Discover generation-specific capabilities rather than prescribe one package universally |
|
||||
| AMD | Supported kernel/userspace combination, commonly amdgpu plus Mesa RADV for Vulkan | RADV and ROCm serve different purposes. ROCm has its own hardware/OS/firmware compatibility matrix; absence of ROCm does not mean ordinary graphics is broken |
|
||||
| NVIDIA | NVIDIA's supported userspace/kernel stack, or Mesa NVK and applicable OpenGL path | NVK is a legitimate Vulkan implementation. NVIDIA's open kernel modules still require matching NVIDIA userspace and GSP firmware; “open module installed” does not establish compatibility |
|
||||
| ARM SoCs | Panfrost/PanVK for supported Mali, Freedreno/Turnip for supported Adreno, other model-specific Mesa/vendor paths | aarch64 names the CPU architecture, not the GPU API capability. Some devices support GLES without Vulkan; experimental support must not be force-enabled automatically |
|
||||
|
||||
Mesa documents RADV's separation from the kernel driver and hardware limitations; Panfrost lists distinct API support by GPU and explicitly warns about experimental PanVK enablement. NVIDIA's inspected `615.71.09` open-module release supports x86_64/aarch64 and Turing-or-later hardware, with corresponding-release userspace/firmware requirements. These examples justify capability probing, not a universal driver recommendation. [7][8][9]
|
||||
|
||||
Qualify **x86_64/glibc, x86_64/musl, aarch64/glibc and aarch64/musl** independently for the chosen executable and transitive libraries. Source availability does not certify a binary across that matrix. Void explicitly states proprietary NVIDIA drivers do not support musl; packaging Odin differently cannot erase that driver limitation. Mesa-based paths can be candidates where that GPU and distribution support them. Current ROCm support is also a specific matrix, not a promise for every Linux distribution or libc. [8][10]
|
||||
|
||||
## 3. Workload candidates and concrete tradeoffs
|
||||
|
||||
| Candidate and inspected version | Measurements, footprint and control | Conditional role |
|
||||
| --- | --- | --- |
|
||||
| **clpeak 2.1.4**, Apache-2.0; released August 27, 2026 | Current code supports Vulkan, OpenCL, CUDA, ROCm/HIP, oneAPI and CPU, among others. CLI has backend/device/test selection and JSON/CSV/XML output. `--max-time` controls each GPU test's timed phase; warmup/calibration add time. C++17/CMake; SDKs/backends are optional but auto-detected by default | Strong first candidate for a deliberately restricted CLI build and selected Vulkan FP32/bandwidth workloads. Pin enabled backends, shaders and compiler; avoid its “run every backend/device/test” default |
|
||||
| **vkpeak 20260527**, MIT; source activity in August 2026 | Vulkan peak scalar/vector/matrix arithmetic and transfer tests using ncnn. Select device and scenarios. Small top-level program, substantial transitive shader/runtime dependency. No user time-budget option is documented in the inspected CLI | Alternative focused compute candidate. Its README explicitly says peak metrics do not represent real-world use. Source returns zero for some unsupported features **and failures**, so zero cannot be interpreted as measured zero performance |
|
||||
| **vkmark 2025.01**, LGPL-2.1-or-later | Configurable Vulkan rendering scenes, dimensions, present mode, duration and device UUID selection. C++17, Vulkan, GLM and Assimp; optional XCB/Wayland/DRM/GBM dependencies | Candidate graphics profile after backend qualification. The released source includes a headless plugin requiring `VK_EXT_headless_surface`; its manpage backend list omits that plugin. Generic Vulkan support alone is insufficient |
|
||||
| **glmark2 2023.01**, GPLv3 | OpenGL 2.0/GLES2 scenes; per-scene duration, off-screen mode, frame-end/swap controls, output validation and CSV/XML results. Build flavors include X11, Wayland, DRM and GBM; GL/EGL/GLES and image libraries/assets | Useful compatibility and rendering candidate; its older API workloads are not a complete modern-GPU assessment. GBM source can use a selected render node; `--off-screen` on an X11 build does not imply display-server independence |
|
||||
|
||||
Primary released READMEs, manuals, licenses and implementation sources support this comparison. glmark2's latest inspected tag remains 2023.01 with main activity in September 2025; vkmark's latest tag is 2025.01 with main activity in September 2025. These are maturity/maintenance observations, not evidence of current hardware certification. clpeak and vkpeak show more recent source/release activity, but still require qualification. [11–14]
|
||||
|
||||
Two implementation traps matter immediately:
|
||||
|
||||
- clpeak's Vulkan instance requests Vulkan 1.0 or 1.1 depending on compiled optional features. Its reported capability floor therefore depends on the build. Its timing code performs warmup and calibration before the timed batch; `--max-time` is not an end-to-end timeout. Its Vulkan backend distinguishes CPU and integrated/discrete GPU device types. [11]
|
||||
- vkpeak adapts work and reports peak results, with memory sizing based partly on device heap information. A selected subset is more controllable than its complete default suite, but a wrapper still needs independent resource and runtime bounds. Neither tool's advertised throughput proves it checks the numerical result required by Odin's correctness stage. [12]
|
||||
|
||||
Licenses above describe inspected project code. Bundling requires a separate manifest for assets, embedded dependencies, modifications and any vendor runtime redistribution terms; these source inspections are not a completed distribution-license audit. Installing a large CUDA/ROCm SDK solely to enable baseline benchmarking would weaken the universal deployment objective. Vendor-specific compute paths are better considered optional capability profiles.
|
||||
|
||||
## 4. Headless, desktop, multiple GPUs and virtualization
|
||||
|
||||
Keep three execution classes distinct: **surface-free compute**, **offscreen rendering**, and **desktop presentation**. Vulkan does not require every physical device or queue to support presentation. Support must be queried for the actual surface. FIFO presentation waits on vertical blanking; an FPS result can therefore reflect display/compositor policy rather than maximum render throughput. Fix and record present mode, resolution and backend. [15]
|
||||
|
||||
vkmark's headless plugin still uses a Vulkan surface/swapchain extension; glmark2's GBM backend opens a render node and creates a GBM surface. These are different requirements and workloads. A KMS/direct-display backend may need display ownership and disturb the session, so it is not an automatic fallback when X11/Wayland fails. An SSH terminal can have usable GPU compute without a display socket; classify presentation as unavailable in that session rather than infer a missing graphics driver. [1][13][14]
|
||||
|
||||
Enumerate all devices, map them to stable identifiers where available, and let the run profile select the display GPU, another named GPU or separate per-GPU runs. Do not treat index zero as “best GPU.” Mesa's selection variables can reorder enumeration; vkmark's UUID selector and NVIDIA's documented UUID/PCI-ID selection illustrate stronger identity mechanisms. Avoid summing overlapping APIs or independently averaging all installed GPUs into one unexplained number. [2][5][9][13]
|
||||
|
||||
Virtual hardware needs its own label. Mesa Venus serializes Vulkan through virtio-gpu to a host renderer and can operate over hardware **or Lavapipe**; guest enumeration does not establish physical passthrough. A VM result measures that guest path, while its host telemetry may be hidden. Passthrough must be established from the recorded environment and device evidence. Containers similarly need device access and compatible userspace libraries; missing exposure is not proof the host has no GPU. [6][16]
|
||||
|
||||
## 5. Reproducibility, safety and diagnosis
|
||||
|
||||
Before scoring, freeze workload version, scene/kernel code, input size, precision/vector width, backend, device, output format, build flags and compiler. Record kernel, userspace driver identity, power source, thermal state, display mode and concurrent load. Warmup/cache policy must be explicit. Compare repeated runs under the same policy; do not compare shader compilation included in one result with warmed execution in another.
|
||||
|
||||
Compute FLOPS, transfer bandwidth and scene FPS answer different questions. A CUDA FP16 matrix peak cannot replace a Vulkan FP32 score; software rendering cannot replace the hardware result; unavailable features are not zeros. Preserve upstream metrics and chosen aggregation rules. glmark2/vkmark aggregate FPS does not automatically supply frame-time percentiles. Small rendering scenes can also be CPU/driver limited, so unexpectedly low throughput is evidence to investigate, not automatic proof of defective hardware.
|
||||
|
||||
Recommended safety boundaries are one GPU workload at a time, bounded memory demand with system/VRAM reserve, an independent wall-clock watchdog, staged process-group cancellation, and refusal of unbounded “run forever” modes. Include initialization/calibration in the overall budget. Stop on device loss, repeated API errors or meaningful driver-reported critical thermal findings; retain partial results. Killing a client cannot guarantee immediate recovery from a kernel/driver hang. Do not automatically overclock, alter fan/power limits, reset a GPU or replace drivers. Resource limits and safe cancellation remain acceptance tests for the selected workload.
|
||||
|
||||
Telemetry is supporting evidence, with provider-specific meaning:
|
||||
|
||||
- NVIDIA `nvidia-smi` documents unsupported values as `N/A`, separate errors for permission denial, unloaded driver and missing NVML, and stable UUID/PCI selection. Its utility success does not test Vulkan/OpenGL presentation. Read-only query adapters must preserve unavailable/error outcomes. [9]
|
||||
- amdgpu exposes temperature, load, power and other sysfs metrics, but support varies. Its APU power reading includes CPU power, so it is not interchangeable with discrete-GPU-only power. [17]
|
||||
- Linux DRM fdinfo defines per-client engine-busy counters, capacities and accounting rules; availability depends on the driver and accessible process descriptors. These can help attribute work without assuming one vendor's utilization meaning applies everywhere. [18]
|
||||
|
||||
Advice should name the evidence: “Vulkan userspace driver could not load,” “render-node access denied,” “software renderer selected,” “this feature is unsupported,” or “workload lost the device.” A package/version mismatch needs concrete loader or vendor evidence; a successful fallback may be intentional. Present a distro-appropriate investigation step with confidence and tradeoffs, rather than an unconditional “install proprietary drivers.”
|
||||
|
||||
## 6. Decisions and limits carried forward
|
||||
|
||||
The workload decision must choose: headless compute and graphics requirements; whether visible GPU presentation is allowed; selected tool/build and minimum API features; treatment of software/virtual/unsupported paths; default multi-GPU selection; memory/runtime budgets; and the correctness check and scoring eligibility rules.
|
||||
|
||||
Before a supported release, qualify the four architecture/libc lanes, real Intel/AMD/NVIDIA hardware, representative ARM SoCs, X11/Wayland/headless sessions, multiple GPUs, denied permissions, software rendering, VM acceleration/passthrough, and cancellation/device-loss fixtures. No such execution evidence was produced here. No package-size estimates or full vendor conformance/redistribution audit were established.
|
||||
|
||||
Context7 resolution preceded documentation lookup. Its glmark2 queries yielded no relevant main-project documentation; clpeak was unindexed; vkpeak resolved to an unrelated speech tool and was rejected; NVIDIA NVML searches produced unrelated/wrapper results. Official tagged source and vendor documentation supplied those gaps. The guessed Mesa Lavapipe page returned 404; software/virtual-path claims use inspected Mesa driver documentation and API/device evidence instead. Current Mesa pages contain evolving and occasionally differing generation summaries, so no exhaustive model support table is inferred from them.
|
||||
|
||||
## Sources
|
||||
|
||||
1. [Linux DRM userspace API, render nodes](https://docs.kernel.org/gpu/drm-uapi.html).
|
||||
2. [Vulkan device/queue specification](https://github.com/KhronosGroup/Vulkan-Docs/blob/main/chapters/devsandqueues.adoc), physical-device, driver, UUID and DRM properties.
|
||||
3. Vulkan Tools SDK tag: [vulkaninfo documentation](https://github.com/KhronosGroup/Vulkan-Tools/blob/vulkan-sdk-1.4.357.0/vulkaninfo/vulkaninfo.md), [license](https://github.com/KhronosGroup/Vulkan-Tools/blob/vulkan-sdk-1.4.357.0/LICENSE.txt).
|
||||
4. Mesa [platforms/drivers](https://docs.mesa3d.org/systems.html), [LLVMpipe](https://docs.mesa3d.org/drivers/llvmpipe.html), [Zink](https://docs.mesa3d.org/drivers/zink.html).
|
||||
5. [Mesa environment variables](https://docs.mesa3d.org/envvars.html), [Vulkan loader driver discovery](https://github.com/KhronosGroup/Vulkan-Loader/blob/main/docs/LoaderDriverInterface.md).
|
||||
6. [OpenCL device enumeration](https://github.com/KhronosGroup/OpenCL-Registry/blob/main/specs/unified/refpages/man/html/clGetDeviceIDs.html), [Vulkan Guide support/null-driver discussion](https://github.com/KhronosGroup/Vulkan-Guide/blob/main/chapters/checking_for_support.adoc), inspected through Context7.
|
||||
7. Mesa [ANV](https://docs.mesa3d.org/drivers/anv.html), [RADV](https://docs.mesa3d.org/drivers/radv.html), [NVK](https://docs.mesa3d.org/drivers/nvk.html), [Panfrost](https://docs.mesa3d.org/drivers/panfrost.html), [Freedreno/Turnip](https://docs.mesa3d.org/drivers/freedreno.html).
|
||||
8. [ROCm current compatibility matrix](https://rocm.docs.amd.com/en/latest/compatibility/compatibility-matrix.html).
|
||||
9. NVIDIA [open-module 615.71.09 README](https://github.com/NVIDIA/open-gpu-kernel-modules/blob/615.71.09/README.md), [nvidia-smi documentation](https://docs.nvidia.com/deploy/nvidia-smi/index.html).
|
||||
10. [Void musl compatibility](https://docs.voidlinux.org/installation/musl.html).
|
||||
11. clpeak 2.1.4: [README](https://github.com/krrishnarraj/clpeak/blob/2.1.4/README.md), [CLI options](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/common/options.cpp), [Vulkan timing/instance implementation](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/vulkan/vk_peak.cpp), [device mapping](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/vulkan/vulkan_device.cpp), [build options](https://github.com/krrishnarraj/clpeak/blob/2.1.4/CMakeLists.txt), [license](https://github.com/krrishnarraj/clpeak/blob/2.1.4/LICENSE).
|
||||
12. vkpeak 20260527: [README](https://github.com/nihui/vkpeak/blob/20260527/README.md), [implementation](https://github.com/nihui/vkpeak/blob/20260527/vkpeak.cpp), [build/dependency configuration](https://github.com/nihui/vkpeak/blob/20260527/CMakeLists.txt), [license](https://github.com/nihui/vkpeak/blob/20260527/LICENSE).
|
||||
13. vkmark 2025.01: [README](https://github.com/vkmark/vkmark/blob/2025.01/README.md), [manual](https://github.com/vkmark/vkmark/blob/2025.01/doc/vkmark.1), [headless implementation and license notice](https://github.com/vkmark/vkmark/blob/2025.01/src/ws/headless_native_system.cpp), [backend build](https://github.com/vkmark/vkmark/blob/2025.01/src/meson.build).
|
||||
14. glmark2 2023.01: [README/license declaration](https://github.com/glmark2/glmark2/blob/2023.01/README), [manual](https://github.com/glmark2/glmark2/blob/2023.01/doc/glmark2.1.in), [GBM implementation](https://github.com/glmark2/glmark2/blob/2023.01/src/native-state-gbm.cpp), [build flavors](https://github.com/glmark2/glmark2/blob/2023.01/meson_options.txt).
|
||||
15. [Vulkan WSI specification](https://github.com/KhronosGroup/Vulkan-Docs/blob/main/chapters/VK_KHR_surface/wsi.adoc), surface support, headless surfaces and present modes.
|
||||
16. [Mesa Virtio-GPU Venus](https://docs.mesa3d.org/drivers/venus.html).
|
||||
17. [Linux amdgpu thermal/power monitoring](https://docs.kernel.org/gpu/amdgpu/thermal.html).
|
||||
18. [Linux DRM usage-statistics ABI](https://docs.kernel.org/gpu/drm-usage-stats.html).
|
||||
@@ -1,139 +0,0 @@
|
||||
# Storage measurements and trustworthy health advice
|
||||
|
||||
Research for [Establish storage measurements and trustworthy health advice](https://git.bongbetic.com/xavierk/odin/issues/4), part of Odin's Wayfinder map. Access date for every source: **2026-09-25**.
|
||||
|
||||
This report establishes evidence and candidate policies. It does not select Odin's final workloads, thresholds, privileged execution design, or score. No benchmark, device query, self-test, installation, or hardware change was performed during this investigation.
|
||||
|
||||
## Findings that shape the decision
|
||||
|
||||
The strongest candidate is **fio for file-based performance measurements, smartmontools for cross-protocol health findings, and an optional nvme-cli adapter for additional NVMe evidence**. These tools cover different responsibilities. A fast benchmark cannot establish drive health; a passing SMART status cannot establish future reliability. Health findings should therefore remain visible independently of the performance score and capability coverage.
|
||||
|
||||
There is a material compatibility change already: nvme-cli **v3.1**, released September 18, 2026, documents `nvme log smart`; `nvme smart-log` is a deprecated compatibility alias. Its default output format version is now 2, with version 1 available. fio **3.43** was released September 23. A bleeding-edge development environment is compatible with reproducible measurements only if Odin records tool versions and keeps workload and parser versions explicit. Package-manager availability alone does not establish supported commands or JSON schemas. [S1][S3]
|
||||
|
||||
## Candidate tools and measurement scope
|
||||
|
||||
| Candidate | Useful responsibility | Limits and recommendation to consider |
|
||||
| --- | --- | --- |
|
||||
| fio | Sequential/random reads and writes, block sizes, queue depths, latency distributions, bounded I/O, optional verification | Best primary workload candidate. Select a small fixed workload vocabulary; do not accept arbitrary user-supplied job files into privileged execution. |
|
||||
| smartctl | ATA, SCSI and NVMe identity, health, existing error/self-test logs; JSON; many bridge/controller adapters | Best baseline health reader. Decode protocol-specific semantics and command status separately. Some transports are unsafe for automatic probing. |
|
||||
| nvme-cli | NVMe-specific identity, SMART and detailed logs | Useful optional supplement. Version 2/3 command and JSON differences need explicit compatibility handling. Avoid duplicating the same controller's health as several independent findings. |
|
||||
| Native `/proc` and `/sys` | I/O pressure, completed I/O, queue activity and available sensors | Low-dependency contextual evidence, not a workload or a substitute for SMART. |
|
||||
| GNU `dd` | Bounded sequential copying, optionally direct I/O and final synchronization | Possible explicitly labelled basic fallback. Its copy-oriented output does not supply fio's workload control or latency distributions; its result must not silently substitute into the same scored workload. |
|
||||
|
||||
Sources: fio HOWTO, smartctl manual, nvme-cli released documentation, kernel PSI/I/O documentation and GNU manual. [S1–S4][S8][S9][S14]
|
||||
|
||||
A compact candidate performance set is:
|
||||
|
||||
| Measurement | Candidate workload, still to be selected | What its result means |
|
||||
| --- | --- | --- |
|
||||
| Sequential read/write | Large blocks, for example 1 MiB, one job, depth 1 | Large-file throughput through the selected filesystem and storage path |
|
||||
| Random read/write | 4 KiB, one job, depth 1 | Small-request responsiveness; report latency and IOPS |
|
||||
| Queued random read | Same block size at a documented higher depth, such as 16 or 32 | Concurrency capability; a different workload from depth 1 |
|
||||
| Durable small writes | A separate small, bounded workload with defined sync frequency | Application-visible cost of requesting persistence |
|
||||
| Optional integrity check | Write and verify Odin-owned file blocks with fio checksums | Whether the tested data path returned those bytes correctly; not a full-surface drive or whole-RAM certification |
|
||||
|
||||
Record read and write throughput in explicit units, IOPS, completed bytes, operation count, errors, elapsed time, and p50/p95/p99 latency where sample counts support them. Distinguish fio completion latency from total latency: total includes submission latency. Record achieved queue-depth distribution; requesting depth greater than one does not make a synchronous engine asynchronous. fio `psync` is a useful depth-1 compatibility candidate; `io_uring` and `libaio` are queued-engine candidates where supported. Engine changes must be visible in results and comparability rules. [S1]
|
||||
|
||||
Measure one storage workload at a time during reference runs. Concurrent CPU or memory stress can instead be an explicitly identified contention experiment. Record filesystem, mount options, device topology, encryption/RAID/virtualization, kernel, selected engine, power/thermal state, background I/O, and pre/post free space. The measurement describes this path under these conditions, not the NVMe/HDD in isolation.
|
||||
|
||||
## Safe operation and comparability constraints
|
||||
|
||||
The following are proposed invariants, rather than finalized profile numbers:
|
||||
|
||||
1. **Own every writable byte.** Create a private run directory on a deliberately selected filesystem and exclusively create its regular files. Validate ownership, type and target identity; reject symlink redirection and raw block/character devices. `O_CREAT|O_EXCL` supplies exclusive creation semantics. Do not use arbitrary existing user files as write targets. Avoid selecting `/tmp` automatically: tmpfs stores files in virtual memory and may use swap. [S7][S11]
|
||||
2. **Budget storage space and cumulative writes separately.** fio `size` defines the working region, while `io_size` can independently bound I/O. `runtime` stops at the earlier of completion or time limit; `time_based` loops the workload. Thus a small file plus a timed loop can write many times its size. Prefer explicit byte and time bounds without `time_based` for ordinary write profiles. Count fixture preparation, repetitions and verification-related writes in a per-run host-write budget. Reserve free space, account for quotas and metadata, recheck during execution, and stop on ENOSPC or I/O errors. Space and byte thresholds remain product decisions. [S1]
|
||||
3. **Do not promise a physical NAND-write limit.** A workload's host bytes are measurable; filesystem/controller write amplification and unrelated host activity are additional. NVMe Data Units Written measures host data in units of 1,000 × 512 bytes, rounded up, excluding metadata; it is not a universal NAND-wear counter. Background activity also prevents attributing its entire delta to Odin. [S5]
|
||||
4. **Make cache and durability modes explicit.** fio `direct=1` normally requests `O_DIRECT`; support and alignment vary by filesystem and kernel, and misaligned requests can fail or fall back to buffered I/O. Direct I/O does not by itself provide `O_SYNC` persistence guarantees, bypass every device cache, or prove sustained media speed. `invalidate` is conditional on platform/file support. Avoid global `drop_caches`: kernel documentation warns of additional I/O and CPU costs. A buffered fallback must be labelled and excluded from direct-I/O comparisons. [S1][S7][S12]
|
||||
5. **Include preparation and flush costs honestly.** Read tests over newly created fixtures still require writes. A user choosing no writes can reuse an identified valid fixture or skip that workload; Odin should not create one silently. Do not measure unwritten sparse-file holes as disk reads. For writes, document `end_fsync` or other synchronization and report end-to-end time including the final flush separately from unsynchronized throughput. Control data compressibility/deduplication using a declared fio buffer policy; generating fresh data adds CPU cost. Preserve normal filesystem settings rather than silently disabling compression or copy-on-write. [S1][S7]
|
||||
6. **Treat cancellation and cleanup as part of the run.** Bound the entire job group; stop launching work on cancellation, retain partial status, reap workers, and remove only proven Odin-owned artifacts. A worker stuck in kernel I/O may not stop immediately. Crash recovery needs a manifest and ownership checks before deletion. Cleanup failure is a reported outcome, never a reason to recursively delete a user-selected directory.
|
||||
7. **Collect health before load and reduce work when evidence is serious.** A candidate policy is to skip storage stress when critical media/reliability findings or current unreadable data are already present. Pause on documented thermal alarms or loss of safety headroom. Display estimated host writes before a write run. Avoid automatic discard/TRIM, formatting, SMART feature changes, firmware updates, cache-policy changes or repair operations as benchmark preparation.
|
||||
|
||||
Short bounded tests cannot establish steady-state SSD performance after exhaustion of a large write cache, or scan every HDD sector. Making test data larger than all caches can conflict with a quick run and a conservative write budget. Report the actual duration and working set; do not extrapolate a short burst into an endurance or sustained-performance guarantee. The tradeoff between low impact and sustained measurements needs an explicit run-profile decision.
|
||||
|
||||
## Health evidence and field interpretation
|
||||
|
||||
Prefer structured output with the original tool version, schema identifier, command outcome, timestamp, device identity and transport. A missing field is unknown, not zero. Preserve large counters losslessly: smartctl JSON can emit string/byte-array companions for integers exceeding JavaScript's safe integer range; `--json=v` requests them consistently. This matters for an Ink/JavaScript consumer. [S2]
|
||||
|
||||
For smartctl NVMe output, the primary object is `nvme_smart_health_information_log`. The inspected source confirms the following keys and conversions. Raw NVMe temperature is Kelvin; smartctl's `temperature` here is already Celsius. Do not convert it twice. [S6]
|
||||
|
||||
| Evidence | Meaning | Candidate interpretation |
|
||||
| --- | --- | --- |
|
||||
| `critical_warning` bit 0; `available_spare` vs `available_spare_threshold` | Spare capacity below the controller's threshold | Urgent preservation/service finding; display the device-provided threshold |
|
||||
| Bit 1; `temperature`, warning/critical temperature time | Above an over-temperature or below an under-temperature threshold | Stop heat-producing tests; investigate cooling/environment. This alone is not proof that replacement is needed |
|
||||
| Bit 2 | NVM subsystem reliability degraded | Urgent backup and replacement/service assessment |
|
||||
| Bit 3 | Media placed read-only for a device reliability condition | Urgent preservation and replacement/service assessment; distinct from user namespace write protection |
|
||||
| Bits 4/5 | Volatile-memory backup failure; persistent-memory region read-only/unreliable | Urgent loss-of-protection/service finding when applicable; explain the specific subsystem |
|
||||
| `percentage_used` | Vendor estimate of endurance consumed | 100 means estimated endurance consumed, **not guaranteed failure**; values may exceed 100. Plan replacement according to manufacturer guidance and workload, without inventing days remaining |
|
||||
| `media_errors` | Unrecovered data-integrity errors, including ECC/CRC/tag errors | Investigate any nonzero history; escalating recent deltas plus failed I/O are much stronger urgent evidence than an isolated old count |
|
||||
| `num_err_log_entries` | Lifetime number of error-information entries | Inspect status/cause and recency. It is not interchangeable with media errors |
|
||||
| `unsafe_shutdowns` | Loss of power without shutdown notification | Investigate shutdown/power history and correlate with errors; not proof of failed media |
|
||||
| `data_units_written`, power-on hours, thermal counters | Usage/history with specified units and reporting limits | Useful trends and context; no universal lifespan formula |
|
||||
|
||||
NVMe warning bits are current state, not persistent event history; zero today does not erase yesterday's finding. Some temperature fields are optional, and zero can mean unsupported. Per-namespace SMART is optional; the global namespace identifier can describe a controller's aggregate. Preserve scope rather than assigning identical controller totals to every namespace. [S3][S5][S6]
|
||||
|
||||
**ATA needs a separate mapping.** Keep attribute ID, raw representation, normalized current/worst value, threshold, type and failure state. smartctl states that these meanings are vendor-specific; SSD meanings can differ and displayed names can be wrong for models absent from its drive database. The label `Pre-fail` by itself does not mean a drive is failing: the current normalized value must cross its threshold. [S2]
|
||||
|
||||
Common drive-database candidates include reallocated sectors (5), pending sectors (197), offline uncorrectable sectors (198), and interface CRC errors (199). Interpret them only with a matching model/firmware/database rule; do not apply a universal raw-count threshold or turn interface errors directly into a disk-replacement recommendation. Preserve lifetime history and recent deltas separately. SMART RETURN STATUS, failed applicable thresholds, existing self-test failures, and observed host I/O errors are stronger when they agree. SCSI health uses its own exception/sense reporting rather than ATA attribute assumptions. [S2][S15]
|
||||
|
||||
smartctl exit status is a bitmask. Bits 0–2 can describe invocation/access/command problems; bits 3–7 describe failing status, thresholds and historical error/self-test evidence. A nonzero exit must not discard usable JSON, and access failure must not become “bad drive.” Reading an existing self-test log is different from starting a test. The manual notes that running self-tests can degrade performance and normal I/O can extend their duration; any future self-test workflow needs a separate user decision and scheduling. [S2]
|
||||
|
||||
## Candidate advice rubric
|
||||
|
||||
This rubric is a proposed interpretation layer over the documented evidence, not a manufacturer's warranty or an adopted Odin policy.
|
||||
|
||||
| Finding class | Evidence sufficient to consider it | Appropriate wording/action |
|
||||
| --- | --- | --- |
|
||||
| **Replace/service now** | Credible ATA failing status/current applicable prefailure threshold; NVMe degraded reliability/read-only media; serious repeated data-integrity failures attributable to the device | “Preserve accessible data now; avoid further stress; arrange replacement or service.” Cite exact flags, device scope and timestamps. Hardware attribution may still need confirmation |
|
||||
| **Investigate urgently** | New media errors, pending/uncorrectable sectors, recent failed self-tests, resets/timeouts, thermal alarms, loss of power-loss protection | Identify the failing path; correlate controller, connection, power and filesystem evidence. Do not automatically blame the medium |
|
||||
| **Monitor / plan replacement** | Stable historical findings or vendor-estimated endurance consumed without current failure evidence | Retain trends, explain wear status and manufacturer limits, and plan according to importance/workload. No invented remaining-life percentage |
|
||||
| **No concerning evidence observed** | Successful supported collection with no relevant current finding | State what was checked and when; keep normal backup advice independent of a performance score |
|
||||
| **Unknown / limited coverage** | Missing permission/tool/field, sleeping drive, unsupported bridge/controller, virtual device, ambiguous identity | Explain the missing capability and a bounded next step. Never convert unavailable evidence into a healthy badge |
|
||||
|
||||
The smartctl manual recommends preserving data promptly when the drive reports failing health. Conversely, Google's primary HDD population study found that SMART-only models were unlikely to predict individual failures reliably. That older HDD result is not a calibrated modern-SSD failure model, but it reinforces the distinction between a useful warning and a guarantee of future health. NVMe's own endurance-field semantics explicitly reject equating 100% usage with failure. [S2][S5][S16]
|
||||
|
||||
## Compatibility, privilege and general health
|
||||
|
||||
USB, SAT and RAID support must follow known transport rules. The smartctl manual documents bridge-specific NVMe adapters and per-physical-disk MegaRAID addressing; a RAID logical volume is not automatically one physical drive. Particularly important: its **JMB39x/JMS56x transport uses READ/WRITE commands to a RAID-volume sector**. It warns that the wrong device can be overwritten and interruption can prevent restoration. Exclude these from routine automated probing; “try every device type” is not a safe compatibility strategy. Even standby-aware queries may wake a disk during autodetection, so record unsupported power-state handling. [S2]
|
||||
|
||||
VMs need an explicit virtual-device classification. QEMU's NVMe implementation constructs SMART data from its emulated controller and block-accounting state. A guest can therefore show valid-looking SMART without revealing the host drive's health. Guest tests establish guest-path performance; actual physical passthrough and device identity require separate verification. VM coverage cannot establish USB, physical RAID, real wear counters or thermal behavior. [S17]
|
||||
|
||||
Keep ordinary file workloads unprivileged. Device queries may require additional device permissions or kernel capabilities; NVMe's Linux passthrough code explicitly gates classes of commands. A future privileged mechanism should allow only validated read operations and selected devices, with no arbitrary shell or passthrough-command forwarding. Permission denial is an expected capability outcome, not an instruction to run the whole TUI as root. [S18]
|
||||
|
||||
For overall system health, useful complementary evidence is:
|
||||
|
||||
- `/proc/pressure/io`: `some` measures time with some stalled tasks, `full` time with all non-idle tasks stalled; use same-window deltas alongside workload latency. This detects pressure, not its sole cause. [S8]
|
||||
- `/proc/diskstats` or per-device sysfs statistics: completed I/O, time and queue context. Counters have concurrency/accounting caveats; busy percentage alone does not establish NVMe saturation. [S9]
|
||||
- Available hwmon readings, limits and alarm flags: retain sensor identity and units; chip-specific alarms and missing sensors preclude a universal hard-coded temperature cutoff. Standard hwmon ABI readings are intended to be readable by unprivileged applications. [S10]
|
||||
- Kernel errors and existing EDAC/RAS evidence: distinguish corrected errors from uncorrected/fatal errors and report available history. EDAC documentation explicitly says corrected errors may, but need not, predict later uncorrected errors. Missing reporting hardware/driver is unknown. Kernel log access can require `CAP_SYSLOG` when `dmesg_restrict=1`; do not assume systemd/journald on Void or other distributions. [S13][S19]
|
||||
|
||||
Kernel or mount optimizations should be suggestions tied to an observed limitation and a documented tradeoff, recorded for subsequent comparable runs. This research supports observing current settings and thermal/power/error evidence; it supplies no evidence for blanket scheduler, write-cache, governor, or filesystem changes.
|
||||
|
||||
## Remaining decisions and evidence gaps
|
||||
|
||||
The next human decisions are the ordinary run's write authorization/budget, minimum free-space reserve, workload lengths and repetitions, required versus optional queued/sync/verification tests, how reduced-capability results affect score eligibility, the supported transport list, the privilege interaction, and the exact advice wording. A sustained-media profile would need a separate impact budget.
|
||||
|
||||
Implementation work will need parser fixtures from supported smartctl/nvme-cli versions; success, partial and denied-permission results; real ATA/NVMe/USB/RAID samples; healthy and failing vendor examples; and proof of cancellation, space reservation, direct-I/O handling and cleanup across filesystems. No such hardware validation occurred here. No calibrated cross-device replacement thresholds or modern SSD remaining-life model were found or claimed.
|
||||
|
||||
Context7 library resolution succeeded for fio, smartmontools, nvme-cli, Linux kernel, GNU Coreutils and QEMU. Both allowed nvme-cli documentation fetches returned “Could not fetch documentation snippets”; its official released documents and source were inspected instead. The NVM Express specifications landing page returned HTTP 403, so NVMe field semantics here are grounded in maintained libnvme definitions and smartmontools implementation rather than a directly retrieved current specification PDF. Those are material evidence limits, not silently filled gaps.
|
||||
|
||||
## Sources inspected
|
||||
|
||||
- **S1:** [fio 3.43 HOWTO](https://github.com/axboe/fio/blob/fio-3.43/HOWTO.rst), relevant workload, size/runtime, buffering, engines, percentile, verification and error sections; [release](https://github.com/axboe/fio/releases/tag/fio-3.43).
|
||||
- **S2:** [smartctl manual source](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/smartctl.8.in), health, attributes, JSON, exit status, device transports, standby and self-tests.
|
||||
- **S3:** nvme-cli v3.1 [SMART log command](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/nvme-log-smart.txt), [global options](https://github.com/linux-nvme/nvme-cli/blob/v3.1/Documentation/global-options.txt), [legacy alias](https://github.com/linux-nvme/nvme-cli/blob/master/Documentation/nvme-smart-log.txt), and [release](https://github.com/linux-nvme/nvme-cli/releases/tag/v3.1).
|
||||
- **S4:** [smartmontools NVMe support examples](https://www.smartmontools.org/wiki/NVMe_Support), inspected through Context7.
|
||||
- **S5:** [libnvme types](https://github.com/linux-nvme/libnvme/blob/master/src/nvme/types.h), `nvme_smart_log` and `nvme_smart_crit` documentation.
|
||||
- **S6:** [smartmontools NVMe JSON implementation](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/nvmeprint.cpp).
|
||||
- **S7:** [Linux man-pages open(2)](https://man7.org/linux/man-pages/man2/open.2.html), exclusive creation and direct/synchronized I/O.
|
||||
- **S8:** [Linux PSI documentation](https://www.kernel.org/doc/html/latest/accounting/psi.html).
|
||||
- **S9:** [Linux I/O statistics documentation](https://www.kernel.org/doc/html/latest/admin-guide/iostats.html).
|
||||
- **S10:** [Linux hwmon sysfs interface](https://www.kernel.org/doc/html/latest/hwmon/sysfs-interface.html).
|
||||
- **S11:** [Linux tmpfs documentation](https://docs.kernel.org/filesystems/tmpfs.html).
|
||||
- **S12:** [Linux VM sysctl documentation](https://docs.kernel.org/admin-guide/sysctl/vm.html), `drop_caches`.
|
||||
- **S13:** [Linux RAS documentation source](https://www.kernel.org/doc/html/latest/_sources/admin-guide/RAS/main.rst.txt), error categories and EDAC.
|
||||
- **S14:** [GNU Coreutils dd manual](https://www.gnu.org/software/coreutils/manual/html_node/dd-invocation.html).
|
||||
- **S15:** [smartmontools drive database](https://github.com/smartmontools/smartmontools/blob/master/smartmontools/drivedb.h), default and model-dependent attribute mappings.
|
||||
- **S16:** [Google, Failure Trends in a Large Disk Drive Population](https://research.google/pubs/failure-trends-in-a-large-disk-drive-population/), primary publication abstract, 2007.
|
||||
- **S17:** [QEMU NVMe implementation](https://github.com/qemu/qemu/blob/master/hw/nvme/ctrl.c), `nvme_smart_info`; [NVMe device documentation](https://github.com/qemu/qemu/blob/master/docs/system/devices/nvme.rst), inspected through Context7.
|
||||
- **S18:** [Linux NVMe ioctl authorization](https://github.com/torvalds/linux/blob/master/drivers/nvme/host/ioctl.c).
|
||||
- **S19:** [Linux kernel sysctl documentation](https://www.kernel.org/doc/html/latest/admin-guide/sysctl/kernel.html), `dmesg_restrict`.
|
||||
Reference in New Issue
Block a user