Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f384c81f0f |
@@ -1,115 +0,0 @@
|
|||||||
# GPU driver evidence and performance workloads
|
|
||||||
|
|
||||||
Research for [Establish GPU driver evidence and performance workloads](https://git.bongbetic.com/xavierk/odin/issues/3). All sources were inspected on **2026-09-25**. These are conditional recommendations for the workload decision, not an adopted suite or support guarantee. No drivers were installed, settings changed, device queries executed, or benchmarks run.
|
|
||||||
|
|
||||||
## Decision summary
|
|
||||||
|
|
||||||
Odin can establish **which device and API worked for a specified operation under the current session**. It cannot certify one universally “correct” driver from a package name, loaded module, advertised API version or benchmark score. Keep discovery, successful execution, output validation, presentation and performance as distinct evidence.
|
|
||||||
|
|
||||||
The strongest initial options are a small **headless Vulkan compute profile** using a qualified clpeak build, plus **API-specific rendering workloads** where supported. vkmark and glmark2 offer useful scenes, but their backend requirements and relatively infrequent releases need qualification. No candidate alone measures compute throughput, rendering, compositor behavior, video acceleration and browser usability.
|
|
||||||
|
|
||||||
Visible GPU windows remain a decision for **Choose Odin’s workload suite and run profiles**. The user's approval of headed/headless browser modes does not settle GPU presentation. Terminal orchestration can support headless workloads without opening a window; it cannot thereby prove the desktop presentation path works.
|
|
||||||
|
|
||||||
## 1. Evidence to collect per device
|
|
||||||
|
|
||||||
| Stage | Evidence | Justified conclusion |
|
|
||||||
| --- | --- | --- |
|
|
||||||
| Hardware and kernel path | DRM/sysfs device, associated PCI or platform identity, bound kernel driver, accessible render node | Device and kernel path are present; userspace API operation is still unproven |
|
|
||||||
| API discovery | Vulkan physical-device properties/features/queues; OpenGL vendor/renderer/version from the actual context; OpenCL platform/device type | This implementation advertises these capabilities in this environment |
|
|
||||||
| Execution | Bounded allocation, command submission, completion and error outcome on the selected device | The tested operation completed; report device loss, allocation failure and timeout distinctly |
|
|
||||||
| Correctness | A known-output compute or rendering check with a defined tolerance | The sampled operation returned an expected result; throughput alone does not supply this proof |
|
|
||||||
| Presentation | Successful creation and presentation to a particular X11/Wayland surface, if this mode is selected | That surface/session path worked; headless success is a different finding |
|
|
||||||
|
|
||||||
Linux render nodes permit non-global rendering without DRM-master authentication, subject to ordinary filesystem permissions. They do not grant modesetting rights. Treat denied access as capability coverage, not “bad GPU,” and do not respond by elevating the entire TUI or changing device permissions. DRM discovery must include platform devices: ARM GPUs need not be PCI devices. [1]
|
|
||||||
|
|
||||||
Vulkan properties include device type, vendor/device identifiers, device UUID, and driver identification. `deviceType=CPU` identifies a typically host-processor implementation; the specification calls device type informational, so combine it with driver/renderer evidence. `driverVersion` is vendor-specified, not universal semantic versioning. `conformanceVersion` describes the implementer's prior conformance testing, not a test of this installation. Where supported, `VK_EXT_physical_device_drm` connects API devices to DRM node major/minor numbers. [2]
|
|
||||||
|
|
||||||
**vulkaninfo** is a useful discovery candidate. Its `--summary` covers enumerated devices; `--json=<index>` writes a Vulkan Profiles JSON file for one device. Plain `--json` defaults to the first device, so one successful invocation does not inventory every GPU. Use a private output directory and record the tool/schema version. The inspected SDK tag is `vulkan-sdk-1.4.357.0`, with Apache-2.0 project licensing. [3]
|
|
||||||
|
|
||||||
Mesa LLVMpipe/Softpipe are software renderers. Zink is an OpenGL implementation over Vulkan and can use a hardware Vulkan driver: the word “Mesa” or “Zink” is not evidence of software rendering. Mesa and the Vulkan loader also expose selection/override variables, including `LIBGL_ALWAYS_SOFTWARE`, `DRI_PRIME`, `MESA_VK_DEVICE_SELECT` and `VK_DRIVER_FILES`. Record relevant effective overrides and selected devices; do not silently change the user's stack during baseline measurement. Vulkan/OpenCL CPU devices and mock drivers must not contribute a hardware-GPU score. [4][5][6]
|
|
||||||
|
|
||||||
## 2. Driver and architecture scope
|
|
||||||
|
|
||||||
| Hardware family | Paths worth supporting conditionally | Boundary |
|
|
||||||
| --- | --- | --- |
|
|
||||||
| Intel | Appropriate Linux kernel driver plus Mesa OpenGL/ANV; an independently available compute runtime | Working OpenGL does not establish Vulkan or OpenCL support. Discover generation-specific capabilities rather than prescribe one package universally |
|
|
||||||
| AMD | Supported kernel/userspace combination, commonly amdgpu plus Mesa RADV for Vulkan | RADV and ROCm serve different purposes. ROCm has its own hardware/OS/firmware compatibility matrix; absence of ROCm does not mean ordinary graphics is broken |
|
|
||||||
| NVIDIA | NVIDIA's supported userspace/kernel stack, or Mesa NVK and applicable OpenGL path | NVK is a legitimate Vulkan implementation. NVIDIA's open kernel modules still require matching NVIDIA userspace and GSP firmware; “open module installed” does not establish compatibility |
|
|
||||||
| ARM SoCs | Panfrost/PanVK for supported Mali, Freedreno/Turnip for supported Adreno, other model-specific Mesa/vendor paths | aarch64 names the CPU architecture, not the GPU API capability. Some devices support GLES without Vulkan; experimental support must not be force-enabled automatically |
|
|
||||||
|
|
||||||
Mesa documents RADV's separation from the kernel driver and hardware limitations; Panfrost lists distinct API support by GPU and explicitly warns about experimental PanVK enablement. NVIDIA's inspected `615.71.09` open-module release supports x86_64/aarch64 and Turing-or-later hardware, with corresponding-release userspace/firmware requirements. These examples justify capability probing, not a universal driver recommendation. [7][8][9]
|
|
||||||
|
|
||||||
Qualify **x86_64/glibc, x86_64/musl, aarch64/glibc and aarch64/musl** independently for the chosen executable and transitive libraries. Source availability does not certify a binary across that matrix. Void explicitly states proprietary NVIDIA drivers do not support musl; packaging Odin differently cannot erase that driver limitation. Mesa-based paths can be candidates where that GPU and distribution support them. Current ROCm support is also a specific matrix, not a promise for every Linux distribution or libc. [8][10]
|
|
||||||
|
|
||||||
## 3. Workload candidates and concrete tradeoffs
|
|
||||||
|
|
||||||
| Candidate and inspected version | Measurements, footprint and control | Conditional role |
|
|
||||||
| --- | --- | --- |
|
|
||||||
| **clpeak 2.1.4**, Apache-2.0; released August 27, 2026 | Current code supports Vulkan, OpenCL, CUDA, ROCm/HIP, oneAPI and CPU, among others. CLI has backend/device/test selection and JSON/CSV/XML output. `--max-time` controls each GPU test's timed phase; warmup/calibration add time. C++17/CMake; SDKs/backends are optional but auto-detected by default | Strong first candidate for a deliberately restricted CLI build and selected Vulkan FP32/bandwidth workloads. Pin enabled backends, shaders and compiler; avoid its “run every backend/device/test” default |
|
|
||||||
| **vkpeak 20260527**, MIT; source activity in August 2026 | Vulkan peak scalar/vector/matrix arithmetic and transfer tests using ncnn. Select device and scenarios. Small top-level program, substantial transitive shader/runtime dependency. No user time-budget option is documented in the inspected CLI | Alternative focused compute candidate. Its README explicitly says peak metrics do not represent real-world use. Source returns zero for some unsupported features **and failures**, so zero cannot be interpreted as measured zero performance |
|
|
||||||
| **vkmark 2025.01**, LGPL-2.1-or-later | Configurable Vulkan rendering scenes, dimensions, present mode, duration and device UUID selection. C++17, Vulkan, GLM and Assimp; optional XCB/Wayland/DRM/GBM dependencies | Candidate graphics profile after backend qualification. The released source includes a headless plugin requiring `VK_EXT_headless_surface`; its manpage backend list omits that plugin. Generic Vulkan support alone is insufficient |
|
|
||||||
| **glmark2 2023.01**, GPLv3 | OpenGL 2.0/GLES2 scenes; per-scene duration, off-screen mode, frame-end/swap controls, output validation and CSV/XML results. Build flavors include X11, Wayland, DRM and GBM; GL/EGL/GLES and image libraries/assets | Useful compatibility and rendering candidate; its older API workloads are not a complete modern-GPU assessment. GBM source can use a selected render node; `--off-screen` on an X11 build does not imply display-server independence |
|
|
||||||
|
|
||||||
Primary released READMEs, manuals, licenses and implementation sources support this comparison. glmark2's latest inspected tag remains 2023.01 with main activity in September 2025; vkmark's latest tag is 2025.01 with main activity in September 2025. These are maturity/maintenance observations, not evidence of current hardware certification. clpeak and vkpeak show more recent source/release activity, but still require qualification. [11–14]
|
|
||||||
|
|
||||||
Two implementation traps matter immediately:
|
|
||||||
|
|
||||||
- clpeak's Vulkan instance requests Vulkan 1.0 or 1.1 depending on compiled optional features. Its reported capability floor therefore depends on the build. Its timing code performs warmup and calibration before the timed batch; `--max-time` is not an end-to-end timeout. Its Vulkan backend distinguishes CPU and integrated/discrete GPU device types. [11]
|
|
||||||
- vkpeak adapts work and reports peak results, with memory sizing based partly on device heap information. A selected subset is more controllable than its complete default suite, but a wrapper still needs independent resource and runtime bounds. Neither tool's advertised throughput proves it checks the numerical result required by Odin's correctness stage. [12]
|
|
||||||
|
|
||||||
Licenses above describe inspected project code. Bundling requires a separate manifest for assets, embedded dependencies, modifications and any vendor runtime redistribution terms; these source inspections are not a completed distribution-license audit. Installing a large CUDA/ROCm SDK solely to enable baseline benchmarking would weaken the universal deployment objective. Vendor-specific compute paths are better considered optional capability profiles.
|
|
||||||
|
|
||||||
## 4. Headless, desktop, multiple GPUs and virtualization
|
|
||||||
|
|
||||||
Keep three execution classes distinct: **surface-free compute**, **offscreen rendering**, and **desktop presentation**. Vulkan does not require every physical device or queue to support presentation. Support must be queried for the actual surface. FIFO presentation waits on vertical blanking; an FPS result can therefore reflect display/compositor policy rather than maximum render throughput. Fix and record present mode, resolution and backend. [15]
|
|
||||||
|
|
||||||
vkmark's headless plugin still uses a Vulkan surface/swapchain extension; glmark2's GBM backend opens a render node and creates a GBM surface. These are different requirements and workloads. A KMS/direct-display backend may need display ownership and disturb the session, so it is not an automatic fallback when X11/Wayland fails. An SSH terminal can have usable GPU compute without a display socket; classify presentation as unavailable in that session rather than infer a missing graphics driver. [1][13][14]
|
|
||||||
|
|
||||||
Enumerate all devices, map them to stable identifiers where available, and let the run profile select the display GPU, another named GPU or separate per-GPU runs. Do not treat index zero as “best GPU.” Mesa's selection variables can reorder enumeration; vkmark's UUID selector and NVIDIA's documented UUID/PCI-ID selection illustrate stronger identity mechanisms. Avoid summing overlapping APIs or independently averaging all installed GPUs into one unexplained number. [2][5][9][13]
|
|
||||||
|
|
||||||
Virtual hardware needs its own label. Mesa Venus serializes Vulkan through virtio-gpu to a host renderer and can operate over hardware **or Lavapipe**; guest enumeration does not establish physical passthrough. A VM result measures that guest path, while its host telemetry may be hidden. Passthrough must be established from the recorded environment and device evidence. Containers similarly need device access and compatible userspace libraries; missing exposure is not proof the host has no GPU. [6][16]
|
|
||||||
|
|
||||||
## 5. Reproducibility, safety and diagnosis
|
|
||||||
|
|
||||||
Before scoring, freeze workload version, scene/kernel code, input size, precision/vector width, backend, device, output format, build flags and compiler. Record kernel, userspace driver identity, power source, thermal state, display mode and concurrent load. Warmup/cache policy must be explicit. Compare repeated runs under the same policy; do not compare shader compilation included in one result with warmed execution in another.
|
|
||||||
|
|
||||||
Compute FLOPS, transfer bandwidth and scene FPS answer different questions. A CUDA FP16 matrix peak cannot replace a Vulkan FP32 score; software rendering cannot replace the hardware result; unavailable features are not zeros. Preserve upstream metrics and chosen aggregation rules. glmark2/vkmark aggregate FPS does not automatically supply frame-time percentiles. Small rendering scenes can also be CPU/driver limited, so unexpectedly low throughput is evidence to investigate, not automatic proof of defective hardware.
|
|
||||||
|
|
||||||
Recommended safety boundaries are one GPU workload at a time, bounded memory demand with system/VRAM reserve, an independent wall-clock watchdog, staged process-group cancellation, and refusal of unbounded “run forever” modes. Include initialization/calibration in the overall budget. Stop on device loss, repeated API errors or meaningful driver-reported critical thermal findings; retain partial results. Killing a client cannot guarantee immediate recovery from a kernel/driver hang. Do not automatically overclock, alter fan/power limits, reset a GPU or replace drivers. Resource limits and safe cancellation remain acceptance tests for the selected workload.
|
|
||||||
|
|
||||||
Telemetry is supporting evidence, with provider-specific meaning:
|
|
||||||
|
|
||||||
- NVIDIA `nvidia-smi` documents unsupported values as `N/A`, separate errors for permission denial, unloaded driver and missing NVML, and stable UUID/PCI selection. Its utility success does not test Vulkan/OpenGL presentation. Read-only query adapters must preserve unavailable/error outcomes. [9]
|
|
||||||
- amdgpu exposes temperature, load, power and other sysfs metrics, but support varies. Its APU power reading includes CPU power, so it is not interchangeable with discrete-GPU-only power. [17]
|
|
||||||
- Linux DRM fdinfo defines per-client engine-busy counters, capacities and accounting rules; availability depends on the driver and accessible process descriptors. These can help attribute work without assuming one vendor's utilization meaning applies everywhere. [18]
|
|
||||||
|
|
||||||
Advice should name the evidence: “Vulkan userspace driver could not load,” “render-node access denied,” “software renderer selected,” “this feature is unsupported,” or “workload lost the device.” A package/version mismatch needs concrete loader or vendor evidence; a successful fallback may be intentional. Present a distro-appropriate investigation step with confidence and tradeoffs, rather than an unconditional “install proprietary drivers.”
|
|
||||||
|
|
||||||
## 6. Decisions and limits carried forward
|
|
||||||
|
|
||||||
The workload decision must choose: headless compute and graphics requirements; whether visible GPU presentation is allowed; selected tool/build and minimum API features; treatment of software/virtual/unsupported paths; default multi-GPU selection; memory/runtime budgets; and the correctness check and scoring eligibility rules.
|
|
||||||
|
|
||||||
Before a supported release, qualify the four architecture/libc lanes, real Intel/AMD/NVIDIA hardware, representative ARM SoCs, X11/Wayland/headless sessions, multiple GPUs, denied permissions, software rendering, VM acceleration/passthrough, and cancellation/device-loss fixtures. No such execution evidence was produced here. No package-size estimates or full vendor conformance/redistribution audit were established.
|
|
||||||
|
|
||||||
Context7 resolution preceded documentation lookup. Its glmark2 queries yielded no relevant main-project documentation; clpeak was unindexed; vkpeak resolved to an unrelated speech tool and was rejected; NVIDIA NVML searches produced unrelated/wrapper results. Official tagged source and vendor documentation supplied those gaps. The guessed Mesa Lavapipe page returned 404; software/virtual-path claims use inspected Mesa driver documentation and API/device evidence instead. Current Mesa pages contain evolving and occasionally differing generation summaries, so no exhaustive model support table is inferred from them.
|
|
||||||
|
|
||||||
## Sources
|
|
||||||
|
|
||||||
1. [Linux DRM userspace API, render nodes](https://docs.kernel.org/gpu/drm-uapi.html).
|
|
||||||
2. [Vulkan device/queue specification](https://github.com/KhronosGroup/Vulkan-Docs/blob/main/chapters/devsandqueues.adoc), physical-device, driver, UUID and DRM properties.
|
|
||||||
3. Vulkan Tools SDK tag: [vulkaninfo documentation](https://github.com/KhronosGroup/Vulkan-Tools/blob/vulkan-sdk-1.4.357.0/vulkaninfo/vulkaninfo.md), [license](https://github.com/KhronosGroup/Vulkan-Tools/blob/vulkan-sdk-1.4.357.0/LICENSE.txt).
|
|
||||||
4. Mesa [platforms/drivers](https://docs.mesa3d.org/systems.html), [LLVMpipe](https://docs.mesa3d.org/drivers/llvmpipe.html), [Zink](https://docs.mesa3d.org/drivers/zink.html).
|
|
||||||
5. [Mesa environment variables](https://docs.mesa3d.org/envvars.html), [Vulkan loader driver discovery](https://github.com/KhronosGroup/Vulkan-Loader/blob/main/docs/LoaderDriverInterface.md).
|
|
||||||
6. [OpenCL device enumeration](https://github.com/KhronosGroup/OpenCL-Registry/blob/main/specs/unified/refpages/man/html/clGetDeviceIDs.html), [Vulkan Guide support/null-driver discussion](https://github.com/KhronosGroup/Vulkan-Guide/blob/main/chapters/checking_for_support.adoc), inspected through Context7.
|
|
||||||
7. Mesa [ANV](https://docs.mesa3d.org/drivers/anv.html), [RADV](https://docs.mesa3d.org/drivers/radv.html), [NVK](https://docs.mesa3d.org/drivers/nvk.html), [Panfrost](https://docs.mesa3d.org/drivers/panfrost.html), [Freedreno/Turnip](https://docs.mesa3d.org/drivers/freedreno.html).
|
|
||||||
8. [ROCm current compatibility matrix](https://rocm.docs.amd.com/en/latest/compatibility/compatibility-matrix.html).
|
|
||||||
9. NVIDIA [open-module 615.71.09 README](https://github.com/NVIDIA/open-gpu-kernel-modules/blob/615.71.09/README.md), [nvidia-smi documentation](https://docs.nvidia.com/deploy/nvidia-smi/index.html).
|
|
||||||
10. [Void musl compatibility](https://docs.voidlinux.org/installation/musl.html).
|
|
||||||
11. clpeak 2.1.4: [README](https://github.com/krrishnarraj/clpeak/blob/2.1.4/README.md), [CLI options](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/common/options.cpp), [Vulkan timing/instance implementation](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/vulkan/vk_peak.cpp), [device mapping](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/vulkan/vulkan_device.cpp), [build options](https://github.com/krrishnarraj/clpeak/blob/2.1.4/CMakeLists.txt), [license](https://github.com/krrishnarraj/clpeak/blob/2.1.4/LICENSE).
|
|
||||||
12. vkpeak 20260527: [README](https://github.com/nihui/vkpeak/blob/20260527/README.md), [implementation](https://github.com/nihui/vkpeak/blob/20260527/vkpeak.cpp), [build/dependency configuration](https://github.com/nihui/vkpeak/blob/20260527/CMakeLists.txt), [license](https://github.com/nihui/vkpeak/blob/20260527/LICENSE).
|
|
||||||
13. vkmark 2025.01: [README](https://github.com/vkmark/vkmark/blob/2025.01/README.md), [manual](https://github.com/vkmark/vkmark/blob/2025.01/doc/vkmark.1), [headless implementation and license notice](https://github.com/vkmark/vkmark/blob/2025.01/src/ws/headless_native_system.cpp), [backend build](https://github.com/vkmark/vkmark/blob/2025.01/src/meson.build).
|
|
||||||
14. glmark2 2023.01: [README/license declaration](https://github.com/glmark2/glmark2/blob/2023.01/README), [manual](https://github.com/glmark2/glmark2/blob/2023.01/doc/glmark2.1.in), [GBM implementation](https://github.com/glmark2/glmark2/blob/2023.01/src/native-state-gbm.cpp), [build flavors](https://github.com/glmark2/glmark2/blob/2023.01/meson_options.txt).
|
|
||||||
15. [Vulkan WSI specification](https://github.com/KhronosGroup/Vulkan-Docs/blob/main/chapters/VK_KHR_surface/wsi.adoc), surface support, headless surfaces and present modes.
|
|
||||||
16. [Mesa Virtio-GPU Venus](https://docs.mesa3d.org/drivers/venus.html).
|
|
||||||
17. [Linux amdgpu thermal/power monitoring](https://docs.kernel.org/gpu/amdgpu/thermal.html).
|
|
||||||
18. [Linux DRM usage-statistics ABI](https://docs.kernel.org/gpu/drm-usage-stats.html).
|
|
||||||
@@ -0,0 +1,140 @@
|
|||||||
|
# Defensible median scoring and comparison rules
|
||||||
|
|
||||||
|
Research for [Establish defensible median scoring and comparison rules](https://git.bongbetic.com/xavierk/odin/issues/7), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
|
||||||
|
|
||||||
|
**Accessed:** 25 September 2026. **Status:** decision evidence and recommendations; no final scoring formula, calibrated reference values, benchmark runs, or implementation. The user's requirement is a **median** score.
|
||||||
|
|
||||||
|
## What the median should mean
|
||||||
|
|
||||||
|
There are three separate choices:
|
||||||
|
|
||||||
|
| Level | Meaning | Main limitation |
|
||||||
|
|---|---|---|
|
||||||
|
| Median of repeated measurements | Typical result for one fixed workload under stated conditions | Hides occasional long stalls; does not combine CPU, storage and browser results |
|
||||||
|
| Median across normalized workloads | Typical relative performance across a fixed test set | Domains with many tests gain more influence; changing references can change rankings |
|
||||||
|
| Median across domain summaries | Typical relative performance across explicitly chosen domains | Domain definitions matter; poor performance in a minority of domains can disappear from the headline |
|
||||||
|
|
||||||
|
NIST defines the sample median as the middle observation, or the arithmetic average of the middle two for an even sample count. Its resistance to extremes is useful, but it is a measure of location, not completeness, reliability, or worst-case response. [NIST location][nist-location]
|
||||||
|
|
||||||
|
**Recommendation for discussion:** use medians to summarize repeated valid measurements, normalize against frozen references, form predefined domain summaries, then use a median across those domains for the requested headline. Preserve each stage and its raw inputs. This proposes an aggregation structure; the domain membership, weighting, reference values, repeat counts and numeric scale remain decisions.
|
||||||
|
|
||||||
|
Established suites demonstrate why the levels must stay explicit. SPEC CPU 2017 takes median execution times from three runs, or the slower of two, then uses a **geometric mean** across ratios. Speedometer uses inverse geometric means across test durations and arithmetic means across iterations. These are methodological precedents, not permission to substitute a geometric mean for Odin's requested median. Preserve a tool's native result under its original name; label Odin's further aggregation separately. [SPEC rules][spec-rules] [Speedometer methodology][speedometer]
|
||||||
|
|
||||||
|
## Normalization and category balance
|
||||||
|
|
||||||
|
Milliseconds, operations/second and GB/s cannot share a meaningful raw median. A candidate approach is a dimensionless ratio against the **same workload's** reference: observation/reference for positive higher-is-better measures, reference/observation for positive lower-is-better measures. SPEC uses the latter for elapsed-time ratios. The metric identity must include its units, workload, size, concurrency, timing boundaries and direction. Reject invalid/nonfinite inputs under declared validity rules; a zero measured duration must not become an infinite score. [SPEC overview][spec-overview]
|
||||||
|
|
||||||
|
**Fictional arithmetic examples throughout this report:** these numbers illustrate consequences only. They are not Odin calibration or measured hardware results. A displayed index of 100 at the reference is an arbitrary illustrative scale, not a recommendation.
|
||||||
|
|
||||||
|
| Fictional metric | Reference | Observed | Illustrative ratio / index |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| Work throughput | 25 operations/s | 50 operations/s | 2.0 / 200 |
|
||||||
|
| Memory bandwidth | 25 GB/s | 20 GB/s | 0.8 / 80 |
|
||||||
|
| Request latency | 4 ms | 2 ms | 2.0 / 200 |
|
||||||
|
|
||||||
|
A median of these indices is 200, despite memory bandwidth being below reference. That is a consequence of the chosen statistic. Show domain detail and slow-tail measurements alongside it. “200 versus 100” describes this index; it does not establish that every application is twice as fast.
|
||||||
|
|
||||||
|
Fix the order of operations. For fictional repeated times `[1, 3]` ms and a 3 ms reference, normalizing the raw median gives `3 / 2 = 1.5`; taking the median of individual ratios `[3, 1]` gives 2. Even-count averaging and reciprocal normalization do not commute. The report format must specify which result it contains. [median definition][nist-location]
|
||||||
|
|
||||||
|
Balance domains before counting metrics. If ten CPU tests each score 160 and four other domains score 40, 80, 100 and 120, the flat fourteen-test median is 160. A median across five domain summaries is 100. Adding CPU subtests should not silently redefine the product's priorities. Similarly, adding Python/Rust/Java variants must not automatically multiply the language domain's influence. A median of domain medians is a deliberate hierarchical index, not the pooled median of all observations.
|
||||||
|
|
||||||
|
Exclude health counters, memory-test pass/fail, driver availability and installed RAM capacity from throughput arithmetic. Multiple correlated outputs from one workload—throughput, IOPS, average latency and several percentiles—also need an explicit selection rule before any becomes an independent scored contribution.
|
||||||
|
|
||||||
|
## Calibration is a substantive decision
|
||||||
|
|
||||||
|
SPEC establishes per-workload reference times on a named machine and publishes the calculation rules. Its documentation explains that reference changes preserve relative overall rankings for its geometric-mean calculation. **That invariance does not generally hold for a median across normalized metrics.** [SPEC reference explanation][spec-overview]
|
||||||
|
|
||||||
|
For three fictional higher-is-better workloads, let machine A produce `[1, 10, 10]` and B produce `[2, 2, 20]` in each workload's own units:
|
||||||
|
|
||||||
|
| Fictional reference vector | A's normalized results → median | B's normalized results → median | Ordering |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `[1, 1, 1]` | `[1, 10, 10]` → 10 | `[2, 2, 20]` → 2 | A higher |
|
||||||
|
| `[1, 10, 10]` | `[1, 1, 1]` → 1 | `[2, 0.2, 2]` → 2 | B higher |
|
||||||
|
|
||||||
|
The measurements did not change. Changing the reference altered the relative scales and which observations occupied the middle. Retain the requested median, make the reference rationale public, and version reference changes rather than treating them as cosmetic rescaling.
|
||||||
|
|
||||||
|
| Reference option | What it supports | Decision cost |
|
||||||
|
|---|---|---|
|
||||||
|
| Named reference configuration | Auditable, fixed anchor with documented per-test measurements | One machine's balance influences the median; configurations and repeatability need validation |
|
||||||
|
| Frozen reference cohort | Per-test references from a documented collection of machines | Cohort selection, sampling bias and revision policy become part of the score |
|
||||||
|
| User's own baseline | Local before/after comparisons | A score relative to oneself cannot rank different machines |
|
||||||
|
|
||||||
|
Recommend evaluating a frozen, locally distributable calibration manifest. Include reference measurements and provenance, reference conditions, workload and artifact digests, units/directions, aggregation order, required domains and calibration identity. Store it with results so calculation remains reproducible offline. Raw observations must survive changes; a recalculated score should identify its new calibration and preserve the original.
|
||||||
|
|
||||||
|
An arbitrary scale factor is acceptable if described as an index. A claim such as “100 is the median Linux machine,” a percentile rank, or a universal poor/good threshold requires representative population evidence that does not exist yet. Separately normalizing each architecture to its own average would also prevent interpreting those numbers as one common cross-architecture scale.
|
||||||
|
|
||||||
|
## Repetitions, warmup and uncertainty
|
||||||
|
|
||||||
|
Google Benchmark documents warmup, repetitions, median, standard deviation and coefficient of variation; it distinguishes user-visible wall time from CPU consumption. NIST recommends examining ordered observations for changing location/spread and says potential outliers should not simply be deleted when their cause is unknown. These support retaining all repeated observations and their conditions. [Google Benchmark][google-guide] [NIST run sequence][nist-runseq] [NIST outliers][nist-outliers]
|
||||||
|
|
||||||
|
Recommended measurement rules:
|
||||||
|
|
||||||
|
- Define warmup separately for each workload and retain its duration. Warm caches/JIT throughput, cold launch time and sustained thermal performance are different questions. Do not discard a slow first run from a declared cold-start test.
|
||||||
|
- Fix repetition and stopping rules before observing scores. A quick run may estimate a median without enough evidence for a useful confidence interval; it should not claim the precision of the standard profile. Calibrate the minimum repeats against the 10–20 minute budget.
|
||||||
|
- Retain run order, warmup, elapsed time, temperatures/power context, competing activity and invalidation reasons. A drifting sequence is not interchangeable independent noise. Repetitions within one process or thermal episode are not automatically independent runs.
|
||||||
|
- Exclude observations only for declared validity failures such as incorrect output, changed workload, cancellation or protocol failure. Preserve them with reasons. A slow but valid run can represent the usability problem Odin is meant to reveal.
|
||||||
|
- Show central spread such as MAD or IQR, plus tails where the workload supplies enough events. NIST defines MAD and IQR as distinct measures of spread; neither is itself a confidence interval. A median across repeated p99 values must not be labelled the p99 of all requests. [NIST scale][nist-scale]
|
||||||
|
|
||||||
|
NIST documents median confidence intervals based on order statistics/binomial probabilities, interpolated methods and bootstrap alternatives. Choose and validate a median-appropriate method; do not apply a mean's standard-error formula to a median. Confidence also depends on sample count and assumptions about the measurements. For an aggregate, account for shared run-level variation and state whether uncertainty in the calibration reference is included. The spread **between different domain scores** is not sampling uncertainty about the headline. [NIST median intervals][nist-median-ci]
|
||||||
|
|
||||||
|
Keep a graph of results in time order. Google documents CPU selection, boost, scheduler contention, SMT, caches and NUMA as variance sources. Its suggestions for controlled laboratory microbenchmarks include changing system settings; Odin's installed-system baseline should record existing conditions and label any tuned experiment separately. The existing [CPU/memory report][odin-cpu] and [portability/UI report][odin-portability] explain TUI interference and qualification needs. Stable repeated numbers alone do not prove that a workload represents real usability. [Google variance][google-variance]
|
||||||
|
|
||||||
|
## Comparability must be attached to every score
|
||||||
|
|
||||||
|
SPEC requires performance-relevant observation conditions and valid workload outputs; its CPU suite intentionally measures processor, memory subsystem **and compilers**. Even a fixed-toolchain comparison describes a defined software/hardware configuration. [SPEC rules][spec-rules] [SPEC overview][spec-overview]
|
||||||
|
|
||||||
|
Recommend two explicit comparison purposes:
|
||||||
|
|
||||||
|
- **Installed-system usability:** the chosen installed browser, shell, runtimes, drivers and kernel are part of what is measured. Version/configuration changes may explain a score change without any hardware change.
|
||||||
|
- **Controlled reference workload:** fixed workload assets, runtime/compiler contracts, flags, input data and execution modes improve comparison across machines. Architecture-specific artifacts must implement equivalent declared work and validate outputs; different ISA policies need disclosure.
|
||||||
|
|
||||||
|
Neither mode needs to masquerade as a pure hardware measurement. Keep their result identities distinct. Kernel/libc/distro differences can be the subject of a comparison, but they must be visible and the workload contract must remain equivalent.
|
||||||
|
|
||||||
|
A comparison identity should include suite/scoring/calibration versions, workload set, run profile, tool/artifact versions, options and data digests, timing/aggregation rules, browser mode, hardware/virtualization context and validity/coverage. Require a documented equivalence decision before combining scores across changed tools or workloads. SPEC warns that scores across different suite generations generally cannot be converted. [SPEC overview][spec-overview]
|
||||||
|
|
||||||
|
Speedometer 3.1 instructs users to use a clean browser profile, close competing programs/tabs, keep its page focused, avoid device interaction, use AC power and allow cooling when needed. Its official UI computes a 95% interval around its **arithmetic mean**; that interval cannot be attached to Odin's median unchanged. Preserve native browser score/uncertainty and label any median of complete runs separately. Headed and headless measurements need separate identities until an equivalence study justifies any shared interpretation; background versus foreground execution is also material. [instructions][speedometer-instructions] [3.1 result code][speedometer-main]
|
||||||
|
|
||||||
|
VM results characterize the guest allocation and virtualization environment. Keep native, virtualized and emulated cohorts identifiable; VM compatibility success does not establish native performance. Storage cache mode, queue depth, engine, filesystem and durability policy similarly belong to the workload identity. A fallback such as buffered I/O cannot silently replace a direct-I/O measurement with the same scoring identity. [CPU/memory report][odin-cpu] [storage report][odin-storage] [portability report][odin-portability]
|
||||||
|
|
||||||
|
## Missing tests and eligibility
|
||||||
|
|
||||||
|
For fictional domain indices `[40, 80, 100, 120, 160]`, the complete median is 100. Omitting 40 produces 110; omitting both 40 and 80 produces 120. Available-only aggregation can reward absent or deliberately skipped weak components.
|
||||||
|
|
||||||
|
**Recommendation:** define a versioned required set for the full score. Permit a clearly named partial median and domain results when the full set is unavailable, with the exact subset identified. Compare partial scores only over the same compatible subset; a pairwise intersection comparison must recompute **both** results and label that narrower scope. An optional pack must not silently change the headline's membership.
|
||||||
|
|
||||||
|
Keep distinct outcomes: completed-valid, completed-with-limitations, unsupported, missing dependency, permission denied, unsafe to run, cancelled, timed out, and failed validation. The eventual validity contract decides whether a limited result remains score-eligible. Never impute missing results as zero, a reference score, or a healthy pass. A required workload failing verification makes the full score ineligible, while preserving completed measurements and the associated finding. Good numbers from other domains should not cancel that failure.
|
||||||
|
|
||||||
|
The user accepted reporting unavailable tests across Linux targets. That does not resolve which domains are mandatory, whether every machine should still display a partial number, or how partial results should look. Those are explicit product decisions.
|
||||||
|
|
||||||
|
## Presentation, recommendations and open decisions
|
||||||
|
|
||||||
|
Recommend a headline containing the median, score identity, full/partial state, and eligible-domain coverage. The next view should show domain values, raw units, repeat count/spread, tail latency, invalidations and reference details. Show health findings beside performance: a fast drive with serious SMART evidence still needs attention. Missing telemetry must stay unknown. The storage and CPU reports establish why speed cannot determine drive replacement or certify memory health. [storage][odin-storage] [CPU/memory][odin-cpu]
|
||||||
|
|
||||||
|
Optimization advice should cite the observation and matching rule: for example, measured foreground stalls plus pressure evidence can support investigating memory contention. A low normalized score alone does not identify its cause. Keep severity of health evidence, completeness of coverage, measurement uncertainty and performance position as separate concepts; avoid one synthetic “confidence/health” percentage that mixes them.
|
||||||
|
|
||||||
|
Before implementation, decide:
|
||||||
|
|
||||||
|
1. The headline's median level, domain membership and balancing rules; whether responsiveness contributes or remains an accompanying measurement.
|
||||||
|
2. The reference configuration/cohort, scale and calibration-release policy; collect actual calibration data before inventing thresholds.
|
||||||
|
3. Required versus optional coverage, partial-score display and exact comparison eligibility.
|
||||||
|
4. Installed-system versus controlled-workload defaults, architecture/ISA policies, browser modes and VM cohorts.
|
||||||
|
5. Repetition/warmup/stopping rules, outlier validity rules, median interval method and honest quick/standard/extended precision claims.
|
||||||
|
6. Evidence requirements for optimization rules and separation of urgent health findings from the score.
|
||||||
|
|
||||||
|
**Evidence limits:** no calibration population, repeatability measurements, TUI-overhead budget or headed/headless equivalence study was produced. Examples are arithmetic demonstrations only. Context7 successfully resolved Google Benchmark; two BrowserBench/Speedometer searches returned unrelated packages, so its official repository and deployed 3.1 documentation were inspected directly. NIST and SPEC sources were inspected directly as statistical and benchmark-methodology references. This report neither adopts SPEC's workloads nor claims that their aggregation rules are Odin's final design.
|
||||||
|
|
||||||
|
[nist-location]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda351.htm
|
||||||
|
[nist-scale]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda356.htm
|
||||||
|
[nist-outliers]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda35h.htm
|
||||||
|
[nist-runseq]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda33p.htm
|
||||||
|
[nist-median-ci]: https://www.itl.nist.gov/div898/software/dataplot/refman1/auxillar/mediancl.htm
|
||||||
|
[spec-rules]: https://www.spec.org/cpu2017/Docs/runrules.html
|
||||||
|
[spec-overview]: https://www.spec.org/cpu2017/Docs/overview.html
|
||||||
|
[google-guide]: https://github.com/google/benchmark/blob/main/docs/user_guide.md
|
||||||
|
[google-variance]: https://github.com/google/benchmark/blob/main/docs/reducing_variance.md
|
||||||
|
[speedometer]: https://github.com/WebKit/Speedometer/blob/main/README.md
|
||||||
|
[speedometer-instructions]: https://browserbench.org/Speedometer3.1/instructions.html
|
||||||
|
[speedometer-main]: https://browserbench.org/Speedometer3.1/resources/main.mjs
|
||||||
|
[odin-cpu]: https://git.bongbetic.com/xavierk/odin/src/commit/90213d7f62cbd118f08ea8ff2f8042e94aa038a7/docs/research/cpu-memory.md
|
||||||
|
[odin-storage]: https://git.bongbetic.com/xavierk/odin/src/commit/6e5af87a64faedd4a8ad31ba10d9be4b499e8349/docs/research/storage-health.md
|
||||||
|
[odin-portability]: https://git.bongbetic.com/xavierk/odin/src/commit/20681cd4f184a9fc0164dd638a252de44ef230f5/docs/research/portability-ui.md
|
||||||
Reference in New Issue
Block a user