Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7ff9d74e8f |
@@ -1,135 +0,0 @@
|
||||
# Establish browser, shell, and language workload validity
|
||||
|
||||
Research for [Establish browser, shell, and language workload validity](https://git.bongbetic.com/xavierk/odin/issues/5), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
|
||||
|
||||
Access date: **2026-09-25**. This is a candidate analysis, not the final suite. No browsers, toolchains, or dependencies were installed; no benchmarks or builds were run.
|
||||
|
||||
## Recommendation for the selection decision
|
||||
|
||||
Preserve five distinct measurement purposes: browser application responsiveness, shell interaction, process startup, compilation, and warmed application execution. A single fast loop cannot represent all five. Retain upstream benchmark names, versions, metrics, and correctness rules; an Odin aggregate is a separate decision.
|
||||
|
||||
A practical initial option is **Speedometer 3.1**, a small Bash latency/throughput set, and optional language packs with pinned workload inputs. JetStream, MotionMark, representative compilation projects, and larger JVM applications fit extended investigation. Installed software answers “how usable is this machine as configured”; pinned reference software improves comparison across machines. Store these as distinct modes rather than treating them as interchangeable.
|
||||
|
||||
## 1. Browser candidates and official measurement rules
|
||||
|
||||
BrowserBench’s live index currently points to **Speedometer 3.1**, **JetStream 3.0**, and **MotionMark 1.3.2**. [1]
|
||||
|
||||
| Candidate | What it measures | Conditional role and interpretation |
|
||||
| --- | --- | --- |
|
||||
| Speedometer 3.1 | End-to-end simulated web-app interactions: TodoMVC frameworks, editors, charts, and news applications. | Strong default browser-responsiveness candidate. Preserve the full default suite and iteration protocol for an upstream-style result. It does not measure browser process startup or network loading. |
|
||||
| JetStream 3.0 | JavaScript/WebAssembly workloads, combining startup, average, and worst-case execution behavior. | Useful extended engine-compute view. “Startup” is workload first-iteration behavior, not opening the browser. Its 77 default workloads and upstream geometric aggregation should remain intact. |
|
||||
| MotionMark 1.3.2 | Browser graphics scenes adjusted toward a target frame rate, with confidence intervals. | Useful optional browser-rendering view where a suitable graphics/display path exists. It is not a standalone GPU benchmark or evidence that the intended hardware driver is active. |
|
||||
|
||||
Speedometer’s release/3.1 source currently resolves to commit `1386415be8fef2f6b6bbdbe1828872471c5d802a`. Its `package.json` still says `3.0.0-alpha`, demonstrating why package metadata alone is not a benchmark identity. The release source defaults to **10 iterations** and an **800×600 workload viewport**; URL parameters can change iterations, suites, warmup, and measurement behavior. The result includes individual measurements and JSON exports. Preserve these settings and raw results; do not silently shorten the suite and report an ordinary “Speedometer 3.1” score. [2]
|
||||
|
||||
Official Speedometer instructions recommend a latest stable browser, a separate clean profile, closed background applications/tabs, a focused benchmark page, AC power, no interaction, and cooling between runs where needed. A pinned historical browser is useful for longitudinal comparison but must be identified as a reference-browser experiment. Speedometer aggregates inverse geometric-mean durations, then averages scores across iterations and reports a 95% confidence interval around that mean. Retain that native statistic and uncertainty. Any Odin median across complete benchmark runs is a separate statistic; do not relabel either or average heterogeneous subtest times. [2][3]
|
||||
|
||||
JetStream explicitly says scores are not comparable across its versions. Its startup, worst-case, and average components are deliberately part of the metric; generic harness warmups must not erase those startup observations. MotionMark depends on drawing-area class, viewport, refresh behavior, and browser rendering. Its instructions recommend maximizing the window at the default resolution and 60 Hz for comparison across browsers; a changed display policy is a recorded experimental condition, not something Odin should silently impose. [4][5]
|
||||
|
||||
### Offline hosting and licensing
|
||||
|
||||
Speedometer documents a local static HTTP server, so a prepared source/assets snapshot can be served on loopback without fetching the public site during measurement. Its root license permits source/binary redistribution with the copyright, conditions, and disclaimer retained. That does not replace the license notices of bundled frameworks, fonts, datasets, or other third-party assets. JetStream and MotionMark’s deployed HTML also carries permissive two-condition notices, while their workload provenance points to multiple projects. Before distributing a pack, inventory the exact selected files and retain their individual notices. [2][4][5][6]
|
||||
|
||||
Recommended preparation contract: content-addressed assets, release/commit identity, dependency/license manifest, complete local resource closure, and a future offline-network audit. Bind the server to loopback and prepare everything before timing. A copied public page without its subresources is not a verified offline benchmark. Keep upstream workload code unchanged where feasible; record every adapter patch. The existence of localhost support does not prove the entire pack is self-contained until that artifact is checked.
|
||||
|
||||
### Automation and the Linux boundary
|
||||
|
||||
Two viable routes serve different purposes:
|
||||
|
||||
1. **Native installed browser plus matching WebDriver.** Speedometer itself documents Selenium testing with ChromeDriver, GeckoDriver, and other browser drivers. This route can use distro-supported native browsers on Void/musl or other systems outside Playwright’s matrix, provided the actual browser/driver pair exists and is qualified. Use explicit browser and driver paths and versions. Selenium Manager can download software by default; its documented offline mode disables network requests/downloads. Its documentation still describes limited Linux architecture support, so automatic management cannot be presumed to cover aarch64 or musl. Native automation remains a capability to prove, not a promise. [7]
|
||||
2. **Pinned browser supplied with Playwright on its supported platforms.** Current documentation lists Debian 12/13 and Ubuntu 22.04/24.04/26.04 on x86-64/arm64. Its bundled Firefox/WebKit builds require glibc; musl distributions are unsupported. Stock Firefox cannot simply be substituted because Playwright requires its patched Firefox. Branded Chrome/Edge support and custom executable paths also do not establish universal compatibility. [8]
|
||||
|
||||
**Odin can launch on a system even when its browser module is unavailable.** On musl, use a qualified native browser/driver pair or report an explicit capability skip. A glibc container/chroot or software-rendered browser is a separate environment and cannot transparently stand in for native desktop usability.
|
||||
|
||||
The user has approved **both headed and headless browser modes**, with terminal controls/results and distinct measurement labels. Chrome’s unified headless mode shares browser code with headed Chrome, but Playwright also provides a distinct headless-shell binary and documents behavioral differences. Sharing code does not establish equal scores, display scheduling, GPU access, or compositor behavior. Label a headless cohort explicitly and avoid claiming compliance with focused-window desktop run conditions. A headed cohort needs a graphical session and the official run conditions. Do not merge their reference distributions or invent a conversion factor. The default mode for each run profile remains a selection decision. [8][9]
|
||||
|
||||
Automation should start the pinned page, observe completion, and collect its full result without repeated polling/instrumentation inside timed work. Correctness failures, missing workloads, or timeouts invalidate a full-suite result. Shortened or filtered runs can be Odin browser probes with their own identities; they are not substitutes for official full-suite statistics.
|
||||
|
||||
Record browser binary/version/channel, driver and automation versions, full flags, profile policy, headed/headless/shell distinction, viewport/device scale, display resolution/refresh, X11/Wayland/virtual display, renderer and acceleration evidence when exposed, kernel/libc/architecture, power state, suite digest, and all query parameters. An unavailable GPU detail is unknown, not proof of hardware acceleration. Browser startup, if desired, should be a separate process-to-ready-page test with its cache/profile policy declared.
|
||||
|
||||
## 2. Meaningful Bash and shell measurements
|
||||
|
||||
Bash startup depends on invocation. Login shells read `/etc/profile` and the first readable user login profile; interactive non-login shells read `.bashrc`; noninteractive shells may execute the file named by `BASH_ENV`. `--noprofile` and `--norc` control different paths. Therefore “time bash” is underspecified. [10]
|
||||
|
||||
| Candidate | Measurement definition | Boundary to preserve |
|
||||
| --- | --- | --- |
|
||||
| Clean process startup | Directly launch a known Bash binary with controlled environment/startup-file policy and a trivial command; measure process launch through exit. | This includes loader/process/exit costs. It is not prompt readiness. Explicitly control `BASH_ENV`, inherited functions, locale, and working directory. |
|
||||
| Interactive readiness | Launch a clean interactive shell under a PTY; measure until a declared prompt-ready marker, then measure prompt return after a builtin command. | Account for PTY setup and marker instrumentation. `bash -i -c exit` does not observe a real prompt. Test login and non-login separately if both are offered. |
|
||||
| Builtin workload throughput | Fixed arithmetic, parameter expansion, arrays, and parsing over seeded in-memory inputs, returning a verified checksum. | Bounded output; no per-iteration `date`, `cat`, `grep`, or other subprocesses. Input size and Bash version are part of the workload identity. |
|
||||
| External-command/pipeline workload | Fixed commands, input corpus, and pipe structure, with exact executable versions. | Measures the combined shell, process-launch, tool, and I/O stack. Report it as such. |
|
||||
| User configuration readiness | Opt-in observation using the user’s real startup files and prompt configuration. | These files execute arbitrary user-configured commands, may contact networks or change state, and may not be repeatable. Report local usability separately from the clean reference result. |
|
||||
|
||||
**Harness trap:** Hyperfine defaults to an intermediate `/bin/sh` on Unix and calibrates/subtracts shell-spawn time. For Bash startup, use its direct-command mode or another direct process timer so the Bash process is the measured payload. Hyperfine also supports warmups and repetitions; warmed filesystem caches are a declared condition, not “cold boot.” Do not globally drop caches or alter system tuning just to manufacture a shell metric. [11]
|
||||
|
||||
Treat time-to-first-prompt and prompt-to-next-prompt as distributions, not a single best sample. Bound captured PTY output, support cancellation, and keep configured-shell timeout/error details. Store configuration fingerprints with care; benchmark records should not copy potentially secret startup-file contents.
|
||||
|
||||
## 3. Language workload candidates
|
||||
|
||||
Each result describes a **workload + implementation + toolchain + environment**, not an intrinsic language ranking. Numeric kernels are useful but should not be the only evidence. Libraries, parsing, allocation, and application behavior matter; Python’s C-backed JSON module and a Rust regex library do not represent the same implementation simply because both are called language tests.
|
||||
|
||||
| Language / candidate | Useful workload selection | Evidence, tradeoffs, and role |
|
||||
| --- | --- | --- |
|
||||
| Python: pyperformance + pyperf | `python_startup`, `python_startup_no_site`, `json_loads`/`json_dumps`, one regex workload, and optionally a pure-Python workload such as `richards`. | Maintained Python project favors real applications. Current PyPI metadata: pyperformance **1.14.0**, Python ≥3.10; pyperf **2.10.0**, Python ≥3.9. Suite docs contain older requirements, so prefer release metadata. MIT project license; inspect selected workload notices. Small subsets fit standard profiles. [12] |
|
||||
| Rust: rustc-perf subsets | Runtime groups include parsing, text search, compression, hash maps, and numeric/graphics kernels; compile candidates include pinned real crates such as `syn`. | Official compiler-performance project clearly separates compilation from generated-program execution. Runtime suite is explicitly **experimental**, with changing workloads; pin a commit and validate chosen outputs. MIT infrastructure, separate compile-benchmark licenses. Avoid adopting the full collector by default: it has profiling/environment dependencies. [13] |
|
||||
| C++: LLVM test-suite subset | A small C++ application/proxy workload such as the serial miniFE variant, plus a smaller verified workload; optionally time compiling the same pinned sources. | Suite compiles/runs whole programs, checks reference output, and records compile and execution times separately. CMake/lit and dependencies increase setup cost. Apache-2.0 with LLVM exceptions for project code; third-party directories have separate terms. Good standard/extended candidate after footprint qualification. [14] |
|
||||
| Java: selected workloads under JMH | A packaged, reviewed data-processing/allocation workload with validated results and configured warmup/forks. | JMH **1.37** is the current published Maven release inspected. It is a harness, not a representative suite by itself; samples explain pitfalls rather than defining an Odin score. GPLv2 with the Classpath exception on designated files. Suitable where a small, controlled JVM workload is desired. [15] |
|
||||
| Java: DaCapo application subset | `lusearch` for search, `h2` for database-style execution, or another justified application. | Latest inspected release **23.11-MR2-chopin**; application suite with nontrivial memory use and output validation. Release notes establish Java 11–21 compatibility, not arbitrary newer JDKs. Harness Apache-2.0; component programs retain their licenses. Better extended candidate than a compulsory quick test. [16] |
|
||||
|
||||
pyperformance explicitly says it is not tuned for PyPy. CPython, PyPy, free-threaded builds, optional JIT builds, and distro build choices require distinct metadata and qualification. Interpreter startup with and without `site` measures different initialization; preserve both names if offered. Some pyperformance dependencies have native extensions, so a Python package being source-available does not prove wheel availability on every musl/architecture combination. [12]
|
||||
|
||||
## 4. Fair measurement protocol
|
||||
|
||||
**Startup, compilation, and warmed execution need separate records.** For Rust/C++, compile once before execution measurements unless compilation is the workload. For Java, bytecode compilation with `javac` differs from runtime JIT work. For Python, process startup/imports differ from repeated execution in an initialized interpreter. If the workload intentionally includes setup, say so and apply the same rule on every system.
|
||||
|
||||
pyperf uses calibration, multiple processes, warmups, values, and metadata. Its fast mode explicitly trades accuracy for speed. JMH provides forks, warmup/measurement iterations, state setup, and result consumption to avoid dead-code elimination; its maintainers warn that a harness does not eliminate benchmarking mistakes. Retain cold/startup samples and warmed samples separately. Use predetermined warmup/measurement rules and report instability; do not keep warming until a favorable result appears. [15][17]
|
||||
|
||||
Proposed compilation policy for selection:
|
||||
|
||||
- Pin sources, datasets, dependency lockfiles, language standard/edition, compiler identity, target triple, optimization flags, linker, libc, and library versions.
|
||||
- Use a documented release optimization configuration. Cargo’s release defaults include `opt-level=3`, incremental off, and 16 codegen units; its bench profile inherits release. Record overrides and dependency profiles. C++ should use an explicitly chosen optimization level and standard; `-Ofast` changes standards-compliance assumptions and cannot silently replace a strict floating-point contract. [18]
|
||||
- Keep baseline-ISA builds and host-tuned builds distinct. `target-cpu=native`/`-march=native`, SIMD dispatch, LTO, PGO, allocator choice, and thread counts can materially change results. A single binary is not portable across x86_64 and aarch64; equal flags do not mean equal generated instructions.
|
||||
- Compilation tests distinguish clean builds, no-op rebuilds, and controlled incremental edits. Pre-fetch dependencies and hold build parallelism/cache policy fixed. Network resolution is not compilation speed; compiler caches must be controlled or explicitly measured.
|
||||
- Validate outputs/checksums and numerical tolerances. Prevent constant folding/dead work, but avoid timing validation if the workload definition excludes it. Do not compare implementations with different precision, algorithms, data sizes, or hidden thread counts under a shared metric name.
|
||||
|
||||
An **installed-toolchain mode** best reveals current developer experience. A **reference-toolchain mode** uses pinned, prepared artifacts for stronger cross-machine comparison, with acquisition size, libc/ISA support, licenses, and security updates owned explicitly. It must not install into or replace the user’s default toolchain during a run. Report compiler/runtime version changes as changes in execution conditions, not unexplained hardware improvements.
|
||||
|
||||
## 5. Profile options, budgets, and unavailable tools
|
||||
|
||||
The following are provisional **application-domain budgets**, not measured runtime promises or the whole Odin run duration:
|
||||
|
||||
| Profile | Candidate scope | Budget policy to evaluate |
|
||||
| --- | --- | --- |
|
||||
| Quick | Browser capability probe; clean Bash startup and one builtin workload; startup/small verified work on selected installed languages. | Approximately 30–60 seconds for this domain. No full browser score if the official suite does not fit. Short observations carry lower-confidence status. |
|
||||
| Standard | Full Speedometer 3.1 on one qualified browser; Bash set; selected language pack(s) with normal warmup/repetitions. | Reserve minutes rather than seconds; a pilot target is 2–5 minutes for browser and 1–3 minutes per selected language pack. Slow hosts may exceed it. |
|
||||
| Extended | JetStream, optional MotionMark, more process forks, compilation workloads, selected DaCapo applications, additional browsers/toolchains. | User-visible per-module time/memory/disk ceilings; an initial 15–30-minute application allocation needs calibration on low-end physical hardware. |
|
||||
|
||||
Preparation/download/compilation costs are shown separately unless compilation is the named workload. Check tools and resource requirements before beginning. Do not shorten an upstream test behind the user’s back to meet a deadline: mark it incomplete and preserve diagnostics.
|
||||
|
||||
Missing interpreter/compiler/browser/driver, unsupported ABI, unavailable display, insufficient resources, timeout, and failed output validation are different outcomes. None is a zero performance measurement. Skips reduce **capability coverage**; the score-design ticket must determine eligibility for an aggregate. Optional-tool absence should leave completed results usable. No hardware/browser/type of VM was executed here to substantiate duration estimates.
|
||||
|
||||
## Remaining decisions and evidence gaps
|
||||
|
||||
Both browser modes are already approved. Select the default mode for each profile; the native-browser/driver and reference-browser support matrix; exact upstream full-suite versus Odin-probe identities; language subsets; installed/reference toolchain policy; fixed compilation/warmup rules; and measured profile budgets. Browser automation overhead, all-assets-offline closure, per-asset redistribution inventory, selected-suite behavior on musl/aarch64 and newer JDKs, and correctness/variance qualification remain future proof work. An application score must not conceal these unresolved cohort boundaries.
|
||||
|
||||
## Sources and method
|
||||
|
||||
Primary sources inspected **2026-09-25**. Context7 resolution preceded queries for Speedometer/WebKit, Bash, Hyperfine, Playwright, pyperformance/pyperf, Rust, LLVM, JMH, GCC, and Selenium. Exact Speedometer and pyperformance lookups returned unrelated entries; those were rejected and their own repositories inspected. Queries stayed within the three-command limit per lookup; no quota error occurred.
|
||||
|
||||
1. [BrowserBench current index](https://browserbench.org/).
|
||||
2. Speedometer release source: [commit](https://github.com/WebKit/Speedometer/commit/1386415be8fef2f6b6bbdbe1828872471c5d802a), [parameters](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/params.mjs), [result client](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/main.mjs), [package metadata](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/package.json), [local hosting](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Development.md).
|
||||
3. [Speedometer 3.1 instructions](https://browserbench.org/Speedometer3.1/instructions.html), [workload descriptions](https://browserbench.org/Speedometer3.1/about.html), [measurement/scoring explanation](https://github.com/WebKit/Speedometer/blob/main/README.md).
|
||||
4. [JetStream 3.0 implementation page](https://browserbench.org/JetStream3.0/), [in-depth methodology and provenance](https://browserbench.org/JetStream3.0/in-depth.html).
|
||||
5. [MotionMark 1.3.2 implementation page](https://browserbench.org/MotionMark1.3.2/), [methodology, display requirements, and version history](https://browserbench.org/MotionMark1.3.2/about.html).
|
||||
6. [Speedometer release license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE).
|
||||
7. [Speedometer browser-driver testing documentation](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Testing.md), [Selenium Manager: explicit drivers, offline mode, and architecture limitations](https://www.selenium.dev/documentation/selenium_manager/).
|
||||
8. [Playwright system requirements](https://playwright.dev/docs/intro), [browser distributions and headless modes](https://github.com/microsoft/playwright/blob/main/docs/src/browsers.md), [musl limitations](https://github.com/microsoft/playwright/blob/main/docs/src/docker.md).
|
||||
9. [Chrome unified headless and headless-shell distinction](https://developer.chrome.com/docs/chromium/headless).
|
||||
10. [GNU Bash startup files](https://www.gnu.org/software/bash/manual/html_node/Bash-Startup-Files.html), [invocation options](https://www.gnu.org/software/bash/manual/html_node/Invoking-Bash.html).
|
||||
11. [Hyperfine README: intermediate-shell correction, direct mode, warmups, repetitions](https://github.com/sharkdp/hyperfine/blob/master/README.md).
|
||||
12. [pyperformance purpose and license](https://github.com/python/pyperformance/blob/main/README.rst), [PyPI release metadata](https://pypi.org/pypi/pyperformance/json), [pyperf release metadata](https://pypi.org/pypi/pyperf/json), [benchmark manifest](https://github.com/python/pyperformance/blob/main/pyperformance/data-files/benchmarks/MANIFEST), [workload descriptions](https://github.com/python/pyperformance/blob/main/doc/benchmarks.rst), [dependency/runtime guidance](https://github.com/python/pyperformance/blob/main/doc/usage.rst), [startup implementation](https://github.com/python/pyperformance/blob/main/pyperformance/data-files/benchmarks/bm_python_startup/run_benchmark.py).
|
||||
13. [rustc-perf purpose/licenses](https://github.com/rust-lang/rustc-perf/blob/master/README.md), [compile suite](https://github.com/rust-lang/rustc-perf/blob/master/collector/compile-benchmarks/README.md), [experimental runtime suite](https://github.com/rust-lang/rustc-perf/blob/master/collector/runtime-benchmarks/README.md), [runtime groups](https://github.com/rust-lang/rustc-perf/tree/master/collector/runtime-benchmarks), [collector requirements](https://github.com/rust-lang/rustc-perf/blob/master/collector/README.md).
|
||||
14. [LLVM test-suite guide](https://llvm.org/docs/TestSuiteGuide.html), [testing/output-validation model](https://llvm.org/docs/TestingGuide.html), [serial miniFE description](https://github.com/llvm/llvm-test-suite/blob/main/MultiSource/Benchmarks/DOE-ProxyApps-C++/miniFE/README), [license and third-party exceptions](https://github.com/llvm/llvm-test-suite/blob/main/LICENSE.TXT).
|
||||
15. [JMH purpose and cautions](https://github.com/openjdk/jmh/blob/master/README.md), [published version metadata](https://repo.maven.apache.org/maven2/org/openjdk/jmh/jmh-core/maven-metadata.xml), [parameter/warmup/fork example](https://github.com/openjdk/jmh/blob/master/jmh-samples/src/main/java/org/openjdk/jmh/samples/JMHSample_27_Params.java), [result consumption and profilers](https://github.com/openjdk/jmh/blob/master/jmh-samples/src/main/java/org/openjdk/jmh/samples/JMHSample_35_Profilers.java), [license](https://github.com/openjdk/jmh/blob/master/LICENSE).
|
||||
16. [DaCapo latest release](https://github.com/dacapobench/dacapobench/releases/tag/v23.11-MR2-chopin), [release compatibility and workload notes](https://github.com/dacapobench/dacapobench/blob/0db32562cf169730c163d88df2eeb28217ca7d03/benchmarks/RELEASE_NOTES.md), [purpose/reporting/license guidance](https://github.com/dacapobench/dacapobench/blob/master/README.md).
|
||||
17. [pyperf runner configuration](https://github.com/psf/pyperf/blob/main/doc/runner.rst), [process/calibration architecture](https://github.com/psf/pyperf/blob/main/doc/run_benchmark.rst).
|
||||
18. [Cargo profile defaults](https://doc.rust-lang.org/cargo/reference/profiles.html), [rustc code-generation options](https://doc.rust-lang.org/rustc/codegen-options/index.html), [GCC optimization semantics](https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html).
|
||||
@@ -0,0 +1,115 @@
|
||||
# GPU driver evidence and performance workloads
|
||||
|
||||
Research for [Establish GPU driver evidence and performance workloads](https://git.bongbetic.com/xavierk/odin/issues/3). All sources were inspected on **2026-09-25**. These are conditional recommendations for the workload decision, not an adopted suite or support guarantee. No drivers were installed, settings changed, device queries executed, or benchmarks run.
|
||||
|
||||
## Decision summary
|
||||
|
||||
Odin can establish **which device and API worked for a specified operation under the current session**. It cannot certify one universally “correct” driver from a package name, loaded module, advertised API version or benchmark score. Keep discovery, successful execution, output validation, presentation and performance as distinct evidence.
|
||||
|
||||
The strongest initial options are a small **headless Vulkan compute profile** using a qualified clpeak build, plus **API-specific rendering workloads** where supported. vkmark and glmark2 offer useful scenes, but their backend requirements and relatively infrequent releases need qualification. No candidate alone measures compute throughput, rendering, compositor behavior, video acceleration and browser usability.
|
||||
|
||||
Visible GPU windows remain a decision for **Choose Odin’s workload suite and run profiles**. The user's approval of headed/headless browser modes does not settle GPU presentation. Terminal orchestration can support headless workloads without opening a window; it cannot thereby prove the desktop presentation path works.
|
||||
|
||||
## 1. Evidence to collect per device
|
||||
|
||||
| Stage | Evidence | Justified conclusion |
|
||||
| --- | --- | --- |
|
||||
| Hardware and kernel path | DRM/sysfs device, associated PCI or platform identity, bound kernel driver, accessible render node | Device and kernel path are present; userspace API operation is still unproven |
|
||||
| API discovery | Vulkan physical-device properties/features/queues; OpenGL vendor/renderer/version from the actual context; OpenCL platform/device type | This implementation advertises these capabilities in this environment |
|
||||
| Execution | Bounded allocation, command submission, completion and error outcome on the selected device | The tested operation completed; report device loss, allocation failure and timeout distinctly |
|
||||
| Correctness | A known-output compute or rendering check with a defined tolerance | The sampled operation returned an expected result; throughput alone does not supply this proof |
|
||||
| Presentation | Successful creation and presentation to a particular X11/Wayland surface, if this mode is selected | That surface/session path worked; headless success is a different finding |
|
||||
|
||||
Linux render nodes permit non-global rendering without DRM-master authentication, subject to ordinary filesystem permissions. They do not grant modesetting rights. Treat denied access as capability coverage, not “bad GPU,” and do not respond by elevating the entire TUI or changing device permissions. DRM discovery must include platform devices: ARM GPUs need not be PCI devices. [1]
|
||||
|
||||
Vulkan properties include device type, vendor/device identifiers, device UUID, and driver identification. `deviceType=CPU` identifies a typically host-processor implementation; the specification calls device type informational, so combine it with driver/renderer evidence. `driverVersion` is vendor-specified, not universal semantic versioning. `conformanceVersion` describes the implementer's prior conformance testing, not a test of this installation. Where supported, `VK_EXT_physical_device_drm` connects API devices to DRM node major/minor numbers. [2]
|
||||
|
||||
**vulkaninfo** is a useful discovery candidate. Its `--summary` covers enumerated devices; `--json=<index>` writes a Vulkan Profiles JSON file for one device. Plain `--json` defaults to the first device, so one successful invocation does not inventory every GPU. Use a private output directory and record the tool/schema version. The inspected SDK tag is `vulkan-sdk-1.4.357.0`, with Apache-2.0 project licensing. [3]
|
||||
|
||||
Mesa LLVMpipe/Softpipe are software renderers. Zink is an OpenGL implementation over Vulkan and can use a hardware Vulkan driver: the word “Mesa” or “Zink” is not evidence of software rendering. Mesa and the Vulkan loader also expose selection/override variables, including `LIBGL_ALWAYS_SOFTWARE`, `DRI_PRIME`, `MESA_VK_DEVICE_SELECT` and `VK_DRIVER_FILES`. Record relevant effective overrides and selected devices; do not silently change the user's stack during baseline measurement. Vulkan/OpenCL CPU devices and mock drivers must not contribute a hardware-GPU score. [4][5][6]
|
||||
|
||||
## 2. Driver and architecture scope
|
||||
|
||||
| Hardware family | Paths worth supporting conditionally | Boundary |
|
||||
| --- | --- | --- |
|
||||
| Intel | Appropriate Linux kernel driver plus Mesa OpenGL/ANV; an independently available compute runtime | Working OpenGL does not establish Vulkan or OpenCL support. Discover generation-specific capabilities rather than prescribe one package universally |
|
||||
| AMD | Supported kernel/userspace combination, commonly amdgpu plus Mesa RADV for Vulkan | RADV and ROCm serve different purposes. ROCm has its own hardware/OS/firmware compatibility matrix; absence of ROCm does not mean ordinary graphics is broken |
|
||||
| NVIDIA | NVIDIA's supported userspace/kernel stack, or Mesa NVK and applicable OpenGL path | NVK is a legitimate Vulkan implementation. NVIDIA's open kernel modules still require matching NVIDIA userspace and GSP firmware; “open module installed” does not establish compatibility |
|
||||
| ARM SoCs | Panfrost/PanVK for supported Mali, Freedreno/Turnip for supported Adreno, other model-specific Mesa/vendor paths | aarch64 names the CPU architecture, not the GPU API capability. Some devices support GLES without Vulkan; experimental support must not be force-enabled automatically |
|
||||
|
||||
Mesa documents RADV's separation from the kernel driver and hardware limitations; Panfrost lists distinct API support by GPU and explicitly warns about experimental PanVK enablement. NVIDIA's inspected `615.71.09` open-module release supports x86_64/aarch64 and Turing-or-later hardware, with corresponding-release userspace/firmware requirements. These examples justify capability probing, not a universal driver recommendation. [7][8][9]
|
||||
|
||||
Qualify **x86_64/glibc, x86_64/musl, aarch64/glibc and aarch64/musl** independently for the chosen executable and transitive libraries. Source availability does not certify a binary across that matrix. Void explicitly states proprietary NVIDIA drivers do not support musl; packaging Odin differently cannot erase that driver limitation. Mesa-based paths can be candidates where that GPU and distribution support them. Current ROCm support is also a specific matrix, not a promise for every Linux distribution or libc. [8][10]
|
||||
|
||||
## 3. Workload candidates and concrete tradeoffs
|
||||
|
||||
| Candidate and inspected version | Measurements, footprint and control | Conditional role |
|
||||
| --- | --- | --- |
|
||||
| **clpeak 2.1.4**, Apache-2.0; released August 27, 2026 | Current code supports Vulkan, OpenCL, CUDA, ROCm/HIP, oneAPI and CPU, among others. CLI has backend/device/test selection and JSON/CSV/XML output. `--max-time` controls each GPU test's timed phase; warmup/calibration add time. C++17/CMake; SDKs/backends are optional but auto-detected by default | Strong first candidate for a deliberately restricted CLI build and selected Vulkan FP32/bandwidth workloads. Pin enabled backends, shaders and compiler; avoid its “run every backend/device/test” default |
|
||||
| **vkpeak 20260527**, MIT; source activity in August 2026 | Vulkan peak scalar/vector/matrix arithmetic and transfer tests using ncnn. Select device and scenarios. Small top-level program, substantial transitive shader/runtime dependency. No user time-budget option is documented in the inspected CLI | Alternative focused compute candidate. Its README explicitly says peak metrics do not represent real-world use. Source returns zero for some unsupported features **and failures**, so zero cannot be interpreted as measured zero performance |
|
||||
| **vkmark 2025.01**, LGPL-2.1-or-later | Configurable Vulkan rendering scenes, dimensions, present mode, duration and device UUID selection. C++17, Vulkan, GLM and Assimp; optional XCB/Wayland/DRM/GBM dependencies | Candidate graphics profile after backend qualification. The released source includes a headless plugin requiring `VK_EXT_headless_surface`; its manpage backend list omits that plugin. Generic Vulkan support alone is insufficient |
|
||||
| **glmark2 2023.01**, GPLv3 | OpenGL 2.0/GLES2 scenes; per-scene duration, off-screen mode, frame-end/swap controls, output validation and CSV/XML results. Build flavors include X11, Wayland, DRM and GBM; GL/EGL/GLES and image libraries/assets | Useful compatibility and rendering candidate; its older API workloads are not a complete modern-GPU assessment. GBM source can use a selected render node; `--off-screen` on an X11 build does not imply display-server independence |
|
||||
|
||||
Primary released READMEs, manuals, licenses and implementation sources support this comparison. glmark2's latest inspected tag remains 2023.01 with main activity in September 2025; vkmark's latest tag is 2025.01 with main activity in September 2025. These are maturity/maintenance observations, not evidence of current hardware certification. clpeak and vkpeak show more recent source/release activity, but still require qualification. [11–14]
|
||||
|
||||
Two implementation traps matter immediately:
|
||||
|
||||
- clpeak's Vulkan instance requests Vulkan 1.0 or 1.1 depending on compiled optional features. Its reported capability floor therefore depends on the build. Its timing code performs warmup and calibration before the timed batch; `--max-time` is not an end-to-end timeout. Its Vulkan backend distinguishes CPU and integrated/discrete GPU device types. [11]
|
||||
- vkpeak adapts work and reports peak results, with memory sizing based partly on device heap information. A selected subset is more controllable than its complete default suite, but a wrapper still needs independent resource and runtime bounds. Neither tool's advertised throughput proves it checks the numerical result required by Odin's correctness stage. [12]
|
||||
|
||||
Licenses above describe inspected project code. Bundling requires a separate manifest for assets, embedded dependencies, modifications and any vendor runtime redistribution terms; these source inspections are not a completed distribution-license audit. Installing a large CUDA/ROCm SDK solely to enable baseline benchmarking would weaken the universal deployment objective. Vendor-specific compute paths are better considered optional capability profiles.
|
||||
|
||||
## 4. Headless, desktop, multiple GPUs and virtualization
|
||||
|
||||
Keep three execution classes distinct: **surface-free compute**, **offscreen rendering**, and **desktop presentation**. Vulkan does not require every physical device or queue to support presentation. Support must be queried for the actual surface. FIFO presentation waits on vertical blanking; an FPS result can therefore reflect display/compositor policy rather than maximum render throughput. Fix and record present mode, resolution and backend. [15]
|
||||
|
||||
vkmark's headless plugin still uses a Vulkan surface/swapchain extension; glmark2's GBM backend opens a render node and creates a GBM surface. These are different requirements and workloads. A KMS/direct-display backend may need display ownership and disturb the session, so it is not an automatic fallback when X11/Wayland fails. An SSH terminal can have usable GPU compute without a display socket; classify presentation as unavailable in that session rather than infer a missing graphics driver. [1][13][14]
|
||||
|
||||
Enumerate all devices, map them to stable identifiers where available, and let the run profile select the display GPU, another named GPU or separate per-GPU runs. Do not treat index zero as “best GPU.” Mesa's selection variables can reorder enumeration; vkmark's UUID selector and NVIDIA's documented UUID/PCI-ID selection illustrate stronger identity mechanisms. Avoid summing overlapping APIs or independently averaging all installed GPUs into one unexplained number. [2][5][9][13]
|
||||
|
||||
Virtual hardware needs its own label. Mesa Venus serializes Vulkan through virtio-gpu to a host renderer and can operate over hardware **or Lavapipe**; guest enumeration does not establish physical passthrough. A VM result measures that guest path, while its host telemetry may be hidden. Passthrough must be established from the recorded environment and device evidence. Containers similarly need device access and compatible userspace libraries; missing exposure is not proof the host has no GPU. [6][16]
|
||||
|
||||
## 5. Reproducibility, safety and diagnosis
|
||||
|
||||
Before scoring, freeze workload version, scene/kernel code, input size, precision/vector width, backend, device, output format, build flags and compiler. Record kernel, userspace driver identity, power source, thermal state, display mode and concurrent load. Warmup/cache policy must be explicit. Compare repeated runs under the same policy; do not compare shader compilation included in one result with warmed execution in another.
|
||||
|
||||
Compute FLOPS, transfer bandwidth and scene FPS answer different questions. A CUDA FP16 matrix peak cannot replace a Vulkan FP32 score; software rendering cannot replace the hardware result; unavailable features are not zeros. Preserve upstream metrics and chosen aggregation rules. glmark2/vkmark aggregate FPS does not automatically supply frame-time percentiles. Small rendering scenes can also be CPU/driver limited, so unexpectedly low throughput is evidence to investigate, not automatic proof of defective hardware.
|
||||
|
||||
Recommended safety boundaries are one GPU workload at a time, bounded memory demand with system/VRAM reserve, an independent wall-clock watchdog, staged process-group cancellation, and refusal of unbounded “run forever” modes. Include initialization/calibration in the overall budget. Stop on device loss, repeated API errors or meaningful driver-reported critical thermal findings; retain partial results. Killing a client cannot guarantee immediate recovery from a kernel/driver hang. Do not automatically overclock, alter fan/power limits, reset a GPU or replace drivers. Resource limits and safe cancellation remain acceptance tests for the selected workload.
|
||||
|
||||
Telemetry is supporting evidence, with provider-specific meaning:
|
||||
|
||||
- NVIDIA `nvidia-smi` documents unsupported values as `N/A`, separate errors for permission denial, unloaded driver and missing NVML, and stable UUID/PCI selection. Its utility success does not test Vulkan/OpenGL presentation. Read-only query adapters must preserve unavailable/error outcomes. [9]
|
||||
- amdgpu exposes temperature, load, power and other sysfs metrics, but support varies. Its APU power reading includes CPU power, so it is not interchangeable with discrete-GPU-only power. [17]
|
||||
- Linux DRM fdinfo defines per-client engine-busy counters, capacities and accounting rules; availability depends on the driver and accessible process descriptors. These can help attribute work without assuming one vendor's utilization meaning applies everywhere. [18]
|
||||
|
||||
Advice should name the evidence: “Vulkan userspace driver could not load,” “render-node access denied,” “software renderer selected,” “this feature is unsupported,” or “workload lost the device.” A package/version mismatch needs concrete loader or vendor evidence; a successful fallback may be intentional. Present a distro-appropriate investigation step with confidence and tradeoffs, rather than an unconditional “install proprietary drivers.”
|
||||
|
||||
## 6. Decisions and limits carried forward
|
||||
|
||||
The workload decision must choose: headless compute and graphics requirements; whether visible GPU presentation is allowed; selected tool/build and minimum API features; treatment of software/virtual/unsupported paths; default multi-GPU selection; memory/runtime budgets; and the correctness check and scoring eligibility rules.
|
||||
|
||||
Before a supported release, qualify the four architecture/libc lanes, real Intel/AMD/NVIDIA hardware, representative ARM SoCs, X11/Wayland/headless sessions, multiple GPUs, denied permissions, software rendering, VM acceleration/passthrough, and cancellation/device-loss fixtures. No such execution evidence was produced here. No package-size estimates or full vendor conformance/redistribution audit were established.
|
||||
|
||||
Context7 resolution preceded documentation lookup. Its glmark2 queries yielded no relevant main-project documentation; clpeak was unindexed; vkpeak resolved to an unrelated speech tool and was rejected; NVIDIA NVML searches produced unrelated/wrapper results. Official tagged source and vendor documentation supplied those gaps. The guessed Mesa Lavapipe page returned 404; software/virtual-path claims use inspected Mesa driver documentation and API/device evidence instead. Current Mesa pages contain evolving and occasionally differing generation summaries, so no exhaustive model support table is inferred from them.
|
||||
|
||||
## Sources
|
||||
|
||||
1. [Linux DRM userspace API, render nodes](https://docs.kernel.org/gpu/drm-uapi.html).
|
||||
2. [Vulkan device/queue specification](https://github.com/KhronosGroup/Vulkan-Docs/blob/main/chapters/devsandqueues.adoc), physical-device, driver, UUID and DRM properties.
|
||||
3. Vulkan Tools SDK tag: [vulkaninfo documentation](https://github.com/KhronosGroup/Vulkan-Tools/blob/vulkan-sdk-1.4.357.0/vulkaninfo/vulkaninfo.md), [license](https://github.com/KhronosGroup/Vulkan-Tools/blob/vulkan-sdk-1.4.357.0/LICENSE.txt).
|
||||
4. Mesa [platforms/drivers](https://docs.mesa3d.org/systems.html), [LLVMpipe](https://docs.mesa3d.org/drivers/llvmpipe.html), [Zink](https://docs.mesa3d.org/drivers/zink.html).
|
||||
5. [Mesa environment variables](https://docs.mesa3d.org/envvars.html), [Vulkan loader driver discovery](https://github.com/KhronosGroup/Vulkan-Loader/blob/main/docs/LoaderDriverInterface.md).
|
||||
6. [OpenCL device enumeration](https://github.com/KhronosGroup/OpenCL-Registry/blob/main/specs/unified/refpages/man/html/clGetDeviceIDs.html), [Vulkan Guide support/null-driver discussion](https://github.com/KhronosGroup/Vulkan-Guide/blob/main/chapters/checking_for_support.adoc), inspected through Context7.
|
||||
7. Mesa [ANV](https://docs.mesa3d.org/drivers/anv.html), [RADV](https://docs.mesa3d.org/drivers/radv.html), [NVK](https://docs.mesa3d.org/drivers/nvk.html), [Panfrost](https://docs.mesa3d.org/drivers/panfrost.html), [Freedreno/Turnip](https://docs.mesa3d.org/drivers/freedreno.html).
|
||||
8. [ROCm current compatibility matrix](https://rocm.docs.amd.com/en/latest/compatibility/compatibility-matrix.html).
|
||||
9. NVIDIA [open-module 615.71.09 README](https://github.com/NVIDIA/open-gpu-kernel-modules/blob/615.71.09/README.md), [nvidia-smi documentation](https://docs.nvidia.com/deploy/nvidia-smi/index.html).
|
||||
10. [Void musl compatibility](https://docs.voidlinux.org/installation/musl.html).
|
||||
11. clpeak 2.1.4: [README](https://github.com/krrishnarraj/clpeak/blob/2.1.4/README.md), [CLI options](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/common/options.cpp), [Vulkan timing/instance implementation](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/vulkan/vk_peak.cpp), [device mapping](https://github.com/krrishnarraj/clpeak/blob/2.1.4/src/vulkan/vulkan_device.cpp), [build options](https://github.com/krrishnarraj/clpeak/blob/2.1.4/CMakeLists.txt), [license](https://github.com/krrishnarraj/clpeak/blob/2.1.4/LICENSE).
|
||||
12. vkpeak 20260527: [README](https://github.com/nihui/vkpeak/blob/20260527/README.md), [implementation](https://github.com/nihui/vkpeak/blob/20260527/vkpeak.cpp), [build/dependency configuration](https://github.com/nihui/vkpeak/blob/20260527/CMakeLists.txt), [license](https://github.com/nihui/vkpeak/blob/20260527/LICENSE).
|
||||
13. vkmark 2025.01: [README](https://github.com/vkmark/vkmark/blob/2025.01/README.md), [manual](https://github.com/vkmark/vkmark/blob/2025.01/doc/vkmark.1), [headless implementation and license notice](https://github.com/vkmark/vkmark/blob/2025.01/src/ws/headless_native_system.cpp), [backend build](https://github.com/vkmark/vkmark/blob/2025.01/src/meson.build).
|
||||
14. glmark2 2023.01: [README/license declaration](https://github.com/glmark2/glmark2/blob/2023.01/README), [manual](https://github.com/glmark2/glmark2/blob/2023.01/doc/glmark2.1.in), [GBM implementation](https://github.com/glmark2/glmark2/blob/2023.01/src/native-state-gbm.cpp), [build flavors](https://github.com/glmark2/glmark2/blob/2023.01/meson_options.txt).
|
||||
15. [Vulkan WSI specification](https://github.com/KhronosGroup/Vulkan-Docs/blob/main/chapters/VK_KHR_surface/wsi.adoc), surface support, headless surfaces and present modes.
|
||||
16. [Mesa Virtio-GPU Venus](https://docs.mesa3d.org/drivers/venus.html).
|
||||
17. [Linux amdgpu thermal/power monitoring](https://docs.kernel.org/gpu/amdgpu/thermal.html).
|
||||
18. [Linux DRM usage-statistics ABI](https://docs.kernel.org/gpu/drm-usage-stats.html).
|
||||
Reference in New Issue
Block a user