2 changed files with 135 additions and 140 deletions
+135
View File
@@ -0,0 +1,135 @@
# Establish browser, shell, and language workload validity
Research for [Establish browser, shell, and language workload validity](https://git.bongbetic.com/xavierk/odin/issues/5), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
Access date: **2026-09-25**. This is a candidate analysis, not the final suite. No browsers, toolchains, or dependencies were installed; no benchmarks or builds were run.
## Recommendation for the selection decision
Preserve five distinct measurement purposes: browser application responsiveness, shell interaction, process startup, compilation, and warmed application execution. A single fast loop cannot represent all five. Retain upstream benchmark names, versions, metrics, and correctness rules; an Odin aggregate is a separate decision.
A practical initial option is **Speedometer 3.1**, a small Bash latency/throughput set, and optional language packs with pinned workload inputs. JetStream, MotionMark, representative compilation projects, and larger JVM applications fit extended investigation. Installed software answers “how usable is this machine as configured”; pinned reference software improves comparison across machines. Store these as distinct modes rather than treating them as interchangeable.
## 1. Browser candidates and official measurement rules
BrowserBench’s live index currently points to **Speedometer 3.1**, **JetStream 3.0**, and **MotionMark 1.3.2**. [1]
| Candidate | What it measures | Conditional role and interpretation |
| --- | --- | --- |
| Speedometer 3.1 | End-to-end simulated web-app interactions: TodoMVC frameworks, editors, charts, and news applications. | Strong default browser-responsiveness candidate. Preserve the full default suite and iteration protocol for an upstream-style result. It does not measure browser process startup or network loading. |
| JetStream 3.0 | JavaScript/WebAssembly workloads, combining startup, average, and worst-case execution behavior. | Useful extended engine-compute view. “Startup” is workload first-iteration behavior, not opening the browser. Its 77 default workloads and upstream geometric aggregation should remain intact. |
| MotionMark 1.3.2 | Browser graphics scenes adjusted toward a target frame rate, with confidence intervals. | Useful optional browser-rendering view where a suitable graphics/display path exists. It is not a standalone GPU benchmark or evidence that the intended hardware driver is active. |
Speedometer’s release/3.1 source currently resolves to commit `1386415be8fef2f6b6bbdbe1828872471c5d802a`. Its `package.json` still says `3.0.0-alpha`, demonstrating why package metadata alone is not a benchmark identity. The release source defaults to **10 iterations** and an **800×600 workload viewport**; URL parameters can change iterations, suites, warmup, and measurement behavior. The result includes individual measurements and JSON exports. Preserve these settings and raw results; do not silently shorten the suite and report an ordinary “Speedometer 3.1” score. [2]
Official Speedometer instructions recommend a latest stable browser, a separate clean profile, closed background applications/tabs, a focused benchmark page, AC power, no interaction, and cooling between runs where needed. A pinned historical browser is useful for longitudinal comparison but must be identified as a reference-browser experiment. Speedometer aggregates inverse geometric-mean durations, then averages scores across iterations and reports a 95% confidence interval around that mean. Retain that native statistic and uncertainty. Any Odin median across complete benchmark runs is a separate statistic; do not relabel either or average heterogeneous subtest times. [2][3]
JetStream explicitly says scores are not comparable across its versions. Its startup, worst-case, and average components are deliberately part of the metric; generic harness warmups must not erase those startup observations. MotionMark depends on drawing-area class, viewport, refresh behavior, and browser rendering. Its instructions recommend maximizing the window at the default resolution and 60 Hz for comparison across browsers; a changed display policy is a recorded experimental condition, not something Odin should silently impose. [4][5]
### Offline hosting and licensing
Speedometer documents a local static HTTP server, so a prepared source/assets snapshot can be served on loopback without fetching the public site during measurement. Its root license permits source/binary redistribution with the copyright, conditions, and disclaimer retained. That does not replace the license notices of bundled frameworks, fonts, datasets, or other third-party assets. JetStream and MotionMark’s deployed HTML also carries permissive two-condition notices, while their workload provenance points to multiple projects. Before distributing a pack, inventory the exact selected files and retain their individual notices. [2][4][5][6]
Recommended preparation contract: content-addressed assets, release/commit identity, dependency/license manifest, complete local resource closure, and a future offline-network audit. Bind the server to loopback and prepare everything before timing. A copied public page without its subresources is not a verified offline benchmark. Keep upstream workload code unchanged where feasible; record every adapter patch. The existence of localhost support does not prove the entire pack is self-contained until that artifact is checked.
### Automation and the Linux boundary
Two viable routes serve different purposes:
1. **Native installed browser plus matching WebDriver.** Speedometer itself documents Selenium testing with ChromeDriver, GeckoDriver, and other browser drivers. This route can use distro-supported native browsers on Void/musl or other systems outside Playwright’s matrix, provided the actual browser/driver pair exists and is qualified. Use explicit browser and driver paths and versions. Selenium Manager can download software by default; its documented offline mode disables network requests/downloads. Its documentation still describes limited Linux architecture support, so automatic management cannot be presumed to cover aarch64 or musl. Native automation remains a capability to prove, not a promise. [7]
2. **Pinned browser supplied with Playwright on its supported platforms.** Current documentation lists Debian 12/13 and Ubuntu 22.04/24.04/26.04 on x86-64/arm64. Its bundled Firefox/WebKit builds require glibc; musl distributions are unsupported. Stock Firefox cannot simply be substituted because Playwright requires its patched Firefox. Branded Chrome/Edge support and custom executable paths also do not establish universal compatibility. [8]
**Odin can launch on a system even when its browser module is unavailable.** On musl, use a qualified native browser/driver pair or report an explicit capability skip. A glibc container/chroot or software-rendered browser is a separate environment and cannot transparently stand in for native desktop usability.
The user has approved **both headed and headless browser modes**, with terminal controls/results and distinct measurement labels. Chrome’s unified headless mode shares browser code with headed Chrome, but Playwright also provides a distinct headless-shell binary and documents behavioral differences. Sharing code does not establish equal scores, display scheduling, GPU access, or compositor behavior. Label a headless cohort explicitly and avoid claiming compliance with focused-window desktop run conditions. A headed cohort needs a graphical session and the official run conditions. Do not merge their reference distributions or invent a conversion factor. The default mode for each run profile remains a selection decision. [8][9]
Automation should start the pinned page, observe completion, and collect its full result without repeated polling/instrumentation inside timed work. Correctness failures, missing workloads, or timeouts invalidate a full-suite result. Shortened or filtered runs can be Odin browser probes with their own identities; they are not substitutes for official full-suite statistics.
Record browser binary/version/channel, driver and automation versions, full flags, profile policy, headed/headless/shell distinction, viewport/device scale, display resolution/refresh, X11/Wayland/virtual display, renderer and acceleration evidence when exposed, kernel/libc/architecture, power state, suite digest, and all query parameters. An unavailable GPU detail is unknown, not proof of hardware acceleration. Browser startup, if desired, should be a separate process-to-ready-page test with its cache/profile policy declared.
## 2. Meaningful Bash and shell measurements
Bash startup depends on invocation. Login shells read `/etc/profile` and the first readable user login profile; interactive non-login shells read `.bashrc`; noninteractive shells may execute the file named by `BASH_ENV`. `--noprofile` and `--norc` control different paths. Therefore “time bash” is underspecified. [10]
| Candidate | Measurement definition | Boundary to preserve |
| --- | --- | --- |
| Clean process startup | Directly launch a known Bash binary with controlled environment/startup-file policy and a trivial command; measure process launch through exit. | This includes loader/process/exit costs. It is not prompt readiness. Explicitly control `BASH_ENV`, inherited functions, locale, and working directory. |
| Interactive readiness | Launch a clean interactive shell under a PTY; measure until a declared prompt-ready marker, then measure prompt return after a builtin command. | Account for PTY setup and marker instrumentation. `bash -i -c exit` does not observe a real prompt. Test login and non-login separately if both are offered. |
| Builtin workload throughput | Fixed arithmetic, parameter expansion, arrays, and parsing over seeded in-memory inputs, returning a verified checksum. | Bounded output; no per-iteration `date`, `cat`, `grep`, or other subprocesses. Input size and Bash version are part of the workload identity. |
| External-command/pipeline workload | Fixed commands, input corpus, and pipe structure, with exact executable versions. | Measures the combined shell, process-launch, tool, and I/O stack. Report it as such. |
| User configuration readiness | Opt-in observation using the user’s real startup files and prompt configuration. | These files execute arbitrary user-configured commands, may contact networks or change state, and may not be repeatable. Report local usability separately from the clean reference result. |
**Harness trap:** Hyperfine defaults to an intermediate `/bin/sh` on Unix and calibrates/subtracts shell-spawn time. For Bash startup, use its direct-command mode or another direct process timer so the Bash process is the measured payload. Hyperfine also supports warmups and repetitions; warmed filesystem caches are a declared condition, not “cold boot.” Do not globally drop caches or alter system tuning just to manufacture a shell metric. [11]
Treat time-to-first-prompt and prompt-to-next-prompt as distributions, not a single best sample. Bound captured PTY output, support cancellation, and keep configured-shell timeout/error details. Store configuration fingerprints with care; benchmark records should not copy potentially secret startup-file contents.
## 3. Language workload candidates
Each result describes a **workload + implementation + toolchain + environment**, not an intrinsic language ranking. Numeric kernels are useful but should not be the only evidence. Libraries, parsing, allocation, and application behavior matter; Python’s C-backed JSON module and a Rust regex library do not represent the same implementation simply because both are called language tests.
| Language / candidate | Useful workload selection | Evidence, tradeoffs, and role |
| --- | --- | --- |
| Python: pyperformance + pyperf | `python_startup`, `python_startup_no_site`, `json_loads`/`json_dumps`, one regex workload, and optionally a pure-Python workload such as `richards`. | Maintained Python project favors real applications. Current PyPI metadata: pyperformance **1.14.0**, Python ≥3.10; pyperf **2.10.0**, Python ≥3.9. Suite docs contain older requirements, so prefer release metadata. MIT project license; inspect selected workload notices. Small subsets fit standard profiles. [12] |
| Rust: rustc-perf subsets | Runtime groups include parsing, text search, compression, hash maps, and numeric/graphics kernels; compile candidates include pinned real crates such as `syn`. | Official compiler-performance project clearly separates compilation from generated-program execution. Runtime suite is explicitly **experimental**, with changing workloads; pin a commit and validate chosen outputs. MIT infrastructure, separate compile-benchmark licenses. Avoid adopting the full collector by default: it has profiling/environment dependencies. [13] |
| C++: LLVM test-suite subset | A small C++ application/proxy workload such as the serial miniFE variant, plus a smaller verified workload; optionally time compiling the same pinned sources. | Suite compiles/runs whole programs, checks reference output, and records compile and execution times separately. CMake/lit and dependencies increase setup cost. Apache-2.0 with LLVM exceptions for project code; third-party directories have separate terms. Good standard/extended candidate after footprint qualification. [14] |
| Java: selected workloads under JMH | A packaged, reviewed data-processing/allocation workload with validated results and configured warmup/forks. | JMH **1.37** is the current published Maven release inspected. It is a harness, not a representative suite by itself; samples explain pitfalls rather than defining an Odin score. GPLv2 with the Classpath exception on designated files. Suitable where a small, controlled JVM workload is desired. [15] |
| Java: DaCapo application subset | `lusearch` for search, `h2` for database-style execution, or another justified application. | Latest inspected release **23.11-MR2-chopin**; application suite with nontrivial memory use and output validation. Release notes establish Java 11–21 compatibility, not arbitrary newer JDKs. Harness Apache-2.0; component programs retain their licenses. Better extended candidate than a compulsory quick test. [16] |
pyperformance explicitly says it is not tuned for PyPy. CPython, PyPy, free-threaded builds, optional JIT builds, and distro build choices require distinct metadata and qualification. Interpreter startup with and without `site` measures different initialization; preserve both names if offered. Some pyperformance dependencies have native extensions, so a Python package being source-available does not prove wheel availability on every musl/architecture combination. [12]
## 4. Fair measurement protocol
**Startup, compilation, and warmed execution need separate records.** For Rust/C++, compile once before execution measurements unless compilation is the workload. For Java, bytecode compilation with `javac` differs from runtime JIT work. For Python, process startup/imports differ from repeated execution in an initialized interpreter. If the workload intentionally includes setup, say so and apply the same rule on every system.
pyperf uses calibration, multiple processes, warmups, values, and metadata. Its fast mode explicitly trades accuracy for speed. JMH provides forks, warmup/measurement iterations, state setup, and result consumption to avoid dead-code elimination; its maintainers warn that a harness does not eliminate benchmarking mistakes. Retain cold/startup samples and warmed samples separately. Use predetermined warmup/measurement rules and report instability; do not keep warming until a favorable result appears. [15][17]
Proposed compilation policy for selection:
- Pin sources, datasets, dependency lockfiles, language standard/edition, compiler identity, target triple, optimization flags, linker, libc, and library versions.
- Use a documented release optimization configuration. Cargo’s release defaults include `opt-level=3`, incremental off, and 16 codegen units; its bench profile inherits release. Record overrides and dependency profiles. C++ should use an explicitly chosen optimization level and standard; `-Ofast` changes standards-compliance assumptions and cannot silently replace a strict floating-point contract. [18]
- Keep baseline-ISA builds and host-tuned builds distinct. `target-cpu=native`/`-march=native`, SIMD dispatch, LTO, PGO, allocator choice, and thread counts can materially change results. A single binary is not portable across x86_64 and aarch64; equal flags do not mean equal generated instructions.
- Compilation tests distinguish clean builds, no-op rebuilds, and controlled incremental edits. Pre-fetch dependencies and hold build parallelism/cache policy fixed. Network resolution is not compilation speed; compiler caches must be controlled or explicitly measured.
- Validate outputs/checksums and numerical tolerances. Prevent constant folding/dead work, but avoid timing validation if the workload definition excludes it. Do not compare implementations with different precision, algorithms, data sizes, or hidden thread counts under a shared metric name.
An **installed-toolchain mode** best reveals current developer experience. A **reference-toolchain mode** uses pinned, prepared artifacts for stronger cross-machine comparison, with acquisition size, libc/ISA support, licenses, and security updates owned explicitly. It must not install into or replace the user’s default toolchain during a run. Report compiler/runtime version changes as changes in execution conditions, not unexplained hardware improvements.
## 5. Profile options, budgets, and unavailable tools
The following are provisional **application-domain budgets**, not measured runtime promises or the whole Odin run duration:
| Profile | Candidate scope | Budget policy to evaluate |
| --- | --- | --- |
| Quick | Browser capability probe; clean Bash startup and one builtin workload; startup/small verified work on selected installed languages. | Approximately 30–60 seconds for this domain. No full browser score if the official suite does not fit. Short observations carry lower-confidence status. |
| Standard | Full Speedometer 3.1 on one qualified browser; Bash set; selected language pack(s) with normal warmup/repetitions. | Reserve minutes rather than seconds; a pilot target is 2–5 minutes for browser and 1–3 minutes per selected language pack. Slow hosts may exceed it. |
| Extended | JetStream, optional MotionMark, more process forks, compilation workloads, selected DaCapo applications, additional browsers/toolchains. | User-visible per-module time/memory/disk ceilings; an initial 15–30-minute application allocation needs calibration on low-end physical hardware. |
Preparation/download/compilation costs are shown separately unless compilation is the named workload. Check tools and resource requirements before beginning. Do not shorten an upstream test behind the user’s back to meet a deadline: mark it incomplete and preserve diagnostics.
Missing interpreter/compiler/browser/driver, unsupported ABI, unavailable display, insufficient resources, timeout, and failed output validation are different outcomes. None is a zero performance measurement. Skips reduce **capability coverage**; the score-design ticket must determine eligibility for an aggregate. Optional-tool absence should leave completed results usable. No hardware/browser/type of VM was executed here to substantiate duration estimates.
## Remaining decisions and evidence gaps
Both browser modes are already approved. Select the default mode for each profile; the native-browser/driver and reference-browser support matrix; exact upstream full-suite versus Odin-probe identities; language subsets; installed/reference toolchain policy; fixed compilation/warmup rules; and measured profile budgets. Browser automation overhead, all-assets-offline closure, per-asset redistribution inventory, selected-suite behavior on musl/aarch64 and newer JDKs, and correctness/variance qualification remain future proof work. An application score must not conceal these unresolved cohort boundaries.
## Sources and method
Primary sources inspected **2026-09-25**. Context7 resolution preceded queries for Speedometer/WebKit, Bash, Hyperfine, Playwright, pyperformance/pyperf, Rust, LLVM, JMH, GCC, and Selenium. Exact Speedometer and pyperformance lookups returned unrelated entries; those were rejected and their own repositories inspected. Queries stayed within the three-command limit per lookup; no quota error occurred.
1. [BrowserBench current index](https://browserbench.org/).
2. Speedometer release source: [commit](https://github.com/WebKit/Speedometer/commit/1386415be8fef2f6b6bbdbe1828872471c5d802a), [parameters](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/params.mjs), [result client](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/main.mjs), [package metadata](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/package.json), [local hosting](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Development.md).
3. [Speedometer 3.1 instructions](https://browserbench.org/Speedometer3.1/instructions.html), [workload descriptions](https://browserbench.org/Speedometer3.1/about.html), [measurement/scoring explanation](https://github.com/WebKit/Speedometer/blob/main/README.md).
4. [JetStream 3.0 implementation page](https://browserbench.org/JetStream3.0/), [in-depth methodology and provenance](https://browserbench.org/JetStream3.0/in-depth.html).
5. [MotionMark 1.3.2 implementation page](https://browserbench.org/MotionMark1.3.2/), [methodology, display requirements, and version history](https://browserbench.org/MotionMark1.3.2/about.html).
6. [Speedometer release license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE).
7. [Speedometer browser-driver testing documentation](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Testing.md), [Selenium Manager: explicit drivers, offline mode, and architecture limitations](https://www.selenium.dev/documentation/selenium_manager/).
8. [Playwright system requirements](https://playwright.dev/docs/intro), [browser distributions and headless modes](https://github.com/microsoft/playwright/blob/main/docs/src/browsers.md), [musl limitations](https://github.com/microsoft/playwright/blob/main/docs/src/docker.md).
9. [Chrome unified headless and headless-shell distinction](https://developer.chrome.com/docs/chromium/headless).
10. [GNU Bash startup files](https://www.gnu.org/software/bash/manual/html_node/Bash-Startup-Files.html), [invocation options](https://www.gnu.org/software/bash/manual/html_node/Invoking-Bash.html).
11. [Hyperfine README: intermediate-shell correction, direct mode, warmups, repetitions](https://github.com/sharkdp/hyperfine/blob/master/README.md).
12. [pyperformance purpose and license](https://github.com/python/pyperformance/blob/main/README.rst), [PyPI release metadata](https://pypi.org/pypi/pyperformance/json), [pyperf release metadata](https://pypi.org/pypi/pyperf/json), [benchmark manifest](https://github.com/python/pyperformance/blob/main/pyperformance/data-files/benchmarks/MANIFEST), [workload descriptions](https://github.com/python/pyperformance/blob/main/doc/benchmarks.rst), [dependency/runtime guidance](https://github.com/python/pyperformance/blob/main/doc/usage.rst), [startup implementation](https://github.com/python/pyperformance/blob/main/pyperformance/data-files/benchmarks/bm_python_startup/run_benchmark.py).
13. [rustc-perf purpose/licenses](https://github.com/rust-lang/rustc-perf/blob/master/README.md), [compile suite](https://github.com/rust-lang/rustc-perf/blob/master/collector/compile-benchmarks/README.md), [experimental runtime suite](https://github.com/rust-lang/rustc-perf/blob/master/collector/runtime-benchmarks/README.md), [runtime groups](https://github.com/rust-lang/rustc-perf/tree/master/collector/runtime-benchmarks), [collector requirements](https://github.com/rust-lang/rustc-perf/blob/master/collector/README.md).
14. [LLVM test-suite guide](https://llvm.org/docs/TestSuiteGuide.html), [testing/output-validation model](https://llvm.org/docs/TestingGuide.html), [serial miniFE description](https://github.com/llvm/llvm-test-suite/blob/main/MultiSource/Benchmarks/DOE-ProxyApps-C++/miniFE/README), [license and third-party exceptions](https://github.com/llvm/llvm-test-suite/blob/main/LICENSE.TXT).
15. [JMH purpose and cautions](https://github.com/openjdk/jmh/blob/master/README.md), [published version metadata](https://repo.maven.apache.org/maven2/org/openjdk/jmh/jmh-core/maven-metadata.xml), [parameter/warmup/fork example](https://github.com/openjdk/jmh/blob/master/jmh-samples/src/main/java/org/openjdk/jmh/samples/JMHSample_27_Params.java), [result consumption and profilers](https://github.com/openjdk/jmh/blob/master/jmh-samples/src/main/java/org/openjdk/jmh/samples/JMHSample_35_Profilers.java), [license](https://github.com/openjdk/jmh/blob/master/LICENSE).
16. [DaCapo latest release](https://github.com/dacapobench/dacapobench/releases/tag/v23.11-MR2-chopin), [release compatibility and workload notes](https://github.com/dacapobench/dacapobench/blob/0db32562cf169730c163d88df2eeb28217ca7d03/benchmarks/RELEASE_NOTES.md), [purpose/reporting/license guidance](https://github.com/dacapobench/dacapobench/blob/master/README.md).
17. [pyperf runner configuration](https://github.com/psf/pyperf/blob/main/doc/runner.rst), [process/calibration architecture](https://github.com/psf/pyperf/blob/main/doc/run_benchmark.rst).
18. [Cargo profile defaults](https://doc.rust-lang.org/cargo/reference/profiles.html), [rustc code-generation options](https://doc.rust-lang.org/rustc/codegen-options/index.html), [GCC optimization semantics](https://gcc.gnu.org/onlinedocs/gcc/Optimize-Options.html).
-140
View File
@@ -1,140 +0,0 @@
# Defensible median scoring and comparison rules
Research for [Establish defensible median scoring and comparison rules](https://git.bongbetic.com/xavierk/odin/issues/7), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
**Accessed:** 25 September 2026. **Status:** decision evidence and recommendations; no final scoring formula, calibrated reference values, benchmark runs, or implementation. The user's requirement is a **median** score.
## What the median should mean
There are three separate choices:
| Level | Meaning | Main limitation |
|---|---|---|
| Median of repeated measurements | Typical result for one fixed workload under stated conditions | Hides occasional long stalls; does not combine CPU, storage and browser results |
| Median across normalized workloads | Typical relative performance across a fixed test set | Domains with many tests gain more influence; changing references can change rankings |
| Median across domain summaries | Typical relative performance across explicitly chosen domains | Domain definitions matter; poor performance in a minority of domains can disappear from the headline |
NIST defines the sample median as the middle observation, or the arithmetic average of the middle two for an even sample count. Its resistance to extremes is useful, but it is a measure of location, not completeness, reliability, or worst-case response. [NIST location][nist-location]
**Recommendation for discussion:** use medians to summarize repeated valid measurements, normalize against frozen references, form predefined domain summaries, then use a median across those domains for the requested headline. Preserve each stage and its raw inputs. This proposes an aggregation structure; the domain membership, weighting, reference values, repeat counts and numeric scale remain decisions.
Established suites demonstrate why the levels must stay explicit. SPEC CPU 2017 takes median execution times from three runs, or the slower of two, then uses a **geometric mean** across ratios. Speedometer uses inverse geometric means across test durations and arithmetic means across iterations. These are methodological precedents, not permission to substitute a geometric mean for Odin's requested median. Preserve a tool's native result under its original name; label Odin's further aggregation separately. [SPEC rules][spec-rules] [Speedometer methodology][speedometer]
## Normalization and category balance
Milliseconds, operations/second and GB/s cannot share a meaningful raw median. A candidate approach is a dimensionless ratio against the **same workload's** reference: observation/reference for positive higher-is-better measures, reference/observation for positive lower-is-better measures. SPEC uses the latter for elapsed-time ratios. The metric identity must include its units, workload, size, concurrency, timing boundaries and direction. Reject invalid/nonfinite inputs under declared validity rules; a zero measured duration must not become an infinite score. [SPEC overview][spec-overview]
**Fictional arithmetic examples throughout this report:** these numbers illustrate consequences only. They are not Odin calibration or measured hardware results. A displayed index of 100 at the reference is an arbitrary illustrative scale, not a recommendation.
| Fictional metric | Reference | Observed | Illustrative ratio / index |
|---|---:|---:|---:|
| Work throughput | 25 operations/s | 50 operations/s | 2.0 / 200 |
| Memory bandwidth | 25 GB/s | 20 GB/s | 0.8 / 80 |
| Request latency | 4 ms | 2 ms | 2.0 / 200 |
A median of these indices is 200, despite memory bandwidth being below reference. That is a consequence of the chosen statistic. Show domain detail and slow-tail measurements alongside it. “200 versus 100” describes this index; it does not establish that every application is twice as fast.
Fix the order of operations. For fictional repeated times `[1, 3]` ms and a 3 ms reference, normalizing the raw median gives `3 / 2 = 1.5`; taking the median of individual ratios `[3, 1]` gives 2. Even-count averaging and reciprocal normalization do not commute. The report format must specify which result it contains. [median definition][nist-location]
Balance domains before counting metrics. If ten CPU tests each score 160 and four other domains score 40, 80, 100 and 120, the flat fourteen-test median is 160. A median across five domain summaries is 100. Adding CPU subtests should not silently redefine the product's priorities. Similarly, adding Python/Rust/Java variants must not automatically multiply the language domain's influence. A median of domain medians is a deliberate hierarchical index, not the pooled median of all observations.
Exclude health counters, memory-test pass/fail, driver availability and installed RAM capacity from throughput arithmetic. Multiple correlated outputs from one workload—throughput, IOPS, average latency and several percentiles—also need an explicit selection rule before any becomes an independent scored contribution.
## Calibration is a substantive decision
SPEC establishes per-workload reference times on a named machine and publishes the calculation rules. Its documentation explains that reference changes preserve relative overall rankings for its geometric-mean calculation. **That invariance does not generally hold for a median across normalized metrics.** [SPEC reference explanation][spec-overview]
For three fictional higher-is-better workloads, let machine A produce `[1, 10, 10]` and B produce `[2, 2, 20]` in each workload's own units:
| Fictional reference vector | A's normalized results → median | B's normalized results → median | Ordering |
|---|---|---|---|
| `[1, 1, 1]` | `[1, 10, 10]` → 10 | `[2, 2, 20]` → 2 | A higher |
| `[1, 10, 10]` | `[1, 1, 1]` → 1 | `[2, 0.2, 2]` → 2 | B higher |
The measurements did not change. Changing the reference altered the relative scales and which observations occupied the middle. Retain the requested median, make the reference rationale public, and version reference changes rather than treating them as cosmetic rescaling.
| Reference option | What it supports | Decision cost |
|---|---|---|
| Named reference configuration | Auditable, fixed anchor with documented per-test measurements | One machine's balance influences the median; configurations and repeatability need validation |
| Frozen reference cohort | Per-test references from a documented collection of machines | Cohort selection, sampling bias and revision policy become part of the score |
| User's own baseline | Local before/after comparisons | A score relative to oneself cannot rank different machines |
Recommend evaluating a frozen, locally distributable calibration manifest. Include reference measurements and provenance, reference conditions, workload and artifact digests, units/directions, aggregation order, required domains and calibration identity. Store it with results so calculation remains reproducible offline. Raw observations must survive changes; a recalculated score should identify its new calibration and preserve the original.
An arbitrary scale factor is acceptable if described as an index. A claim such as “100 is the median Linux machine,” a percentile rank, or a universal poor/good threshold requires representative population evidence that does not exist yet. Separately normalizing each architecture to its own average would also prevent interpreting those numbers as one common cross-architecture scale.
## Repetitions, warmup and uncertainty
Google Benchmark documents warmup, repetitions, median, standard deviation and coefficient of variation; it distinguishes user-visible wall time from CPU consumption. NIST recommends examining ordered observations for changing location/spread and says potential outliers should not simply be deleted when their cause is unknown. These support retaining all repeated observations and their conditions. [Google Benchmark][google-guide] [NIST run sequence][nist-runseq] [NIST outliers][nist-outliers]
Recommended measurement rules:
- Define warmup separately for each workload and retain its duration. Warm caches/JIT throughput, cold launch time and sustained thermal performance are different questions. Do not discard a slow first run from a declared cold-start test.
- Fix repetition and stopping rules before observing scores. A quick run may estimate a median without enough evidence for a useful confidence interval; it should not claim the precision of the standard profile. Calibrate the minimum repeats against the 10–20 minute budget.
- Retain run order, warmup, elapsed time, temperatures/power context, competing activity and invalidation reasons. A drifting sequence is not interchangeable independent noise. Repetitions within one process or thermal episode are not automatically independent runs.
- Exclude observations only for declared validity failures such as incorrect output, changed workload, cancellation or protocol failure. Preserve them with reasons. A slow but valid run can represent the usability problem Odin is meant to reveal.
- Show central spread such as MAD or IQR, plus tails where the workload supplies enough events. NIST defines MAD and IQR as distinct measures of spread; neither is itself a confidence interval. A median across repeated p99 values must not be labelled the p99 of all requests. [NIST scale][nist-scale]
NIST documents median confidence intervals based on order statistics/binomial probabilities, interpolated methods and bootstrap alternatives. Choose and validate a median-appropriate method; do not apply a mean's standard-error formula to a median. Confidence also depends on sample count and assumptions about the measurements. For an aggregate, account for shared run-level variation and state whether uncertainty in the calibration reference is included. The spread **between different domain scores** is not sampling uncertainty about the headline. [NIST median intervals][nist-median-ci]
Keep a graph of results in time order. Google documents CPU selection, boost, scheduler contention, SMT, caches and NUMA as variance sources. Its suggestions for controlled laboratory microbenchmarks include changing system settings; Odin's installed-system baseline should record existing conditions and label any tuned experiment separately. The existing [CPU/memory report][odin-cpu] and [portability/UI report][odin-portability] explain TUI interference and qualification needs. Stable repeated numbers alone do not prove that a workload represents real usability. [Google variance][google-variance]
## Comparability must be attached to every score
SPEC requires performance-relevant observation conditions and valid workload outputs; its CPU suite intentionally measures processor, memory subsystem **and compilers**. Even a fixed-toolchain comparison describes a defined software/hardware configuration. [SPEC rules][spec-rules] [SPEC overview][spec-overview]
Recommend two explicit comparison purposes:
- **Installed-system usability:** the chosen installed browser, shell, runtimes, drivers and kernel are part of what is measured. Version/configuration changes may explain a score change without any hardware change.
- **Controlled reference workload:** fixed workload assets, runtime/compiler contracts, flags, input data and execution modes improve comparison across machines. Architecture-specific artifacts must implement equivalent declared work and validate outputs; different ISA policies need disclosure.
Neither mode needs to masquerade as a pure hardware measurement. Keep their result identities distinct. Kernel/libc/distro differences can be the subject of a comparison, but they must be visible and the workload contract must remain equivalent.
A comparison identity should include suite/scoring/calibration versions, workload set, run profile, tool/artifact versions, options and data digests, timing/aggregation rules, browser mode, hardware/virtualization context and validity/coverage. Require a documented equivalence decision before combining scores across changed tools or workloads. SPEC warns that scores across different suite generations generally cannot be converted. [SPEC overview][spec-overview]
Speedometer 3.1 instructs users to use a clean browser profile, close competing programs/tabs, keep its page focused, avoid device interaction, use AC power and allow cooling when needed. Its official UI computes a 95% interval around its **arithmetic mean**; that interval cannot be attached to Odin's median unchanged. Preserve native browser score/uncertainty and label any median of complete runs separately. Headed and headless measurements need separate identities until an equivalence study justifies any shared interpretation; background versus foreground execution is also material. [instructions][speedometer-instructions] [3.1 result code][speedometer-main]
VM results characterize the guest allocation and virtualization environment. Keep native, virtualized and emulated cohorts identifiable; VM compatibility success does not establish native performance. Storage cache mode, queue depth, engine, filesystem and durability policy similarly belong to the workload identity. A fallback such as buffered I/O cannot silently replace a direct-I/O measurement with the same scoring identity. [CPU/memory report][odin-cpu] [storage report][odin-storage] [portability report][odin-portability]
## Missing tests and eligibility
For fictional domain indices `[40, 80, 100, 120, 160]`, the complete median is 100. Omitting 40 produces 110; omitting both 40 and 80 produces 120. Available-only aggregation can reward absent or deliberately skipped weak components.
**Recommendation:** define a versioned required set for the full score. Permit a clearly named partial median and domain results when the full set is unavailable, with the exact subset identified. Compare partial scores only over the same compatible subset; a pairwise intersection comparison must recompute **both** results and label that narrower scope. An optional pack must not silently change the headline's membership.
Keep distinct outcomes: completed-valid, completed-with-limitations, unsupported, missing dependency, permission denied, unsafe to run, cancelled, timed out, and failed validation. The eventual validity contract decides whether a limited result remains score-eligible. Never impute missing results as zero, a reference score, or a healthy pass. A required workload failing verification makes the full score ineligible, while preserving completed measurements and the associated finding. Good numbers from other domains should not cancel that failure.
The user accepted reporting unavailable tests across Linux targets. That does not resolve which domains are mandatory, whether every machine should still display a partial number, or how partial results should look. Those are explicit product decisions.
## Presentation, recommendations and open decisions
Recommend a headline containing the median, score identity, full/partial state, and eligible-domain coverage. The next view should show domain values, raw units, repeat count/spread, tail latency, invalidations and reference details. Show health findings beside performance: a fast drive with serious SMART evidence still needs attention. Missing telemetry must stay unknown. The storage and CPU reports establish why speed cannot determine drive replacement or certify memory health. [storage][odin-storage] [CPU/memory][odin-cpu]
Optimization advice should cite the observation and matching rule: for example, measured foreground stalls plus pressure evidence can support investigating memory contention. A low normalized score alone does not identify its cause. Keep severity of health evidence, completeness of coverage, measurement uncertainty and performance position as separate concepts; avoid one synthetic “confidence/health” percentage that mixes them.
Before implementation, decide:
1. The headline's median level, domain membership and balancing rules; whether responsiveness contributes or remains an accompanying measurement.
2. The reference configuration/cohort, scale and calibration-release policy; collect actual calibration data before inventing thresholds.
3. Required versus optional coverage, partial-score display and exact comparison eligibility.
4. Installed-system versus controlled-workload defaults, architecture/ISA policies, browser modes and VM cohorts.
5. Repetition/warmup/stopping rules, outlier validity rules, median interval method and honest quick/standard/extended precision claims.
6. Evidence requirements for optimization rules and separation of urgent health findings from the score.
**Evidence limits:** no calibration population, repeatability measurements, TUI-overhead budget or headed/headless equivalence study was produced. Examples are arithmetic demonstrations only. Context7 successfully resolved Google Benchmark; two BrowserBench/Speedometer searches returned unrelated packages, so its official repository and deployed 3.1 documentation were inspected directly. NIST and SPEC sources were inspected directly as statistical and benchmark-methodology references. This report neither adopts SPEC's workloads nor claims that their aggregation rules are Odin's final design.
[nist-location]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda351.htm
[nist-scale]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda356.htm
[nist-outliers]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda35h.htm
[nist-runseq]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda33p.htm
[nist-median-ci]: https://www.itl.nist.gov/div898/software/dataplot/refman1/auxillar/mediancl.htm
[spec-rules]: https://www.spec.org/cpu2017/Docs/runrules.html
[spec-overview]: https://www.spec.org/cpu2017/Docs/overview.html
[google-guide]: https://github.com/google/benchmark/blob/main/docs/user_guide.md
[google-variance]: https://github.com/google/benchmark/blob/main/docs/reducing_variance.md
[speedometer]: https://github.com/WebKit/Speedometer/blob/main/README.md
[speedometer-instructions]: https://browserbench.org/Speedometer3.1/instructions.html
[speedometer-main]: https://browserbench.org/Speedometer3.1/resources/main.mjs
[odin-cpu]: https://git.bongbetic.com/xavierk/odin/src/commit/90213d7f62cbd118f08ea8ff2f8042e94aa038a7/docs/research/cpu-memory.md
[odin-storage]: https://git.bongbetic.com/xavierk/odin/src/commit/6e5af87a64faedd4a8ad31ba10d9be4b499e8349/docs/research/storage-health.md
[odin-portability]: https://git.bongbetic.com/xavierk/odin/src/commit/20681cd4f184a9fc0164dd638a252de44ef230f5/docs/research/portability-ui.md