Author SHA1 Message Date
Codex f384c81f0f Research median scoring and comparison rules 2026-09-25 23:44:11 +05:30
2 changed files with 140 additions and 139 deletions
-139
View File
@@ -1,139 +0,0 @@
# Safe CPU, memory, and kernel measurements for Odin
Research for [Establish safe CPU, memory, and kernel measurements](https://git.bongbetic.com/xavierk/odin/issues/2), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
**Accessed:** 25 September 2026. **Status:** recommendations for later decisions, not a selected workload suite. No benchmarks, stress tests, privilege changes, or host configuration changes were performed.
## Findings that shape the decision
Odin can reuse established workloads and Linux interfaces, but no single tool measures system usability, peak performance, and hardware health. Keep four outcomes distinct:
- **Performance measurement:** completed work per second or elapsed time under a specified workload.
- **Responsiveness:** foreground request or wakeup latency, including its tail, under a specified competing load.
- **Pressure observation:** time lost to CPU, memory, or I/O contention during that interval.
- **Health finding:** a detected verification failure or reported hardware error, with the observation's coverage.
This separation follows the tools' actual contracts: sysbench benchmarks operations; schbench measures artificial requests and scheduling delays; PSI measures stalls; memtester checks memory contents. Stress-ng explicitly says it was **not intended as a precise benchmark suite**. Successful stress completion establishes only that the chosen work completed without detected failures under those conditions. [sysbench][sysbench] [schbench][schbench] [PSI][psi] [memtester][memtester-man] [stress-ng][stress-readme]
## Candidate comparison
Maintenance observations describe inspected releases or repository activity, not a guarantee of future support. GPL notices and third-party dependencies need checking again for the exact distributable artifacts. Perf is distributed in the Linux source tree, whose COPYING specifies GPL-2.0-only with the syscall exception and notes that other licenses may also apply; preserve the selected tool and dependency notices. [Linux COPYING][kernel-license]
| Candidate | Useful measurement and limitations | Controls, output, portability, maintenance |
|---|---|---|
| **sysbench CPU / memory** | CPU events are a small prime-search workload, not a general application-performance model. The memory test defaults to a **1 KiB block**; its `100G` default is cumulative transfer size, not RAM allocation. A default run therefore cannot stand for DRAM bandwidth. | Thread, event, duration, warmup and percentile controls; human-readable built-in reports require a pinned parser. Upstream advertises x86_64 and aarch64 packages. GPL-2.0-or-later. Latest published release inspected: 1.0.20, April 2020; repository push activity March 2025. Void still packages 1.0.20. [README][sysbench] [CPU source][sysbench-cpu] [memory source][sysbench-memory] [release][sysbench-release] [metadata][sysbench-meta] [Void][void-sysbench] |
| **schbench** | Reports synthetic request latency, wakeup latency, and requests/second. Closer to responsiveness than peak throughput. Its server-inspired matrix workload deliberately penalizes preemption using per-CPU locks; this is not a desktop interaction model. | Runtime, worker/message counts, work size, request rate, affinity and JSON percentiles. Source has x86 and aarch64 paths; Linux pthread/futex interfaces, no normal root requirement found. Musl behavior remains unverified. GPLv2. Inspected upstream commit `6300b8f`, June 2025. [methodology][schbench] [source][schbench-source] |
| **stress-ng** | Appropriate candidate for controlled background load and selected verification diagnostics. Bogo operations are unsuitable as Odin's cross-test score foundation. | Explicit byte/worker/time caps, verification, YAML and differentiated exit codes. Upstream documents musl builds and testing on ARM64/x86-64. GPL-2.0-or-later; release 0.22.01, September 2026. Void's recipe has explicit musl handling. [README][stress-readme] [manual][stress-man] [release][stress-release] [Void][void-stress] |
| **STREAM** | Established sustainable memory-bandwidth kernels: Copy, Scale, Add, Triad. No memory-latency or comprehensive error-detection result. Each array must exceed cache requirements; a small cache-resident run is not a compliant STREAM result. | Portable C; multicore uses OpenMP and its runtime. Array size is a build parameter in reference 5.10. Text output and numerical validation. Reference source dates to 2013: stable method, not evidence of a modern portability test matrix. Custom license permits use/redistribution but imposes result-naming/run-rule conditions. [source and license][stream-source] [run rules][stream-rules] |
| **lmbench `lat_mem_rd`** | Pointer-chain latency over sizes/strides exposes cache, memory and TLB behavior. Its manual acknowledges vulnerability to stride-sensitive prefetchers; do not present this as an architecture-independent “true RAM latency.” | Warmup, repetitions, bounded size, text pairs. GPLv2 COPYING inspected; Intel's repository is active but contains old documentation. Portability and selected-file licensing need validation before adoption; not a recommended mandatory dependency yet. [manual][lmbench-memory] [README][lmbench-readme] [COPYING][lmbench-license] [metadata][lmbench-meta] |
| **cyclictest / perf** | Cyclictest measures timer wakeup latency; perf supplies diagnostic counters. Neither is a substitute for foreground application response. | Cyclictest offers duration, histogram and JSON; source is GPL-2.0-only, current rt-tests release 2.11. Its startup tests permission to enter SCHED_FIFO even when ordinary policy is selected. Perf depends on kernel/PMU access and can emit JSON. Treat both as optional diagnostic coverage. [cyclictest][cyclictest] [privilege check][rt-utils] [rt-tests release][rt-release] [perf][perf-stat] |
| **memtester** | Online checking of allocated memory, not all installed RAM. Failures can involve memory, CPU, temperature or power; the result does not identify a replaceable DIMM by itself. | Byte size and finite iteration count; default iterations are infinite. Text plus exit-bit mask. Record actual allocation and locking, not only exit status. GPL-2.0-only. Version 4.7.1's December 2024 fix addresses stricter C23/GCC 15 compilation; Void recipe inspected still selects 4.6.0. [manual][memtester-man] [source][memtester-source] [changelog][memtester-changelog] [Void][void-memtester] |
| **Memtest86+** | Offline, bootable diagnostics reach almost all memory without the resident OS. They cannot run as an ordinary in-terminal stage. | GPLv2. Stable v8.10, May 2026, lists x86, x86-64 and LoongArch64. Current main additionally lists AArch64 with UEFI boot. This is a **stable/development difference**, not certified aarch64 coverage. [stable README][memtest-stable] [development README][memtest-main] [release][memtest-release] |
| **EDAC / rasdaemon** | Hardware-error telemetry, complementary to active tests. Availability depends on hardware, firmware, drivers and exposed events. | Read EDAC counters where accessible; optionally consume an existing rasdaemon history. Rasdaemon monitors kernel trace events and has database backends; its repository shows September 2026 activity and GPLv2 metadata. Starting it is a separate privileged monitoring action, not necessary to read available counters. [EDAC ABI][edac-abi] [rasdaemon][rasdaemon] [metadata][ras-meta] |
## Measuring usability under contention
**Recommendation:** evaluate paired idle and loaded foreground-request measurements as a first-class candidate. Run the same bounded foreground work alone, with a fixed CPU background load, and with separately controlled memory pressure. Preserve normal scheduling policy, work size, thread count, placement, requested arrival rate, actual throughput, sample count, and latency distribution. Report absolute latency and degradation relative to idle; a fast idle result can coexist with poor responsiveness under contention.
Existing tools can supply the observations. Schbench records both request and wakeup latency, supports a fixed request rate, and can coexist with an independently bounded stress-ng background worker. Sysbench's rate-limited engine is another candidate: its source adds measured queue time to event duration and reports queue length. These remain synthetic proxies for foreground work; neither measures keyboard-to-pixel delay, terminal rendering, browser interaction, or application launch by itself. Those need their own workload definitions. [schbench methodology][schbench] [schbench source][schbench-source] [sysbench queue][sysbench-core] [sysbench timer][sysbench-timer]
Schbench needs qualification before selection. Its README says warmup defaults to five seconds, while the inspected source defaults to zero and bypasses its warmup reset in request-rate mode. It uses `gettimeofday`, so clock adjustments can contaminate timing. Set options explicitly, retain the revision, validate timing conditions, and decide whether its deliberate preemption penalty fits Odin's goal. Do not adapt work size independently on every system and then compare the resulting latency as equal work. [source][schbench-source]
Cyclictest is a useful separate scheduler diagnostic. Its default behavior can hold `/dev/cpu_dma_latency` at zero and suppress deep idle states; `--default-system` avoids that tuning. The source's unconditional real-time privilege check prevents assuming that an ordinary-policy configuration is universally unprivileged. Its timer latency, especially under SCHED_FIFO, is a different measurement from a normal foreground application's response. [manual][cyclictest] [source][cyclictest-source] [privileges][rt-utils]
## Kernel telemetry and compatibility
**PSI:** Read system and, where available, workload-cgroup `cpu`, `memory`, and `io` pressure. `some` is time when at least some tasks are stalled; `full` is time when all non-idle tasks are stalled together. The cumulative `total` counter permits interval deltas; rolling 10/60/300-second averages can smear a short benchmark across adjacent phases. **System-wide CPU `full` is undefined and exposed as zero for compatibility**—zero there cannot mean perfect responsiveness. PSI exists in the inspected Linux 4.20 source, but requires `CONFIG_PSI` and may be disabled by default pending `psi=1`. Probe actual files and readability, not only kernel version. [PSI][psi] [4.20 documentation][psi-420] [Kconfig][kconfig]
**Capacity and pressure:** Record usable RAM, `MemAvailable`, swap capacity/usage, process or cgroup memory, major faults, and swap/reclaim activity where exposed. `MemAvailable` is an estimate of memory available without swapping, not an allocation guarantee. Swap occupancy alone does not establish current pressure. `/proc/stat` supplies CPU time and steal time, but kernel documentation explicitly warns that `iowait` is unreliable. Correlate these observations with latency and PSI; do not derive a definitive bottleneck from CPU utilization or a single counter. [proc documentation][proc]
**Cgroups:** Observe effective CPU affinity/cpuset, CPU quota, memory/swap limits and relevant ancestors. Host RAM and online CPU counts may exceed what the benchmark is allowed to use. Cgroup v2 `cpu.stat` records throttling; `memory.events` separates high-limit reclaim, OOM conditions and kills. Parse by key: the kernel explicitly allows new `memory.stat` entries in the middle. These interfaces are independent of a particular init system, but writable delegation and enabled controllers are not guaranteed on Void, deb, rpm, containers, or user sessions. [cgroup v2][cgroup]
**Perf:** Treat hardware counters as enrichment. `CONFIG_PERF_EVENTS`, CPU PMU support, virtualization, `perf_event_paranoid`, capabilities and distribution policy can limit access. Kernel documentation recommends `CAP_PERFMON` over broad `CAP_SYS_ADMIN`; Odin should describe missing access rather than lowering system security settings. Record event identity and time-running percentage when multiplexing occurs. PMU-specific cache and pipeline events are not universal normalized scores. Perf's manual also warns of overhead at short sampling intervals, particularly below 100 ms. [security][perf-security] [Kconfig][kconfig] [perf stat][perf-stat]
**Architecture/libc:** Proc/sysfs/cgroup interfaces offer the strongest common layer across x86_64/aarch64 and glibc/musl. Workload binaries still need distinct, verified artifacts and recorded toolchain flags. Upstream stress-ng explicitly documents musl; sysbench's advertised architectures and Void recipes are useful evidence, but none of the inspected material certifies Odin's whole four-way architecture/libc matrix. STREAM additionally needs a compatible OpenMP runtime; rt-tests has library dependencies. A package recipe proves availability intent, not successful operation. Missing checks should carry reasons such as unsupported, permission denied, unavailable dependency, or insufficient safe resources. [stress-ng][stress-readme] [sysbench][sysbench] [Void recipes][void-sysbench] [STREAM][stream-source] [rt-tests Makefile][rt-makefile]
## Resource safety and cancellation
The following is a proposed execution contract, with exact caps left to the safety decision:
1. Calculate a conservative working budget from current `MemAvailable`, effective cgroup/ancestor headroom, expected tool/runtime overhead, and a retained reserve. Recheck while running. There is no sourced universal percentage that guarantees safety when other programs allocate concurrently.
2. Where delegated cgroup v2 control exists, put disposable workers in their own subtree and keep the supervisor outside that subtree. `memory.high` induces reclaim/throttling and **is not a hard cap**; `memory.max` bounds charged memory and can invoke OOM inside the cgroup. `memory.swap.max` separately controls swap. Use deliberate limits and record them because they change results. Caps reduce risk; they cannot guarantee that an unrelated system-wide shortage never kills a process. [cgroup v2][cgroup]
3. Reserve intentional pressure for explicitly selected, isolated work. Without reliable containment or enough reserve, recommend pressure **observation** and small bounded workloads, and report unavailable active-pressure coverage. Allocation success alone is insufficient: memtester's own manual warns about overcommit, swapping and OOM affecting other programs. [memtester][memtester-man]
4. Never expose memtester's physical-address/device modes in the normal benchmark path. They overwrite the mapped region and can crash the system when it belongs to another process or the kernel. For ordinary allocations, verify locked bytes and completed patterns/iterations. Linux permits unprivileged locking up to `RLIMIT_MEMLOCK`; larger locking requires suitable privilege, commonly `CAP_IPC_LOCK`. The tool's “run as root” advice should not force the entire TUI to run as root. [manual][memtester-man] [Linux mlock][mlock]
5. Use finite work/time limits and a supervisor deadline. For stress-ng, enable only reviewed stressors, verification where supported, and no OOM respawn (`--oomable`). Its `--oom-avoid` is a heuristic with measurement overhead, not containment. In 0.22.01, `--vm-bytes` describes a total across VM workers; other stressors have different allocation semantics, so retain the exact version and options. [stress-ng manual][stress-man]
6. Cancel the worker process group gracefully, then terminate remaining descendants after a defined grace period; use `cgroup.kill` when accessible. Preserve a cancelled/partial result. Stress-ng documents SIGINT cleanup, but its timeout can overrun during uninterruptible calls or cleanup. No userspace deadline guarantees immediate cancellation of an uninterruptible kernel task. [manual][stress-man] [cgroup kill][cgroup]
## Repetition, kernel settings, and TUI overhead
**Recommendation:** record warmup separately, repeat bounded measurements, retain all repetitions and dispersion, and report a median only at a clearly defined level. STREAM's official report takes the **best** iteration after discarding the first; a median of repeated STREAM run results is a different statistic. Do not silently relabel its internal minimum as a median, combine raw milliseconds with MB/s, or replace an unavailable result with zero. Overall score normalization belongs to the scoring decision. [STREAM source][stream-source]
Record kernel/build identity, visible preemption/scheduler settings, CPU topology and allowed CPUs, NUMA placement, THP policy, libc, workload/compiler version and flags, governor/driver, boost, power source and temperature observations. NUMA placement and THP policy affect what memory workload is actually measured. Kernel CPUFreq documentation explains that `scaling_cur_freq` can be a requested state rather than measured frequency, and boost depends on thermal/power conditions and package load. A governor name or falling frequency alone does not prove thermal throttling. Correlate sustained performance with available temperatures, thermal trip/cooling states, and power/frequency evidence; unavailable sensors remain unavailable. [CPUFreq][cpufreq] [NUMA][numa] [THP][thp] [thermal interfaces][thermal]
For “single core,” specify whether the worker is pinned and how the core is chosen on heterogeneous CPUs. For “multicore,” specify workers relative to allowed CPUs, SMT and quota; do not silently change those rules between systems. Preserve the machine's existing configuration for the baseline. Potential governor, scheduler, THP, affinity or kernel changes should be advice or separately labelled experiments, not automatic optimization before measuring.
The TUI competes for CPU time, memory bandwidth, cache and terminal I/O. Recommend throttled graph updates, buffered logs, no expensive animation during timed sections, and a quiet measurement mode that preserves cancellation. Avoid hiding this by reserving a core without recording it: that reduces tested capacity. Later validation should compare quiet versus normal rendering on the slowest supported machines and establish an overhead budget. This is a proposed qualification experiment, not evidence that a particular redraw rate is already safe. Lmbench explicitly warns about competing cache/CPU work; PSI's own Kconfig notes overhead can show up in synthetic scheduler stress tests. [lmbench][lmbench-readme] [Kconfig][kconfig]
## Memory errors, VM coverage, and remaining decisions
Online memtester cannot touch RAM occupied by the kernel or other processes. It may allocate less than requested and may continue unlocked; inspected 4.7.1 source can then still exit zero if its pattern checks succeed. Parse allocation/locking evidence alongside exit bits and report “no errors detected in the tested allocation,” with size, iterations and duration. A mismatch warrants investigation, not an automatic RAM-replacement diagnosis. [manual][memtester-man] [source][memtester-source]
EDAC counters reset at driver initialization or explicit reset; preserve counter baselines and `seconds_since_reset` without resetting them. Corrected errors merit attention, but uncorrected errors may cause a panic before a counter increments. DIMM labels can depend on board-specific userspace mapping. Therefore missing EDAC nodes, zero observed deltas, and an empty rasdaemon history cannot certify error-free RAM. Offer offline follow-up where supported; decide how development-only AArch64 Memtest86+ support should be presented. [EDAC ABI][edac-abi] [EDAC model][edac] [Memtest86+ stable][memtest-stable] [development][memtest-main]
VMs can validate packaging, libc/architecture execution, permissions, telemetry fallbacks, cgroup containment and result handling. Guest CPU/memory scores describe the guest allocation and host scheduling conditions; steal time is useful context. They do not certify the host's DIMMs, ECC pipeline, cooling, physical memory-channel bandwidth, or representative bare-metal scheduler tails. Rasdaemon's upstream QEMU tests intentionally inject virtual nonfatal events: useful for exercising decoding, not proving physical hardware health. [proc][proc] [rasdaemon CI description][rasdaemon]
**Recommended next decision:** shortlist sysbench for a narrow CPU baseline, STREAM for bandwidth, schbench for responsiveness qualification, selected stress-ng workers for bounded load, and memtester plus available EDAC/RAS for diagnostics. Keep perf/cyclictest optional; defer mandatory memory-latency scoring until a candidate is validated. Final selection remains open.
The human-facing decisions still needed are the foreground workload's meaning; fixed versus relative background load; inclusion of responsiveness in the median score; safe resource reserves and privileges; repetition/time allocation within the 10–20 minute standard run; treatment of heterogeneous cores and missing coverage; and whether offline/development-tool guidance belongs in the first release.
**Evidence limits:** no candidate was built or executed across the target matrix. Context7 resolved Linux kernel, sysbench, memtester and rt-tests documentation. Stress-ng, STREAM and schbench searches returned unrelated libraries, so no false library match was used; their owning sources were inspected directly. Memtester's upstream HTTPS site failed certificate validation; the report uses the original source/manpage/changelog preserved by Debian, cross-checked against the Void 4.6.0 source archive checksum, and identifies the version difference. Exact artifact compatibility, parser contracts, resource budgets, and score repeatability require later qualification.
[sysbench]: https://github.com/akopytov/sysbench/blob/master/README.md
[sysbench-cpu]: https://github.com/akopytov/sysbench/blob/master/src/tests/cpu/sb_cpu.c
[sysbench-memory]: https://github.com/akopytov/sysbench/blob/master/src/tests/memory/sb_memory.c
[sysbench-core]: https://github.com/akopytov/sysbench/blob/master/src/sysbench.c
[sysbench-timer]: https://github.com/akopytov/sysbench/blob/master/src/sb_timer.h
[sysbench-release]: https://github.com/akopytov/sysbench/releases/tag/1.0.20
[sysbench-meta]: https://api.github.com/repos/akopytov/sysbench
[void-sysbench]: https://github.com/void-linux/void-packages/blob/master/srcpkgs/sysbench/template
[schbench]: https://kernel.googlesource.com/pub/scm/linux/kernel/git/mason/schbench/+/6300b8f3a8922c61ea6bb2cdfa1901a42c0cc6fc/README.md
[schbench-source]: https://kernel.googlesource.com/pub/scm/linux/kernel/git/mason/schbench/+/6300b8f3a8922c61ea6bb2cdfa1901a42c0cc6fc/schbench.c
[stress-readme]: https://github.com/ColinIanKing/stress-ng/blob/V0.22.01/README.md
[stress-man]: https://github.com/ColinIanKing/stress-ng/blob/V0.22.01/stress-ng.1
[stress-release]: https://github.com/ColinIanKing/stress-ng/releases/tag/V0.22.01
[void-stress]: https://github.com/void-linux/void-packages/blob/master/srcpkgs/stress-ng/template
[stream-source]: https://www.cs.virginia.edu/stream/FTP/Code/stream.c
[stream-rules]: https://www.cs.virginia.edu/stream/ref.html
[lmbench-memory]: https://github.com/intel/lmbench/blob/master/doc/lat_mem_rd.8
[lmbench-readme]: https://github.com/intel/lmbench/blob/master/README
[lmbench-license]: https://github.com/intel/lmbench/blob/master/COPYING
[lmbench-meta]: https://api.github.com/repos/intel/lmbench
[cyclictest]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/src/cyclictest/cyclictest.8
[cyclictest-source]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/src/cyclictest/cyclictest.c
[rt-utils]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/src/lib/rt-utils.c
[rt-release]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6
[rt-makefile]: https://kernel.googlesource.com/pub/scm/utils/rt-tests/rt-tests/+/62da2befac98f811af8e56f2b7992fb09faa33d6/Makefile
[perf-stat]: https://github.com/torvalds/linux/blob/master/tools/perf/Documentation/perf-stat.txt
[kernel-license]: https://github.com/torvalds/linux/blob/master/COPYING
[perf-security]: https://docs.kernel.org/admin-guide/perf-security.html
[memtester-man]: https://sources.debian.org/data/main/m/memtester/4.7.1-1/memtester.8
[memtester-source]: https://sources.debian.org/data/main/m/memtester/4.7.1-1/memtester.c
[memtester-changelog]: https://sources.debian.org/data/main/m/memtester/4.7.1-1/CHANGELOG
[void-memtester]: https://github.com/void-linux/void-packages/blob/master/srcpkgs/memtester/template
[mlock]: https://man7.org/linux/man-pages/man2/mlock.2.html
[memtest-stable]: https://github.com/memtest86plus/memtest86plus/blob/v8.10/README.md
[memtest-main]: https://github.com/memtest86plus/memtest86plus/blob/main/README.md
[memtest-release]: https://github.com/memtest86plus/memtest86plus/releases/tag/v8.10
[edac-abi]: https://github.com/torvalds/linux/blob/master/Documentation/ABI/testing/sysfs-devices-edac
[edac]: https://docs.kernel.org/driver-api/edac.html
[rasdaemon]: https://github.com/mchehab/rasdaemon/blob/master/README.rst
[ras-meta]: https://api.github.com/repos/mchehab/rasdaemon
[psi]: https://docs.kernel.org/accounting/psi.html
[psi-420]: https://github.com/torvalds/linux/blob/v4.20/Documentation/accounting/psi.txt
[kconfig]: https://github.com/torvalds/linux/blob/master/init/Kconfig
[proc]: https://docs.kernel.org/filesystems/proc.html
[cgroup]: https://docs.kernel.org/admin-guide/cgroup-v2.html
[cpufreq]: https://docs.kernel.org/admin-guide/pm/cpufreq.html
[numa]: https://docs.kernel.org/admin-guide/mm/numa_memory_policy.html
[thp]: https://docs.kernel.org/admin-guide/mm/transhuge.html
[thermal]: https://docs.kernel.org/driver-api/thermal/sysfs-api.html
+140
View File
@@ -0,0 +1,140 @@
# Defensible median scoring and comparison rules
Research for [Establish defensible median scoring and comparison rules](https://git.bongbetic.com/xavierk/odin/issues/7), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
**Accessed:** 25 September 2026. **Status:** decision evidence and recommendations; no final scoring formula, calibrated reference values, benchmark runs, or implementation. The user's requirement is a **median** score.
## What the median should mean
There are three separate choices:
| Level | Meaning | Main limitation |
|---|---|---|
| Median of repeated measurements | Typical result for one fixed workload under stated conditions | Hides occasional long stalls; does not combine CPU, storage and browser results |
| Median across normalized workloads | Typical relative performance across a fixed test set | Domains with many tests gain more influence; changing references can change rankings |
| Median across domain summaries | Typical relative performance across explicitly chosen domains | Domain definitions matter; poor performance in a minority of domains can disappear from the headline |
NIST defines the sample median as the middle observation, or the arithmetic average of the middle two for an even sample count. Its resistance to extremes is useful, but it is a measure of location, not completeness, reliability, or worst-case response. [NIST location][nist-location]
**Recommendation for discussion:** use medians to summarize repeated valid measurements, normalize against frozen references, form predefined domain summaries, then use a median across those domains for the requested headline. Preserve each stage and its raw inputs. This proposes an aggregation structure; the domain membership, weighting, reference values, repeat counts and numeric scale remain decisions.
Established suites demonstrate why the levels must stay explicit. SPEC CPU 2017 takes median execution times from three runs, or the slower of two, then uses a **geometric mean** across ratios. Speedometer uses inverse geometric means across test durations and arithmetic means across iterations. These are methodological precedents, not permission to substitute a geometric mean for Odin's requested median. Preserve a tool's native result under its original name; label Odin's further aggregation separately. [SPEC rules][spec-rules] [Speedometer methodology][speedometer]
## Normalization and category balance
Milliseconds, operations/second and GB/s cannot share a meaningful raw median. A candidate approach is a dimensionless ratio against the **same workload's** reference: observation/reference for positive higher-is-better measures, reference/observation for positive lower-is-better measures. SPEC uses the latter for elapsed-time ratios. The metric identity must include its units, workload, size, concurrency, timing boundaries and direction. Reject invalid/nonfinite inputs under declared validity rules; a zero measured duration must not become an infinite score. [SPEC overview][spec-overview]
**Fictional arithmetic examples throughout this report:** these numbers illustrate consequences only. They are not Odin calibration or measured hardware results. A displayed index of 100 at the reference is an arbitrary illustrative scale, not a recommendation.
| Fictional metric | Reference | Observed | Illustrative ratio / index |
|---|---:|---:|---:|
| Work throughput | 25 operations/s | 50 operations/s | 2.0 / 200 |
| Memory bandwidth | 25 GB/s | 20 GB/s | 0.8 / 80 |
| Request latency | 4 ms | 2 ms | 2.0 / 200 |
A median of these indices is 200, despite memory bandwidth being below reference. That is a consequence of the chosen statistic. Show domain detail and slow-tail measurements alongside it. “200 versus 100” describes this index; it does not establish that every application is twice as fast.
Fix the order of operations. For fictional repeated times `[1, 3]` ms and a 3 ms reference, normalizing the raw median gives `3 / 2 = 1.5`; taking the median of individual ratios `[3, 1]` gives 2. Even-count averaging and reciprocal normalization do not commute. The report format must specify which result it contains. [median definition][nist-location]
Balance domains before counting metrics. If ten CPU tests each score 160 and four other domains score 40, 80, 100 and 120, the flat fourteen-test median is 160. A median across five domain summaries is 100. Adding CPU subtests should not silently redefine the product's priorities. Similarly, adding Python/Rust/Java variants must not automatically multiply the language domain's influence. A median of domain medians is a deliberate hierarchical index, not the pooled median of all observations.
Exclude health counters, memory-test pass/fail, driver availability and installed RAM capacity from throughput arithmetic. Multiple correlated outputs from one workload—throughput, IOPS, average latency and several percentiles—also need an explicit selection rule before any becomes an independent scored contribution.
## Calibration is a substantive decision
SPEC establishes per-workload reference times on a named machine and publishes the calculation rules. Its documentation explains that reference changes preserve relative overall rankings for its geometric-mean calculation. **That invariance does not generally hold for a median across normalized metrics.** [SPEC reference explanation][spec-overview]
For three fictional higher-is-better workloads, let machine A produce `[1, 10, 10]` and B produce `[2, 2, 20]` in each workload's own units:
| Fictional reference vector | A's normalized results → median | B's normalized results → median | Ordering |
|---|---|---|---|
| `[1, 1, 1]` | `[1, 10, 10]` → 10 | `[2, 2, 20]` → 2 | A higher |
| `[1, 10, 10]` | `[1, 1, 1]` → 1 | `[2, 0.2, 2]` → 2 | B higher |
The measurements did not change. Changing the reference altered the relative scales and which observations occupied the middle. Retain the requested median, make the reference rationale public, and version reference changes rather than treating them as cosmetic rescaling.
| Reference option | What it supports | Decision cost |
|---|---|---|
| Named reference configuration | Auditable, fixed anchor with documented per-test measurements | One machine's balance influences the median; configurations and repeatability need validation |
| Frozen reference cohort | Per-test references from a documented collection of machines | Cohort selection, sampling bias and revision policy become part of the score |
| User's own baseline | Local before/after comparisons | A score relative to oneself cannot rank different machines |
Recommend evaluating a frozen, locally distributable calibration manifest. Include reference measurements and provenance, reference conditions, workload and artifact digests, units/directions, aggregation order, required domains and calibration identity. Store it with results so calculation remains reproducible offline. Raw observations must survive changes; a recalculated score should identify its new calibration and preserve the original.
An arbitrary scale factor is acceptable if described as an index. A claim such as “100 is the median Linux machine,” a percentile rank, or a universal poor/good threshold requires representative population evidence that does not exist yet. Separately normalizing each architecture to its own average would also prevent interpreting those numbers as one common cross-architecture scale.
## Repetitions, warmup and uncertainty
Google Benchmark documents warmup, repetitions, median, standard deviation and coefficient of variation; it distinguishes user-visible wall time from CPU consumption. NIST recommends examining ordered observations for changing location/spread and says potential outliers should not simply be deleted when their cause is unknown. These support retaining all repeated observations and their conditions. [Google Benchmark][google-guide] [NIST run sequence][nist-runseq] [NIST outliers][nist-outliers]
Recommended measurement rules:
- Define warmup separately for each workload and retain its duration. Warm caches/JIT throughput, cold launch time and sustained thermal performance are different questions. Do not discard a slow first run from a declared cold-start test.
- Fix repetition and stopping rules before observing scores. A quick run may estimate a median without enough evidence for a useful confidence interval; it should not claim the precision of the standard profile. Calibrate the minimum repeats against the 10–20 minute budget.
- Retain run order, warmup, elapsed time, temperatures/power context, competing activity and invalidation reasons. A drifting sequence is not interchangeable independent noise. Repetitions within one process or thermal episode are not automatically independent runs.
- Exclude observations only for declared validity failures such as incorrect output, changed workload, cancellation or protocol failure. Preserve them with reasons. A slow but valid run can represent the usability problem Odin is meant to reveal.
- Show central spread such as MAD or IQR, plus tails where the workload supplies enough events. NIST defines MAD and IQR as distinct measures of spread; neither is itself a confidence interval. A median across repeated p99 values must not be labelled the p99 of all requests. [NIST scale][nist-scale]
NIST documents median confidence intervals based on order statistics/binomial probabilities, interpolated methods and bootstrap alternatives. Choose and validate a median-appropriate method; do not apply a mean's standard-error formula to a median. Confidence also depends on sample count and assumptions about the measurements. For an aggregate, account for shared run-level variation and state whether uncertainty in the calibration reference is included. The spread **between different domain scores** is not sampling uncertainty about the headline. [NIST median intervals][nist-median-ci]
Keep a graph of results in time order. Google documents CPU selection, boost, scheduler contention, SMT, caches and NUMA as variance sources. Its suggestions for controlled laboratory microbenchmarks include changing system settings; Odin's installed-system baseline should record existing conditions and label any tuned experiment separately. The existing [CPU/memory report][odin-cpu] and [portability/UI report][odin-portability] explain TUI interference and qualification needs. Stable repeated numbers alone do not prove that a workload represents real usability. [Google variance][google-variance]
## Comparability must be attached to every score
SPEC requires performance-relevant observation conditions and valid workload outputs; its CPU suite intentionally measures processor, memory subsystem **and compilers**. Even a fixed-toolchain comparison describes a defined software/hardware configuration. [SPEC rules][spec-rules] [SPEC overview][spec-overview]
Recommend two explicit comparison purposes:
- **Installed-system usability:** the chosen installed browser, shell, runtimes, drivers and kernel are part of what is measured. Version/configuration changes may explain a score change without any hardware change.
- **Controlled reference workload:** fixed workload assets, runtime/compiler contracts, flags, input data and execution modes improve comparison across machines. Architecture-specific artifacts must implement equivalent declared work and validate outputs; different ISA policies need disclosure.
Neither mode needs to masquerade as a pure hardware measurement. Keep their result identities distinct. Kernel/libc/distro differences can be the subject of a comparison, but they must be visible and the workload contract must remain equivalent.
A comparison identity should include suite/scoring/calibration versions, workload set, run profile, tool/artifact versions, options and data digests, timing/aggregation rules, browser mode, hardware/virtualization context and validity/coverage. Require a documented equivalence decision before combining scores across changed tools or workloads. SPEC warns that scores across different suite generations generally cannot be converted. [SPEC overview][spec-overview]
Speedometer 3.1 instructs users to use a clean browser profile, close competing programs/tabs, keep its page focused, avoid device interaction, use AC power and allow cooling when needed. Its official UI computes a 95% interval around its **arithmetic mean**; that interval cannot be attached to Odin's median unchanged. Preserve native browser score/uncertainty and label any median of complete runs separately. Headed and headless measurements need separate identities until an equivalence study justifies any shared interpretation; background versus foreground execution is also material. [instructions][speedometer-instructions] [3.1 result code][speedometer-main]
VM results characterize the guest allocation and virtualization environment. Keep native, virtualized and emulated cohorts identifiable; VM compatibility success does not establish native performance. Storage cache mode, queue depth, engine, filesystem and durability policy similarly belong to the workload identity. A fallback such as buffered I/O cannot silently replace a direct-I/O measurement with the same scoring identity. [CPU/memory report][odin-cpu] [storage report][odin-storage] [portability report][odin-portability]
## Missing tests and eligibility
For fictional domain indices `[40, 80, 100, 120, 160]`, the complete median is 100. Omitting 40 produces 110; omitting both 40 and 80 produces 120. Available-only aggregation can reward absent or deliberately skipped weak components.
**Recommendation:** define a versioned required set for the full score. Permit a clearly named partial median and domain results when the full set is unavailable, with the exact subset identified. Compare partial scores only over the same compatible subset; a pairwise intersection comparison must recompute **both** results and label that narrower scope. An optional pack must not silently change the headline's membership.
Keep distinct outcomes: completed-valid, completed-with-limitations, unsupported, missing dependency, permission denied, unsafe to run, cancelled, timed out, and failed validation. The eventual validity contract decides whether a limited result remains score-eligible. Never impute missing results as zero, a reference score, or a healthy pass. A required workload failing verification makes the full score ineligible, while preserving completed measurements and the associated finding. Good numbers from other domains should not cancel that failure.
The user accepted reporting unavailable tests across Linux targets. That does not resolve which domains are mandatory, whether every machine should still display a partial number, or how partial results should look. Those are explicit product decisions.
## Presentation, recommendations and open decisions
Recommend a headline containing the median, score identity, full/partial state, and eligible-domain coverage. The next view should show domain values, raw units, repeat count/spread, tail latency, invalidations and reference details. Show health findings beside performance: a fast drive with serious SMART evidence still needs attention. Missing telemetry must stay unknown. The storage and CPU reports establish why speed cannot determine drive replacement or certify memory health. [storage][odin-storage] [CPU/memory][odin-cpu]
Optimization advice should cite the observation and matching rule: for example, measured foreground stalls plus pressure evidence can support investigating memory contention. A low normalized score alone does not identify its cause. Keep severity of health evidence, completeness of coverage, measurement uncertainty and performance position as separate concepts; avoid one synthetic “confidence/health” percentage that mixes them.
Before implementation, decide:
1. The headline's median level, domain membership and balancing rules; whether responsiveness contributes or remains an accompanying measurement.
2. The reference configuration/cohort, scale and calibration-release policy; collect actual calibration data before inventing thresholds.
3. Required versus optional coverage, partial-score display and exact comparison eligibility.
4. Installed-system versus controlled-workload defaults, architecture/ISA policies, browser modes and VM cohorts.
5. Repetition/warmup/stopping rules, outlier validity rules, median interval method and honest quick/standard/extended precision claims.
6. Evidence requirements for optimization rules and separation of urgent health findings from the score.
**Evidence limits:** no calibration population, repeatability measurements, TUI-overhead budget or headed/headless equivalence study was produced. Examples are arithmetic demonstrations only. Context7 successfully resolved Google Benchmark; two BrowserBench/Speedometer searches returned unrelated packages, so its official repository and deployed 3.1 documentation were inspected directly. NIST and SPEC sources were inspected directly as statistical and benchmark-methodology references. This report neither adopts SPEC's workloads nor claims that their aggregation rules are Odin's final design.
[nist-location]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda351.htm
[nist-scale]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda356.htm
[nist-outliers]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda35h.htm
[nist-runseq]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda33p.htm
[nist-median-ci]: https://www.itl.nist.gov/div898/software/dataplot/refman1/auxillar/mediancl.htm
[spec-rules]: https://www.spec.org/cpu2017/Docs/runrules.html
[spec-overview]: https://www.spec.org/cpu2017/Docs/overview.html
[google-guide]: https://github.com/google/benchmark/blob/main/docs/user_guide.md
[google-variance]: https://github.com/google/benchmark/blob/main/docs/reducing_variance.md
[speedometer]: https://github.com/WebKit/Speedometer/blob/main/README.md
[speedometer-instructions]: https://browserbench.org/Speedometer3.1/instructions.html
[speedometer-main]: https://browserbench.org/Speedometer3.1/resources/main.mjs
[odin-cpu]: https://git.bongbetic.com/xavierk/odin/src/commit/90213d7f62cbd118f08ea8ff2f8042e94aa038a7/docs/research/cpu-memory.md
[odin-storage]: https://git.bongbetic.com/xavierk/odin/src/commit/6e5af87a64faedd4a8ad31ba10d9be4b499e8349/docs/research/storage-health.md
[odin-portability]: https://git.bongbetic.com/xavierk/odin/src/commit/20681cd4f184a9fc0164dd638a252de44ef230f5/docs/research/portability-ui.md