Establish defensible median scoring and comparison rules #7

Closed
opened 2026-09-25 17:48:49 +00:00 by xavierk · 1 comment
Owner

Part of Find the way to Odin’s build-ready specification.

Question

How can Odin produce the requested single median performance score without combining incompatible units, hiding missing coverage, or presenting health risk as speed?

Research established benchmark measurement/statistics practices from primary sources. Compare per-trial medians, medians across normalized tests/domains, category balance, lower/higher-is-better metrics, calibration/reference baselines, uncertainty/repetition/warmup/outliers, system noise and thermal drift. Analyze what a median hides and how to retain diagnostic detail. Address installed-system usability versus fixed-toolchain hardware comparisons, browser execution modes, changing suites/versions, VM versus physical results, missing tests, full/partial score eligibility and locally reproducible calibration. User explicitly means median; do not silently substitute geometric mean. Investigate how to present separately the evidence, coverage, confidence and health/optimization rubric and what must be decided by the human.

Deliver sourced alternatives and worked hypothetical examples clearly labeled as illustrative, not measured data. Do not invent a calibrated baseline or final numeric formula.

Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). <!-- wayfinder-map: 1 --> ## Question How can Odin produce the requested single median performance score without combining incompatible units, hiding missing coverage, or presenting health risk as speed? Research established benchmark measurement/statistics practices from primary sources. Compare per-trial medians, medians across normalized tests/domains, category balance, lower/higher-is-better metrics, calibration/reference baselines, uncertainty/repetition/warmup/outliers, system noise and thermal drift. Analyze what a median hides and how to retain diagnostic detail. Address installed-system usability versus fixed-toolchain hardware comparisons, browser execution modes, changing suites/versions, VM versus physical results, missing tests, full/partial score eligibility and locally reproducible calibration. User explicitly means median; do not silently substitute geometric mean. Investigate how to present separately the evidence, coverage, confidence and health/optimization rubric and what must be decided by the human. Deliver sourced alternatives and worked hypothetical examples clearly labeled as illustrative, not measured data. Do not invent a calibrated baseline or final numeric formula.
xavierk added the wayfinder:research label 2026-09-25 17:48:49 +00:00
xavierk added a new dependency 2026-09-25 17:49:52 +00:00
xavierk added a new dependency 2026-09-25 17:49:57 +00:00
xavierk self-assigned this 2026-09-25 18:05:54 +00:00
Author
Owner

Research resolution

The report preserves the requested median while distinguishing three levels: typical repeated result for one workload, median across normalized workloads, and median across domain summaries. A staged median-based design is a candidate, not an adopted formula.

  • Raw milliseconds, throughput and bandwidth cannot be combined directly. Each metric needs a defined workload, direction, units, validity rules and normalization reference; aggregation order matters, including for even-count medians and reciprocal normalization.
  • A flat median can be dominated by domains with more subtests. Fixed domain membership and balancing rules are substantive product choices.
  • Fictional arithmetic examples demonstrate that changing per-test reference baselines can reverse machine rankings under median aggregation. Keep calibration rationale, measurements and identity immutable and versioned; no real baseline was invented here.
  • Removing weak domains can raise an available-only median. The report recommends a fixed required set for a full score, explicit partial coverage, and comparing partial scores only over identical compatible subsets.
  • Repetition, warmup, validity and stopping rules need to precede observed results. Preserve slow valid observations, time order, spread and tails. Median uncertainty requires an appropriate method; a mean’s interval or between-domain spread is not interchangeable with it.
  • Retain native suite results, including Speedometer’s own mean and interval, under their official identities. Any Odin median of complete runs is a separately labeled statistic.
  • Installed-system usability and controlled-reference comparisons have different meanings. Workload/tool/profile/calibration versions, browser mode, VM context, configuration and fallback behavior must travel with each result.
  • Health findings, capability coverage and uncertainty should remain visible beside performance. A low index alone cannot diagnose its cause; no other domain’s speed can certify memory integrity or cancel urgent drive evidence.

Read the cited research report. Evidence is recorded at commit f384c81f0fe3 on research/scoring.

Still for the human decision tickets: median level and domain balance; reference configuration or cohort; scale and calibration release policy; required/optional domains; partial-score presentation; comparison cohorts; repetition/warmup/interval methods; and the diagnostic evidence rubric.

Evidence limits: this report supplies inspected methodology and clearly fictional arithmetic examples, not calibration data, repeatability measurements, a UI-overhead budget or headed/headless equivalence proof. Context7 resolved Google Benchmark but returned unrelated BrowserBench matches; official BrowserBench, NIST and SPEC sources were inspected directly. No benchmarks were run.

<!-- wayfinder-research-resolution: scoring --> ## Research resolution The report preserves the requested median while distinguishing three levels: typical repeated result for one workload, median across normalized workloads, and median across domain summaries. A staged median-based design is a candidate, not an adopted formula. - Raw milliseconds, throughput and bandwidth cannot be combined directly. Each metric needs a defined workload, direction, units, validity rules and normalization reference; aggregation order matters, including for even-count medians and reciprocal normalization. - A flat median can be dominated by domains with more subtests. Fixed domain membership and balancing rules are substantive product choices. - Fictional arithmetic examples demonstrate that changing per-test reference baselines can reverse machine rankings under median aggregation. Keep calibration rationale, measurements and identity immutable and versioned; no real baseline was invented here. - Removing weak domains can raise an available-only median. The report recommends a fixed required set for a full score, explicit partial coverage, and comparing partial scores only over identical compatible subsets. - Repetition, warmup, validity and stopping rules need to precede observed results. Preserve slow valid observations, time order, spread and tails. Median uncertainty requires an appropriate method; a mean’s interval or between-domain spread is not interchangeable with it. - Retain native suite results, including Speedometer’s own mean and interval, under their official identities. Any Odin median of complete runs is a separately labeled statistic. - Installed-system usability and controlled-reference comparisons have different meanings. Workload/tool/profile/calibration versions, browser mode, VM context, configuration and fallback behavior must travel with each result. - Health findings, capability coverage and uncertainty should remain visible beside performance. A low index alone cannot diagnose its cause; no other domain’s speed can certify memory integrity or cancel urgent drive evidence. [Read the cited research report](https://git.bongbetic.com/xavierk/odin/src/commit/f384c81f0fe3a8cb464ff33bc7d2acd15f3e3025/docs/research/scoring.md). Evidence is recorded at commit `f384c81f0fe3` on `research/scoring`. **Still for the human decision tickets:** median level and domain balance; reference configuration or cohort; scale and calibration release policy; required/optional domains; partial-score presentation; comparison cohorts; repetition/warmup/interval methods; and the diagnostic evidence rubric. **Evidence limits:** this report supplies inspected methodology and clearly fictional arithmetic examples, not calibration data, repeatability measurements, a UI-overhead budget or headed/headless equivalence proof. Context7 resolved Google Benchmark but returned unrelated BrowserBench matches; official BrowserBench, NIST and SPEC sources were inspected directly. No benchmarks were run.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/odin#7