Establish defensible median scoring and comparison rules #7
Notifications
Due Date
No due date set.
Blocks
Reference: xavierk/odin#7
Reference in New Issue
Block a user
Part of Find the way to Odin’s build-ready specification.
Question
How can Odin produce the requested single median performance score without combining incompatible units, hiding missing coverage, or presenting health risk as speed?
Research established benchmark measurement/statistics practices from primary sources. Compare per-trial medians, medians across normalized tests/domains, category balance, lower/higher-is-better metrics, calibration/reference baselines, uncertainty/repetition/warmup/outliers, system noise and thermal drift. Analyze what a median hides and how to retain diagnostic detail. Address installed-system usability versus fixed-toolchain hardware comparisons, browser execution modes, changing suites/versions, VM versus physical results, missing tests, full/partial score eligibility and locally reproducible calibration. User explicitly means median; do not silently substitute geometric mean. Investigate how to present separately the evidence, coverage, confidence and health/optimization rubric and what must be decided by the human.
Deliver sourced alternatives and worked hypothetical examples clearly labeled as illustrative, not measured data. Do not invent a calibrated baseline or final numeric formula.
Research resolution
The report preserves the requested median while distinguishing three levels: typical repeated result for one workload, median across normalized workloads, and median across domain summaries. A staged median-based design is a candidate, not an adopted formula.
Read the cited research report. Evidence is recorded at commit
f384c81f0fe3onresearch/scoring.Still for the human decision tickets: median level and domain balance; reference configuration or cohort; scale and calibration release policy; required/optional domains; partial-score presentation; comparison cohorts; repetition/warmup/interval methods; and the diagnostic evidence rubric.
Evidence limits: this report supplies inspected methodology and clearly fictional arithmetic examples, not calibration data, repeatability measurements, a UI-overhead budget or headed/headless equivalence proof. Context7 resolved Google Benchmark but returned unrelated BrowserBench matches; official BrowserBench, NIST and SPEC sources were inspected directly. No benchmarks were run.