Define Odin’s median score and comparability contract #10
Notifications
Due Date
No due date set.
Blocks
Depends on
#13 Evaluate the Bongbetic TUI workflow with InkUI
xavierk/odin
#15 Approve Odin’s build-ready specification
xavierk/odin
#8 Choose Odin’s workload suite and run profiles
xavierk/odin
Reference: xavierk/odin#10
Reference in New Issue
Block a user
Part of Find the way to Odin’s build-ready specification.
Question
What precisely is the headline median score: which comparable normalized observations enter it, what baseline/scale is used, and what does it claim about the system? Set per-test repetitions, category balance, calibration and validation method, complete/partial eligibility, missing-data policy, score/version compatibility, confidence indicators, and whether results describe installed-system usability or a fixed reference environment. Preserve the user’s median requirement and expose bottlenecks that a median can hide. Define how health findings remain visible and whether they affect score eligibility.
Resolve with the human using sourced alternatives and concrete example runs; do not fabricate a real calibration corpus.
Resolution
The human accepted all recommendations in three live rounds. This is the scoring contract; it does not claim that calibration or implementation has occurred.
Meaning and aggregation
The headline describes installed-system performance by default. Prepared reference-software runs are separately labelled; neither isolates hardware performance. The headline is the median of seven equally represented domain scores: CPU, memory bandwidth, GPU, storage, browser, Bash, and languages. Loaded responsiveness and health findings remain prominently visible alongside the score, outside its aggregation.
Normalize each workload against a fixed, versioned reference value. For higher-is-better measurements, score = 100 × measurement/reference; for lower-is-better measurements, score = 100 × reference/measurement. Thus 100 means reference performance, twice the throughput scores 200, and half the startup time scores 200. Real calibration measurements must establish reference values before numerical Odin scores ship; there is no moving population average or invented baseline.
For repeated tests, take the median in native units first, then normalize. At every median stage, an even-sized population uses the arithmetic average of the middle two values. Retain the approved workload definitions and native statistics from Choose Odin’s workload suite and run profiles.
Fixed domain aggregation:
Always show the domain breakdown: illustrative scores 20, 80, 90, 100, 110, 120, 130 yield a headline of 100 despite one severe weakness. These are fictional arithmetic values, not calibration evidence.
Coverage and profiles
A missing required measurement makes its domain score unavailable. A missing domain withholds the full headline. Never substitute zero, silently reduce membership, or emit a partial aggregate in v1. Show completed domain scores, raw measurements, and explicit coverage instead.
Only a completed standard measurement set can produce the headline. Quick provides diagnostics and separately identified probes. Extended modules add detail without changing headline membership; an extended run can show a headline only when it also completes the unchanged standard set.
Calibration and comparisons
Use a frozen calibration release backed by published raw runs and a reproducible procedure. Each workload has one reference value shared across supported architectures, without architecture-specific rescaling. Record reference hardware and software, and validate scores on multiple physical systems before release. Until this evidence exists, display raw results without numerical Odin scores.
Direct score comparisons require matching scoring rules, calibration, workload protocols, and execution modes. Keep headed/headless browser, installed/reference software, direct/buffered storage, and virtual/physical results separate. Installed software versions may differ for useful before/after upgrade comparisons, but identify the change and do not attribute it solely to hardware. Calibration or scoring-rule changes do not silently preserve comparability.
The concrete calibration acquisition procedure and release acceptance criteria need a follow-on planning decision. This ticket establishes the policy, not a real reference corpus; running calibration belongs to the later implementation/validation effort.
Repetition, uncertainty, and validity
Retain approved repetitions: three trials for most workloads, 30 startup samples, and one complete Speedometer suite in standard. Show sample counts, ranges, and native workload statistics. Do not invent a headline confidence interval from three trials or from between-domain variation. Preserve slow valid observations; slowness alone does not invalidate a result. Preserve warmup/protocol rules already agreed in the workload decision.
Health findings never numerically penalize speed. A drive warning stays prominently visible beside an otherwise valid score. Detected memory errors or failed output verification withhold the headline because measurement integrity is uncertain. Safety stops or unfinished required workloads also withhold it. Missing health telemetry means health unknown, not healthy, and does not alone suppress a complete performance score. Detailed evidence thresholds and stop rules remain owned by Define health findings and the optimization advice rubric.
Evidence and remaining planning
This decision builds on Establish defensible median scoring and comparison rules and the approved workload contract. No benchmark or calibration run was performed. Physical repeatability and UI-overhead gates remain in Set distro, VM, and physical-hardware validation gates. The follow-on calibration ticket must settle the reference selection, exact collection/warmup protocol, and acceptance procedure before specification approval.
The Speedometer-specific browser statistic and reference are superseded by Choose a legally distributable v1 browser workload. Specify Odin’s v1 browser interaction workload will set the replacement statistic and identity. Seven-domain aggregation and coverage rules here stand.