Define Odin’s median score and comparability contract #10

Closed
opened 2026-09-25 17:48:51 +00:00 by xavierk · 2 comments
Owner

Part of Find the way to Odin’s build-ready specification.

Question

What precisely is the headline median score: which comparable normalized observations enter it, what baseline/scale is used, and what does it claim about the system? Set per-test repetitions, category balance, calibration and validation method, complete/partial eligibility, missing-data policy, score/version compatibility, confidence indicators, and whether results describe installed-system usability or a fixed reference environment. Preserve the user’s median requirement and expose bottlenecks that a median can hide. Define how health findings remain visible and whether they affect score eligibility.

Resolve with the human using sourced alternatives and concrete example runs; do not fabricate a real calibration corpus.

Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). <!-- wayfinder-map: 1 --> ## Question What precisely is the headline median score: which comparable normalized observations enter it, what baseline/scale is used, and what does it claim about the system? Set per-test repetitions, category balance, calibration and validation method, complete/partial eligibility, missing-data policy, score/version compatibility, confidence indicators, and whether results describe installed-system usability or a fixed reference environment. Preserve the user’s median requirement and expose bottlenecks that a median can hide. Define how health findings remain visible and whether they affect score eligibility. Resolve with the human using sourced alternatives and concrete example runs; do not fabricate a real calibration corpus.
xavierk added the wayfinder:grilling label 2026-09-25 17:48:51 +00:00
xavierk added a new dependency 2026-09-25 17:49:52 +00:00
xavierk added a new dependency 2026-09-25 17:49:54 +00:00
xavierk added a new dependency 2026-09-25 17:49:58 +00:00
xavierk added a new dependency 2026-09-25 17:50:02 +00:00
xavierk added a new dependency 2026-09-25 17:50:06 +00:00
xavierk added a new dependency 2026-09-25 17:50:09 +00:00
xavierk self-assigned this 2026-09-27 05:26:48 +00:00
Author
Owner

Resolution

The human accepted all recommendations in three live rounds. This is the scoring contract; it does not claim that calibration or implementation has occurred.

Meaning and aggregation

The headline describes installed-system performance by default. Prepared reference-software runs are separately labelled; neither isolates hardware performance. The headline is the median of seven equally represented domain scores: CPU, memory bandwidth, GPU, storage, browser, Bash, and languages. Loaded responsiveness and health findings remain prominently visible alongside the score, outside its aggregation.

Normalize each workload against a fixed, versioned reference value. For higher-is-better measurements, score = 100 × measurement/reference; for lower-is-better measurements, score = 100 × reference/measurement. Thus 100 means reference performance, twice the throughput scores 200, and half the startup time scores 200. Real calibration measurements must establish reference values before numerical Odin scores ship; there is no moving population average or invented baseline.

For repeated tests, take the median in native units first, then normalize. At every median stage, an even-sized population uses the arithmetic average of the middle two values. Retain the approved workload definitions and native statistics from Choose Odin’s workload suite and run profiles.

Fixed domain aggregation:

  • CPU: median of normalized single-worker and parallel throughput.
  • Memory bandwidth: median of normalized STREAM Copy and Triad measurements, preserving STREAM's native within-run statistic before taking the cross-run median.
  • GPU: equally represented compute and rendering summaries. Compute summarizes the selected FP32 and global-memory-bandwidth measurements; rendering summarizes the selected shading and texture scenes. Use medians within each summary, then the median of the two summaries.
  • Storage: median of normalized sequential read, sequential write, random read, and random write performance; sequential bandwidth and random IOPS represent the four fixed workloads. Latency distributions remain diagnostic detail rather than additional votes.
  • Browser: normalize the complete native Speedometer score. Standard runs one complete suite; do not replace Speedometer's native aggregation with a median of its internal subtests.
  • Bash: median of normalized direct startup, interactive prompt readiness, and arithmetic throughput.
  • Languages: each of Python, Rust, C++, and Java has a score combining normalized startup and warmed throughput by median; the domain is the median of those four equally represented language scores.

Always show the domain breakdown: illustrative scores 20, 80, 90, 100, 110, 120, 130 yield a headline of 100 despite one severe weakness. These are fictional arithmetic values, not calibration evidence.

Coverage and profiles

A missing required measurement makes its domain score unavailable. A missing domain withholds the full headline. Never substitute zero, silently reduce membership, or emit a partial aggregate in v1. Show completed domain scores, raw measurements, and explicit coverage instead.

Only a completed standard measurement set can produce the headline. Quick provides diagnostics and separately identified probes. Extended modules add detail without changing headline membership; an extended run can show a headline only when it also completes the unchanged standard set.

Calibration and comparisons

Use a frozen calibration release backed by published raw runs and a reproducible procedure. Each workload has one reference value shared across supported architectures, without architecture-specific rescaling. Record reference hardware and software, and validate scores on multiple physical systems before release. Until this evidence exists, display raw results without numerical Odin scores.

Direct score comparisons require matching scoring rules, calibration, workload protocols, and execution modes. Keep headed/headless browser, installed/reference software, direct/buffered storage, and virtual/physical results separate. Installed software versions may differ for useful before/after upgrade comparisons, but identify the change and do not attribute it solely to hardware. Calibration or scoring-rule changes do not silently preserve comparability.

The concrete calibration acquisition procedure and release acceptance criteria need a follow-on planning decision. This ticket establishes the policy, not a real reference corpus; running calibration belongs to the later implementation/validation effort.

Repetition, uncertainty, and validity

Retain approved repetitions: three trials for most workloads, 30 startup samples, and one complete Speedometer suite in standard. Show sample counts, ranges, and native workload statistics. Do not invent a headline confidence interval from three trials or from between-domain variation. Preserve slow valid observations; slowness alone does not invalidate a result. Preserve warmup/protocol rules already agreed in the workload decision.

Health findings never numerically penalize speed. A drive warning stays prominently visible beside an otherwise valid score. Detected memory errors or failed output verification withhold the headline because measurement integrity is uncertain. Safety stops or unfinished required workloads also withhold it. Missing health telemetry means health unknown, not healthy, and does not alone suppress a complete performance score. Detailed evidence thresholds and stop rules remain owned by Define health findings and the optimization advice rubric.

Evidence and remaining planning

This decision builds on Establish defensible median scoring and comparison rules and the approved workload contract. No benchmark or calibration run was performed. Physical repeatability and UI-overhead gates remain in Set distro, VM, and physical-hardware validation gates. The follow-on calibration ticket must settle the reference selection, exact collection/warmup protocol, and acceptance procedure before specification approval.

## Resolution The human accepted all recommendations in three live rounds. This is the scoring contract; it does not claim that calibration or implementation has occurred. ### Meaning and aggregation The headline describes installed-system performance by default. Prepared reference-software runs are separately labelled; neither isolates hardware performance. The headline is the median of seven equally represented domain scores: CPU, memory bandwidth, GPU, storage, browser, Bash, and languages. Loaded responsiveness and health findings remain prominently visible alongside the score, outside its aggregation. Normalize each workload against a fixed, versioned reference value. For higher-is-better measurements, score = 100 × measurement/reference; for lower-is-better measurements, score = 100 × reference/measurement. Thus 100 means reference performance, twice the throughput scores 200, and half the startup time scores 200. Real calibration measurements must establish reference values before numerical Odin scores ship; there is no moving population average or invented baseline. For repeated tests, take the median in native units first, then normalize. At every median stage, an even-sized population uses the arithmetic average of the middle two values. Retain the approved workload definitions and native statistics from [Choose Odin’s workload suite and run profiles](https://git.bongbetic.com/xavierk/odin/issues/8). Fixed domain aggregation: - CPU: median of normalized single-worker and parallel throughput. - Memory bandwidth: median of normalized STREAM Copy and Triad measurements, preserving STREAM's native within-run statistic before taking the cross-run median. - GPU: equally represented compute and rendering summaries. Compute summarizes the selected FP32 and global-memory-bandwidth measurements; rendering summarizes the selected shading and texture scenes. Use medians within each summary, then the median of the two summaries. - Storage: median of normalized sequential read, sequential write, random read, and random write performance; sequential bandwidth and random IOPS represent the four fixed workloads. Latency distributions remain diagnostic detail rather than additional votes. - Browser: normalize the complete native Speedometer score. Standard runs one complete suite; do not replace Speedometer's native aggregation with a median of its internal subtests. - Bash: median of normalized direct startup, interactive prompt readiness, and arithmetic throughput. - Languages: each of Python, Rust, C++, and Java has a score combining normalized startup and warmed throughput by median; the domain is the median of those four equally represented language scores. Always show the domain breakdown: illustrative scores 20, 80, 90, 100, 110, 120, 130 yield a headline of 100 despite one severe weakness. These are fictional arithmetic values, not calibration evidence. ### Coverage and profiles A missing required measurement makes its domain score unavailable. A missing domain withholds the full headline. Never substitute zero, silently reduce membership, or emit a partial aggregate in v1. Show completed domain scores, raw measurements, and explicit coverage instead. Only a completed standard measurement set can produce the headline. Quick provides diagnostics and separately identified probes. Extended modules add detail without changing headline membership; an extended run can show a headline only when it also completes the unchanged standard set. ### Calibration and comparisons Use a frozen calibration release backed by published raw runs and a reproducible procedure. Each workload has one reference value shared across supported architectures, without architecture-specific rescaling. Record reference hardware and software, and validate scores on multiple physical systems before release. Until this evidence exists, display raw results without numerical Odin scores. Direct score comparisons require matching scoring rules, calibration, workload protocols, and execution modes. Keep headed/headless browser, installed/reference software, direct/buffered storage, and virtual/physical results separate. Installed software versions may differ for useful before/after upgrade comparisons, but identify the change and do not attribute it solely to hardware. Calibration or scoring-rule changes do not silently preserve comparability. The concrete calibration acquisition procedure and release acceptance criteria need a follow-on planning decision. This ticket establishes the policy, not a real reference corpus; running calibration belongs to the later implementation/validation effort. ### Repetition, uncertainty, and validity Retain approved repetitions: three trials for most workloads, 30 startup samples, and one complete Speedometer suite in standard. Show sample counts, ranges, and native workload statistics. Do not invent a headline confidence interval from three trials or from between-domain variation. Preserve slow valid observations; slowness alone does not invalidate a result. Preserve warmup/protocol rules already agreed in the workload decision. Health findings never numerically penalize speed. A drive warning stays prominently visible beside an otherwise valid score. Detected memory errors or failed output verification withhold the headline because measurement integrity is uncertain. Safety stops or unfinished required workloads also withhold it. Missing health telemetry means health unknown, not healthy, and does not alone suppress a complete performance score. Detailed evidence thresholds and stop rules remain owned by [Define health findings and the optimization advice rubric](https://git.bongbetic.com/xavierk/odin/issues/11). ### Evidence and remaining planning This decision builds on [Establish defensible median scoring and comparison rules](https://git.bongbetic.com/xavierk/odin/issues/7) and the approved workload contract. No benchmark or calibration run was performed. Physical repeatability and UI-overhead gates remain in [Set distro, VM, and physical-hardware validation gates](https://git.bongbetic.com/xavierk/odin/issues/14). The follow-on calibration ticket must settle the reference selection, exact collection/warmup protocol, and acceptance procedure before specification approval.
xavierk added a new dependency 2026-09-27 05:45:47 +00:00
Author
Owner

The Speedometer-specific browser statistic and reference are superseded by Choose a legally distributable v1 browser workload. Specify Odin’s v1 browser interaction workload will set the replacement statistic and identity. Seven-domain aggregation and coverage rules here stand.

The Speedometer-specific browser statistic and reference are superseded by [Choose a legally distributable v1 browser workload](https://git.bongbetic.com/xavierk/odin/issues/19#issuecomment-6899). [Specify Odin’s v1 browser interaction workload](https://git.bongbetic.com/xavierk/odin/issues/20) will set the replacement statistic and identity. Seven-domain aggregation and coverage rules here stand.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/odin#10