Decide with the human: reference hardware/software selection and rationale; how workload reference values are derived from repeated complete runs; exact warmup, preparation, ordering, cooldown, and environment controls not already fixed by the workload contract; required raw evidence and artifact identity; validation across multiple physical systems; sensitivity to baseline choice; numerical validity and stability acceptance criteria; and how fixes or later recalibration create a new comparison identity. Specify supported mode/protocol reference coverage without introducing architecture-specific rescaling. Assign concrete physical repeatability gates to the existing validation ticket rather than duplicating them.
The accepted policy is fixed: 100 means reference performance; one reference value per workload is shared across supported architectures; no numerical Odin scores ship until real calibration and validation evidence exists. Do not invent that evidence.
This ticket chooses a build-ready procedure and release gates, not an actual calibration corpus. Running benchmarks and publishing measured reference values belong to the later implementation/validation effort, outside this planning map. Surface access gaps as planning prerequisites only where they prevent choosing a viable procedure.
Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
<!-- wayfinder-map: 1 -->
## Question
Which reproducible reference-selection, collection, and acceptance procedure will produce Odin’s first frozen calibration release under [Define Odin’s median score and comparability contract](https://git.bongbetic.com/xavierk/odin/issues/10)?
Decide with the human: reference hardware/software selection and rationale; how workload reference values are derived from repeated complete runs; exact warmup, preparation, ordering, cooldown, and environment controls not already fixed by the workload contract; required raw evidence and artifact identity; validation across multiple physical systems; sensitivity to baseline choice; numerical validity and stability acceptance criteria; and how fixes or later recalibration create a new comparison identity. Specify supported mode/protocol reference coverage without introducing architecture-specific rescaling. Assign concrete physical repeatability gates to the existing validation ticket rather than duplicating them.
The accepted policy is fixed: 100 means reference performance; one reference value per workload is shared across supported architectures; no numerical Odin scores ship until real calibration and validation evidence exists. Do not invent that evidence.
This ticket chooses a build-ready procedure and release gates, not an actual calibration corpus. Running benchmarks and publishing measured reference values belong to the later implementation/validation effort, outside this planning map. Surface access gaps as planning prerequisites only where they prevent choosing a viable procedure.
The human accepted the recommendations across the live decision rounds. This is a calibration procedure and release policy, not a measured corpus. No numerical Odin score may ship until the applicable corpus and validation evidence exist.
Reference and scope
Use one fixed, documented physical x86_64 reference system for each v1 calibration release. Record why its CPU, memory, GPU, storage, and power configuration were selected; it is a reference point, not a claim that its hardware represents a population. One measured reference value per workload applies across x86_64 and aarch64. Do not rescale by architecture. Validate on at least two independent physical systems, including one aarch64 system.
Each scored protocol and execution mode needs real reference evidence and a distinct identity. This includes installed versus prepared reference browser/language software, headed versus headless browser, and direct versus buffered storage. Keep any other protocol differences that change measurement meaning distinct. A mode without a qualified reference corpus shows raw results only. Physical reference evidence does not certify VM scores: v1 records VM measurements and context but withholds numerical Odin scores in VMs until a separate virtual calibration protocol and evidence pass release gates.
Collection procedure
For each candidate scored mode, collect at least five valid complete standard runs on the reference system, spread across at least two calendar days and two cold boots. Use the fixed workload pack, inputs, tool builds, parser revisions, profile, and per-workload repetition/warmup rules from the approved workload contract. Retain the complete time-ordered set of attempts, including failures, timeouts, cancellations, and slow valid trials. A failed or ineligible attempt remains in the evidence but does not enter the reference calculation. Repeat collection until five eligible values exist for every required workload; never trim a valid value merely because it is slow.
Before each run, check capability, permissions, storage headroom, required artifacts, output verification, power state, and sensor availability. Capture a two-minute idle baseline. Record firmware, hardware identifiers, kernel, CPU entitlement and affinity, governor, scheduler, thermal and power state, mount and filesystem settings, GPU driver and API, browser/toolchain versions, runtime/tool versions, and available health telemetry. Do not tune the host automatically. Rotate domain order across the five complete runs and record the exact order; preserve each workload's specified internal order and warmup. Between runs, require two minutes with CPU utilization below 5% and measured temperature within 3 °C of the idle baseline. Stop waiting after 20 minutes and mark the next run ineligible for calibration if readiness is unmet. When temperature sensors are unavailable, require five idle minutes and flag reduced thermal evidence. Every deviation and interruption is recorded.
For each workload, use its approved native within-run statistic, then take the median of the five or more eligible run-level values in native units. For browser standard, one full Speedometer suite contributes its official score per run. The median across complete runs becomes the frozen reference value; the established scoring formula and even-count median rule remain unchanged. Do not substitute a shortened suite, a partial run, or an alternate tool version.
Freeze gates and sensitivity
A calibration candidate passes numerical validity only when every required reference workload has at least five eligible complete run values, every output check passes, and values are finite, positive, in the declared unit and direction, with matching protocol identity. Preserve all raw samples and exclusions with explicit reasons. Across-run median absolute deviation divided by the workload median must be at most 5%; for the complete browser suite, at most 10%. Investigate and rerun unstable candidates rather than discard valid outliers or relax limits after seeing results. These are corpus stability gates, not claims about physical-system repeatability; Set distro, VM, and physical-hardware validation gates owns the latter.
Run the same collection protocol on at least two independent physical validation systems, including one aarch64 system, to test workload validity and the shared-reference interpretation. With the candidate reference frozen, publish their domain and headline results where full coverage permits. Recompute scores using each validation system's eligible values as an alternate reference to expose baseline sensitivity. Flag any domain or headline shift greater than 10%, or any rank reversal, for documented review. A flag is not automatic rejection; unresolved protocol bias blocks the freeze. Do not infer architecture parity from these systems alone. The validation ticket sets the separate physical repeatability and UI-overhead thresholds.
Evidence and identity
Publish an immutable calibration manifest and supporting raw, timestamped run records. Include successful and failed outcomes; per-trial values and native statistics; output checks; environment snapshots; exact source commits, binary and asset SHA-256 hashes, licenses, parser and protocol revisions; run/order/mode IDs; reference calculations; exclusions; cross-system validation; sensitivity analysis; and reviewer rationale for flagged cases. Record an explicit calibration release ID and score-comparison identity. Publish the corpus before enabling numerical scores for its qualified modes.
A correction to reference values, workload protocol, score rule, or calculation creates a new calibration release and comparison identity. Old run records retain their original identity and score. Recalculation for a newer release, if offered, must be explicit and preserve the old result. No silent cross-release score comparison. Prepared versus installed software and other mode boundaries from the scoring contract remain separate.
This resolution introduces no actual measured reference value. Production collection and release validation occur after this planning map. No new child ticket is needed: physical repeatability, UI overhead, distro cells, and hardware access belong to the existing validation ticket.
## Resolution
The human accepted the recommendations across the live decision rounds. This is a calibration procedure and release policy, not a measured corpus. No numerical Odin score may ship until the applicable corpus and validation evidence exist.
### Reference and scope
Use one fixed, documented physical x86_64 reference system for each v1 calibration release. Record why its CPU, memory, GPU, storage, and power configuration were selected; it is a reference point, not a claim that its hardware represents a population. One measured reference value per workload applies across x86_64 and aarch64. Do not rescale by architecture. Validate on at least two independent physical systems, including one aarch64 system.
Each scored protocol and execution mode needs real reference evidence and a distinct identity. This includes installed versus prepared reference browser/language software, headed versus headless browser, and direct versus buffered storage. Keep any other protocol differences that change measurement meaning distinct. A mode without a qualified reference corpus shows raw results only. Physical reference evidence does not certify VM scores: v1 records VM measurements and context but withholds numerical Odin scores in VMs until a separate virtual calibration protocol and evidence pass release gates.
### Collection procedure
For each candidate scored mode, collect at least five valid complete standard runs on the reference system, spread across at least two calendar days and two cold boots. Use the fixed workload pack, inputs, tool builds, parser revisions, profile, and per-workload repetition/warmup rules from the approved workload contract. Retain the complete time-ordered set of attempts, including failures, timeouts, cancellations, and slow valid trials. A failed or ineligible attempt remains in the evidence but does not enter the reference calculation. Repeat collection until five eligible values exist for every required workload; never trim a valid value merely because it is slow.
Before each run, check capability, permissions, storage headroom, required artifacts, output verification, power state, and sensor availability. Capture a two-minute idle baseline. Record firmware, hardware identifiers, kernel, CPU entitlement and affinity, governor, scheduler, thermal and power state, mount and filesystem settings, GPU driver and API, browser/toolchain versions, runtime/tool versions, and available health telemetry. Do not tune the host automatically. Rotate domain order across the five complete runs and record the exact order; preserve each workload's specified internal order and warmup. Between runs, require two minutes with CPU utilization below 5% and measured temperature within 3 °C of the idle baseline. Stop waiting after 20 minutes and mark the next run ineligible for calibration if readiness is unmet. When temperature sensors are unavailable, require five idle minutes and flag reduced thermal evidence. Every deviation and interruption is recorded.
For each workload, use its approved native within-run statistic, then take the median of the five or more eligible run-level values in native units. For browser standard, one full Speedometer suite contributes its official score per run. The median across complete runs becomes the frozen reference value; the established scoring formula and even-count median rule remain unchanged. Do not substitute a shortened suite, a partial run, or an alternate tool version.
### Freeze gates and sensitivity
A calibration candidate passes numerical validity only when every required reference workload has at least five eligible complete run values, every output check passes, and values are finite, positive, in the declared unit and direction, with matching protocol identity. Preserve all raw samples and exclusions with explicit reasons. Across-run median absolute deviation divided by the workload median must be at most 5%; for the complete browser suite, at most 10%. Investigate and rerun unstable candidates rather than discard valid outliers or relax limits after seeing results. These are corpus stability gates, not claims about physical-system repeatability; [Set distro, VM, and physical-hardware validation gates](https://git.bongbetic.com/xavierk/odin/issues/14) owns the latter.
Run the same collection protocol on at least two independent physical validation systems, including one aarch64 system, to test workload validity and the shared-reference interpretation. With the candidate reference frozen, publish their domain and headline results where full coverage permits. Recompute scores using each validation system's eligible values as an alternate reference to expose baseline sensitivity. Flag any domain or headline shift greater than 10%, or any rank reversal, for documented review. A flag is not automatic rejection; unresolved protocol bias blocks the freeze. Do not infer architecture parity from these systems alone. The validation ticket sets the separate physical repeatability and UI-overhead thresholds.
### Evidence and identity
Publish an immutable calibration manifest and supporting raw, timestamped run records. Include successful and failed outcomes; per-trial values and native statistics; output checks; environment snapshots; exact source commits, binary and asset SHA-256 hashes, licenses, parser and protocol revisions; run/order/mode IDs; reference calculations; exclusions; cross-system validation; sensitivity analysis; and reviewer rationale for flagged cases. Record an explicit calibration release ID and score-comparison identity. Publish the corpus before enabling numerical scores for its qualified modes.
A correction to reference values, workload protocol, score rule, or calculation creates a new calibration release and comparison identity. Old run records retain their original identity and score. Recalculation for a newer release, if offered, must be explicit and preserve the old result. No silent cross-release score comparison. Prepared versus installed software and other mode boundaries from the scoring contract remain separate.
This resolution introduces no actual measured reference value. Production collection and release validation occur after this planning map. No new child ticket is needed: physical repeatability, UI overhead, distro cells, and hardware access belong to the existing validation ticket.
The Speedometer-specific browser corpus is superseded by [Choose a legally distributable v1 browser workload](https://git.bongbetic.com/xavierk/odin/issues/19#issuecomment-6899). [Specify Odin’s v1 browser interaction workload](https://git.bongbetic.com/xavierk/odin/issues/20) will define the new protocol before mode-specific calibration. General calibration and release gates here stand.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Part of Find the way to Odin’s build-ready specification.
Question
Which reproducible reference-selection, collection, and acceptance procedure will produce Odin’s first frozen calibration release under Define Odin’s median score and comparability contract?
Decide with the human: reference hardware/software selection and rationale; how workload reference values are derived from repeated complete runs; exact warmup, preparation, ordering, cooldown, and environment controls not already fixed by the workload contract; required raw evidence and artifact identity; validation across multiple physical systems; sensitivity to baseline choice; numerical validity and stability acceptance criteria; and how fixes or later recalibration create a new comparison identity. Specify supported mode/protocol reference coverage without introducing architecture-specific rescaling. Assign concrete physical repeatability gates to the existing validation ticket rather than duplicating them.
The accepted policy is fixed: 100 means reference performance; one reference value per workload is shared across supported architectures; no numerical Odin scores ship until real calibration and validation evidence exists. Do not invent that evidence.
This ticket chooses a build-ready procedure and release gates, not an actual calibration corpus. Running benchmarks and publishing measured reference values belong to the later implementation/validation effort, outside this planning map. Surface access gaps as planning prerequisites only where they prevent choosing a viable procedure.
Resolution
The human accepted the recommendations across the live decision rounds. This is a calibration procedure and release policy, not a measured corpus. No numerical Odin score may ship until the applicable corpus and validation evidence exist.
Reference and scope
Use one fixed, documented physical x86_64 reference system for each v1 calibration release. Record why its CPU, memory, GPU, storage, and power configuration were selected; it is a reference point, not a claim that its hardware represents a population. One measured reference value per workload applies across x86_64 and aarch64. Do not rescale by architecture. Validate on at least two independent physical systems, including one aarch64 system.
Each scored protocol and execution mode needs real reference evidence and a distinct identity. This includes installed versus prepared reference browser/language software, headed versus headless browser, and direct versus buffered storage. Keep any other protocol differences that change measurement meaning distinct. A mode without a qualified reference corpus shows raw results only. Physical reference evidence does not certify VM scores: v1 records VM measurements and context but withholds numerical Odin scores in VMs until a separate virtual calibration protocol and evidence pass release gates.
Collection procedure
For each candidate scored mode, collect at least five valid complete standard runs on the reference system, spread across at least two calendar days and two cold boots. Use the fixed workload pack, inputs, tool builds, parser revisions, profile, and per-workload repetition/warmup rules from the approved workload contract. Retain the complete time-ordered set of attempts, including failures, timeouts, cancellations, and slow valid trials. A failed or ineligible attempt remains in the evidence but does not enter the reference calculation. Repeat collection until five eligible values exist for every required workload; never trim a valid value merely because it is slow.
Before each run, check capability, permissions, storage headroom, required artifacts, output verification, power state, and sensor availability. Capture a two-minute idle baseline. Record firmware, hardware identifiers, kernel, CPU entitlement and affinity, governor, scheduler, thermal and power state, mount and filesystem settings, GPU driver and API, browser/toolchain versions, runtime/tool versions, and available health telemetry. Do not tune the host automatically. Rotate domain order across the five complete runs and record the exact order; preserve each workload's specified internal order and warmup. Between runs, require two minutes with CPU utilization below 5% and measured temperature within 3 °C of the idle baseline. Stop waiting after 20 minutes and mark the next run ineligible for calibration if readiness is unmet. When temperature sensors are unavailable, require five idle minutes and flag reduced thermal evidence. Every deviation and interruption is recorded.
For each workload, use its approved native within-run statistic, then take the median of the five or more eligible run-level values in native units. For browser standard, one full Speedometer suite contributes its official score per run. The median across complete runs becomes the frozen reference value; the established scoring formula and even-count median rule remain unchanged. Do not substitute a shortened suite, a partial run, or an alternate tool version.
Freeze gates and sensitivity
A calibration candidate passes numerical validity only when every required reference workload has at least five eligible complete run values, every output check passes, and values are finite, positive, in the declared unit and direction, with matching protocol identity. Preserve all raw samples and exclusions with explicit reasons. Across-run median absolute deviation divided by the workload median must be at most 5%; for the complete browser suite, at most 10%. Investigate and rerun unstable candidates rather than discard valid outliers or relax limits after seeing results. These are corpus stability gates, not claims about physical-system repeatability; Set distro, VM, and physical-hardware validation gates owns the latter.
Run the same collection protocol on at least two independent physical validation systems, including one aarch64 system, to test workload validity and the shared-reference interpretation. With the candidate reference frozen, publish their domain and headline results where full coverage permits. Recompute scores using each validation system's eligible values as an alternate reference to expose baseline sensitivity. Flag any domain or headline shift greater than 10%, or any rank reversal, for documented review. A flag is not automatic rejection; unresolved protocol bias blocks the freeze. Do not infer architecture parity from these systems alone. The validation ticket sets the separate physical repeatability and UI-overhead thresholds.
Evidence and identity
Publish an immutable calibration manifest and supporting raw, timestamped run records. Include successful and failed outcomes; per-trial values and native statistics; output checks; environment snapshots; exact source commits, binary and asset SHA-256 hashes, licenses, parser and protocol revisions; run/order/mode IDs; reference calculations; exclusions; cross-system validation; sensitivity analysis; and reviewer rationale for flagged cases. Record an explicit calibration release ID and score-comparison identity. Publish the corpus before enabling numerical scores for its qualified modes.
A correction to reference values, workload protocol, score rule, or calculation creates a new calibration release and comparison identity. Old run records retain their original identity and score. Recalculation for a newer release, if offered, must be explicit and preserve the old result. No silent cross-release score comparison. Prepared versus installed software and other mode boundaries from the scoring contract remain separate.
This resolution introduces no actual measured reference value. Production collection and release validation occur after this planning map. No new child ticket is needed: physical repeatability, UI overhead, distro cells, and hardware access belong to the existing validation ticket.
The Speedometer-specific browser corpus is superseded by Choose a legally distributable v1 browser workload. Specify Odin’s v1 browser interaction workload will define the new protocol before mode-specific calibration. General calibration and release gates here stand.