Set distro, VM, and physical-hardware validation gates #14

Closed
opened 2026-09-25 17:48:54 +00:00 by xavierk · 1 comment
Owner

Part of Find the way to Odin’s build-ready specification.

Question

Which exact distro/libc/architecture/kernel cells and physical devices must be exercised before Odin can claim support and score comparability? Define reproducible VM coverage for Void, deb and rpm families; minimal/container/headless/SSH/desktop modes as applicable; driver and SMART passthrough limitations; physical evidence for GPU, storage health, thermals and memory errors; failure/permission/missing-dependency cases; score repeatability and UI-overhead criteria. Identify access or hardware gaps as new task tickets if needed.

Resolve an actionable acceptance matrix with the human. A VM pass must not imply untested physical-hardware support.

Calibration validation handoff

Define Odin’s calibration procedure and release gates fixed the reference-corpus stability gate, but this ticket must set concrete physical repeatability and UI-overhead acceptance thresholds. Define number and spacing of independent full standard runs on the reference and at least two validation systems, including one aarch64; between-boot and within-system spread limits by workload/domain; headed/headless browser and storage-mode coverage where scored; TUI versus non-TUI measurement overhead limits; and evidence needed to accept or reject failures. Do not reuse the corpus median-absolute-deviation limit as proof of cross-system repeatability. VMs may supply compatibility evidence, but v1 VM results have no numerical Odin score.

Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). <!-- wayfinder-map: 1 --> ## Question Which exact distro/libc/architecture/kernel cells and physical devices must be exercised before Odin can claim support and score comparability? Define reproducible VM coverage for Void, deb and rpm families; minimal/container/headless/SSH/desktop modes as applicable; driver and SMART passthrough limitations; physical evidence for GPU, storage health, thermals and memory errors; failure/permission/missing-dependency cases; score repeatability and UI-overhead criteria. Identify access or hardware gaps as new task tickets if needed. Resolve an actionable acceptance matrix with the human. A VM pass must not imply untested physical-hardware support. ## Calibration validation handoff [Define Odin’s calibration procedure and release gates](https://git.bongbetic.com/xavierk/odin/issues/17#issuecomment-6776) fixed the reference-corpus stability gate, but this ticket must set concrete physical repeatability and UI-overhead acceptance thresholds. Define number and spacing of independent full standard runs on the reference and at least two validation systems, including one aarch64; between-boot and within-system spread limits by workload/domain; headed/headless browser and storage-mode coverage where scored; TUI versus non-TUI measurement overhead limits; and evidence needed to accept or reject failures. Do not reuse the corpus median-absolute-deviation limit as proof of cross-system repeatability. VMs may supply compatibility evidence, but v1 VM results have no numerical Odin score.
xavierk added the wayfinder:grilling label 2026-09-25 17:48:54 +00:00
xavierk added a new dependency 2026-09-25 17:50:04 +00:00
xavierk added a new dependency 2026-09-25 17:50:05 +00:00
xavierk added a new dependency 2026-09-25 17:50:06 +00:00
xavierk added a new dependency 2026-09-25 17:50:06 +00:00
xavierk added a new dependency 2026-09-25 17:50:07 +00:00
xavierk added a new dependency 2026-09-25 17:50:12 +00:00
xavierk added a new dependency 2026-09-27 05:45:48 +00:00
xavierk added a new dependency 2026-09-27 20:12:31 +00:00
xavierk added a new dependency 2026-09-27 21:18:14 +00:00
xavierk self-assigned this 2026-09-28 03:44:01 +00:00
Author
Owner

Resolution — validation and support gates

The product owner approved this acceptance matrix in live decision rounds. It is a release contract, not a claim that the systems have been acquired or tests run. Preserve the workload, score, health, runtime, and calibration decisions linked from the map.

Certified platform cells

Odin v1's eight required cells are Void glibc and Void musl on each of x86_64 and aarch64, Debian 13 glibc on each architecture, and Fedora 44 glibc on each architecture. Void is rolling: the release manifest must pin the exact tested repository snapshot, image/rootfs digest, package versions, kernel build/configuration, and Odin/helper builds. Debian and Fedora point releases, kernel builds, and image digests must likewise be pinned at qualification. Each tested build needs recorded CPU feature, cgroup, PSI, GPU, filesystem, terminal, sensor, and permission probes. An upstream Node ABI minimum is not an Odin support guarantee. Older and newer kernel builds remain unverified until requalified; do not extrapolate a universal kernel floor from one cell.

Every cell must boot and run on physical hardware and in a reproducible VM before a bare-metal support claim. Preserve VM image and package digests, virtual device model, QEMU/libvirt version, vCPU/RAM/storage configuration, kernel build, and test log. VM evidence proves guest compatibility only. v1 VM run records show raw values and coverage but no numerical Odin score. A cell validated only in a VM can carry a VM-only claim, never bare-metal certification. A core install or CLI failure in any of the eight required cells blocks the eight-cell v1 release.

In all eight VM and physical cells, exercise native package install/upgrade/removal, CLI/help/plain output, terminal restore, Quick and Standard profiles, cancellation, persisted partial outcomes, and explicit capability skips. Run minimal headless and SSH sessions in every cell. Exercise headed desktop, keyboard and mouse parity, resize, weak/no-color/ASCII terminal behavior, screen-reader-friendly output, and headed browser in the four glibc cells. Musl desktop behavior is claimed only if separately exercised; otherwise report headed capability as unverified or unavailable. Use a focused failure suite on representative cells: missing or mismatched browser driver/tool, missing display or focus, GPU API/driver failure, permission denial, insufficient safe RAM/disk headroom, bad workload output, thermal/device alarm, timeout, cancellation during each worker class, cleanup failure, and low-capability terminal. Confirm each failure produces the established outcome and preserves evidence without inventing a score. Smoke-test Debian 13 and Void musl x86_64 containers for cgroup/permission detection; v1 makes no container support or score claim.

These distro choices reflect current official release and architecture information: Debian releases, Fedora 44, Void architectures and libc variants. Void's musl limitations prohibit assuming proprietary NVIDIA support. QEMU's Arm documentation describes an emulated platform, not physical Arm hardware.

Physical systems and device evidence

Use one documented physical x86_64 calibration reference, one independent lower-end x86_64 validation system, and one physical aarch64 validation system. The second x86_64 system needs at least 8 GiB RAM, a hardware Vulkan/GL GPU from a different driver family than the reference, a usable display, and enough free filesystem space for the approved storage workload. Select an Arm system with a real GPU and display that can complete all seven scored domains in at least one qualified OS cell. Both validation systems are current access gaps: Secure independent x86_64 validation system for Odin v1 and Secure physical aarch64 validation system for Odin v1. Securing them is required before the relevant score and support claims.

Across the physical rigs, exercise native NVMe and SATA HDD paths, real SMART and thermal evidence, two x86 GPU driver families, and one Arm hardware GPU. Verify device discovery, compute/render correctness, permissions, storage direct-I/O behavior, health interpretation, and physical attribution. A virtual disk, virtual renderer, or generic USB bridge does not stand in for native physical-device evidence. Test missing or unreadable EDAC/RAS data and error classification with fixtures when no real error exists; do not induce hardware faults or overstate a clean memory check. Record which hardware and modes each result actually covers.

Repeatability and TUI overhead

On the reference and both validation systems, collect six independent complete Standard runs in one fixed scored OS/mode configuration: three after each of two cold boots on different calendar days. Retain all attempted and slow valid runs, preflight and environment snapshots, per-trial values, failed outputs, exclusions with reasons, and cooling/readiness conditions from the calibration procedure. The reference's six eligible runs may also satisfy the existing five-run calibration minimum, but its corpus median-absolute-deviation gate does not replace these physical repeatability checks. Qualify another OS or scored mode separately before claiming its numerical score; functional support alone does not imply score qualification.

For each required workload, compute run-level native values according to its fixed protocol, then six-run spread as (maximum - minimum) / median × 100%. It must be at most 10% for CPU, memory bandwidth, Bash, and language workloads, and at most 20% for GPU, storage, and browser workloads. Each available domain-score spread must be at most 10%; a complete headline-score spread at most 5%. Compare the median of the three runs from each cold boot with absolute difference / median of the two boot medians × 100%: at most 5% for CPU, memory, Bash, and language workloads and every available domain, and at most 10% for GPU, storage, and browser workloads. Preserve valid outliers; investigate and rerun after a documented correction instead of trimming them. Missing required measurements yield no domain/headline and cannot pass a full-score gate.

On reference and Arm systems, perform three matched, alternating TUI and plain-terminal Standard-run pairs under the same workload and browser mode, with readiness restored between runs. For each workload, calculate the absolute TUI/plain difference divided by the plain value in each pair; the median of the three percentages must be at most 5%. A larger difference blocks that score identity until interference is resolved. Keep the measurement runner and environment identical aside from UI rendering. Record order, system load, and raw pairs. This gate measures UI interference, not calibration-corpus stability.

Modes, claims, and release result

Qualify installed-browser headed and headless identities on all three physical systems with compatible browser/driver pairs. Record refresh rate, device-pixel ratio, acceleration path, and focus. Qualify direct-I/O storage first. Buffered storage and prepared-reference software modes can score only after their distinct physical reference corpus and validation evidence pass the existing calibration gates. One complete seven-domain physical aarch64 run, qualified against the shared reference, is required before claiming cross-architecture headline comparability. Physical bare-metal support for each cell requires that cell's physical boot and functional suite; numerical score eligibility additionally requires its applicable physical mode, repeatability, workload correctness, and calibration gates. Other configurations retain raw measurements and explicit coverage gaps.

A failed workload or repeatability gate blocks its score identity and dependent headline, while preserving partial records and reasoned outcomes. It need not block a core package release when all eight cells meet the install/CLI/support gates and the unsupported score is withheld. A core failure in any required cell blocks the eight-cell v1 release. No score ships before the immutable physical calibration corpus and approved mode manifest exist. This policy does not claim that a 10–20 minute target is always met: validate it on the lower-end physical rig, and preserve timeouts rather than shortening protocols.

## Resolution — validation and support gates The product owner approved this acceptance matrix in live decision rounds. It is a release contract, not a claim that the systems have been acquired or tests run. Preserve the workload, score, health, runtime, and calibration decisions linked from the map. ### Certified platform cells Odin v1's eight required cells are Void glibc and Void musl on each of x86_64 and aarch64, Debian 13 glibc on each architecture, and Fedora 44 glibc on each architecture. Void is rolling: the release manifest must pin the exact tested repository snapshot, image/rootfs digest, package versions, kernel build/configuration, and Odin/helper builds. Debian and Fedora point releases, kernel builds, and image digests must likewise be pinned at qualification. Each tested build needs recorded CPU feature, cgroup, PSI, GPU, filesystem, terminal, sensor, and permission probes. An upstream Node ABI minimum is not an Odin support guarantee. Older and newer kernel builds remain unverified until requalified; do not extrapolate a universal kernel floor from one cell. Every cell must boot and run on physical hardware and in a reproducible VM before a bare-metal support claim. Preserve VM image and package digests, virtual device model, QEMU/libvirt version, vCPU/RAM/storage configuration, kernel build, and test log. VM evidence proves guest compatibility only. v1 VM run records show raw values and coverage but no numerical Odin score. A cell validated only in a VM can carry a VM-only claim, never bare-metal certification. A core install or CLI failure in any of the eight required cells blocks the eight-cell v1 release. In all eight VM and physical cells, exercise native package install/upgrade/removal, CLI/help/plain output, terminal restore, Quick and Standard profiles, cancellation, persisted partial outcomes, and explicit capability skips. Run minimal headless and SSH sessions in every cell. Exercise headed desktop, keyboard and mouse parity, resize, weak/no-color/ASCII terminal behavior, screen-reader-friendly output, and headed browser in the four glibc cells. Musl desktop behavior is claimed only if separately exercised; otherwise report headed capability as unverified or unavailable. Use a focused failure suite on representative cells: missing or mismatched browser driver/tool, missing display or focus, GPU API/driver failure, permission denial, insufficient safe RAM/disk headroom, bad workload output, thermal/device alarm, timeout, cancellation during each worker class, cleanup failure, and low-capability terminal. Confirm each failure produces the established outcome and preserves evidence without inventing a score. Smoke-test Debian 13 and Void musl x86_64 containers for cgroup/permission detection; v1 makes no container support or score claim. These distro choices reflect current official release and architecture information: [Debian releases](https://www.debian.org/releases/index.html), [Fedora 44](https://fedoraproject.org/workstation/download/), [Void architectures and libc variants](https://docs.voidlinux.org/installation/). Void's [musl limitations](https://docs.voidlinux.org/installation/musl.html) prohibit assuming proprietary NVIDIA support. [QEMU's Arm documentation](https://www.qemu.org/docs/master/system/target-arm) describes an emulated platform, not physical Arm hardware. ### Physical systems and device evidence Use one documented physical x86_64 calibration reference, one independent lower-end x86_64 validation system, and one physical aarch64 validation system. The second x86_64 system needs at least 8 GiB RAM, a hardware Vulkan/GL GPU from a different driver family than the reference, a usable display, and enough free filesystem space for the approved storage workload. Select an Arm system with a real GPU and display that can complete all seven scored domains in at least one qualified OS cell. Both validation systems are current access gaps: [Secure independent x86_64 validation system for Odin v1](https://git.bongbetic.com/xavierk/odin/issues/22) and [Secure physical aarch64 validation system for Odin v1](https://git.bongbetic.com/xavierk/odin/issues/21). Securing them is required before the relevant score and support claims. Across the physical rigs, exercise native NVMe and SATA HDD paths, real SMART and thermal evidence, two x86 GPU driver families, and one Arm hardware GPU. Verify device discovery, compute/render correctness, permissions, storage direct-I/O behavior, health interpretation, and physical attribution. A virtual disk, virtual renderer, or generic USB bridge does not stand in for native physical-device evidence. Test missing or unreadable EDAC/RAS data and error classification with fixtures when no real error exists; do not induce hardware faults or overstate a clean memory check. Record which hardware and modes each result actually covers. ### Repeatability and TUI overhead On the reference and both validation systems, collect six independent complete Standard runs in one fixed scored OS/mode configuration: three after each of two cold boots on different calendar days. Retain all attempted and slow valid runs, preflight and environment snapshots, per-trial values, failed outputs, exclusions with reasons, and cooling/readiness conditions from the calibration procedure. The reference's six eligible runs may also satisfy the existing five-run calibration minimum, but its corpus median-absolute-deviation gate does not replace these physical repeatability checks. Qualify another OS or scored mode separately before claiming its numerical score; functional support alone does not imply score qualification. For each required workload, compute run-level native values according to its fixed protocol, then six-run spread as `(maximum - minimum) / median × 100%`. It must be at most 10% for CPU, memory bandwidth, Bash, and language workloads, and at most 20% for GPU, storage, and browser workloads. Each available domain-score spread must be at most 10%; a complete headline-score spread at most 5%. Compare the median of the three runs from each cold boot with `absolute difference / median of the two boot medians × 100%`: at most 5% for CPU, memory, Bash, and language workloads and every available domain, and at most 10% for GPU, storage, and browser workloads. Preserve valid outliers; investigate and rerun after a documented correction instead of trimming them. Missing required measurements yield no domain/headline and cannot pass a full-score gate. On reference and Arm systems, perform three matched, alternating TUI and plain-terminal Standard-run pairs under the same workload and browser mode, with readiness restored between runs. For each workload, calculate the absolute TUI/plain difference divided by the plain value in each pair; the median of the three percentages must be at most 5%. A larger difference blocks that score identity until interference is resolved. Keep the measurement runner and environment identical aside from UI rendering. Record order, system load, and raw pairs. This gate measures UI interference, not calibration-corpus stability. ### Modes, claims, and release result Qualify installed-browser headed and headless identities on all three physical systems with compatible browser/driver pairs. Record refresh rate, device-pixel ratio, acceleration path, and focus. Qualify direct-I/O storage first. Buffered storage and prepared-reference software modes can score only after their distinct physical reference corpus and validation evidence pass the existing calibration gates. One complete seven-domain physical aarch64 run, qualified against the shared reference, is required before claiming cross-architecture headline comparability. Physical bare-metal support for each cell requires that cell's physical boot and functional suite; numerical score eligibility additionally requires its applicable physical mode, repeatability, workload correctness, and calibration gates. Other configurations retain raw measurements and explicit coverage gaps. A failed workload or repeatability gate blocks its score identity and dependent headline, while preserving partial records and reasoned outcomes. It need not block a core package release when all eight cells meet the install/CLI/support gates and the unsupported score is withheld. A core failure in any required cell blocks the eight-cell v1 release. No score ships before the immutable physical calibration corpus and approved mode manifest exist. This policy does not claim that a 10–20 minute target is always met: validate it on the lower-end physical rig, and preserve timeouts rather than shortening protocols.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/odin#14