Linux users lack one terminal tool that measures their installed system across CPU, memory, GPU, storage, browser, shell, and common language workloads while explaining missing capability and credible health concerns. Isolated benchmark numbers do not show the measurement conditions, cannot establish hardware health, and can mislead when compared across different workloads or environments. Odin needs a reproducible, locally stored, accessible benchmark and diagnostic experience with a defensible median score and evidence-based advice.
Solution
Build Odin by Bongbetic, a Linux terminal suite with Quick, Standard, and explicitly selected Extended run profiles. Standard attempts a fixed seven-domain workload set and publishes a headline median score only when every required measurement is valid under a qualified calibration release. All runs retain native measurements, coverage, health evidence, and execution context locally. The terminal interface provides themed InkUI dashboards, keyboard and qualified mouse controls, plain output, history, comparison, and private-by-default export/import. The supervised runner bounds work and records partial outcomes. Native packages target the approved Void, Debian, and Fedora certification cells. Numerical claims ship only after physical calibration and validation gates pass.
The approved source decisions are indexed by Find the way to Odin’s build-ready specification. Approve Odin’s build-ready specification records the product owner's handoff approval. Where this synthesis compresses a byte-level protocol, the linked resolution comment remains normative. Later browser, editor, headed-validation, and Extended-browser decisions supersede only the identified older text; no Speedometer result or asset enters Odin v1.
User Stories
As a Linux user, I want to launch Odin as a normal terminal command, so that I can inspect my system without a desktop application.
As a Bash user, I want documented commands and exit statuses, so that I can use Odin in scripts.
As a user in a pipe or weak terminal, I want plain machine-usable output, so that results remain readable without terminal graphics.
As a user with a capable terminal, I want Bongbetic branding, themed InkUI components, and colorful graphs, so that measurements and warnings are easy to scan.
As a keyboard user, I want every action reachable without a mouse, so that the interface remains operable over SSH.
As a mouse user, I want clickable controls where my terminal supports mouse reporting, so that I can navigate naturally.
As a user whose terminal lacks mouse reporting, I want explicit keyboard parity, so that no action becomes inaccessible.
As a user with low color or ASCII output, I want measurements and warnings preserved, so that appearance settings do not hide facts.
As a screen-reader user, I want a text-friendly output path, so that measurements and warnings are understandable.
As a user who resizes or interrupts the terminal, I want the display and terminal state restored, so that my shell remains usable.
As a new user, I want a Quick profile with bounded read-only discovery, so that I can triage capability and health in a few minutes.
As a user seeking a broad result, I want a Standard profile targeting 10–20 minutes, so that the requested performance areas are measured together.
As a user seeking deeper evidence, I want separately chosen Extended modules with time and resource estimates, so that I control additional stress and writes.
As a user on a limited system, I want unavailable workloads named with reasons, so that missing measurements are not mistaken for zero performance.
As a user with several GPUs, I want the measured device identified and ambiguous selection resolved, so that a result refers to the intended GPU.
As a user with several drives, I want to choose the filesystem path for storage performance while safe discovery can inspect other devices, so that the result has clear scope.
As a user concerned about storage wear, I want a clear write estimate and consent prompt, so that a benchmark does not write unexpectedly.
As a user who declines storage writes, I want a preserved coverage outcome, so that the run still reports useful evidence.
As a user, I want bounded CPU throughput measurements at one worker and at my effective CPU entitlement, so that I can see narrow single-thread and parallel behavior.
As a user, I want memory bandwidth and loaded-responsiveness measurements shown separately, so that throughput does not disguise latency under load.
As a user, I want memory checks limited to tested bytes and patterns, so that a passing finite check is not presented as proof of fault-free RAM.
As a user, I want available EDAC/RAS and pressure evidence retained, so that memory errors and contention can be investigated.
As a user, I want GPU discovery, API execution, known-output correctness, and presentation reported as distinct claims, so that driver status is not overstated.
As a user without a qualified hardware GPU path, I want software or virtual rendering labelled separately, so that it is not counted as hardware performance.
As a user, I want bounded sequential and random storage measurements with mode and device mapping, so that speed results have an interpretable scope.
As a user with an unsupported or unsafe health transport, I want an unknown outcome instead of guessed device commands, so that diagnosis does not create risk.
As a user, I want an offline browser interaction suite using my installed browser, so that browser responsiveness reflects my installed system.
As a user, I want headed and headless browser results identified separately, so that display cadence and execution mode are not conflated.
As a user without a compatible browser or driver, I want a precise capability skip, so that Odin does not silently install or substitute software.
As a user, I want Bash direct startup, interactive prompt readiness, and arithmetic throughput reported separately, so that shell latency and computation remain distinct.
As a user, I want Python, Rust, C++, and Java startup and warmed throughput measurements, so that I can assess representative installed toolchains.
As a user, I want language output checks, so that a fast wrong result is never treated as performance.
As a user, I want kernel, power, thermal, driver, filesystem, and software context recorded, so that I can interpret a surprising measurement later.
As a user, I want health findings separated from speed scores, so that a fast failing drive does not look healthy.
As a user, I want health severity and evidence confidence shown separately, so that urgency is not confused with certainty.
As a user with unavailable telemetry, I want current health marked unknown, so that missing evidence is not treated as clean evidence.
As a user, I want drive replacement advice only when device-attributable evidence supports it, so that low speed alone does not trigger replacement advice.
As a user, I want optimization advice tied to observed limitations, tradeoffs, and a verification step, so that I can judge whether a change helped.
As a user, I want Odin to stop affected work on qualified thermal, device, integrity, or resource alarms, so that the run respects safety limits.
As a user with detected memory corruption, I want the whole run stopped and its partial evidence retained, so that invalid measurements are not scored.
As a user who cancels a run, I want worker termination and cleanup attempted with outcomes recorded, so that I know what remains uncertain.
As a user who declines privilege, I want the run retained with diagnostic coverage gaps, so that elevation is optional and explicit.
As a user, I want a seven-domain headline median only after complete valid Standard coverage, so that missing areas cannot inflate the score.
As a user, I want domain breakdowns beside the headline, so that a median does not hide a weak domain.
As a user, I want raw measurements before calibration is available, so that Odin never invents a reference score.
As a user, I want scores normalized against an immutable calibration release, so that the meaning of 100 and score identity stay stable.
As a user, I want Quick and incomplete runs preserved without a partial headline, so that useful evidence survives without a misleading aggregate.
As a user, I want loaded responsiveness and health findings visible beside the score, so that unscored diagnostics still inform decisions.
As a user, I want installed and prepared-reference software modes labelled separately, so that comparisons retain their software context.
As a user, I want VM measurements marked raw-only in v1, so that guest compatibility is not misrepresented as physical score qualification.
As a user, I want every benchmark run saved locally, including cancelled and failed runs, so that the evidence remains auditable.
As a user, I want raw trials, timestamps, validity reasons, and calculation inputs retained, so that a summary can be checked later.
As a user, I want run notes stored separately from original evidence, so that annotation does not rewrite measurements.
As a user, I want history to persist until I delete it, so that old results do not disappear unexpectedly.
As a user, I want saved history to survive interruption or save failure, so that earlier run records remain intact.
As a user, I want to list, filter, inspect, and compare runs in the terminal, so that I can study changes without a hosted service.
As a user, I want incompatibility reasons shown for unlike runs, so that a combined change is not calculated across different identities.
As a user, I want a portable machine-readable export with private fields redacted by default, so that I can share results without leaking stable identifiers.
As a user requesting full fidelity, I want an explicit preview of identifying fields, so that disclosure is deliberate.
As a user, I want validated import to preserve original run identity and mark provenance, so that imported results cannot masquerade as local measurements.
As a user, I want a clear confirmation before deleting a run, so that history is not removed by accident.
As a package user, I want native xbps, deb, or rpm delivery on certified cells, so that installation matches my Linux distribution.
As a package user, I want optional workload preparation explicit and recorded, so that normal runs never download or upgrade tools.
As a user on musl or another limited cell, I want capability detection and scoped support claims, so that package support does not promise unavailable drivers.
As a release reviewer, I want reproducible build, license, and asset manifests, so that distributed packs have verifiable provenance.
As a release reviewer, I want physical and VM certification evidence for each required cell, so that a support claim matches actual tests.
As a release reviewer, I want physical repeatability, browser qualification, and TUI overhead measured, so that numerical scores have evidence beyond a successful launch.
As a release reviewer, I want failed gates to withhold their score identity while retaining raw evidence, so that a package can report capability honestly.
As a release reviewer, I want the independent x86_64 and physical aarch64 validation systems documented, so that cross-system and cross-architecture claims have real hardware evidence.
As a user, I want no automatic kernel tuning, driver installation, drive repair, or destructive raw-device writes, so that Odin remains a measurement and advice tool.
Implementation Decisions
System boundary and delivery
The v1 terminal application uses TypeScript, React 19, Ink 6, Node 24, and pinned source-copied InkUI components. Bun and an Ink 7 migration are outside this v1 stack. Ship native xbps, deb, and rpm packages for qualified cells. Use a maintained distro Node 24 where available; otherwise bundle a pinned Node 24 runtime and own its updates and notices. Void musl requires a qualified distro-built Node and qualified helpers. No unofficial runtime ABI floor becomes a general Odin support promise. Source: Choose Odin’s runtime and Linux compatibility contract.
Keep the terminal UI separate from an unprivileged, supervised measurement runner. The runner owns immutable run parameters, explicit helper paths and arguments, controlled environment, process groups, deadlines, resource and output bounds, cancellation, cleanup attempts, and durable partial outcomes. Do not construct workload shell commands. Privileged diagnostics expose narrowly validated operation-and-target requests, not a root UI or generic command runner. Quiet decorative updates during timed measurements.
Build around a public CLI boundary that reaches capability probing, run planning, supervised execution, scoring, health interpretation, and local run-record persistence. This is the primary observable behavior seam. The browser pack and physical validation add separate release evidence for timing and hardware claims.
Core packages contain only rights-cleared versioned assets. Optional browser, language, and helper preparation is explicit before a run; record origin, version, digest, size, licenses, notices, and cache location. An ordinary run never downloads, installs, upgrades, repairs, or tunes the host. Missing preparation is a capability outcome.
Profiles, workload identities, and limits
Quick is a roughly 2–3 minute read-only triage profile: safe health and capability discovery, short CPU/Bash/language probes, memory/PSI context, GPU discovery and bounded correctness, and browser automation availability. Read-only prohibits storage writes, device-changing commands, and self-tests; finite RAM/GPU computation remains possible. Quick has its own shortened probe identities and no full headline.
Standard targets 10–20 minutes and attempts the fixed suite where capability and prepared assets permit. One filesystem path and one GPU are selected for performance per run. Each worker has a fixed end-to-end deadline including setup and cleanup; a slow system can produce an incomplete run. Never shorten a protocol or trim a slow valid observation to meet the target.
Extended modules are separately selected with duration, memory, and write estimates. They include contained memory pressure and longer scheduler work; visible GPU presentation and additional devices; queued/durable/integrity or larger file storage work under separate write consent; three full Odin Browser Interaction v1 suites; configured-shell and pipeline work; and compilation/longer language work. JetStream and MotionMark are not v1 modules. Offline Memtest86+ and long drive self-tests remain separate follow-up flows. Sources: Choose Odin’s workload suite and run profiles, Resolve Odin v1 Extended browser workload scope.
Outcomes distinguish completed, not scheduled, unavailable by reason, invalid output, execution failure, timeout, cancellation, and safety stop. Preserve partial samples and cleanup status. A missing capability never becomes numerical zero or a substitute workload. Version each workload, protocol, parser, helper, asset, execution mode, and software mode.
Retain at least the greater of 1 GiB and 25% of effective available memory outside workers. STREAM's three arrays each exceed the relevant last-level cache and are at least the greater of 128 MiB and four times reported last-level cache; total arrays are at most 1.5 GiB. Skip safely when the cache or headroom cannot be established. Online memtester is one finite pass over at most 256 MiB when safe. GPU workers require a qualified VRAM cap no greater than 256 MiB or 25% of known free VRAM. Bounded CPU load targets about half the effective CPU entitlement. The selected storage fixture and all repeated writes stay below the agreed 1 GiB host-write ceiling.
Standard performance and diagnostic suite
CPU: pinned sysbench 1.0.20 prime search with cpu-max-prime=20000; one-worker and effective-entitlement worker counts; three ten-second timed runs each. The parallel count is the ceiling of the lesser of allowed logical CPUs and effective cgroup quota, with a minimum of one. Preserve events/s, workers, affinity, quota, and throttling. This is a narrow integer workload.
Memory bandwidth and responsiveness: STREAM 5.10 Copy and Triad preserve their native best-iteration statistic, with three independent runs for cross-run medians. Pinned schbench revision 6300b8f3a8922c61ea6bb2cdfa1901a42c0cc6fc measures idle and bounded stress-ng 0.22.01 loaded latency in three 15-second pairs. Preserve p50/p95/p99, achieved load, throughput, samples, and PSI deltas. Schbench's v1 parameters and untimed preparation are normative in Choose Odin’s workload suite and run profiles. Stress-ng bogo-ops are not scored.
Memory diagnostics: one finite memtester 4.7.1 pass over at most 256 MiB with actual allocated and locked bytes, patterns, errors, and timeout recorded; available EDAC/RAS and pressure readings before and after. A pass describes only tested scope. Do not infer error-free memory from missing counters.
GPU: enumerate the selected hardware device, kernel driver, renderer/API, and permissions. Use qualified Vulkan-only clpeak 2.1.4 FP32 and global bandwidth tests plus glmark2 2023.01 fixed 800×600 offscreen GL/GLES shading and texture scenes where supported. Repeat short tests three times and run a separate known-output correctness operation. Keep compute, rendering, and presentation claims distinct. Standard opens no presentation window. Software or virtual rendering has a distinct identity.
Storage: fio 3.43 operates only on a private 256 MiB fixture in a user-selected filesystem. At queue depth one with psync, run 1 MiB sequential read over 256 MiB, sequential write over 128 MiB, and 4 KiB random read and write over 4 MiB each, with three repetitions. Prefer aligned direct I/O; buffered fallback is distinct. Record bandwidth/IOPS, latency distributions, bytes, engine, cache/flush mode, filesystem, and device mapping. Fixture preparation plus repeated writes total 652 MiB before metadata; enforce a 1 GiB ceiling and leave at least 2 GiB plus 10% of filesystem capacity free. Show the estimate and require consent. Never use existing user files or a raw device. This is burst-path performance, not sustained-media speed.
Browser: use the Odin-owned offline Browser Interaction v1 pack and installed Chromium-family or Firefox browser with a compatible local WebDriver. The pack has deterministic grid, board, and Markdown editor fixtures and eight scripted actions in each scenario. Two untimed warmup cycles precede ten measured cycles, producing 240 valid action-to-visible-update proxy samples. A trusted page input-event timestamp to the second consecutive animation frame after the expected UI state is the native latency measure; this is neither physical photon timing nor the browser's INP metric. The suite statistic is nearest-rank p95, sorted sample 228 of 240, in milliseconds. Standard runs one full suite with a five-minute end-to-end stage deadline. Preserve every sample, action identity, state check, and context. A missing sample, wrong state, external request, focus loss, bad timing, crash, or timeout withholds the browser domain. Run on loopback with a fresh profile, no extensions, clean initial cache, 800×600 CSS viewport, and 100% zoom. Headed and headless, installed and prepared-reference modes, and a separate headless-shell binary have distinct identities. Exact fixture generation, action order, validity, and pack requirements are normative in Specify Odin’s v1 browser interaction workload; Reconcile browser editor fixture and scripted action replaces the conflicting editor action text with one-based lines, exact states, and digests. Board lane one is numeric lane ID 1, so moving C000 from lane ID 0 is an actual change. The final shipped app, fixtures, selectors, checker expectations, and assets define the pack identity.
Bash: 30 clean direct startup samples, 30 clean interactive PTY prompt-readiness samples, and three timed runs of the pinned builtin arithmetic work, verifying sum 43824450000. Startup files are excluded by default. Keep startup, prompt, and throughput distinct.
Languages: Python, Rust, C++, and Java run Odin-owned text-parse-v1, prepared before timing from the same 1,048,576-record, 12 MiB ASCII corpus. Each verifies record count and weighted sum 266529503798272; the corpus SHA-256 is 1f7707f5f5e19ec9292a0aecdb67fc241176e58b86dccc04b8cc13acf421633b. Report completed passes/s from three fixed five-second windows after declared warmup, plus 30 clean process-to-marker startup trials per language. Rust/C++ startup launches a prepared binary, not a compiler. No claim ranks languages intrinsically. Exact corpus formula and source protocol are normative in Choose Odin’s workload suite and run profiles.
Overall health: before and after load, collect available temperature/alarm, CPU/memory/I/O PSI, EDAC/RAS, kernel/device I/O error, and safely supported SMART evidence. Record tool/parser versions, scope, timestamps, and permission gaps. Use qualified read-only protocol operations; never guess a transport, start self-tests, or interpret unknown telemetry as healthy. Source: Define health findings and the optimization advice rubric.
Score, health, and advice
The headline median score is the median of exactly seven equally represented domain scores: CPU, memory bandwidth, GPU, storage, browser, Bash, and languages. For higher-is-better native values use 100 × measured/reference; for lower-is-better use 100 × reference/measured. Take repeated-trial medians in native units first; even-sized medians average their middle two values. Within domains: CPU uses single and parallel throughput; memory uses Copy and Triad; GPU uses compute and rendering summaries; storage uses four fixed read/write results; browser uses one complete p95 result; Bash uses startup, prompt, and arithmetic; languages take a startup/throughput median per language then the median of four languages. Loaded responsiveness and health remain beside, not inside, the numerical aggregate. Source: Define Odin’s median score and comparability contract, with the browser substitution from Specify Odin’s v1 browser interaction workload.
No full headline appears unless a Standard set has every required valid measurement, the corresponding physical mode has qualified reference evidence, and integrity/safety rules permit it. Missing data, detected memory errors, failed output verification, incomplete required work, or a safety stop withhold the headline. A slow valid result remains valid. Missing health telemetry alone leaves health unknown but does not suppress a complete valid score. Quick and incomplete runs show raw values, available domains, and coverage without a partial aggregate. VM runs are raw-only in v1. Never invent a reference or confidence interval.
Health finding severity is informational, warning, or critical; finding confidence is directly observed, corroborated, or tentative. Apply confidence to the claim actually supported. Missing, stale, unreadable, or uninterpretable evidence is unknown. Separate current observation, during-run change, and historical evidence. A passing limited check is not overall health certification. Drive replacement advice requires credible device-reported failure, applicable degradation, read-only device failure, or repeated device-attributable integrity errors; endurance consumption, low speed, temperature, or ambiguous interface errors alone need their own narrower advice. Use documented alarms and applicable limits, never invented universal thresholds.
Stop affected work and later work on the same component for active thermal alarms, device loss, integrity error, or breached limits. Stop the entire run for memory corruption or inseparable safety/integrity risk. Do not automatically retry. Preserve evidence, partial samples, cleanup failures, and unresolved I/O. Advice requires corroborating workload-specific evidence before asserting a bottleneck; hypotheses remain labelled, and recommendations state tradeoffs plus a verification step. Tuning, repair, and driver changes remain advice only.
Run records, history, and comparison
A run record preserves every attempt, including invalid, cancelled, incomplete, and safety-stopped runs: raw trials, native statistics, timestamps/order, units, summary inputs, validity and coverage reasons, environment, workload/tool/parser/protocol versions, calibration and score identities, permissions, thermal/power context, health source and confidence, advice evidence, and cleanup outcome. Original evidence and scores are immutable; user notes are separate. Version schema and calculations. Source: Define local results, history, and export behavior.
Save by default in a per-user local data directory distinct from the storage benchmark target. Permit choosing/opening that directory. Retain history until explicit deletion; show storage use and support deliberate cleanup. Interrupted or failed writes preserve committed history and report whether the current run was saved or needs recovery. Terminal history lists, filters, inspects, compares, annotates, exports, and confirms deletion.
Direct comparison requires matching scoring rules, calibration release, workload protocol, browser execution/software mode, storage direct/buffered mode, and physical/virtual context. Show unlike runs side by side with precise incompatibility reasons, without a combined change or unsupported causal claim. Installed software versions may differ within a compatible cohort but remain visible as environmental changes.
v1 exports a versioned portable machine-readable record with all measurement and interpretation evidence. Default export redacts hostname, username, home paths, serials, and stable hardware IDs while retaining non-identifying comparison facts. Full-fidelity export requires explicit selection and preview. If privacy rules conflict with comparison-critical fields, explain the conflict rather than silently leaking or deleting data. Import only after schema and integrity checks, mark provenance, and preserve original identity.
TUI and support contract
Offer the three product-owner-approved layouts: Mission Control default, Guided Run, and Run Ledger. Preserve plan, live-run, summary, history, advice, profile, theme, ASCII, cancellation, and layout navigation described by Evaluate the Bongbetic TUI workflow with InkUI. Keyboard parity is mandatory; qualify mouse only where the terminal reports it. Handle focus, scroll, resize, weak/no-color/ASCII paths, screen-reader-friendly output, and restoration of raw input, mouse mode, cursor, and screen on normal exit and handled interruption. Do not claim restoration after SIGKILL.
The eight required certification cells are Void glibc and Void musl on each of x86_64 and aarch64, plus Debian 13 glibc and Fedora 44 glibc on each architecture. Pin exact tested image/rootfs, repository snapshot or point release, package builds, kernel configuration, Odin and helper builds. Test each cell physically and in a reproducible VM before a bare-metal claim; VM evidence proves guest compatibility only. All eight cells need package, CLI/plain output, terminal restore, Quick/Standard, cancellation, partial persistence, capability skips, minimal/SSH checks. All six glibc cells additionally require headed UI and browser functional checks on physical hardware and in VMs; required headed UI failure blocks the eight-cell release. Mouse support remains terminal-capability-scoped. Void musl headed behavior is claimed only if separately qualified. Sources: Set distro, VM, and physical-hardware validation gates, Clarify headed validation for glibc certification cells.
Testing Decisions
The primary automated seam is the public CLI with controlled workload, capability, sensor, and device fixtures. Test externally visible exit status, plain/TUI information, worker outcome, score eligibility, saved run record, cancellation, and cleanup reporting. Test behavior rather than private functions, rendering trees, or implementation-specific file layout. No production code or existing test suite is present yet; the approved TUI prototype supplies interaction examples but simulated measurements, not benchmark prior art.
Exercise each workload adapter with known inputs, wrong-output fixtures, version mismatch, missing permission, insufficient safe resources, timeout, cancellation, cleanup failure, and slow-but-valid output. Recompute pinned Bash and language checks independently. Check native units, repetitions, result identities, and score formulas at the CLI/run-record boundary. Include complete and missing-domain cases, memory-error and GPU/browser verification failures, unknown health telemetry, and privileged-check denial.
Browser acceptance uses real supported Chromium-family and Firefox plus compatible drivers for headed and headless modes. Verify deterministic fixture bytes; corrected editor expected states and SHA-256 after each action; independent grid/board/editor state checks; exactly 240 positive finite samples; p95 rank; focus and animation-frame timing; loopback-only assets; no outbound requests; deadline; partial evidence; and mode identity. Run unscored offline smoke tests with outbound network blocked for each qualified engine/mode. The pack manifest uses RFC 8785 canonical JSON, sorted normalized paths, per-file byte length, SHA-256, provenance, rights/notice evidence, and canonical-manifest SHA-256 as pack ID. Reject unlisted assets, symlinks, path traversal, and missing notices. Changing final pack bytes changes identity and requires requalification. Source: Specify Odin’s v1 browser interaction workload.
Test health rules with protocol-specific fixtures, including historical versus current alarms, missing/stale data, ambiguous I/O attribution, smartctl nonzero bitmask with usable output, parser version mismatch, GPU discovery/execution/correctness/presentation distinctions, VM guest scope, drive transport safety, and uncertain advice. Test safety-stop scope and preserved partial records. No test deliberately creates real hardware faults.
Test local history through interrupted writes, recovery, immutable evidence with editable notes, compatible and incompatible comparisons, default-redacted and explicit full-fidelity exports, schema/integrity failures on import, imported provenance, deletion confirmation, and history retention.
Calibrate every scored protocol/mode on one documented physical x86_64 reference with at least five complete eligible Standard runs across two days and two cold boots. Capture a two-minute idle baseline, readiness/cooling conditions, all failed attempts, exact environment and pack/tool hashes. Each reference workload uses the median of run-level native values. Corpus median absolute deviation divided by median must be at most 5%, or 10% for the browser suite. Validate the shared reference on at least two independent physical systems including aarch64, review alternate-reference sensitivity greater than 10% or rank reversal, and publish immutable raw corpus and manifest before numerical scores. New browser p95 requires fresh reference evidence; old Speedometer values never transfer. Source: Define Odin’s calibration procedure and release gates.
On the reference and both validation systems, run six complete Standard runs in one fixed scored OS/mode: three after each of two cold boots on different days. Apply the approved workload, domain, headline, and cross-boot spread limits; preserve valid outliers and every failed attempt. On reference and Arm, run three alternating matched TUI/plain pairs; median per-workload TUI overhead must be at most 5%. Physically exercise native NVMe and SATA HDD, two x86 GPU driver families, one Arm GPU, real SMART/thermal paths, device mapping, and storage direct I/O. VMs prove compatibility, not physical hardware or numerical score validity. Exact thresholds and matrix are normative in Set distro, VM, and physical-hardware validation gates.
A core install, CLI, or required headed UI failure in a required cell blocks the eight-cell release. A failed workload, repeatability, or calibration gate withholds that score identity and dependent headline while retaining raw evidence; it need not block a functioning core package. A scored mode ships only after its physical reference, correctness, repeatability, pack, and applicable cell gates pass. Validate the Standard time target on the lower-end physical x86_64 system without shortening protocols.
Out of Scope
Official Speedometer 3.1 assets, scores, or references; the uncleared full archive is not a v1 distribution dependency. JetStream 3.0 and MotionMark 1.3.2 are deferred beyond v1.
Hosted leaderboards, accounts, mandatory uploads, and cloud result storage.
Default destructive raw-device write tests, automatic tuning, repairs, driver installation, kernel changes, or self-tests. Long drive self-tests and offline Memtest86+ are separate follow-up flows.
Claims that a benchmark score proves hardware health, that a finite memory pass certifies RAM, that VM results certify bare metal, or that a package install certifies every workload or numerical score.
Production implementation, calibration collection, or release evidence performed by this specification issue itself; they belong to subsequent build and validation tickets.
Further Notes
This spec is the synthesis for implementation ticketing. Find the way to Odin’s build-ready specification remains the decision index; linked resolution comments retain byte-level and threshold details. Use those links when splitting tracer-bullet implementation and validation tickets. Do not silently change a fixed workload, checker, score rule, reference, safety rule, or support claim.
A failed candidate qualification or new driver/tool incompatibility may require a new explicit decision. Until then, record the precise coverage gap rather than substituting a workload or claiming support.
## Problem Statement
Linux users lack one terminal tool that measures their installed system across CPU, memory, GPU, storage, browser, shell, and common language workloads while explaining missing capability and credible health concerns. Isolated benchmark numbers do not show the measurement conditions, cannot establish hardware health, and can mislead when compared across different workloads or environments. Odin needs a reproducible, locally stored, accessible benchmark and diagnostic experience with a defensible median score and evidence-based advice.
## Solution
Build **Odin by Bongbetic**, a Linux terminal suite with Quick, Standard, and explicitly selected Extended run profiles. Standard attempts a fixed seven-domain workload set and publishes a headline median score only when every required measurement is valid under a qualified calibration release. All runs retain native measurements, coverage, health evidence, and execution context locally. The terminal interface provides themed InkUI dashboards, keyboard and qualified mouse controls, plain output, history, comparison, and private-by-default export/import. The supervised runner bounds work and records partial outcomes. Native packages target the approved Void, Debian, and Fedora certification cells. Numerical claims ship only after physical calibration and validation gates pass.
The approved source decisions are indexed by [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). [Approve Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/15#issuecomment-6999) records the product owner's handoff approval. Where this synthesis compresses a byte-level protocol, the linked resolution comment remains normative. Later browser, editor, headed-validation, and Extended-browser decisions supersede only the identified older text; no Speedometer result or asset enters Odin v1.
## User Stories
1. As a Linux user, I want to launch Odin as a normal terminal command, so that I can inspect my system without a desktop application.
2. As a Bash user, I want documented commands and exit statuses, so that I can use Odin in scripts.
3. As a user in a pipe or weak terminal, I want plain machine-usable output, so that results remain readable without terminal graphics.
4. As a user with a capable terminal, I want Bongbetic branding, themed InkUI components, and colorful graphs, so that measurements and warnings are easy to scan.
5. As a keyboard user, I want every action reachable without a mouse, so that the interface remains operable over SSH.
6. As a mouse user, I want clickable controls where my terminal supports mouse reporting, so that I can navigate naturally.
7. As a user whose terminal lacks mouse reporting, I want explicit keyboard parity, so that no action becomes inaccessible.
8. As a user with low color or ASCII output, I want measurements and warnings preserved, so that appearance settings do not hide facts.
9. As a screen-reader user, I want a text-friendly output path, so that measurements and warnings are understandable.
10. As a user who resizes or interrupts the terminal, I want the display and terminal state restored, so that my shell remains usable.
11. As a new user, I want a Quick profile with bounded read-only discovery, so that I can triage capability and health in a few minutes.
12. As a user seeking a broad result, I want a Standard profile targeting 10–20 minutes, so that the requested performance areas are measured together.
13. As a user seeking deeper evidence, I want separately chosen Extended modules with time and resource estimates, so that I control additional stress and writes.
14. As a user on a limited system, I want unavailable workloads named with reasons, so that missing measurements are not mistaken for zero performance.
15. As a user with several GPUs, I want the measured device identified and ambiguous selection resolved, so that a result refers to the intended GPU.
16. As a user with several drives, I want to choose the filesystem path for storage performance while safe discovery can inspect other devices, so that the result has clear scope.
17. As a user concerned about storage wear, I want a clear write estimate and consent prompt, so that a benchmark does not write unexpectedly.
18. As a user who declines storage writes, I want a preserved coverage outcome, so that the run still reports useful evidence.
19. As a user, I want bounded CPU throughput measurements at one worker and at my effective CPU entitlement, so that I can see narrow single-thread and parallel behavior.
20. As a user, I want memory bandwidth and loaded-responsiveness measurements shown separately, so that throughput does not disguise latency under load.
21. As a user, I want memory checks limited to tested bytes and patterns, so that a passing finite check is not presented as proof of fault-free RAM.
22. As a user, I want available EDAC/RAS and pressure evidence retained, so that memory errors and contention can be investigated.
23. As a user, I want GPU discovery, API execution, known-output correctness, and presentation reported as distinct claims, so that driver status is not overstated.
24. As a user without a qualified hardware GPU path, I want software or virtual rendering labelled separately, so that it is not counted as hardware performance.
25. As a user, I want bounded sequential and random storage measurements with mode and device mapping, so that speed results have an interpretable scope.
26. As a user with an unsupported or unsafe health transport, I want an unknown outcome instead of guessed device commands, so that diagnosis does not create risk.
27. As a user, I want an offline browser interaction suite using my installed browser, so that browser responsiveness reflects my installed system.
28. As a user, I want headed and headless browser results identified separately, so that display cadence and execution mode are not conflated.
29. As a user without a compatible browser or driver, I want a precise capability skip, so that Odin does not silently install or substitute software.
30. As a user, I want Bash direct startup, interactive prompt readiness, and arithmetic throughput reported separately, so that shell latency and computation remain distinct.
31. As a user, I want Python, Rust, C++, and Java startup and warmed throughput measurements, so that I can assess representative installed toolchains.
32. As a user, I want language output checks, so that a fast wrong result is never treated as performance.
33. As a user, I want kernel, power, thermal, driver, filesystem, and software context recorded, so that I can interpret a surprising measurement later.
34. As a user, I want health findings separated from speed scores, so that a fast failing drive does not look healthy.
35. As a user, I want health severity and evidence confidence shown separately, so that urgency is not confused with certainty.
36. As a user with unavailable telemetry, I want current health marked unknown, so that missing evidence is not treated as clean evidence.
37. As a user, I want drive replacement advice only when device-attributable evidence supports it, so that low speed alone does not trigger replacement advice.
38. As a user, I want optimization advice tied to observed limitations, tradeoffs, and a verification step, so that I can judge whether a change helped.
39. As a user, I want Odin to stop affected work on qualified thermal, device, integrity, or resource alarms, so that the run respects safety limits.
40. As a user with detected memory corruption, I want the whole run stopped and its partial evidence retained, so that invalid measurements are not scored.
41. As a user who cancels a run, I want worker termination and cleanup attempted with outcomes recorded, so that I know what remains uncertain.
42. As a user who declines privilege, I want the run retained with diagnostic coverage gaps, so that elevation is optional and explicit.
43. As a user, I want a seven-domain headline median only after complete valid Standard coverage, so that missing areas cannot inflate the score.
44. As a user, I want domain breakdowns beside the headline, so that a median does not hide a weak domain.
45. As a user, I want raw measurements before calibration is available, so that Odin never invents a reference score.
46. As a user, I want scores normalized against an immutable calibration release, so that the meaning of 100 and score identity stay stable.
47. As a user, I want Quick and incomplete runs preserved without a partial headline, so that useful evidence survives without a misleading aggregate.
48. As a user, I want loaded responsiveness and health findings visible beside the score, so that unscored diagnostics still inform decisions.
49. As a user, I want installed and prepared-reference software modes labelled separately, so that comparisons retain their software context.
50. As a user, I want VM measurements marked raw-only in v1, so that guest compatibility is not misrepresented as physical score qualification.
51. As a user, I want every benchmark run saved locally, including cancelled and failed runs, so that the evidence remains auditable.
52. As a user, I want raw trials, timestamps, validity reasons, and calculation inputs retained, so that a summary can be checked later.
53. As a user, I want run notes stored separately from original evidence, so that annotation does not rewrite measurements.
54. As a user, I want history to persist until I delete it, so that old results do not disappear unexpectedly.
55. As a user, I want saved history to survive interruption or save failure, so that earlier run records remain intact.
56. As a user, I want to list, filter, inspect, and compare runs in the terminal, so that I can study changes without a hosted service.
57. As a user, I want incompatibility reasons shown for unlike runs, so that a combined change is not calculated across different identities.
58. As a user, I want a portable machine-readable export with private fields redacted by default, so that I can share results without leaking stable identifiers.
59. As a user requesting full fidelity, I want an explicit preview of identifying fields, so that disclosure is deliberate.
60. As a user, I want validated import to preserve original run identity and mark provenance, so that imported results cannot masquerade as local measurements.
61. As a user, I want a clear confirmation before deleting a run, so that history is not removed by accident.
62. As a package user, I want native xbps, deb, or rpm delivery on certified cells, so that installation matches my Linux distribution.
63. As a package user, I want optional workload preparation explicit and recorded, so that normal runs never download or upgrade tools.
64. As a user on musl or another limited cell, I want capability detection and scoped support claims, so that package support does not promise unavailable drivers.
65. As a release reviewer, I want reproducible build, license, and asset manifests, so that distributed packs have verifiable provenance.
66. As a release reviewer, I want physical and VM certification evidence for each required cell, so that a support claim matches actual tests.
67. As a release reviewer, I want physical repeatability, browser qualification, and TUI overhead measured, so that numerical scores have evidence beyond a successful launch.
68. As a release reviewer, I want failed gates to withhold their score identity while retaining raw evidence, so that a package can report capability honestly.
69. As a release reviewer, I want the independent x86_64 and physical aarch64 validation systems documented, so that cross-system and cross-architecture claims have real hardware evidence.
70. As a user, I want no automatic kernel tuning, driver installation, drive repair, or destructive raw-device writes, so that Odin remains a measurement and advice tool.
## Implementation Decisions
### System boundary and delivery
- The v1 terminal application uses TypeScript, React 19, Ink 6, Node 24, and pinned source-copied InkUI components. Bun and an Ink 7 migration are outside this v1 stack. Ship native xbps, deb, and rpm packages for qualified cells. Use a maintained distro Node 24 where available; otherwise bundle a pinned Node 24 runtime and own its updates and notices. Void musl requires a qualified distro-built Node and qualified helpers. No unofficial runtime ABI floor becomes a general Odin support promise. Source: [Choose Odin’s runtime and Linux compatibility contract](https://git.bongbetic.com/xavierk/odin/issues/9#issuecomment-6845).
- Keep the terminal UI separate from an unprivileged, supervised measurement runner. The runner owns immutable run parameters, explicit helper paths and arguments, controlled environment, process groups, deadlines, resource and output bounds, cancellation, cleanup attempts, and durable partial outcomes. Do not construct workload shell commands. Privileged diagnostics expose narrowly validated operation-and-target requests, not a root UI or generic command runner. Quiet decorative updates during timed measurements.
- Build around a public CLI boundary that reaches capability probing, run planning, supervised execution, scoring, health interpretation, and local run-record persistence. This is the primary observable behavior seam. The browser pack and physical validation add separate release evidence for timing and hardware claims.
- Core packages contain only rights-cleared versioned assets. Optional browser, language, and helper preparation is explicit before a run; record origin, version, digest, size, licenses, notices, and cache location. An ordinary run never downloads, installs, upgrades, repairs, or tunes the host. Missing preparation is a capability outcome.
### Profiles, workload identities, and limits
- Quick is a roughly 2–3 minute read-only triage profile: safe health and capability discovery, short CPU/Bash/language probes, memory/PSI context, GPU discovery and bounded correctness, and browser automation availability. Read-only prohibits storage writes, device-changing commands, and self-tests; finite RAM/GPU computation remains possible. Quick has its own shortened probe identities and no full headline.
- Standard targets 10–20 minutes and attempts the fixed suite where capability and prepared assets permit. One filesystem path and one GPU are selected for performance per run. Each worker has a fixed end-to-end deadline including setup and cleanup; a slow system can produce an incomplete run. Never shorten a protocol or trim a slow valid observation to meet the target.
- Extended modules are separately selected with duration, memory, and write estimates. They include contained memory pressure and longer scheduler work; visible GPU presentation and additional devices; queued/durable/integrity or larger file storage work under separate write consent; three full Odin Browser Interaction v1 suites; configured-shell and pipeline work; and compilation/longer language work. JetStream and MotionMark are not v1 modules. Offline Memtest86+ and long drive self-tests remain separate follow-up flows. Sources: [Choose Odin’s workload suite and run profiles](https://git.bongbetic.com/xavierk/odin/issues/8#issuecomment-6705), [Resolve Odin v1 Extended browser workload scope](https://git.bongbetic.com/xavierk/odin/issues/25#issuecomment-6981).
- Outcomes distinguish completed, not scheduled, unavailable by reason, invalid output, execution failure, timeout, cancellation, and safety stop. Preserve partial samples and cleanup status. A missing capability never becomes numerical zero or a substitute workload. Version each workload, protocol, parser, helper, asset, execution mode, and software mode.
- Retain at least the greater of 1 GiB and 25% of effective available memory outside workers. STREAM's three arrays each exceed the relevant last-level cache and are at least the greater of 128 MiB and four times reported last-level cache; total arrays are at most 1.5 GiB. Skip safely when the cache or headroom cannot be established. Online memtester is one finite pass over at most 256 MiB when safe. GPU workers require a qualified VRAM cap no greater than 256 MiB or 25% of known free VRAM. Bounded CPU load targets about half the effective CPU entitlement. The selected storage fixture and all repeated writes stay below the agreed 1 GiB host-write ceiling.
### Standard performance and diagnostic suite
- CPU: pinned sysbench 1.0.20 prime search with `cpu-max-prime=20000`; one-worker and effective-entitlement worker counts; three ten-second timed runs each. The parallel count is the ceiling of the lesser of allowed logical CPUs and effective cgroup quota, with a minimum of one. Preserve events/s, workers, affinity, quota, and throttling. This is a narrow integer workload.
- Memory bandwidth and responsiveness: STREAM 5.10 Copy and Triad preserve their native best-iteration statistic, with three independent runs for cross-run medians. Pinned schbench revision `6300b8f3a8922c61ea6bb2cdfa1901a42c0cc6fc` measures idle and bounded stress-ng 0.22.01 loaded latency in three 15-second pairs. Preserve p50/p95/p99, achieved load, throughput, samples, and PSI deltas. Schbench's v1 parameters and untimed preparation are normative in [Choose Odin’s workload suite and run profiles](https://git.bongbetic.com/xavierk/odin/issues/8#issuecomment-6705). Stress-ng bogo-ops are not scored.
- Memory diagnostics: one finite memtester 4.7.1 pass over at most 256 MiB with actual allocated and locked bytes, patterns, errors, and timeout recorded; available EDAC/RAS and pressure readings before and after. A pass describes only tested scope. Do not infer error-free memory from missing counters.
- GPU: enumerate the selected hardware device, kernel driver, renderer/API, and permissions. Use qualified Vulkan-only clpeak 2.1.4 FP32 and global bandwidth tests plus glmark2 2023.01 fixed 800×600 offscreen GL/GLES shading and texture scenes where supported. Repeat short tests three times and run a separate known-output correctness operation. Keep compute, rendering, and presentation claims distinct. Standard opens no presentation window. Software or virtual rendering has a distinct identity.
- Storage: fio 3.43 operates only on a private 256 MiB fixture in a user-selected filesystem. At queue depth one with `psync`, run 1 MiB sequential read over 256 MiB, sequential write over 128 MiB, and 4 KiB random read and write over 4 MiB each, with three repetitions. Prefer aligned direct I/O; buffered fallback is distinct. Record bandwidth/IOPS, latency distributions, bytes, engine, cache/flush mode, filesystem, and device mapping. Fixture preparation plus repeated writes total 652 MiB before metadata; enforce a 1 GiB ceiling and leave at least 2 GiB plus 10% of filesystem capacity free. Show the estimate and require consent. Never use existing user files or a raw device. This is burst-path performance, not sustained-media speed.
- Browser: use the Odin-owned offline Browser Interaction v1 pack and installed Chromium-family or Firefox browser with a compatible local WebDriver. The pack has deterministic grid, board, and Markdown editor fixtures and eight scripted actions in each scenario. Two untimed warmup cycles precede ten measured cycles, producing 240 valid action-to-visible-update proxy samples. A trusted page input-event timestamp to the second consecutive animation frame after the expected UI state is the native latency measure; this is neither physical photon timing nor the browser's INP metric. The suite statistic is nearest-rank p95, sorted sample 228 of 240, in milliseconds. Standard runs one full suite with a five-minute end-to-end stage deadline. Preserve every sample, action identity, state check, and context. A missing sample, wrong state, external request, focus loss, bad timing, crash, or timeout withholds the browser domain. Run on loopback with a fresh profile, no extensions, clean initial cache, 800×600 CSS viewport, and 100% zoom. Headed and headless, installed and prepared-reference modes, and a separate headless-shell binary have distinct identities. Exact fixture generation, action order, validity, and pack requirements are normative in [Specify Odin’s v1 browser interaction workload](https://git.bongbetic.com/xavierk/odin/issues/20#issuecomment-6927); [Reconcile browser editor fixture and scripted action](https://git.bongbetic.com/xavierk/odin/issues/23#issuecomment-6960) replaces the conflicting editor action text with one-based lines, exact states, and digests. Board lane one is numeric lane ID 1, so moving C000 from lane ID 0 is an actual change. The final shipped app, fixtures, selectors, checker expectations, and assets define the pack identity.
- Bash: 30 clean direct startup samples, 30 clean interactive PTY prompt-readiness samples, and three timed runs of the pinned builtin arithmetic work, verifying sum `43824450000`. Startup files are excluded by default. Keep startup, prompt, and throughput distinct.
- Languages: Python, Rust, C++, and Java run Odin-owned `text-parse-v1`, prepared before timing from the same 1,048,576-record, 12 MiB ASCII corpus. Each verifies record count and weighted sum `266529503798272`; the corpus SHA-256 is `1f7707f5f5e19ec9292a0aecdb67fc241176e58b86dccc04b8cc13acf421633b`. Report completed passes/s from three fixed five-second windows after declared warmup, plus 30 clean process-to-marker startup trials per language. Rust/C++ startup launches a prepared binary, not a compiler. No claim ranks languages intrinsically. Exact corpus formula and source protocol are normative in [Choose Odin’s workload suite and run profiles](https://git.bongbetic.com/xavierk/odin/issues/8#issuecomment-6705).
- Overall health: before and after load, collect available temperature/alarm, CPU/memory/I/O PSI, EDAC/RAS, kernel/device I/O error, and safely supported SMART evidence. Record tool/parser versions, scope, timestamps, and permission gaps. Use qualified read-only protocol operations; never guess a transport, start self-tests, or interpret unknown telemetry as healthy. Source: [Define health findings and the optimization advice rubric](https://git.bongbetic.com/xavierk/odin/issues/11#issuecomment-6740).
### Score, health, and advice
- The headline median score is the median of exactly seven equally represented domain scores: CPU, memory bandwidth, GPU, storage, browser, Bash, and languages. For higher-is-better native values use `100 × measured/reference`; for lower-is-better use `100 × reference/measured`. Take repeated-trial medians in native units first; even-sized medians average their middle two values. Within domains: CPU uses single and parallel throughput; memory uses Copy and Triad; GPU uses compute and rendering summaries; storage uses four fixed read/write results; browser uses one complete p95 result; Bash uses startup, prompt, and arithmetic; languages take a startup/throughput median per language then the median of four languages. Loaded responsiveness and health remain beside, not inside, the numerical aggregate. Source: [Define Odin’s median score and comparability contract](https://git.bongbetic.com/xavierk/odin/issues/10#issuecomment-6723), with the browser substitution from [Specify Odin’s v1 browser interaction workload](https://git.bongbetic.com/xavierk/odin/issues/20#issuecomment-6927).
- No full headline appears unless a Standard set has every required valid measurement, the corresponding physical mode has qualified reference evidence, and integrity/safety rules permit it. Missing data, detected memory errors, failed output verification, incomplete required work, or a safety stop withhold the headline. A slow valid result remains valid. Missing health telemetry alone leaves health unknown but does not suppress a complete valid score. Quick and incomplete runs show raw values, available domains, and coverage without a partial aggregate. VM runs are raw-only in v1. Never invent a reference or confidence interval.
- Health finding severity is informational, warning, or critical; finding confidence is directly observed, corroborated, or tentative. Apply confidence to the claim actually supported. Missing, stale, unreadable, or uninterpretable evidence is unknown. Separate current observation, during-run change, and historical evidence. A passing limited check is not overall health certification. Drive replacement advice requires credible device-reported failure, applicable degradation, read-only device failure, or repeated device-attributable integrity errors; endurance consumption, low speed, temperature, or ambiguous interface errors alone need their own narrower advice. Use documented alarms and applicable limits, never invented universal thresholds.
- Stop affected work and later work on the same component for active thermal alarms, device loss, integrity error, or breached limits. Stop the entire run for memory corruption or inseparable safety/integrity risk. Do not automatically retry. Preserve evidence, partial samples, cleanup failures, and unresolved I/O. Advice requires corroborating workload-specific evidence before asserting a bottleneck; hypotheses remain labelled, and recommendations state tradeoffs plus a verification step. Tuning, repair, and driver changes remain advice only.
### Run records, history, and comparison
- A run record preserves every attempt, including invalid, cancelled, incomplete, and safety-stopped runs: raw trials, native statistics, timestamps/order, units, summary inputs, validity and coverage reasons, environment, workload/tool/parser/protocol versions, calibration and score identities, permissions, thermal/power context, health source and confidence, advice evidence, and cleanup outcome. Original evidence and scores are immutable; user notes are separate. Version schema and calculations. Source: [Define local results, history, and export behavior](https://git.bongbetic.com/xavierk/odin/issues/12#issuecomment-6750).
- Save by default in a per-user local data directory distinct from the storage benchmark target. Permit choosing/opening that directory. Retain history until explicit deletion; show storage use and support deliberate cleanup. Interrupted or failed writes preserve committed history and report whether the current run was saved or needs recovery. Terminal history lists, filters, inspects, compares, annotates, exports, and confirms deletion.
- Direct comparison requires matching scoring rules, calibration release, workload protocol, browser execution/software mode, storage direct/buffered mode, and physical/virtual context. Show unlike runs side by side with precise incompatibility reasons, without a combined change or unsupported causal claim. Installed software versions may differ within a compatible cohort but remain visible as environmental changes.
- v1 exports a versioned portable machine-readable record with all measurement and interpretation evidence. Default export redacts hostname, username, home paths, serials, and stable hardware IDs while retaining non-identifying comparison facts. Full-fidelity export requires explicit selection and preview. If privacy rules conflict with comparison-critical fields, explain the conflict rather than silently leaking or deleting data. Import only after schema and integrity checks, mark provenance, and preserve original identity.
### TUI and support contract
- Offer the three product-owner-approved layouts: Mission Control default, Guided Run, and Run Ledger. Preserve plan, live-run, summary, history, advice, profile, theme, ASCII, cancellation, and layout navigation described by [Evaluate the Bongbetic TUI workflow with InkUI](https://git.bongbetic.com/xavierk/odin/issues/13#issuecomment-6895). Keyboard parity is mandatory; qualify mouse only where the terminal reports it. Handle focus, scroll, resize, weak/no-color/ASCII paths, screen-reader-friendly output, and restoration of raw input, mouse mode, cursor, and screen on normal exit and handled interruption. Do not claim restoration after SIGKILL.
- The eight required certification cells are Void glibc and Void musl on each of x86_64 and aarch64, plus Debian 13 glibc and Fedora 44 glibc on each architecture. Pin exact tested image/rootfs, repository snapshot or point release, package builds, kernel configuration, Odin and helper builds. Test each cell physically and in a reproducible VM before a bare-metal claim; VM evidence proves guest compatibility only. All eight cells need package, CLI/plain output, terminal restore, Quick/Standard, cancellation, partial persistence, capability skips, minimal/SSH checks. All six glibc cells additionally require headed UI and browser functional checks on physical hardware and in VMs; required headed UI failure blocks the eight-cell release. Mouse support remains terminal-capability-scoped. Void musl headed behavior is claimed only if separately qualified. Sources: [Set distro, VM, and physical-hardware validation gates](https://git.bongbetic.com/xavierk/odin/issues/14#issuecomment-6939), [Clarify headed validation for glibc certification cells](https://git.bongbetic.com/xavierk/odin/issues/24#issuecomment-6965).
## Testing Decisions
- The primary automated seam is the public CLI with controlled workload, capability, sensor, and device fixtures. Test externally visible exit status, plain/TUI information, worker outcome, score eligibility, saved run record, cancellation, and cleanup reporting. Test behavior rather than private functions, rendering trees, or implementation-specific file layout. No production code or existing test suite is present yet; the approved TUI prototype supplies interaction examples but simulated measurements, not benchmark prior art.
- Exercise each workload adapter with known inputs, wrong-output fixtures, version mismatch, missing permission, insufficient safe resources, timeout, cancellation, cleanup failure, and slow-but-valid output. Recompute pinned Bash and language checks independently. Check native units, repetitions, result identities, and score formulas at the CLI/run-record boundary. Include complete and missing-domain cases, memory-error and GPU/browser verification failures, unknown health telemetry, and privileged-check denial.
- Browser acceptance uses real supported Chromium-family and Firefox plus compatible drivers for headed and headless modes. Verify deterministic fixture bytes; corrected editor expected states and SHA-256 after each action; independent grid/board/editor state checks; exactly 240 positive finite samples; p95 rank; focus and animation-frame timing; loopback-only assets; no outbound requests; deadline; partial evidence; and mode identity. Run unscored offline smoke tests with outbound network blocked for each qualified engine/mode. The pack manifest uses RFC 8785 canonical JSON, sorted normalized paths, per-file byte length, SHA-256, provenance, rights/notice evidence, and canonical-manifest SHA-256 as pack ID. Reject unlisted assets, symlinks, path traversal, and missing notices. Changing final pack bytes changes identity and requires requalification. Source: [Specify Odin’s v1 browser interaction workload](https://git.bongbetic.com/xavierk/odin/issues/20#issuecomment-6927).
- Test health rules with protocol-specific fixtures, including historical versus current alarms, missing/stale data, ambiguous I/O attribution, smartctl nonzero bitmask with usable output, parser version mismatch, GPU discovery/execution/correctness/presentation distinctions, VM guest scope, drive transport safety, and uncertain advice. Test safety-stop scope and preserved partial records. No test deliberately creates real hardware faults.
- Test local history through interrupted writes, recovery, immutable evidence with editable notes, compatible and incompatible comparisons, default-redacted and explicit full-fidelity exports, schema/integrity failures on import, imported provenance, deletion confirmation, and history retention.
- Calibrate every scored protocol/mode on one documented physical x86_64 reference with at least five complete eligible Standard runs across two days and two cold boots. Capture a two-minute idle baseline, readiness/cooling conditions, all failed attempts, exact environment and pack/tool hashes. Each reference workload uses the median of run-level native values. Corpus median absolute deviation divided by median must be at most 5%, or 10% for the browser suite. Validate the shared reference on at least two independent physical systems including aarch64, review alternate-reference sensitivity greater than 10% or rank reversal, and publish immutable raw corpus and manifest before numerical scores. New browser p95 requires fresh reference evidence; old Speedometer values never transfer. Source: [Define Odin’s calibration procedure and release gates](https://git.bongbetic.com/xavierk/odin/issues/17#issuecomment-6776).
- On the reference and both validation systems, run six complete Standard runs in one fixed scored OS/mode: three after each of two cold boots on different days. Apply the approved workload, domain, headline, and cross-boot spread limits; preserve valid outliers and every failed attempt. On reference and Arm, run three alternating matched TUI/plain pairs; median per-workload TUI overhead must be at most 5%. Physically exercise native NVMe and SATA HDD, two x86 GPU driver families, one Arm GPU, real SMART/thermal paths, device mapping, and storage direct I/O. VMs prove compatibility, not physical hardware or numerical score validity. Exact thresholds and matrix are normative in [Set distro, VM, and physical-hardware validation gates](https://git.bongbetic.com/xavierk/odin/issues/14#issuecomment-6939).
- A core install, CLI, or required headed UI failure in a required cell blocks the eight-cell release. A failed workload, repeatability, or calibration gate withholds that score identity and dependent headline while retaining raw evidence; it need not block a functioning core package. A scored mode ships only after its physical reference, correctness, repeatability, pack, and applicable cell gates pass. Validate the Standard time target on the lower-end physical x86_64 system without shortening protocols.
## Out of Scope
- Official Speedometer 3.1 assets, scores, or references; the uncleared full archive is not a v1 distribution dependency. JetStream 3.0 and MotionMark 1.3.2 are deferred beyond v1.
- Hosted leaderboards, accounts, mandatory uploads, and cloud result storage.
- Default destructive raw-device write tests, automatic tuning, repairs, driver installation, kernel changes, or self-tests. Long drive self-tests and offline Memtest86+ are separate follow-up flows.
- Claims that a benchmark score proves hardware health, that a finite memory pass certifies RAM, that VM results certify bare metal, or that a package install certifies every workload or numerical score.
- Production implementation, calibration collection, or release evidence performed by this specification issue itself; they belong to subsequent build and validation tickets.
## Further Notes
- This spec is the synthesis for implementation ticketing. [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1) remains the decision index; linked resolution comments retain byte-level and threshold details. Use those links when splitting tracer-bullet implementation and validation tickets. Do not silently change a fixed workload, checker, score rule, reference, safety rule, or support claim.
- Access to [Secure physical aarch64 validation system for Odin v1](https://git.bongbetic.com/xavierk/odin/issues/21) and [Secure independent x86_64 validation system for Odin v1](https://git.bongbetic.com/xavierk/odin/issues/22) remains open. These are release-validation prerequisites, not duplicate implementation slices. Link them as native blockers for physical validation and scored-release tickets that need them; unrelated build work can proceed.
- A failed candidate qualification or new driver/tool incompatibility may require a new explicit decision. Until then, record the precise coverage gap rather than substituting a workload or claiming support.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Problem Statement
Linux users lack one terminal tool that measures their installed system across CPU, memory, GPU, storage, browser, shell, and common language workloads while explaining missing capability and credible health concerns. Isolated benchmark numbers do not show the measurement conditions, cannot establish hardware health, and can mislead when compared across different workloads or environments. Odin needs a reproducible, locally stored, accessible benchmark and diagnostic experience with a defensible median score and evidence-based advice.
Solution
Build Odin by Bongbetic, a Linux terminal suite with Quick, Standard, and explicitly selected Extended run profiles. Standard attempts a fixed seven-domain workload set and publishes a headline median score only when every required measurement is valid under a qualified calibration release. All runs retain native measurements, coverage, health evidence, and execution context locally. The terminal interface provides themed InkUI dashboards, keyboard and qualified mouse controls, plain output, history, comparison, and private-by-default export/import. The supervised runner bounds work and records partial outcomes. Native packages target the approved Void, Debian, and Fedora certification cells. Numerical claims ship only after physical calibration and validation gates pass.
The approved source decisions are indexed by Find the way to Odin’s build-ready specification. Approve Odin’s build-ready specification records the product owner's handoff approval. Where this synthesis compresses a byte-level protocol, the linked resolution comment remains normative. Later browser, editor, headed-validation, and Extended-browser decisions supersede only the identified older text; no Speedometer result or asset enters Odin v1.
User Stories
Implementation Decisions
System boundary and delivery
Profiles, workload identities, and limits
Standard performance and diagnostic suite
cpu-max-prime=20000; one-worker and effective-entitlement worker counts; three ten-second timed runs each. The parallel count is the ceiling of the lesser of allowed logical CPUs and effective cgroup quota, with a minimum of one. Preserve events/s, workers, affinity, quota, and throttling. This is a narrow integer workload.6300b8f3a8922c61ea6bb2cdfa1901a42c0cc6fcmeasures idle and bounded stress-ng 0.22.01 loaded latency in three 15-second pairs. Preserve p50/p95/p99, achieved load, throughput, samples, and PSI deltas. Schbench's v1 parameters and untimed preparation are normative in Choose Odin’s workload suite and run profiles. Stress-ng bogo-ops are not scored.psync, run 1 MiB sequential read over 256 MiB, sequential write over 128 MiB, and 4 KiB random read and write over 4 MiB each, with three repetitions. Prefer aligned direct I/O; buffered fallback is distinct. Record bandwidth/IOPS, latency distributions, bytes, engine, cache/flush mode, filesystem, and device mapping. Fixture preparation plus repeated writes total 652 MiB before metadata; enforce a 1 GiB ceiling and leave at least 2 GiB plus 10% of filesystem capacity free. Show the estimate and require consent. Never use existing user files or a raw device. This is burst-path performance, not sustained-media speed.43824450000. Startup files are excluded by default. Keep startup, prompt, and throughput distinct.text-parse-v1, prepared before timing from the same 1,048,576-record, 12 MiB ASCII corpus. Each verifies record count and weighted sum266529503798272; the corpus SHA-256 is1f7707f5f5e19ec9292a0aecdb67fc241176e58b86dccc04b8cc13acf421633b. Report completed passes/s from three fixed five-second windows after declared warmup, plus 30 clean process-to-marker startup trials per language. Rust/C++ startup launches a prepared binary, not a compiler. No claim ranks languages intrinsically. Exact corpus formula and source protocol are normative in Choose Odin’s workload suite and run profiles.Score, health, and advice
100 × measured/reference; for lower-is-better use100 × reference/measured. Take repeated-trial medians in native units first; even-sized medians average their middle two values. Within domains: CPU uses single and parallel throughput; memory uses Copy and Triad; GPU uses compute and rendering summaries; storage uses four fixed read/write results; browser uses one complete p95 result; Bash uses startup, prompt, and arithmetic; languages take a startup/throughput median per language then the median of four languages. Loaded responsiveness and health remain beside, not inside, the numerical aggregate. Source: Define Odin’s median score and comparability contract, with the browser substitution from Specify Odin’s v1 browser interaction workload.Run records, history, and comparison
TUI and support contract
Testing Decisions
Out of Scope
Further Notes