Author SHA1 Message Date
Codex 20681cd4f1 Research InkUI architecture and Linux delivery options 2026-09-25 23:30:37 +05:30
2 changed files with 141 additions and 140 deletions
+141
View File
@@ -0,0 +1,141 @@
# Establish InkUI architecture and Linux delivery options
Research for [Establish InkUI architecture and Linux delivery options](https://git.bongbetic.com/xavierk/odin/issues/6), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
Access date: **2026-09-25**. This report records evidence and conditional recommendations, not an architecture decision. No dependencies were installed, application code built, or benchmarks run. InkUI is a user requirement throughout.
## Decision summary
The strongest initial candidate is a TypeScript/React terminal application using **copied InkUI components**, a maintained Node runtime, and separately supervised workload processes. Node 24 provides a maintained baseline today. Ink 7.1.1 offers useful mouse-layout and terminal-lifecycle primitives, but moving InkUI beyond its declared Ink 6 range requires a focused compatibility proof. Staying on Ink 6.8.0 avoids that version change while leaving more application work around positioning and terminal ownership.
Bun is a credible alternative delivery/runtime candidate because it publishes glibc and musl binaries for both requested architectures and supports compiled executables. Its documented Node compatibility differences mean it needs an explicit subprocess, terminal, and asset-loading qualification before adoption. Neither runtime choice makes every benchmark available on every machine. **Capability coverage**, operating-system support, and comparable **performance measurements** need separate contracts.
## 1. Exact InkUI identity and maturity
The requested project is **kamlesh723/InkUI**, not the separate `vadimdemedes/ink-ui` library. Its installer is `@inkui-cli/inkui`; npm reports **0.5.0**, published May 4, 2026. GitHub’s latest release is still **v0.4.0**, published April 12. The inspected main commit is `e3110d89b3f0933bcb297a6af33318124c889f36`, also May 4. Pin the actual source commit and installer version rather than treating those release signals as interchangeable. [1][2]
InkUI copies `.tsx` source and shared theme code into the consuming project. Individual component packages also exist. The source-copy route fits adapting mouse handling, accessible output, and Bongbetic themes, but transfers responsibility for reviewing upstream fixes and maintaining local changes. Its MIT notice must accompany copied/substantial source. Upstream documentation specifies Node ≥20, React `^19.0.0`, and Ink `^6.0.0`; its CI covers Ubuntu with Node 20, not a Linux/libc/terminal matrix. These facts establish an early component project with useful source, not broad deployment certification. [1][3]
Verified component capabilities:
- **Themes:** semantic color tokens, dark/light defaults, custom colors, and border styles. A global theme context is application code in the guide, not automatic global behavior.
- **Keyboard:** input components use Ink input handlers; hooks provide indexed focus cycling and key bindings. Odin still owns modal focus, consistent shortcuts, and cancellation.
- **Graphs:** Sparkline uses Unicode blocks, averages samples when downsampling, and can display latest/min/max values. Its declared `height` parameter is unused in the inspected implementation. Gauge supports bar, ring, and arc forms with thresholds. These are small displays, not a complete plotting/interaction system.
- **Fallback:** ASCII borders exist, but graph glyphs remain Unicode. Terminal sizing tracks resize events and defaults to 80×24.
- **Mouse/accessibility:** inspection found no terminal mouse implementation or ARIA annotations in InkUI components; mouse references in the repository belong to its website. The README’s accessibility description is not evidence of complete accessible interactions. [4]
## 2. Runtime choices and support floors
| Candidate | Evidence and fit | Decision condition |
| --- | --- | --- |
| Node 24 + React 19 + Ink 6.8.0 | Ink 6.8 requires Node ≥20 and React ≥19; it satisfies InkUI’s declared major range. Node 24 remains supported through April 2028. | Prefer if preserving the declared InkUI range matters most and the interaction proof can supply the missing mouse/terminal behavior cleanly. |
| Node 24 + React ≥19.2 + Ink 7.1.1 | The released Ink 7.1.1 manifest requires Node ≥22 and React ≥19.2. It adds public element coordinates and terminal controls useful to Odin. | Prefer if copied InkUI components pass adaptation tests; do not force installation past peer constraints and call that compatibility. |
| Node 26 | Scheduled Current release today, with LTS due October 28, 2026. Official Linux builds also require `libatomic`, unlike the Node 24 build documentation. | Consider as a forward-compatibility lane or release baseline after qualification; “newest” alone is insufficient. |
| Bun 1.4.2 | Latest inspected release; official x64/arm64 glibc/musl artifacts and executable compilation. Node `tty` is documented as implemented, with behavioral differences; `child_process` and `process` remain partially compatible. | Candidate only after the same terminal, process-tree cancellation, signals, and dependency-asset tests as Node. No InkUI-on-Bun certification was established here. |
Node 20 reached its scheduled EOL on April 30, 2026; Node 22 is maintained until April 2027, while Node 25 is already EOL. React’s current registry version is 19.3.0, but an actual runtime/renderer pair must be pinned together. The released Ink 7.1.1 manifest differs from its evolving main branch, so use release-tag evidence for implementation. [5][6][7]
**Node 24’s official Linux baseline** is kernel ≥4.18, glibc ≥2.28, and libstdc++ providing `GLIBCXX_3.4.25`, for x64 and arm64 Tier 1 platforms. Older kernels may run but are not the documented official-binary baseline. Upstream also excludes vendor-EOL operating systems. x64 musl appears as Experimental; arm64 musl is absent from that platform table. Unofficial musl builds exist for both architectures but explicitly disclaim rigorous testing and guarantees. [8]
Void supplies downstream Node builds: its inspected package recipe is Node 24.18.0 with distribution library dependencies and cross-build handling. That is a viable source of a musl runtime, not proof that an arbitrary downloaded Node binary works on musl. Void itself supports glibc and musl variants. The handbook specifically states that proprietary NVIDIA drivers do not support musl; Odin must explain that limitation rather than attempting to hide it with packaging. [9]
Bun’s current documentation specifies glibc ≥2.17, x64 SSE4.2/Nehalem or newer, official musl alternatives, and both requested architectures. It recommends kernel ≥5.6 while claiming operation down to 3.10 with degraded newer syscalls. Current x64 documentation says one binary selects AVX paths at runtime; old advice about choosing separate baseline/modern builds is stale. The inspected installation page does not establish a precise musl version floor. Certify the chosen release and artifact; do not interpret a runtime boot floor as Odin’s complete feature floor. [7]
## 3. Mouse, terminal behavior, and accessibility
**Mouse is the main UI feasibility gate.** Xterm defines reporting modes and SGR mouse encoding; enabling reporting is only the transport. Odin also needs event decoding, click/wheel semantics, clipping, focus transfer, hit testing, and restoration of modes. A small established input adapter is worth assessing before custom parsing, but this investigation does not endorse an unexamined package. InkUI remains the component source. [10]
Ink 6’s public measurement API returns only dimensions. Ink 7.1.1 returns `x`, `y`, width, and height after layout. Its documentation explicitly warns that these coordinates are relative to the live layout, not the terminal viewport. Static output above the live region, scroll offsets, borders, and resizing must be accounted for before comparing mouse coordinates. Neither inspected Ink API supplies a complete mouse interaction system. The proof should cover clicking a scrolled table row, scrolling inside a panel, overlapping dialogs, resizing, and keyboard parity. [5][11]
Ink 7.1.1 documents alternate-screen ownership, noninteractive output, render-rate control, terminal suspension, and screen-reader mode. Noninteractive rendering emits only the final non-static frame at unmount, which is not a sufficient progress/reporting policy by itself. Ink’s basic screen-reader mode supports a subset of ARIA and `INK_SCREEN_READER=true`; Odin must add meaningful labels, numeric graph alternatives, state descriptions, and restrained updates. Release-tag documentation says raw-mode setters can throw when unsupported; indexed main-branch documentation described newer behavior. Do not mix these contracts. [11]
Recommended terminal contract for the subsequent UX decision:
- Full interactive layout in certified terminals; keyboard access to every action; mouse where reporting is available. Mouse absence should not make a run unusable.
- An explicit plain mode for redirected output, unavailable raw input, `TERM=dumb`, and accessibility preferences. Keep readable progress, result summaries, and noninteractive arguments for Bash automation. Bash support means a normal executable and stable exit behavior; orchestration should not depend on users’ shell aliases or startup files.
- Independent controls for color, Unicode decoration, animation, and interactive layout. `NO_COLOR` concerns ANSI color and explicitly does not disable bold/underline; it is not an ASCII or screen-reader switch. Graph numbers and health states must remain understandable without color. [12]
- No unsupported claim that `$TERM` alone proves capability. Qualify an xterm-compatible/VTE terminal, a modern enhanced-protocol terminal, SSH with a PTY, tmux with mouse on/off, Linux console, and pipes. tmux mediates events and terminal features; its configuration and remote terminfo are part of the compatibility environment. [13]
- On normal exit, cancellation, exceptions, SIGTERM, and terminal handoff: restore raw mode, mouse reporting, cursor, and screen state. SIGKILL cannot execute application cleanup; qualify recovery behavior and a readable rerun, without promising impossible restoration guarantees.
## 4. Distribution and offline execution
Treat **xbps, deb, and rpm as delivery formats**, not interchangeable dependency namespaces or compatibility guarantees. A native payload still carries architecture, libc, shared-library, and CPU constraints. Debian’s policies distinguish absolute dependencies from recommendations and require generated shared-library dependencies; RPM records requirements and package signatures; Void separates architecture/libc repositories and requires signed remote repositories. [14]
Three delivery options merit a final choice:
| Delivery | Advantage | Ownership cost / qualification |
| --- | --- | --- |
| Native packages using a maintained distro Node | Distro security updates, conventional dependency ownership, natural Void musl integration. | Supported distro releases must provide the selected runtime range; package names and versions vary. JS-only files can be architecture-independent, but included runtime/helper binaries cannot. |
| Portable archive containing pinned Node plus built JS/assets | Predictable runtime without npm on the target; simple inspection and Bash launcher. | Separate x64/arm64 and glibc/musl artifacts; Bongbetic owns runtime security updates, signatures, licenses, and musl build provenance. A bundled runtime still has system-library requirements. |
| Bun compiled executable | Documented targets cover the four architecture/libc combinations and can embed assets. | A separate binary per target still needs qualification. Current docs enable `.env`/`bunfig.toml` autoloading by default; deterministic execution should disable or explicitly control that behavior. Subprocess compatibility remains a gate. |
Node 24’s single-executable feature is in active development and embeds a CommonJS script; ESM dependencies, bundled assets, and cross-platform code-cache restrictions add work. It should not be the default merely to produce one file. A portable directory can be a complete product. [15]
Recommend separating acquisition from a **benchmark run**. Core UI, manifests, fixtures, and chosen baseline workloads should run offline once prepared. Optional tools, browser engines, language toolchains, and large datasets need declared versions, origins, digests, licenses, sizes, and cache locations. Distribution packages or a verified offline bundle can supply them. Installation should be an explicit preparation action; a timed measurement must not fetch “latest,” upgrade drivers, or change kernel settings.
When a dependency is unavailable, record a precise capability-coverage reason. Do not replace a workload silently and reuse its score identity. Decide later whether the default installer acquires the full standard run profile or a smaller core with optional packs. Keep package-manager scripts and network access out of privileged measurement operations.
## 5. Measurement and privilege boundaries
A UI runtime does not need to be the workload implementation language. Node’s asynchronous child-process API can orchestrate existing executables without a shell, so the evidence does not require a second application language now. A native helper becomes justified if a selected workload needs precise memory access, direct kernel APIs, or a small enforceable watchdog/privilege boundary that existing tools do not provide. That choice belongs after workload requirements are known. [16]
Proposed ownership boundaries:
1. An unprivileged InkUI process owns navigation, branding, readable progress, and reports.
2. A runner owns the run manifest, executable identity, timing, output limits, process lifecycle, and persisted measurements. Use explicit arguments and controlled environment variables, not shell command construction.
3. Privileged operations, when required, receive narrowly validated operation/target requests. Do not elevate the complete React/Bun/Node application or expose a generic root command executor. Validate device identity, paths, ownership, and temporary-file boundaries; retain original-user ownership of local results. Read-only diagnostics and storage writes require distinct authorization/safety rules.
UI graphs, animation timers, garbage collection, logging, and synchronous parsing can perturb CPU, memory, and I/O measurements. Separate processes prevent event-loop coupling but still share hardware. Recommend pausing decorative updates during timed windows, buffering bounded observations, and reporting afterward. If live charts remain necessary, cap their sampling/render rate and quantify UI-on versus quiet-mode overhead on low-end hardware. Do not silently reserve a core or change affinity/governors: these alter the benchmark’s execution conditions. Record any chosen isolation policy in the run manifest.
Cancellation is more than calling `child.kill()`: Node documents that killing a parent does not kill Linux descendants. A runner needs process-group ownership, staged termination, exit collection, and orphan handling after UI failure. cgroup v2 can provide stronger group limits and termination, but controller availability and delegation permissions must be probed. Kernel documentation states that controller exposure depends on configuration and hierarchy ownership; `cgroup.kill` and memory limits cannot be presumed available on every system. No systemd dependency is necessary for the core model. [16][17]
For dangerous pressure profiles, refuse execution when required containment or the independent watchdog cannot be established; explain the unavailable capability. A plain performance run can use a simpler supervision path if its workload contract allows it. Discover actual kernel interfaces such as PSI files rather than assuming them from version numbers; record missing features without inventing zero pressure. [17][18]
## 6. Proposed certification matrix
These are future release gates, not completed tests. Use exact image/artifact digests and record kernel, architecture, libc, runtime, terminal, and relevant permissions.
| Lane | Suggested coverage | What it establishes |
| --- | --- | --- |
| Native packages | Void glibc + musl; maintained Debian/Ubuntu; maintained Fedora and enterprise RPM family. | Install, upgrade/remove, dependency declarations, offline start, paths/ownership, normal-user behavior. |
| Architecture/libc | x86_64/glibc, x86_64/musl, aarch64/glibc, aarch64/musl. At least Void and Alpine as distinct musl environments if both are claimed. | Loader/runtime and transitive-asset compatibility. Native ARM hardware is required before treating emulation success as ARM performance evidence. |
| Kernel/permissions | Oldest promised supported ABI; supported LTS and current distro kernels; runit and systemd; cgroup v2 delegated/denied/absent; restricted proc/sys access. | Capability probing, fallback, supervision, explicit refusal where safe containment is unavailable. |
| Terminal/input | Local VTE/xterm-compatible terminal, a modern terminal, SSH PTY, tmux, Linux console, 80×24 and narrow resize, Unicode/ASCII, no color, screen reader, piped output. | Input, readable numeric alternatives, output contract, cleanup, no raw-mode crash. |
| Failures | Cancel each workload phase, terminate UI/runner, missing tool, permission denial, low disk space, invalid persisted result. | Process/resource cleanup, bounded logs, durable result state, useful failure explanations. |
| Physical measurement | Intel/AMD x86_64 and ARM; AMD/Intel/NVIDIA graphics where supported; NVMe/SATA/HDD; laptops with thermal/power transitions. | Device visibility, real drivers, SMART interpretation paths, thermal effects, representative performance and UI-overhead measurement. |
VMs are suitable for package/ABI/terminal/failure-path qualification and measuring the VM itself. Virtual storage, virtual GPUs, host scheduling, and hidden hardware telemetry cannot certify physical drive replacement advice, native GPU performance, or host memory health. Passthrough creates a separate recorded environment, not an exemption from those distinctions. Real faults should be covered with parser fixtures and controlled validation; do not deliberately damage hardware to prove a health rule.
## 7. Decisions still required
1. Node 24 versus another maintained runtime; Ink 6 compatibility versus adapting copied components to Ink 7; exact version pin and upgrade policy.
2. Native-package dependencies, portable runtime archives, or both; who maintains and certifies musl builds; minimum OS/kernel/libc/CPU contract.
3. A mouse/focus/terminal-lifecycle prototype retaining InkUI, plus the accessible/plain output acceptance criteria.
4. Core versus optional offline workload packs, licensing and size limits, and dependency acquisition policy.
5. Per-workload privilege/containment requirements, behavior after UI death, and the quantitative UI-overhead acceptance threshold.
6. Exact release certification matrix and which capabilities can be certified only on physical hardware.
Evidence limits: no compatibility or performance claim was established by execution; no universal Node/Bun/InkUI bundle was produced; installed Void package availability on every target was not checked; Bun’s precise musl floor and InkUI-on-Ink-7/Bun behavior remain unverified. Browser/driver/workload selection and scoring are separate investigations.
## Sources and method
All sources below were inspected on **2026-09-25**. Context7 library resolution preceded documentation queries for Node, Ink, Bun, XBPS/Void, Debian Policy, RPM, and tmux. Two earlier exact-InkUI searches returned unrelated libraries and were rejected; the requested repository was inspected directly. Context7 results tracking `master` were checked against released source where version behavior mattered. No quota error occurred.
1. [InkUI README at inspected commit](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/README.md), [installer manifest](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/apps/cli/package.json), [npm registry](https://registry.npmjs.org/@inkui-cli%2Finkui).
2. [InkUI v0.4.0 release](https://github.com/kamlesh723/InkUI/releases/tag/v0.4.0), [inspected commit](https://github.com/kamlesh723/InkUI/commit/e3110d89b3f0933bcb297a6af33318124c889f36).
3. [InkUI MIT license](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/LICENSE), [installation guide](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/apps/docs/content/getting-started/installation.mdx), [CI](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/.github/workflows/ci.yml).
4. InkUI inspected source: [themes](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/packages/core/src/theme.ts), [hooks](https://github.com/kamlesh723/InkUI/tree/e3110d89b3f0933bcb297a6af33318124c889f36/packages/hooks/src), [Sparkline](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/packages/sparkline/src/Sparkline.tsx), [Gauge](https://github.com/kamlesh723/InkUI/blob/e3110d89b3f0933bcb297a6af33318124c889f36/packages/gauge/src/Gauge.tsx).
5. [Ink 6.8.0 README](https://github.com/vadimdemedes/ink/blob/v6.8.0/readme.md), [Ink 7.1.1 manifest](https://github.com/vadimdemedes/ink/blob/v7.1.1/package.json), [Ink npm metadata](https://registry.npmjs.org/ink), [React registry](https://registry.npmjs.org/react/latest).
6. [Node official release schedule](https://github.com/nodejs/Release/blob/main/schedule.json), [Node 26 platform/build document](https://github.com/nodejs/node/blob/v26.x/BUILDING.md).
7. [Bun 1.4.2 release](https://github.com/oven-sh/bun/releases/tag/bun-v1.4.2), [installation/platform requirements](https://bun.com/docs/installation), [Node API compatibility](https://bun.com/docs/runtime/nodejs-compat).
8. [Node 24 supported platforms and official binary requirements](https://github.com/nodejs/node/blob/v24.x/BUILDING.md), [unofficial-builds limitations and targets](https://github.com/nodejs/unofficial-builds/blob/main/README.md).
9. [Void Node package recipe](https://github.com/void-linux/void-packages/blob/master/srcpkgs/nodejs/template), [Void musl support and incompatible software](https://docs.voidlinux.org/installation/musl.html).
10. [Xterm control sequences: mouse tracking and SGR encoding](https://invisible-island.net/xterm/ctlseqs/ctlseqs.html#h2-Mouse-Tracking).
11. [Ink 7.1.1 released README](https://github.com/vadimdemedes/ink/blob/v7.1.1/readme.md), [released measurement implementation](https://github.com/vadimdemedes/ink/blob/v7.1.1/src/measure-element.ts).
12. [NO_COLOR informal standard and FAQ](https://no-color.org/).
13. [tmux getting started](https://github.com/tmux/tmux/wiki/Getting-Started), [modifier keys and terminfo](https://github.com/tmux/tmux/wiki/Modifier-Keys), [tmux manual source](https://github.com/tmux/tmux/blob/master/tmux.1).
14. [Void repositories](https://docs.voidlinux.org/xbps/repositories/index.html), [Void signing](https://docs.voidlinux.org/xbps/repositories/signing.html), [Debian package relationships](https://www.debian.org/doc/debian-policy/ch-relationships.html), [Debian shared-library dependencies](https://www.debian.org/doc/debian-policy/ch-sharedlibs.html), [RPM dependency tags](https://github.com/rpm-software-management/rpm/blob/master/docs/manual/tags.md), [RPM package signatures](https://github.com/rpm-software-management/rpm/blob/master/docs/manual/format_v4.md).
15. [Bun executable targets, embedding, and autoload controls](https://bun.com/docs/bundler/executables), [Node 24 single-executable applications](https://github.com/nodejs/node/blob/v24.x/doc/api/single-executable-applications.md).
16. [Node child-process API: spawn, detached processes, signals, and descendant caveat](https://nodejs.org/docs/latest-v24.x/api/child_process.html).
17. [Linux cgroup v2: delegation, controllers, memory limits, and cgroup.kill](https://docs.kernel.org/admin-guide/cgroup-v2.html).
18. [Linux pressure stall information](https://docs.kernel.org/accounting/psi.html).
-140
View File
@@ -1,140 +0,0 @@
# Defensible median scoring and comparison rules
Research for [Establish defensible median scoring and comparison rules](https://git.bongbetic.com/xavierk/odin/issues/7), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
**Accessed:** 25 September 2026. **Status:** decision evidence and recommendations; no final scoring formula, calibrated reference values, benchmark runs, or implementation. The user's requirement is a **median** score.
## What the median should mean
There are three separate choices:
| Level | Meaning | Main limitation |
|---|---|---|
| Median of repeated measurements | Typical result for one fixed workload under stated conditions | Hides occasional long stalls; does not combine CPU, storage and browser results |
| Median across normalized workloads | Typical relative performance across a fixed test set | Domains with many tests gain more influence; changing references can change rankings |
| Median across domain summaries | Typical relative performance across explicitly chosen domains | Domain definitions matter; poor performance in a minority of domains can disappear from the headline |
NIST defines the sample median as the middle observation, or the arithmetic average of the middle two for an even sample count. Its resistance to extremes is useful, but it is a measure of location, not completeness, reliability, or worst-case response. [NIST location][nist-location]
**Recommendation for discussion:** use medians to summarize repeated valid measurements, normalize against frozen references, form predefined domain summaries, then use a median across those domains for the requested headline. Preserve each stage and its raw inputs. This proposes an aggregation structure; the domain membership, weighting, reference values, repeat counts and numeric scale remain decisions.
Established suites demonstrate why the levels must stay explicit. SPEC CPU 2017 takes median execution times from three runs, or the slower of two, then uses a **geometric mean** across ratios. Speedometer uses inverse geometric means across test durations and arithmetic means across iterations. These are methodological precedents, not permission to substitute a geometric mean for Odin's requested median. Preserve a tool's native result under its original name; label Odin's further aggregation separately. [SPEC rules][spec-rules] [Speedometer methodology][speedometer]
## Normalization and category balance
Milliseconds, operations/second and GB/s cannot share a meaningful raw median. A candidate approach is a dimensionless ratio against the **same workload's** reference: observation/reference for positive higher-is-better measures, reference/observation for positive lower-is-better measures. SPEC uses the latter for elapsed-time ratios. The metric identity must include its units, workload, size, concurrency, timing boundaries and direction. Reject invalid/nonfinite inputs under declared validity rules; a zero measured duration must not become an infinite score. [SPEC overview][spec-overview]
**Fictional arithmetic examples throughout this report:** these numbers illustrate consequences only. They are not Odin calibration or measured hardware results. A displayed index of 100 at the reference is an arbitrary illustrative scale, not a recommendation.
| Fictional metric | Reference | Observed | Illustrative ratio / index |
|---|---:|---:|---:|
| Work throughput | 25 operations/s | 50 operations/s | 2.0 / 200 |
| Memory bandwidth | 25 GB/s | 20 GB/s | 0.8 / 80 |
| Request latency | 4 ms | 2 ms | 2.0 / 200 |
A median of these indices is 200, despite memory bandwidth being below reference. That is a consequence of the chosen statistic. Show domain detail and slow-tail measurements alongside it. “200 versus 100” describes this index; it does not establish that every application is twice as fast.
Fix the order of operations. For fictional repeated times `[1, 3]` ms and a 3 ms reference, normalizing the raw median gives `3 / 2 = 1.5`; taking the median of individual ratios `[3, 1]` gives 2. Even-count averaging and reciprocal normalization do not commute. The report format must specify which result it contains. [median definition][nist-location]
Balance domains before counting metrics. If ten CPU tests each score 160 and four other domains score 40, 80, 100 and 120, the flat fourteen-test median is 160. A median across five domain summaries is 100. Adding CPU subtests should not silently redefine the product's priorities. Similarly, adding Python/Rust/Java variants must not automatically multiply the language domain's influence. A median of domain medians is a deliberate hierarchical index, not the pooled median of all observations.
Exclude health counters, memory-test pass/fail, driver availability and installed RAM capacity from throughput arithmetic. Multiple correlated outputs from one workload—throughput, IOPS, average latency and several percentiles—also need an explicit selection rule before any becomes an independent scored contribution.
## Calibration is a substantive decision
SPEC establishes per-workload reference times on a named machine and publishes the calculation rules. Its documentation explains that reference changes preserve relative overall rankings for its geometric-mean calculation. **That invariance does not generally hold for a median across normalized metrics.** [SPEC reference explanation][spec-overview]
For three fictional higher-is-better workloads, let machine A produce `[1, 10, 10]` and B produce `[2, 2, 20]` in each workload's own units:
| Fictional reference vector | A's normalized results → median | B's normalized results → median | Ordering |
|---|---|---|---|
| `[1, 1, 1]` | `[1, 10, 10]` → 10 | `[2, 2, 20]` → 2 | A higher |
| `[1, 10, 10]` | `[1, 1, 1]` → 1 | `[2, 0.2, 2]` → 2 | B higher |
The measurements did not change. Changing the reference altered the relative scales and which observations occupied the middle. Retain the requested median, make the reference rationale public, and version reference changes rather than treating them as cosmetic rescaling.
| Reference option | What it supports | Decision cost |
|---|---|---|
| Named reference configuration | Auditable, fixed anchor with documented per-test measurements | One machine's balance influences the median; configurations and repeatability need validation |
| Frozen reference cohort | Per-test references from a documented collection of machines | Cohort selection, sampling bias and revision policy become part of the score |
| User's own baseline | Local before/after comparisons | A score relative to oneself cannot rank different machines |
Recommend evaluating a frozen, locally distributable calibration manifest. Include reference measurements and provenance, reference conditions, workload and artifact digests, units/directions, aggregation order, required domains and calibration identity. Store it with results so calculation remains reproducible offline. Raw observations must survive changes; a recalculated score should identify its new calibration and preserve the original.
An arbitrary scale factor is acceptable if described as an index. A claim such as “100 is the median Linux machine,” a percentile rank, or a universal poor/good threshold requires representative population evidence that does not exist yet. Separately normalizing each architecture to its own average would also prevent interpreting those numbers as one common cross-architecture scale.
## Repetitions, warmup and uncertainty
Google Benchmark documents warmup, repetitions, median, standard deviation and coefficient of variation; it distinguishes user-visible wall time from CPU consumption. NIST recommends examining ordered observations for changing location/spread and says potential outliers should not simply be deleted when their cause is unknown. These support retaining all repeated observations and their conditions. [Google Benchmark][google-guide] [NIST run sequence][nist-runseq] [NIST outliers][nist-outliers]
Recommended measurement rules:
- Define warmup separately for each workload and retain its duration. Warm caches/JIT throughput, cold launch time and sustained thermal performance are different questions. Do not discard a slow first run from a declared cold-start test.
- Fix repetition and stopping rules before observing scores. A quick run may estimate a median without enough evidence for a useful confidence interval; it should not claim the precision of the standard profile. Calibrate the minimum repeats against the 10–20 minute budget.
- Retain run order, warmup, elapsed time, temperatures/power context, competing activity and invalidation reasons. A drifting sequence is not interchangeable independent noise. Repetitions within one process or thermal episode are not automatically independent runs.
- Exclude observations only for declared validity failures such as incorrect output, changed workload, cancellation or protocol failure. Preserve them with reasons. A slow but valid run can represent the usability problem Odin is meant to reveal.
- Show central spread such as MAD or IQR, plus tails where the workload supplies enough events. NIST defines MAD and IQR as distinct measures of spread; neither is itself a confidence interval. A median across repeated p99 values must not be labelled the p99 of all requests. [NIST scale][nist-scale]
NIST documents median confidence intervals based on order statistics/binomial probabilities, interpolated methods and bootstrap alternatives. Choose and validate a median-appropriate method; do not apply a mean's standard-error formula to a median. Confidence also depends on sample count and assumptions about the measurements. For an aggregate, account for shared run-level variation and state whether uncertainty in the calibration reference is included. The spread **between different domain scores** is not sampling uncertainty about the headline. [NIST median intervals][nist-median-ci]
Keep a graph of results in time order. Google documents CPU selection, boost, scheduler contention, SMT, caches and NUMA as variance sources. Its suggestions for controlled laboratory microbenchmarks include changing system settings; Odin's installed-system baseline should record existing conditions and label any tuned experiment separately. The existing [CPU/memory report][odin-cpu] and [portability/UI report][odin-portability] explain TUI interference and qualification needs. Stable repeated numbers alone do not prove that a workload represents real usability. [Google variance][google-variance]
## Comparability must be attached to every score
SPEC requires performance-relevant observation conditions and valid workload outputs; its CPU suite intentionally measures processor, memory subsystem **and compilers**. Even a fixed-toolchain comparison describes a defined software/hardware configuration. [SPEC rules][spec-rules] [SPEC overview][spec-overview]
Recommend two explicit comparison purposes:
- **Installed-system usability:** the chosen installed browser, shell, runtimes, drivers and kernel are part of what is measured. Version/configuration changes may explain a score change without any hardware change.
- **Controlled reference workload:** fixed workload assets, runtime/compiler contracts, flags, input data and execution modes improve comparison across machines. Architecture-specific artifacts must implement equivalent declared work and validate outputs; different ISA policies need disclosure.
Neither mode needs to masquerade as a pure hardware measurement. Keep their result identities distinct. Kernel/libc/distro differences can be the subject of a comparison, but they must be visible and the workload contract must remain equivalent.
A comparison identity should include suite/scoring/calibration versions, workload set, run profile, tool/artifact versions, options and data digests, timing/aggregation rules, browser mode, hardware/virtualization context and validity/coverage. Require a documented equivalence decision before combining scores across changed tools or workloads. SPEC warns that scores across different suite generations generally cannot be converted. [SPEC overview][spec-overview]
Speedometer 3.1 instructs users to use a clean browser profile, close competing programs/tabs, keep its page focused, avoid device interaction, use AC power and allow cooling when needed. Its official UI computes a 95% interval around its **arithmetic mean**; that interval cannot be attached to Odin's median unchanged. Preserve native browser score/uncertainty and label any median of complete runs separately. Headed and headless measurements need separate identities until an equivalence study justifies any shared interpretation; background versus foreground execution is also material. [instructions][speedometer-instructions] [3.1 result code][speedometer-main]
VM results characterize the guest allocation and virtualization environment. Keep native, virtualized and emulated cohorts identifiable; VM compatibility success does not establish native performance. Storage cache mode, queue depth, engine, filesystem and durability policy similarly belong to the workload identity. A fallback such as buffered I/O cannot silently replace a direct-I/O measurement with the same scoring identity. [CPU/memory report][odin-cpu] [storage report][odin-storage] [portability report][odin-portability]
## Missing tests and eligibility
For fictional domain indices `[40, 80, 100, 120, 160]`, the complete median is 100. Omitting 40 produces 110; omitting both 40 and 80 produces 120. Available-only aggregation can reward absent or deliberately skipped weak components.
**Recommendation:** define a versioned required set for the full score. Permit a clearly named partial median and domain results when the full set is unavailable, with the exact subset identified. Compare partial scores only over the same compatible subset; a pairwise intersection comparison must recompute **both** results and label that narrower scope. An optional pack must not silently change the headline's membership.
Keep distinct outcomes: completed-valid, completed-with-limitations, unsupported, missing dependency, permission denied, unsafe to run, cancelled, timed out, and failed validation. The eventual validity contract decides whether a limited result remains score-eligible. Never impute missing results as zero, a reference score, or a healthy pass. A required workload failing verification makes the full score ineligible, while preserving completed measurements and the associated finding. Good numbers from other domains should not cancel that failure.
The user accepted reporting unavailable tests across Linux targets. That does not resolve which domains are mandatory, whether every machine should still display a partial number, or how partial results should look. Those are explicit product decisions.
## Presentation, recommendations and open decisions
Recommend a headline containing the median, score identity, full/partial state, and eligible-domain coverage. The next view should show domain values, raw units, repeat count/spread, tail latency, invalidations and reference details. Show health findings beside performance: a fast drive with serious SMART evidence still needs attention. Missing telemetry must stay unknown. The storage and CPU reports establish why speed cannot determine drive replacement or certify memory health. [storage][odin-storage] [CPU/memory][odin-cpu]
Optimization advice should cite the observation and matching rule: for example, measured foreground stalls plus pressure evidence can support investigating memory contention. A low normalized score alone does not identify its cause. Keep severity of health evidence, completeness of coverage, measurement uncertainty and performance position as separate concepts; avoid one synthetic “confidence/health” percentage that mixes them.
Before implementation, decide:
1. The headline's median level, domain membership and balancing rules; whether responsiveness contributes or remains an accompanying measurement.
2. The reference configuration/cohort, scale and calibration-release policy; collect actual calibration data before inventing thresholds.
3. Required versus optional coverage, partial-score display and exact comparison eligibility.
4. Installed-system versus controlled-workload defaults, architecture/ISA policies, browser modes and VM cohorts.
5. Repetition/warmup/stopping rules, outlier validity rules, median interval method and honest quick/standard/extended precision claims.
6. Evidence requirements for optimization rules and separation of urgent health findings from the score.
**Evidence limits:** no calibration population, repeatability measurements, TUI-overhead budget or headed/headless equivalence study was produced. Examples are arithmetic demonstrations only. Context7 successfully resolved Google Benchmark; two BrowserBench/Speedometer searches returned unrelated packages, so its official repository and deployed 3.1 documentation were inspected directly. NIST and SPEC sources were inspected directly as statistical and benchmark-methodology references. This report neither adopts SPEC's workloads nor claims that their aggregation rules are Odin's final design.
[nist-location]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda351.htm
[nist-scale]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda356.htm
[nist-outliers]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda35h.htm
[nist-runseq]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda33p.htm
[nist-median-ci]: https://www.itl.nist.gov/div898/software/dataplot/refman1/auxillar/mediancl.htm
[spec-rules]: https://www.spec.org/cpu2017/Docs/runrules.html
[spec-overview]: https://www.spec.org/cpu2017/Docs/overview.html
[google-guide]: https://github.com/google/benchmark/blob/main/docs/user_guide.md
[google-variance]: https://github.com/google/benchmark/blob/main/docs/reducing_variance.md
[speedometer]: https://github.com/WebKit/Speedometer/blob/main/README.md
[speedometer-instructions]: https://browserbench.org/Speedometer3.1/instructions.html
[speedometer-main]: https://browserbench.org/Speedometer3.1/resources/main.mjs
[odin-cpu]: https://git.bongbetic.com/xavierk/odin/src/commit/90213d7f62cbd118f08ea8ff2f8042e94aa038a7/docs/research/cpu-memory.md
[odin-storage]: https://git.bongbetic.com/xavierk/odin/src/commit/6e5af87a64faedd4a8ad31ba10d9be4b499e8349/docs/research/storage-health.md
[odin-portability]: https://git.bongbetic.com/xavierk/odin/src/commit/20681cd4f184a9fc0164dd638a252de44ef230f5/docs/research/portability-ui.md