Establish safe CPU, memory, and kernel measurements #2

Closed
opened 2026-09-25 17:48:45 +00:00 by xavierk · 1 comment
Owner

Part of Find the way to Odin’s build-ready specification.

Question

Which existing, maintained workloads and Linux telemetry can reliably characterize CPU throughput, latency/bottlenecks, memory pressure, and memory errors within Odin’s profiles?

Investigate CPU single/multicore and scheduler behavior, warmup, thermal/power/governor effects, PSI/cgroups/perf access and kernel configuration; assess existing options for foreground response latency under bounded CPU/memory contention, distinguishing usability under load from peak throughput and timer wakeup jitter; memory capacity/bandwidth/latency/swap/pressure; online memory-test coverage versus offline tests, ECC/EDAC visibility and limits. Compare established tools and their licenses/maintenance, output stability, resource caps, cancellation, OOM avoidance, permissions, architectures/libcs and fallback behavior. Distinguish stress, benchmarking, pressure observation, and error detection. Include what a VM can and cannot validate, and how live TUI overhead could contaminate measurements.

Deliver cited facts, a candidate matrix, recommended options and confidence limits. Do not run stress tests or select the final workload suite.

Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). <!-- wayfinder-map: 1 --> ## Question Which existing, maintained workloads and Linux telemetry can reliably characterize CPU throughput, latency/bottlenecks, memory pressure, and memory errors within Odin’s profiles? Investigate CPU single/multicore and scheduler behavior, warmup, thermal/power/governor effects, PSI/cgroups/perf access and kernel configuration; assess existing options for foreground response latency under bounded CPU/memory contention, distinguishing usability under load from peak throughput and timer wakeup jitter; memory capacity/bandwidth/latency/swap/pressure; online memory-test coverage versus offline tests, ECC/EDAC visibility and limits. Compare established tools and their licenses/maintenance, output stability, resource caps, cancellation, OOM avoidance, permissions, architectures/libcs and fallback behavior. Distinguish stress, benchmarking, pressure observation, and error detection. Include what a VM can and cannot validate, and how live TUI overhead could contaminate measurements. Deliver cited facts, a candidate matrix, recommended options and confidence limits. Do not run stress tests or select the final workload suite.
xavierk added the wayfinder:research label 2026-09-25 17:48:45 +00:00
xavierk added a new dependency 2026-09-25 17:49:47 +00:00
xavierk added a new dependency 2026-09-25 17:49:55 +00:00
xavierk self-assigned this 2026-09-25 17:50:13 +00:00
Author
Owner

Research resolution

The investigation separates throughput, responsiveness under load, pressure observations, and health findings. No single candidate establishes all four.

  • The conditional shortlist is sysbench for a narrow CPU baseline, STREAM for memory bandwidth, schbench for responsiveness qualification, selected stress-ng workers for controlled load, and memtester plus EDAC/RAS for diagnostics.
  • Schbench measures synthetic request and wakeup latency under configurable load. It can help compare idle versus contended response, but does not itself measure terminal or browser interaction.
  • stress-ng explicitly disclaims precise benchmarking; its bogo-operation values should not become Odin’s general performance score. STREAM’s internal best-iteration statistic must remain distinguishable from any median across repeated runs.
  • PSI, /proc, and cgroup telemetry need capability and permission detection. A missing metric, or a system-wide CPU PSI full value of zero, cannot establish healthy responsiveness.
  • Safe pressure work requires explicit worker/allocation budgets, retained headroom, applicable memory/swap controls, and supervised cancellation. memory.high is not a hard cap, and a userspace timeout cannot guarantee immediate termination of a task blocked in the kernel.
  • Online memtester only checks its actual allocation and can continue without locking while still exiting successfully. Allocation, locking, iterations, and detected errors must accompany any result. EDAC absence or zero observed deltas cannot certify error-free RAM.
  • Stable Memtest86+ and development-main AArch64 support differ; offline tests cannot be presented as an ordinary in-terminal stage. VM success does not certify physical DIMMs or cooling.
  • Kernel, topology, allowed CPUs, NUMA/THP, compiler flags, governor/power/thermal context, and TUI activity affect the meaning of a measurement. The report proposes qualification of foreground latency and quiet-versus-live UI overhead.

Read the cited research report. Evidence is recorded at commit 90213d7f62cb on research/cpu-memory.

Still for the human decision tickets: exact foreground workloads and background intensity; score inclusion; resource/headroom/privilege policies; repetition and time allocation; heterogeneous-core behavior; missing coverage; offline diagnostic scope.

Evidence limits: no candidates were built or run on the target matrix and no overhead or repeatability budget is yet calibrated. Context7 lacked suitable stress-ng/STREAM/schbench matches, so owning sources were inspected directly. Memtester’s upstream TLS failed; original sources preserved by Debian and the Void package archive were used with version differences recorded. No benchmarks or host changes were performed.

<!-- wayfinder-research-resolution: cpu_memory --> ## Research resolution The investigation separates **throughput, responsiveness under load, pressure observations, and health findings**. No single candidate establishes all four. - The conditional shortlist is sysbench for a narrow CPU baseline, STREAM for memory bandwidth, schbench for responsiveness qualification, selected stress-ng workers for controlled load, and memtester plus EDAC/RAS for diagnostics. - Schbench measures synthetic request and wakeup latency under configurable load. It can help compare idle versus contended response, but does not itself measure terminal or browser interaction. - stress-ng explicitly disclaims precise benchmarking; its bogo-operation values should not become Odin’s general performance score. STREAM’s internal best-iteration statistic must remain distinguishable from any median across repeated runs. - PSI, `/proc`, and cgroup telemetry need capability and permission detection. A missing metric, or a system-wide CPU PSI full value of zero, cannot establish healthy responsiveness. - Safe pressure work requires explicit worker/allocation budgets, retained headroom, applicable memory/swap controls, and supervised cancellation. `memory.high` is not a hard cap, and a userspace timeout cannot guarantee immediate termination of a task blocked in the kernel. - Online memtester only checks its actual allocation and can continue without locking while still exiting successfully. Allocation, locking, iterations, and detected errors must accompany any result. EDAC absence or zero observed deltas cannot certify error-free RAM. - Stable Memtest86+ and development-main AArch64 support differ; offline tests cannot be presented as an ordinary in-terminal stage. VM success does not certify physical DIMMs or cooling. - Kernel, topology, allowed CPUs, NUMA/THP, compiler flags, governor/power/thermal context, and TUI activity affect the meaning of a measurement. The report proposes qualification of foreground latency and quiet-versus-live UI overhead. [Read the cited research report](https://git.bongbetic.com/xavierk/odin/src/commit/90213d7f62cbd118f08ea8ff2f8042e94aa038a7/docs/research/cpu-memory.md). Evidence is recorded at commit `90213d7f62cb` on `research/cpu-memory`. **Still for the human decision tickets:** exact foreground workloads and background intensity; score inclusion; resource/headroom/privilege policies; repetition and time allocation; heterogeneous-core behavior; missing coverage; offline diagnostic scope. **Evidence limits:** no candidates were built or run on the target matrix and no overhead or repeatability budget is yet calibrated. Context7 lacked suitable stress-ng/STREAM/schbench matches, so owning sources were inspected directly. Memtester’s upstream TLS failed; original sources preserved by Debian and the Void package archive were used with version differences recorded. No benchmarks or host changes were performed.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/odin#2