Establish browser, shell, and language workload validity #5

Closed
opened 2026-09-25 17:48:48 +00:00 by xavierk · 1 comment
Owner

Part of Find the way to Odin’s build-ready specification.

Question

Which browser, shell, and programming-language workloads can produce useful and reproducible Odin measurements in quick, standard, and extended runs?

Evaluate official Speedometer 3/BrowserBench-family candidates, licenses and redistribution/offline hosting, current versions, browser automation, headed versus headless comparability, GPU acceleration/environment metadata and startup versus steady-state behavior. Define meaningful Bash/shell latency and throughput candidates without confusing external command performance with shell performance. For Python, Rust, C++, and Java investigate maintained representative workloads, compiler/interpreter/JIT versions, optimization flags, warmup, compile versus execution time, installed versus reference toolchains, startup, and cross-architecture fairness. Avoid unsupported intrinsic language rankings. Identify runtime budgets and skipped-tool semantics.

Deliver primary-source comparisons and recommendations, leaving final selections to the human decision ticket. Do not install toolchains or run benchmarks.

Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). <!-- wayfinder-map: 1 --> ## Question Which browser, shell, and programming-language workloads can produce useful and reproducible Odin measurements in quick, standard, and extended runs? Evaluate official Speedometer 3/BrowserBench-family candidates, licenses and redistribution/offline hosting, current versions, browser automation, headed versus headless comparability, GPU acceleration/environment metadata and startup versus steady-state behavior. Define meaningful Bash/shell latency and throughput candidates without confusing external command performance with shell performance. For Python, Rust, C++, and Java investigate maintained representative workloads, compiler/interpreter/JIT versions, optimization flags, warmup, compile versus execution time, installed versus reference toolchains, startup, and cross-architecture fairness. Avoid unsupported intrinsic language rankings. Identify runtime budgets and skipped-tool semantics. Deliver primary-source comparisons and recommendations, leaving final selections to the human decision ticket. Do not install toolchains or run benchmarks.
xavierk added the wayfinder:research label 2026-09-25 17:48:48 +00:00
xavierk added a new dependency 2026-09-25 17:49:49 +00:00
xavierk self-assigned this 2026-09-25 18:01:32 +00:00
Author
Owner

Research resolution

The candidate analysis separates browser responsiveness, shell interaction, process startup, compilation, and warmed execution. Results describe a workload/toolchain/environment, not an intrinsic ranking of programming languages.

  • Speedometer 3.1 is a strong browser-responsiveness candidate; JetStream 3.0 and MotionMark 1.3.2 add different engine/graphics views. Preserve official workload settings, native statistics and raw results; an Odin median across complete runs is separately labeled.
  • Release/commit/assets identify a workload more reliably than package metadata alone. Prepared local hosting is possible, but full offline resource closure and third-party license inventory require qualification of the selected pack.
  • Native installed browsers with matching WebDriver are one route; pinned Playwright browsers are another supported-platform cohort. Playwright’s bundled browsers are not a universal musl solution. Use a qualified native pair or an explicit capability skip, not a silent substitute environment.
  • Headed, headless and headless-shell results need distinct environment identities. Shared browser code does not establish equal scores or compliance with focused-window run conditions.
  • Bash startup, true PTY prompt readiness, builtin throughput, and external-command pipelines are separate workloads. Hyperfine’s default shell-spawn subtraction is unsuitable for measuring that same startup overhead; user startup-file execution must be an explicit local-usability mode.
  • Language candidates include pyperformance/pyperf, selected rustc-perf compile/runtime workloads, LLVM test-suite C++ subsets, and JMH-based or DaCapo Java workloads. Their maintenance, licensing, footprint, version and portability limitations are documented; JMH is a harness rather than a representative suite by itself.
  • Startup, compilation and warmed/JIT execution remain separate records. Installed versus reference toolchains, optimization flags, output correctness, warmup, repetitions, workload versions and failures/skips need declared rules.

Read the cited research report. Evidence is recorded at commit 80be5a7ba92e on research/application-workloads.

Still for the human decision tickets: exact default/extended workloads and budgets; required versus optional language packs; native versus reference browser/toolchain policy; automation route; shell profile scope; final scoring inclusion and comparability cohorts.

Evidence limits: no tools were installed and no benchmarks/builds were run. Exact artifacts still need offline-network, licensing, output-correctness, libc/architecture and runtime-budget qualification. Experimental Rust runtime workloads and documented Java-version limits must not be treated as universal language support.

<!-- wayfinder-research-resolution: application_workloads --> ## Research resolution The candidate analysis separates browser responsiveness, shell interaction, process startup, compilation, and warmed execution. Results describe a workload/toolchain/environment, not an intrinsic ranking of programming languages. - Speedometer 3.1 is a strong browser-responsiveness candidate; JetStream 3.0 and MotionMark 1.3.2 add different engine/graphics views. Preserve official workload settings, native statistics and raw results; an Odin median across complete runs is separately labeled. - Release/commit/assets identify a workload more reliably than package metadata alone. Prepared local hosting is possible, but full offline resource closure and third-party license inventory require qualification of the selected pack. - Native installed browsers with matching WebDriver are one route; pinned Playwright browsers are another supported-platform cohort. Playwright’s bundled browsers are not a universal musl solution. Use a qualified native pair or an explicit capability skip, not a silent substitute environment. - Headed, headless and headless-shell results need distinct environment identities. Shared browser code does not establish equal scores or compliance with focused-window run conditions. - Bash startup, true PTY prompt readiness, builtin throughput, and external-command pipelines are separate workloads. Hyperfine’s default shell-spawn subtraction is unsuitable for measuring that same startup overhead; user startup-file execution must be an explicit local-usability mode. - Language candidates include pyperformance/pyperf, selected rustc-perf compile/runtime workloads, LLVM test-suite C++ subsets, and JMH-based or DaCapo Java workloads. Their maintenance, licensing, footprint, version and portability limitations are documented; JMH is a harness rather than a representative suite by itself. - Startup, compilation and warmed/JIT execution remain separate records. Installed versus reference toolchains, optimization flags, output correctness, warmup, repetitions, workload versions and failures/skips need declared rules. [Read the cited research report](https://git.bongbetic.com/xavierk/odin/src/commit/80be5a7ba92ecdb9a303bd9cad42d2866742c5f9/docs/research/application-workloads.md). Evidence is recorded at commit `80be5a7ba92e` on `research/application-workloads`. **Still for the human decision tickets:** exact default/extended workloads and budgets; required versus optional language packs; native versus reference browser/toolchain policy; automation route; shell profile scope; final scoring inclusion and comparability cohorts. **Evidence limits:** no tools were installed and no benchmarks/builds were run. Exact artifacts still need offline-network, licensing, output-correctness, libc/architecture and runtime-budget qualification. Experimental Rust runtime workloads and documented Java-version limits must not be treated as universal language support.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/odin#5