Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f384c81f0f |
@@ -0,0 +1,140 @@
|
||||
# Defensible median scoring and comparison rules
|
||||
|
||||
Research for [Establish defensible median scoring and comparison rules](https://git.bongbetic.com/xavierk/odin/issues/7), part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1).
|
||||
|
||||
**Accessed:** 25 September 2026. **Status:** decision evidence and recommendations; no final scoring formula, calibrated reference values, benchmark runs, or implementation. The user's requirement is a **median** score.
|
||||
|
||||
## What the median should mean
|
||||
|
||||
There are three separate choices:
|
||||
|
||||
| Level | Meaning | Main limitation |
|
||||
|---|---|---|
|
||||
| Median of repeated measurements | Typical result for one fixed workload under stated conditions | Hides occasional long stalls; does not combine CPU, storage and browser results |
|
||||
| Median across normalized workloads | Typical relative performance across a fixed test set | Domains with many tests gain more influence; changing references can change rankings |
|
||||
| Median across domain summaries | Typical relative performance across explicitly chosen domains | Domain definitions matter; poor performance in a minority of domains can disappear from the headline |
|
||||
|
||||
NIST defines the sample median as the middle observation, or the arithmetic average of the middle two for an even sample count. Its resistance to extremes is useful, but it is a measure of location, not completeness, reliability, or worst-case response. [NIST location][nist-location]
|
||||
|
||||
**Recommendation for discussion:** use medians to summarize repeated valid measurements, normalize against frozen references, form predefined domain summaries, then use a median across those domains for the requested headline. Preserve each stage and its raw inputs. This proposes an aggregation structure; the domain membership, weighting, reference values, repeat counts and numeric scale remain decisions.
|
||||
|
||||
Established suites demonstrate why the levels must stay explicit. SPEC CPU 2017 takes median execution times from three runs, or the slower of two, then uses a **geometric mean** across ratios. Speedometer uses inverse geometric means across test durations and arithmetic means across iterations. These are methodological precedents, not permission to substitute a geometric mean for Odin's requested median. Preserve a tool's native result under its original name; label Odin's further aggregation separately. [SPEC rules][spec-rules] [Speedometer methodology][speedometer]
|
||||
|
||||
## Normalization and category balance
|
||||
|
||||
Milliseconds, operations/second and GB/s cannot share a meaningful raw median. A candidate approach is a dimensionless ratio against the **same workload's** reference: observation/reference for positive higher-is-better measures, reference/observation for positive lower-is-better measures. SPEC uses the latter for elapsed-time ratios. The metric identity must include its units, workload, size, concurrency, timing boundaries and direction. Reject invalid/nonfinite inputs under declared validity rules; a zero measured duration must not become an infinite score. [SPEC overview][spec-overview]
|
||||
|
||||
**Fictional arithmetic examples throughout this report:** these numbers illustrate consequences only. They are not Odin calibration or measured hardware results. A displayed index of 100 at the reference is an arbitrary illustrative scale, not a recommendation.
|
||||
|
||||
| Fictional metric | Reference | Observed | Illustrative ratio / index |
|
||||
|---|---:|---:|---:|
|
||||
| Work throughput | 25 operations/s | 50 operations/s | 2.0 / 200 |
|
||||
| Memory bandwidth | 25 GB/s | 20 GB/s | 0.8 / 80 |
|
||||
| Request latency | 4 ms | 2 ms | 2.0 / 200 |
|
||||
|
||||
A median of these indices is 200, despite memory bandwidth being below reference. That is a consequence of the chosen statistic. Show domain detail and slow-tail measurements alongside it. “200 versus 100” describes this index; it does not establish that every application is twice as fast.
|
||||
|
||||
Fix the order of operations. For fictional repeated times `[1, 3]` ms and a 3 ms reference, normalizing the raw median gives `3 / 2 = 1.5`; taking the median of individual ratios `[3, 1]` gives 2. Even-count averaging and reciprocal normalization do not commute. The report format must specify which result it contains. [median definition][nist-location]
|
||||
|
||||
Balance domains before counting metrics. If ten CPU tests each score 160 and four other domains score 40, 80, 100 and 120, the flat fourteen-test median is 160. A median across five domain summaries is 100. Adding CPU subtests should not silently redefine the product's priorities. Similarly, adding Python/Rust/Java variants must not automatically multiply the language domain's influence. A median of domain medians is a deliberate hierarchical index, not the pooled median of all observations.
|
||||
|
||||
Exclude health counters, memory-test pass/fail, driver availability and installed RAM capacity from throughput arithmetic. Multiple correlated outputs from one workload—throughput, IOPS, average latency and several percentiles—also need an explicit selection rule before any becomes an independent scored contribution.
|
||||
|
||||
## Calibration is a substantive decision
|
||||
|
||||
SPEC establishes per-workload reference times on a named machine and publishes the calculation rules. Its documentation explains that reference changes preserve relative overall rankings for its geometric-mean calculation. **That invariance does not generally hold for a median across normalized metrics.** [SPEC reference explanation][spec-overview]
|
||||
|
||||
For three fictional higher-is-better workloads, let machine A produce `[1, 10, 10]` and B produce `[2, 2, 20]` in each workload's own units:
|
||||
|
||||
| Fictional reference vector | A's normalized results → median | B's normalized results → median | Ordering |
|
||||
|---|---|---|---|
|
||||
| `[1, 1, 1]` | `[1, 10, 10]` → 10 | `[2, 2, 20]` → 2 | A higher |
|
||||
| `[1, 10, 10]` | `[1, 1, 1]` → 1 | `[2, 0.2, 2]` → 2 | B higher |
|
||||
|
||||
The measurements did not change. Changing the reference altered the relative scales and which observations occupied the middle. Retain the requested median, make the reference rationale public, and version reference changes rather than treating them as cosmetic rescaling.
|
||||
|
||||
| Reference option | What it supports | Decision cost |
|
||||
|---|---|---|
|
||||
| Named reference configuration | Auditable, fixed anchor with documented per-test measurements | One machine's balance influences the median; configurations and repeatability need validation |
|
||||
| Frozen reference cohort | Per-test references from a documented collection of machines | Cohort selection, sampling bias and revision policy become part of the score |
|
||||
| User's own baseline | Local before/after comparisons | A score relative to oneself cannot rank different machines |
|
||||
|
||||
Recommend evaluating a frozen, locally distributable calibration manifest. Include reference measurements and provenance, reference conditions, workload and artifact digests, units/directions, aggregation order, required domains and calibration identity. Store it with results so calculation remains reproducible offline. Raw observations must survive changes; a recalculated score should identify its new calibration and preserve the original.
|
||||
|
||||
An arbitrary scale factor is acceptable if described as an index. A claim such as “100 is the median Linux machine,” a percentile rank, or a universal poor/good threshold requires representative population evidence that does not exist yet. Separately normalizing each architecture to its own average would also prevent interpreting those numbers as one common cross-architecture scale.
|
||||
|
||||
## Repetitions, warmup and uncertainty
|
||||
|
||||
Google Benchmark documents warmup, repetitions, median, standard deviation and coefficient of variation; it distinguishes user-visible wall time from CPU consumption. NIST recommends examining ordered observations for changing location/spread and says potential outliers should not simply be deleted when their cause is unknown. These support retaining all repeated observations and their conditions. [Google Benchmark][google-guide] [NIST run sequence][nist-runseq] [NIST outliers][nist-outliers]
|
||||
|
||||
Recommended measurement rules:
|
||||
|
||||
- Define warmup separately for each workload and retain its duration. Warm caches/JIT throughput, cold launch time and sustained thermal performance are different questions. Do not discard a slow first run from a declared cold-start test.
|
||||
- Fix repetition and stopping rules before observing scores. A quick run may estimate a median without enough evidence for a useful confidence interval; it should not claim the precision of the standard profile. Calibrate the minimum repeats against the 10–20 minute budget.
|
||||
- Retain run order, warmup, elapsed time, temperatures/power context, competing activity and invalidation reasons. A drifting sequence is not interchangeable independent noise. Repetitions within one process or thermal episode are not automatically independent runs.
|
||||
- Exclude observations only for declared validity failures such as incorrect output, changed workload, cancellation or protocol failure. Preserve them with reasons. A slow but valid run can represent the usability problem Odin is meant to reveal.
|
||||
- Show central spread such as MAD or IQR, plus tails where the workload supplies enough events. NIST defines MAD and IQR as distinct measures of spread; neither is itself a confidence interval. A median across repeated p99 values must not be labelled the p99 of all requests. [NIST scale][nist-scale]
|
||||
|
||||
NIST documents median confidence intervals based on order statistics/binomial probabilities, interpolated methods and bootstrap alternatives. Choose and validate a median-appropriate method; do not apply a mean's standard-error formula to a median. Confidence also depends on sample count and assumptions about the measurements. For an aggregate, account for shared run-level variation and state whether uncertainty in the calibration reference is included. The spread **between different domain scores** is not sampling uncertainty about the headline. [NIST median intervals][nist-median-ci]
|
||||
|
||||
Keep a graph of results in time order. Google documents CPU selection, boost, scheduler contention, SMT, caches and NUMA as variance sources. Its suggestions for controlled laboratory microbenchmarks include changing system settings; Odin's installed-system baseline should record existing conditions and label any tuned experiment separately. The existing [CPU/memory report][odin-cpu] and [portability/UI report][odin-portability] explain TUI interference and qualification needs. Stable repeated numbers alone do not prove that a workload represents real usability. [Google variance][google-variance]
|
||||
|
||||
## Comparability must be attached to every score
|
||||
|
||||
SPEC requires performance-relevant observation conditions and valid workload outputs; its CPU suite intentionally measures processor, memory subsystem **and compilers**. Even a fixed-toolchain comparison describes a defined software/hardware configuration. [SPEC rules][spec-rules] [SPEC overview][spec-overview]
|
||||
|
||||
Recommend two explicit comparison purposes:
|
||||
|
||||
- **Installed-system usability:** the chosen installed browser, shell, runtimes, drivers and kernel are part of what is measured. Version/configuration changes may explain a score change without any hardware change.
|
||||
- **Controlled reference workload:** fixed workload assets, runtime/compiler contracts, flags, input data and execution modes improve comparison across machines. Architecture-specific artifacts must implement equivalent declared work and validate outputs; different ISA policies need disclosure.
|
||||
|
||||
Neither mode needs to masquerade as a pure hardware measurement. Keep their result identities distinct. Kernel/libc/distro differences can be the subject of a comparison, but they must be visible and the workload contract must remain equivalent.
|
||||
|
||||
A comparison identity should include suite/scoring/calibration versions, workload set, run profile, tool/artifact versions, options and data digests, timing/aggregation rules, browser mode, hardware/virtualization context and validity/coverage. Require a documented equivalence decision before combining scores across changed tools or workloads. SPEC warns that scores across different suite generations generally cannot be converted. [SPEC overview][spec-overview]
|
||||
|
||||
Speedometer 3.1 instructs users to use a clean browser profile, close competing programs/tabs, keep its page focused, avoid device interaction, use AC power and allow cooling when needed. Its official UI computes a 95% interval around its **arithmetic mean**; that interval cannot be attached to Odin's median unchanged. Preserve native browser score/uncertainty and label any median of complete runs separately. Headed and headless measurements need separate identities until an equivalence study justifies any shared interpretation; background versus foreground execution is also material. [instructions][speedometer-instructions] [3.1 result code][speedometer-main]
|
||||
|
||||
VM results characterize the guest allocation and virtualization environment. Keep native, virtualized and emulated cohorts identifiable; VM compatibility success does not establish native performance. Storage cache mode, queue depth, engine, filesystem and durability policy similarly belong to the workload identity. A fallback such as buffered I/O cannot silently replace a direct-I/O measurement with the same scoring identity. [CPU/memory report][odin-cpu] [storage report][odin-storage] [portability report][odin-portability]
|
||||
|
||||
## Missing tests and eligibility
|
||||
|
||||
For fictional domain indices `[40, 80, 100, 120, 160]`, the complete median is 100. Omitting 40 produces 110; omitting both 40 and 80 produces 120. Available-only aggregation can reward absent or deliberately skipped weak components.
|
||||
|
||||
**Recommendation:** define a versioned required set for the full score. Permit a clearly named partial median and domain results when the full set is unavailable, with the exact subset identified. Compare partial scores only over the same compatible subset; a pairwise intersection comparison must recompute **both** results and label that narrower scope. An optional pack must not silently change the headline's membership.
|
||||
|
||||
Keep distinct outcomes: completed-valid, completed-with-limitations, unsupported, missing dependency, permission denied, unsafe to run, cancelled, timed out, and failed validation. The eventual validity contract decides whether a limited result remains score-eligible. Never impute missing results as zero, a reference score, or a healthy pass. A required workload failing verification makes the full score ineligible, while preserving completed measurements and the associated finding. Good numbers from other domains should not cancel that failure.
|
||||
|
||||
The user accepted reporting unavailable tests across Linux targets. That does not resolve which domains are mandatory, whether every machine should still display a partial number, or how partial results should look. Those are explicit product decisions.
|
||||
|
||||
## Presentation, recommendations and open decisions
|
||||
|
||||
Recommend a headline containing the median, score identity, full/partial state, and eligible-domain coverage. The next view should show domain values, raw units, repeat count/spread, tail latency, invalidations and reference details. Show health findings beside performance: a fast drive with serious SMART evidence still needs attention. Missing telemetry must stay unknown. The storage and CPU reports establish why speed cannot determine drive replacement or certify memory health. [storage][odin-storage] [CPU/memory][odin-cpu]
|
||||
|
||||
Optimization advice should cite the observation and matching rule: for example, measured foreground stalls plus pressure evidence can support investigating memory contention. A low normalized score alone does not identify its cause. Keep severity of health evidence, completeness of coverage, measurement uncertainty and performance position as separate concepts; avoid one synthetic “confidence/health” percentage that mixes them.
|
||||
|
||||
Before implementation, decide:
|
||||
|
||||
1. The headline's median level, domain membership and balancing rules; whether responsiveness contributes or remains an accompanying measurement.
|
||||
2. The reference configuration/cohort, scale and calibration-release policy; collect actual calibration data before inventing thresholds.
|
||||
3. Required versus optional coverage, partial-score display and exact comparison eligibility.
|
||||
4. Installed-system versus controlled-workload defaults, architecture/ISA policies, browser modes and VM cohorts.
|
||||
5. Repetition/warmup/stopping rules, outlier validity rules, median interval method and honest quick/standard/extended precision claims.
|
||||
6. Evidence requirements for optimization rules and separation of urgent health findings from the score.
|
||||
|
||||
**Evidence limits:** no calibration population, repeatability measurements, TUI-overhead budget or headed/headless equivalence study was produced. Examples are arithmetic demonstrations only. Context7 successfully resolved Google Benchmark; two BrowserBench/Speedometer searches returned unrelated packages, so its official repository and deployed 3.1 documentation were inspected directly. NIST and SPEC sources were inspected directly as statistical and benchmark-methodology references. This report neither adopts SPEC's workloads nor claims that their aggregation rules are Odin's final design.
|
||||
|
||||
[nist-location]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda351.htm
|
||||
[nist-scale]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda356.htm
|
||||
[nist-outliers]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda35h.htm
|
||||
[nist-runseq]: https://www.itl.nist.gov/div898/handbook/eda/section3/eda33p.htm
|
||||
[nist-median-ci]: https://www.itl.nist.gov/div898/software/dataplot/refman1/auxillar/mediancl.htm
|
||||
[spec-rules]: https://www.spec.org/cpu2017/Docs/runrules.html
|
||||
[spec-overview]: https://www.spec.org/cpu2017/Docs/overview.html
|
||||
[google-guide]: https://github.com/google/benchmark/blob/main/docs/user_guide.md
|
||||
[google-variance]: https://github.com/google/benchmark/blob/main/docs/reducing_variance.md
|
||||
[speedometer]: https://github.com/WebKit/Speedometer/blob/main/README.md
|
||||
[speedometer-instructions]: https://browserbench.org/Speedometer3.1/instructions.html
|
||||
[speedometer-main]: https://browserbench.org/Speedometer3.1/resources/main.mjs
|
||||
[odin-cpu]: https://git.bongbetic.com/xavierk/odin/src/commit/90213d7f62cbd118f08ea8ff2f8042e94aa038a7/docs/research/cpu-memory.md
|
||||
[odin-storage]: https://git.bongbetic.com/xavierk/odin/src/commit/6e5af87a64faedd4a8ad31ba10d9be4b499e8349/docs/research/storage-health.md
|
||||
[odin-portability]: https://git.bongbetic.com/xavierk/odin/src/commit/20681cd4f184a9fc0164dd638a252de44ef230f5/docs/research/portability-ui.md
|
||||
@@ -1,38 +0,0 @@
|
||||
# Speedometer 3.1 pinned pack: redistribution audit
|
||||
|
||||
## Decision
|
||||
|
||||
**Do not redistribute the complete pinned archive yet.** The top-level BSD-style license permits redistribution of Speedometer's own work when its notice, conditions, and disclaimer travel with it. It does not establish rights for every embedded work. The pinned archive contains identifiable third-party material whose grant or compliance path is not yet established. This is a source audit, not a legal opinion or a benchmark run.
|
||||
|
||||
Source: [WebKit/Speedometer commit `1386415be8fef2f6b6bbdbe1828872471c5d802a`](https://github.com/WebKit/Speedometer/tree/1386415be8fef2f6b6bbdbe1828872471c5d802a), [root license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE).
|
||||
|
||||
## Exact inventory
|
||||
|
||||
The companion [per-file manifest](speedometer-license-manifest.tsv) records every relative path, byte count, SHA-256, and **review group** in the full GitHub commit archive: **1,118 files; 61,022,359 uncompressed bytes**. Its SHA-256 is `78edcd45649128fbd4c8659c1d054a2b9bb76a184b58ce59b37895c19f2bdfb6`. The archive download SHA-256 was `cfefa818d7789ed2f2b9f3f1bf608cc35e4e8241489eef47049296ce92791313`. Review groups flag provenance work; they are **not license determinations**. The complete source archive is a conservative candidate pack, not a selected minimum runtime file set.
|
||||
|
||||
Reproduce the manifest from that commit's GitHub `.tar.gz`: strip its single leading directory; sort regular-file paths by UTF-8 path; for each file write `path`, decimal byte count, lowercase SHA-256, and review group as tab-separated fields. Group rules: exact notice names and `*.LICENSE.txt`/`3rdpartylicenses.txt` first; then Gutenberg HTML, chart datasets, all news-site files, Adobe icon SVGs, then TodoMVC, React Stockcharts, charts, editors, and remainder in that order. The manifest itself is not an upstream artifact.
|
||||
|
||||
Nine separate notice files appear in the archive: `LICENSE`; `resources/todomvc/license.md`; `resources/react-stockcharts/build/static/js/2.8e539c84.chunk.js.LICENSE.txt`; Angular and Angular Complex `dist/3rdpartylicenses.txt`; and React, React Complex, React Redux, React Redux Complex `dist/app.bundle.js.LICENSE.txt`. Exact paths and hashes are in the manifest. Inline notices in JavaScript, CSS, HTML, SVG, and source maps also need preservation. A filename scan cannot prove all component notices were extracted.
|
||||
|
||||
## Established obligations and specific gaps
|
||||
|
||||
| Material | Evidence and current finding | Required action before shipping |
|
||||
| --- | --- | --- |
|
||||
| Speedometer-owned source | [Root license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE) requires retention of copyright, conditions, and disclaimer for source; reproduction in documentation or other materials for binary form. | Include root `LICENSE` in pack and distribution materials; keep original headers. |
|
||||
| TodoMVC implementations | [TodoMVC subtree license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/license.md) says MIT **unless otherwise specified**. Generated React and Angular notices identify further components. | Carry subtree license and all bundled notices. Audit each runtime bundle against its dependency versions; fill missing texts/attributions. |
|
||||
| Angular bundles | Both [`3rdpartylicenses.txt` files](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/architecture-examples/angular/dist/3rdpartylicenses.txt) include MIT, Apache-2.0, and CC-BY-4.0 material. | Keep both files with their matching bundles; map each named component to bundle; satisfy Apache notice/change rules and CC attribution requirements as applicable. |
|
||||
| React Stockcharts | Its [README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/react-stockcharts/README.md) claims MIT and points to upstream; [bundle notice](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/react-stockcharts/build/static/js/2.8e539c84.chunk.js.LICENSE.txt) lists dependencies. | Include upstream MIT text plus bundled notice; verify the actual bundled versions. |
|
||||
| News-site template and CSS | Both [Next README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/newssite/news-next/README.md) and [Nuxt README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/newssite/news-nuxt/README.md) credit [`flashdesignory/news-site-template`](https://github.com/flashdesignory/news-site-template). Its repository has no visible license file, and [its package metadata](https://github.com/flashdesignory/news-site-template/blob/main/package.json) declares none. Pinned [news-site-css metadata](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/newssite/news-site-css/package.json) says ISC but does not provide that license text. The source and exported `dist` include adapted template material. **No redistribution grant established for the template.** | Get written permission or a verifiable applicable license from its rights holder, or remove/replace the NewsSite suites and all dependent files. Obtain and carry correct ISC text/notice for news-site-css. Reassess suite set and scoring if suites removed. |
|
||||
| Chart datasets | [Dataset README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/charts/datasets/README) identifies two CSVs copied from an older D3 path. It does not state their data source or rights. `airports.csv`, `flights-airports.csv`, and the README are the three manifest paths. Generated `resources/charts/dist/assets/flights-airports-9a9e6422.js` embeds data. | Trace dataset origin and grant, then include required attribution. Otherwise replace with licensed/synthetic data and rebuild the affected chart assets, or omit that workload. |
|
||||
| Adobe Spectrum icons and CSS | [Big DOM README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/big-dom-generator/README.md) says static shell uses Adobe `@spectrum-css`. Its source contains 22 `Smock_*.svg` icons plus three other SVGs, with no per-file notice. Generated `dist` and complex TodoMVC pages may embed this material. TodoMVC's MIT default alone cannot establish rights over Adobe assets. | Identify exact Adobe package/source and applicable icon/CSS license, preserve its notice, and verify built copies. If unavailable, replace assets and rebuild/verify affected pages. |
|
||||
| Editor text | [`longtext.html`](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/editors/assets/longtext.html) is Project Gutenberg eBook 2650 per [asset README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/editors/assets/README.md). It includes Gutenberg branding and full license. [Gutenberg terms](https://www.gutenberg.org/policy/license) impose access, notice, format, fee, and territorial conditions while branding remains. | Preserve full embedded license and prominent notice; assess distribution geography and price. Simpler path: replace with separately cleared text of equivalent workload shape. |
|
||||
| Other generated bundles and imagery | Charts, Editors, NewsSite, TodoMVC, React Stockcharts, images, maps, and CSS are generated or embedded works. Package lock entries identify dependencies but do not themselves provide all license texts; absence of a notice file is not proof of permission. | Build a component-to-file bill of materials from pinned locks/source maps and licenses. Review all shipped binaries/images individually; acquire missing grants or exclude/rebuild. |
|
||||
|
||||
## Pack gate
|
||||
|
||||
1. Define exact runtime file set; leave development files out only after proving every enabled suite resolves locally. Record each shipped file in a final manifest, with source commit and byte hash. The full-archive manifest here remains comparison baseline.
|
||||
2. Map every final file to originating project or generated bundle components. Record SPDX identifier, copyright holder, evidence URL, required notice text, and fulfillment location. Mark unknown explicitly; no implicit root-license inheritance for third-party work.
|
||||
3. Resolve the specific gaps above. Carry all nine existing notice files when their associated files ship; carry missing upstream notices and keep inline notices. Build a top-level `THIRD_PARTY_NOTICES` index with bundled texts or direct accompanying files.
|
||||
4. Recheck after any exclusion, replacement, or rebuild. Any changed byte needs a new manifest digest and runtime validation. License clearance and offline functional validation are separate gates.
|
||||
|
||||
The earlier [offline pack research](speedometer-offline-pack.md) established static-path plausibility and proposed blocked-network validation; it did not clear redistribution rights. No benchmark, browser installation, or host change was made here.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -1,33 +0,0 @@
|
||||
# Speedometer 3.1 offline pack at the pinned upstream commit
|
||||
|
||||
## Decision-ready finding
|
||||
|
||||
The pinned [WebKit/Speedometer commit `1386415be8fef2f6b6bbdbe1828872471c5d802a`](https://github.com/WebKit/Speedometer/commit/1386415be8fef2f6b6bbdbe1828872471c5d802a) is a plausible source for an Odin **versioned browser workload** served from localhost. It contains built static applications and the benchmark runner. Its page identifies itself as Speedometer **3.1**, although `package.json` still says `3.0.0-alpha`; identify the pack by the full commit and a content digest, not that package version. The [about page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/about.html) says workloads are built as static files and cannot depend on server infrastructure. The [runner](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/benchmark-runner.mjs) loads each suite in an iframe under `resources/`.
|
||||
|
||||
This is a **conditional yes** for an offline pack. Source inspection and static link checks support completeness, but they do not prove that a browser makes no external requests or that every workload succeeds without internet. Require a blocked-network browser smoke test before calling the pack offline-ready. This research did not execute a benchmark or install a browser.
|
||||
|
||||
## Pack contents and static checks
|
||||
|
||||
The [suite list](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/tests.mjs) declares 32 suites, 20 enabled by default. Local inspection of the pinned source archive found every declared suite entry path. The archive contained 1,118 files totaling 61,022,359 uncompressed bytes. For the root page plus the 20 enabled suite entry pages, a static HTML parser checked 194 `script`, asset `link`, and `img` references: none was external or missing. These numbers describe the checked archive, not a run result. The parser did not resolve dynamic JavaScript imports, CSS URLs, route requests, or user navigation.
|
||||
|
||||
The [main page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/index.html) loads local CSS and `resources/main.mjs`; the [runner](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/benchmark-runner.mjs) constructs `resources/${suite.url}`. The Perf Dashboard workload deserves special attention: its [static page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/perf.webkit.org/public/v3/index.html) replaces its API method with `mockAPIs()` and fetches 13 specified local JSON paths. All 13 files exist in the pinned archive. The ordinary [dashboard remote API](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/perf.webkit.org/public/v3/remote.js) supports XHR, so verify the mock remains active in the actual packaged page.
|
||||
|
||||
No `npm install` is needed to serve the already built workload assets. Upstream [development instructions](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Development.md) use `http-server` for local development. Odin can serve the frozen file tree with its own loopback-only static server; do not rebuild application assets as part of a benchmark run. Use an HTTP origin rather than `file://`, since the suite uses modules, iframe paths, and fetches. The [upstream test harness](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/tests/run.mjs) is Selenium based and requires an installed browser and matching driver, per [Testing.md](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Testing.md); it is not a prerequisite for serving the built pack.
|
||||
|
||||
## Redistribution boundary
|
||||
|
||||
The root [LICENSE](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE) permits source and binary redistribution with or without changes if its copyright notice, conditions, and disclaimer are retained or reproduced as specified. This is not a blanket license for all included third-party work. The [TodoMVC subtree license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/license.md) states MIT unless otherwise specified and requires inclusion of its notice in copies or substantial portions. Built bundles also carry license files, including [React bundle notices](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/architecture-examples/react/dist/app.bundle.js.LICENSE.txt), [Angular third-party notices](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/architecture-examples/angular/dist/3rdpartylicenses.txt), and [React Stockcharts notices](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/react-stockcharts/build/static/js/2.8e539c84.chunk.js.LICENSE.txt).
|
||||
|
||||
Pack the complete upstream notice files alongside their assets, preserve inline notices, and record a notice inventory with the pack manifest. A complete third-party license audit remains unresolved: the archive includes many generated bundles and assets, and finding a root license plus named notice files does not establish licensing for every individual asset. Review the final redistributed file set and notices before shipping. This is a licensing assessment from primary source text, not legal advice.
|
||||
|
||||
## Browser mode and result identity
|
||||
|
||||
Upstream [test instructions](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/instructions.html) call for a latest stable browser, clean profile, focused page, closed competing tabs, and no interaction during the run. The [page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/index.html) warns when its visible viewport is below 850 × 650. The [parameters](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/params.mjs) separately default the suite iframe to 800 × 600 and expose iteration count, suite selection, and timing method. Record these exact conditions in Odin's run record.
|
||||
|
||||
Use a headed, focused browser session for the comparable browser workload. Upstream provides no headless equivalence claim in the cited instructions or [Selenium runner](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/tests/run.mjs). Headless may be useful for a smoke test, but treat any headless measurement as a distinct software mode until equivalence is demonstrated. Do not silently mix browser versions, profiles, window/iframe sizes, suite selections, or headed/headless results in one calibrated measurement.
|
||||
|
||||
## Manifest and offline acceptance proposal
|
||||
|
||||
Create a deterministic, content-addressed pack manifest. Record: upstream repository URL and full commit; Speedometer 3.1 display version; every shipped relative path with byte length and SHA-256; a sorted inventory of license/notice paths; default suite names; and packaging schema version. Hash canonical serialized manifest bytes for the pack identifier. Verify each file hash before serving; reject missing, extra, or changed files. Keep the complete source snapshot or an auditable mapping from snapshot to shipped subset. This is an Odin design recommendation based on the pinned [runner paths](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/benchmark-runner.mjs), [suite list](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/tests.mjs), and [license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE); upstream does not prescribe this manifest.
|
||||
|
||||
Before a benchmark run, validate offline behavior without collecting a score: serve the frozen tree on loopback, open the landing page and each default suite page in a disposable browser profile, disable outside network at the browser or sandbox boundary, log attempted requests, and check that all required assets load with no external request or console error. Include dynamic imports, CSS fonts/images, redirects, worker requests, and the Perf Dashboard JSON paths. Do not click Start Test during this gate. If the gate fails, report the missing/remote URL and mark the browser workload unavailable; do not substitute an online fetch. This acceptance procedure is proposed, not claimed as completed.
|
||||
@@ -1,39 +0,0 @@
|
||||
# Speedometer 3.1 offline pack: packaging decision
|
||||
|
||||
## Answer
|
||||
|
||||
**Do not redistribute an unchanged Speedometer 3.1 pack from commit `1386415be8fef2f6b6bbdbe1828872471c5d802a` yet.** The pinned archive is a conservative, precisely inventoried source candidate, but its full redistribution rights are not established. The companion [per-file source inventory](speedometer-license-manifest.tsv) contains all 1,118 archive files, each with byte count and SHA-256 (61,022,359 bytes total; inventory SHA-256 `78edcd45649128fbd4c8659c1d054a2b9bb76a184b58ce59b37895c19f2bdfb6`). It identifies review groups, **not** individual license clearance or a minimal runtime set. The [prior license audit](speedometer-license-audit.md) documents the evidence and gaps. Source: [pinned WebKit/Speedometer tree](https://github.com/WebKit/Speedometer/tree/1386415be8fef2f6b6bbdbe1828872471c5d802a).
|
||||
|
||||
The acquisition decision is concrete: obtain a verifiable grant and required notices for every unresolved work in the selected pack, or replace those works and rebuild/validate the affected assets. If either route cannot clear every **default** suite, do not call a reduced suite set the unchanged official Speedometer 3.1 workload. Select a different browser workload or explicitly specify a derivative suite and scoring identity. The [pinned suite list](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/tests.mjs) enables NewsSite Next/Nuxt, both Charts suites, both Editor suites, and complex TodoMVC variants by default, so simply deleting those assets changes the measured workload. No permission or replacement is established by this report.
|
||||
|
||||
## Asset and notice boundary
|
||||
|
||||
The [runner](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/benchmark-runner.mjs) loads `resources/${suite.url}`; the [suite list](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/tests.mjs) declares 32 suite entries, 20 enabled by default. All declared entry files exist in the pinned archive. Previous static inspection found 194 landing-page and default-entry HTML asset references, all local and present, but did not resolve dynamic imports, CSS URLs, route requests, or interactions. Source: [prior offline-pack research](speedometer-offline-pack.md), [pinned archive](https://github.com/WebKit/Speedometer/tree/1386415be8fef2f6b6bbdbe1828872471c5d802a). No exact smaller **runtime** file set has therefore been proved. Use the full archive as the conservative candidate until a dependency trace and offline gate justify exclusions.
|
||||
|
||||
| Asset group in candidate archive | Evidence and obligation or gap |
|
||||
| --- | --- |
|
||||
| Speedometer-authored files | [Root BSD-style license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/LICENSE): retain its copyright, conditions, and disclaimer in source copies; reproduce them in documentation or other binary-distribution materials. Its terms do not establish rights to embedded third-party works. |
|
||||
| TodoMVC implementations and generated framework bundles | [TodoMVC license](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/license.md) says MIT unless otherwise specified. Keep that text, matching generated-bundle notices, and inline notices; match bundled component versions to source/lockfiles. [Angular notice](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/architecture-examples/angular/dist/3rdpartylicenses.txt) includes MIT, Apache-2.0, and CC-BY-4.0 entries. |
|
||||
| NewsSite Next/Nuxt, template, CSS, imagery | Both [Next](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/newssite/news-next/README.md) and [Nuxt](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/newssite/news-nuxt/README.md) credit the [source template](https://github.com/flashdesignory/news-site-template), whose repository exposes no license file or package license. No grant for adapted template content is established. [Pinned CSS package metadata](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/newssite/news-site-css/package.json) says ISC, but the matching license text and rights for bundled images still need confirmation. |
|
||||
| Charts code and data | [Dataset README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/charts/datasets/README) identifies copied CSVs without source rights. The archive also contains generated chart bundles, including data embedded in `resources/charts/dist/assets/flights-airports-9a9e6422.js`; the [manifest](speedometer-license-manifest.tsv) records their hashes. Trace dataset grants or replace with cleared equivalent data, rebuild, then recheck the result. |
|
||||
| Complex TodoMVC Adobe shell | [Big DOM README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/todomvc/big-dom-generator/README.md) names Adobe `@spectrum-css`; the archive has `Smock_*.svg` icons. Exact package versions, rights, and built-copy notices remain unresolved. |
|
||||
| Editor fixtures and bundles | [`longtext.html`](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/editors/assets/longtext.html) identifies Project Gutenberg eBook 2650 and contains its license. [Project Gutenberg's terms](https://www.gutenberg.org/policy/license) attach redistribution and trademark conditions while its marks remain; geographic rights need checking. Replace with independently cleared text of comparable workload shape if those conditions are unsuitable. Editor bundles also need component/notice mapping. |
|
||||
| React Stockcharts and Perf Dashboard | [Stockcharts README](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/react-stockcharts/README.md) cites MIT; [generated bundle notice](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/react-stockcharts/build/static/js/2.8e539c84.chunk.js.LICENSE.txt) lists components. Keep notices and verify component versions. Audit generated dashboard JavaScript and imagery against their source rights. |
|
||||
|
||||
The candidate archive has nine separate notice files: root `LICENSE`; `resources/todomvc/license.md`; one React Stockcharts `*.LICENSE.txt`; two Angular `3rdpartylicenses.txt`; and four React-family `app.bundle.js.LICENSE.txt`. Their exact paths and hashes are in the [source inventory](speedometer-license-manifest.tsv). Preserve applicable files and inline headers. A top-level notice index must map each shipped component or asset to origin, copyright holder, license/permission evidence, required notice, and fulfillment location. A notice file alone does not cure an absent grant. Source: [pinned tree](https://github.com/WebKit/Speedometer/tree/1386415be8fef2f6b6bbdbe1828872471c5d802a), [license audit](speedometer-license-audit.md).
|
||||
|
||||
## Offline fetch boundary
|
||||
|
||||
The built pack can be served over loopback HTTP without installing its development dependencies; the [about page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/about.html) says suites are static applications without server infrastructure. The [main page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/index.html) loads local runner assets. The Perf Dashboard [static page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/perf.webkit.org/public/v3/index.html) overrides `RemoteAPI.sendHttpRequest`, prefetches 13 local JSON fixtures, and reports unexpected paths. Those 13 paths exist in the archive; an unmocked [remote API implementation](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/perf.webkit.org/public/v3/remote.js) also exists, so packaging must retain the override and verify it executes. Static source inspection has **not** established closure for all dynamic requests, imports, CSS fonts/images, redirects, workers, or navigation. Source: [prior offline-pack research](speedometer-offline-pack.md).
|
||||
|
||||
Prepared native automation is outside the asset pack. The pinned [Selenium harness](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/tests/run.mjs) starts a server and connects to an installed browser/driver; [Testing.md](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/Testing.md) states those installation requirements for upstream tests. Odin's browser binary, driver, profile, window state, headed/headless mode, and automation adapter belong to the **run environment record**, not the frozen workload pack. Upstream [instructions](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/instructions.html) require a focused browser page. The [parameters](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/params.mjs) default the suite iframe to 800 × 600; the [landing page](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/index.html) warns below an 850 × 650 visible viewport. Record both sizes; keep headed and headless measurements separate unless equivalence is demonstrated.
|
||||
|
||||
## Final manifest and acceptance gate
|
||||
|
||||
After rights are cleared and the exact shipped file set is selected, create a canonical, content-addressed manifest **for that set**. Include schema version; upstream URL/full commit; Speedometer display version; immutable suite names, entry URLs, enabled state, and runner parameters; every normalized relative path with byte length and SHA-256; provenance and license evidence identifiers; notice paths and fulfillment mapping; and all replacements/build inputs. Sort paths by UTF-8 bytes, use one specified JSON serialization, exclude the manifest's own digest from hashed content, and publish the SHA-256 of those bytes as pack identity. Reject duplicate or traversal paths, symlinks, missing/extra files, and hash mismatches before serving. Any altered asset, suite list, or notice changes pack identity. This is an Odin packaging proposal; upstream provides no such manifest. Source for source identity and suite paths: [pinned tree](https://github.com/WebKit/Speedometer/tree/1386415be8fef2f6b6bbdbe1828872471c5d802a), [suite list](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/tests.mjs).
|
||||
|
||||
Run a separate **non-benchmark** acceptance gate before any scored use: verify the manifest; serve read-only files on loopback; start a disposable browser profile with outbound network blocked; open the landing page and each enabled suite entry without starting `Start Test`; capture all page/frame/worker requests and response status, redirects, console errors, and service-worker activity; inspect dynamic imports, CSS images/fonts, and the Perf Dashboard fixture requests. If a suite needs interaction to reveal a resource, use an unscored smoke action outside the runner and record it. Fail on any outside request, missing local resource, unexpected remote API path, or console error that affects loading. This tests reachability and fetch closure, **not** benchmark behavior or timing. A full workload execution later remains a separate validation before results are trusted. Source for runner behavior: [runner](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/benchmark-runner.mjs), [dashboard mock](https://github.com/WebKit/Speedometer/blob/1386415be8fef2f6b6bbdbe1828872471c5d802a/resources/perf.webkit.org/public/v3/index.html).
|
||||
|
||||
## State of evidence
|
||||
|
||||
No browser was installed, no benchmark was run, and no host setting was changed for this research. No pack has passed the legal or offline gate. The exact runtime subset, component-to-file rights matrix, successful fetch closure, and permission or replacement choices remain to be established before redistribution.
|
||||
Reference in New Issue
Block a user