Specify Odin’s v1 browser interaction workload #20
Notifications
Due Date
No due date set.
Blocks
#15 Approve Odin’s build-ready specification
xavierk/odin
Reference: xavierk/odin#20
Reference in New Issue
Block a user
Part of Find the way to Odin’s build-ready specification.
Question
What exact Odin-owned, offline browser interaction workload will v1 run and score? Specify representative web-application scenarios and their deterministic fixtures, scripted interactions, measured phases and native statistic, warmup and repetition rules, correctness and invalidation checks, browser setup, headed/headless identities, and the full standard-run time budget. Define the fixed browser-domain membership and normalization direction so Define Odin’s median score and comparability contract can apply without Speedometer’s native score or references. State which results are comparable, which capabilities cause a reasoned skip, and what the validation matrix must exercise.
Set the source and dependency boundary for a distributable pack. Require per-file provenance, license, and notice evidence for every shipped file, a canonical content-addressed manifest, and an unscored smoke test with outbound network blocked. Decide the acceptance criteria and fallback outcome if the pack or reference corpus fails. Produce a build-ready protocol with the human; implementation, actual rights audit, smoke execution, and calibration data collection follow the planning map.
This ticket implements the workload route chosen in Choose a legally distributable v1 browser workload. The changed identity must update the standard profile, browser domain, calibration plan, run-time budget, and distro/physical validation gates. No Speedometer 3.1 asset, score, or reference value silently carries over.
Resolution — product owner approved the protocol in four live rounds
This is the build-ready browser workload policy. No pack has been built, rights audit completed, smoke test run, or calibration value measured in this planning ticket.
Workload identity and fixtures
Odin Browser Interaction v1 is one offline, Odin-owned suite with three fixed web-application scenarios. Its application, including the fixed Markdown subset and renderer, is authored in plain HTML, CSS, and JavaScript without third-party page libraries, images, or fonts. Odin serves the complete pack on loopback. The benchmark page uses no external runtime asset or service.
Build the fixtures once as static UTF-8 JSON and Markdown files from these deterministic rules; never generate runtime random values, dates, or user data:
The ordered actions are fixed:
The script pins target selectors, prepared input values, one triggering browser input per action, and the expected visible state after every action. Field preparation and text selection occur before each measured trigger. A cycle resets all scenarios to the frozen fixtures before replay. The independent checker compares observed states and final scenario digests against expected values shipped in the pack; a no-op or unexpected state fails the suite. The fixed scenario order is grid, board, editor within each cycle. The manifest pins the exact fixture bytes, app source, script, and expected values. Any change to these artifacts creates a new workload identity and requires new qualification and calibration.
Measurement and browser domain
A full suite has two untimed warmup cycles followed by ten measured cycles. Each measured cycle has eight actions in each of the three scenarios, giving exactly 240 valid latency samples. Reset, warmup, setup, assertion work, and cleanup are outside individual action timers but inside the stage deadline. Standard runs one full suite. Quick checks browser and automation capability without a browser score. The explicitly selected extended browser module runs three independent full suites for diagnostic spread; it does not alter standard browser-domain membership.
The page records the trusted input event timestamp on its own monotonic time line. Once the expected UI state is present, it waits for two consecutive requestAnimationFrame callbacks and records the second callback time. Their difference in milliseconds is Odin's action-to-visible-update proxy. Browser-side timing excludes WebDriver command transport. This is a frame-paced proxy, not physical photon timing or the browser's Interaction to Next Paint metric. Clock-origin mismatch, absent frame callbacks, or a missing action sample makes the suite ineligible. The Event Timing API is not the source of scored samples: its minimum 16 ms exposure threshold can omit fast events, and its duration is rounded to 8 ms (MDN Event Timing, W3C Event Timing). requestAnimationFrame cadence follows the display refresh rate (MDN requestAnimationFrame).
Sort all 240 valid action latencies ascending. The native suite statistic is nearest-rank p95: item ceil(0.95 × 240) = 228, in milliseconds. Retain every action value, action identity, per-scenario summaries, sample count, and range; do not trim slow valid values. The browser performance domain has exactly one required measurement: this complete native p95. Normalize lower-is-better as browser domain score = 100 × qualified reference p95 / observed p95. Installed mode uses one frozen reference value per headed/headless mode, measured on its declared physical reference browser and shared across qualified architectures and browser engines. Each prepared reference-browser configuration needs separate reference evidence and a distinct comparison identity. The full headline remains the median of seven domain scores under Define Odin’s median score and comparability contract. No partial browser suite, substitute statistic, or old Speedometer score enters that median.
Browser setup, modes, and comparisons
Standard uses one user-selected installed Chromium-family or Firefox browser with an existing compatible local WebDriver. Verify browser and driver versions and action/timing capability before launch. The normal run never installs a driver, downloads a browser, or substitutes a prepared browser. Missing or incompatible browser/driver is a reasoned browser capability skip. A prepared reference-browser run is an explicit separate software mode and needs its own pinned browser and qualified reference evidence.
Launch the selected browser with a fresh temporary profile, no extensions, a clean cache at start, a fixed 800 × 600 CSS-pixel viewport, and 100% zoom. Use a focused headed window by default when a usable display exists; permit an explicit headless choice, and use headless when no usable display exists. Use the selected browser binary for both modes; a separate headless-shell binary is a different workload identity. The browser page loads only loopback resources. Record browser and driver versions, engine, installed/reference software mode, headed/headless mode, window focus, viewport, device-pixel ratio, display refresh rate when available, GPU acceleration/renderer path, OS/kernel, pack digest, protocol revision, and calibration release.
Direct score comparison requires the same pack, browser protocol, scoring rule, calibration release, and execution/software modes. Installed browser version and engine changes remain visible as environmental differences; no difference is attributed to one cause without evidence. Headed and headless, or installed and prepared-reference results, are not merged. A display refresh difference is flagged in comparison rather than silently changing or suppressing the score. The browser score describes installed-system response, including display cadence in headed mode. Virtual-machine measurements remain raw-only under Define Odin’s calibration procedure and release gates.
Time, validity, and outcomes
The standard browser stage has a five-minute end-to-end deadline, including preflight, local server and browser launch, loading, warmup, all measured cycles, verification, and cleanup. It retains the 10–20 minute overall standard-run target of Choose Odin’s workload suite and run profiles. A slow valid action remains in the sample set. Odin never shortens cycles, drops a slow sample, or silently retries to meet the deadline.
All 240 expected actions must have one finite positive timestamped sample and the correct independently checked UI state. Wrong or missing output, duplicate/missing action, hidden or unfocused headed page, browser crash, failed local asset, workload-page request outside loopback, unusable clock or animation-frame timing, safety stop, or deadline failure prevents a valid browser-domain score. Classify missing capability as unavailable; wrong output as invalid; deadline as timeout/incomplete; and user cancellation as cancelled. Preserve partial measurements, environment, error, and cleanup evidence in the run record. Without the required browser-domain measurement, withhold the full seven-domain headline. Health findings stay separately visible and obey the established integrity and safety-stop rules.
Pack, calibration, and release gates
The runtime browser app, fixture files, action scripts, automation adapter, styles, and any bundled dependencies are pinned as a distributable pack. Every shipped payload file must have path, byte length, SHA-256, source/provenance, license or documented rights basis, and notice evidence, including transitive assets. The pack manifest uses RFC 8785 canonical JSON, with file entries sorted by normalized relative path. The SHA-256 of the canonical manifest is the pack ID; the release record identifies the generated manifest itself. Reject path traversal, symlinks, missing notices, and unlisted runtime assets. No Speedometer 3.1 asset or reference is included.
Before packaging, run an unscored smoke test of every qualified browser engine and headed/headless mode with outbound network blocked and loopback allowed. Confirm complete asset closure, expected UI states, all 240 samples, and cleanup. Before enabling numerical browser or headline scores, collect a new physical, mode-specific reference corpus using Define Odin’s calibration procedure and release gates. Its existing five-complete-run, two-day, two-boot, multi-system, and aarch64 gates apply to the new p95 statistic. The qualified reference is shared across architectures; no Speedometer reference or score transfers. A missing scored mode corpus yields raw browser measurements only and no full headline. A failed rights, manifest, or offline-smoke gate blocks the pack; browser coverage is unavailable and the full headline is withheld.
Set distro, VM, and physical-hardware validation gates must exercise: supported Chromium-family and Firefox driver pairs; headed and headless identities; installed and any prepared-reference software modes; x86_64/aarch64 and glibc/musl cells where claimed; Void, deb, and rpm families; desktop, SSH/headless, and VM compatibility; low-end physical completion within the five-minute browser stage and 10–20 minute standard target; display refresh and device-pixel-ratio effects; GPU acceleration/renderer changes; missing/mismatched drivers, unavailable display, focus loss, bad fixture/output, external requests, timeout, cancellation, and cleanup. VMs establish compatibility, not numerical score validity. Exact distro cells and physical repeatability thresholds remain owned by that validation decision ticket.
These are acceptance gates for implementation and release, not work performed during this planning map.