Define local results, history, and export behavior #12

Closed
opened 2026-09-25 17:48:53 +00:00 by xavierk · 1 comment
Owner

Part of Find the way to Odin’s build-ready specification.

Question

What must a local result preserve so anyone can revisit their run, understand warnings, and make valid comparisons across systems or upgrades? Decide schema/version metadata, units/raw measurements versus summaries, workload/tool/kernel/driver/profile provenance, coverage states, confidence, atomic persistence/recovery, storage location/retention, history UX, export portability and sensitive hardware identity handling. No hosted service is required.

Resolve the product/data contract with the human before choosing incidental persistence details.

Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). <!-- wayfinder-map: 1 --> ## Question What must a local result preserve so anyone can revisit their run, understand warnings, and make valid comparisons across systems or upgrades? Decide schema/version metadata, units/raw measurements versus summaries, workload/tool/kernel/driver/profile provenance, coverage states, confidence, atomic persistence/recovery, storage location/retention, history UX, export portability and sensitive hardware identity handling. No hosted service is required. Resolve the product/data contract with the human before choosing incidental persistence details.
xavierk added the wayfinder:grilling label 2026-09-25 17:48:53 +00:00
xavierk added a new dependency 2026-09-25 17:49:58 +00:00
xavierk added a new dependency 2026-09-25 17:49:59 +00:00
xavierk added a new dependency 2026-09-25 17:50:00 +00:00
xavierk added a new dependency 2026-09-25 17:50:03 +00:00
xavierk added a new dependency 2026-09-25 17:50:11 +00:00
xavierk self-assigned this 2026-09-27 18:35:34 +00:00
Author
Owner

Resolution

The user confirmed the recommendations across five live rounds and confirmed the completed contract. This defines product behavior for local results; it does not select a database engine or claim that persistence has been implemented.

Run record and provenance

  • Keep a full run record for every execution, including completed, incomplete, cancelled, invalid, and safety-stopped runs. Preserve partial samples and cleanup status with their actual outcomes; never present them as completed measurements or scores.
  • Preserve raw trials and native statistics, units, timestamps and time order, sample counts and ranges, validity and coverage outcomes, summary calculations, and score inputs. Retain the exact reason for each unavailable or unfinished workload. Preserve the scoring rule and calibration release used for any published score. A missing required measurement still withholds the full headline under Define Odin’s median score and comparability contract.
  • Preserve enough context to interpret each result: workload definitions and versions, helper and parser versions, profile, protocol, scoring and schema versions, calibration release, execution mode, software mode, kernel and relevant configuration, driver, selected device and filesystem mapping, permissions, and measured thermal or power context. Record observed values and unavailable context distinctly; do not infer a stable condition from absent evidence.
  • Keep the source, observation time, scope, severity, confidence, stale/unknown status, and supporting evidence for health findings and optimization advice. Preserve the separation between health evidence and performance scoring agreed in Define health findings and the optimization advice rubric.
  • Version the record schema and calculation identity. Preserve the original measurement and calculation record as recorded; a later reader may derive a new view or migrate a copy, but must not silently rewrite historical evidence or scores. User notes live separately from the original record.

Local history and recovery

  • Save by default in a per-user local data directory, separate from the path selected for storage benchmarking. Let users choose or open the result location. Keep runs until explicit deletion; show storage use and offer deliberate cleanup. No automatic age or size eviction in v1.
  • A failed or interrupted save must leave previously committed history intact. Recover or flag any incomplete record and state whether the current run was saved. The implementation mechanism is open, but this observable atomicity and recovery behavior is required.
  • The v1 terminal history must list and filter runs, show full evidence, compare two runs, export, and delete a selected run after confirmation. Users may add or edit a note such as an upgrade marker without changing the original evidence.
  • Compare only directly compatible run identities under the existing scoring contract. When scoring rules, calibration, workload protocol, browser mode, software mode, storage mode, or physical/virtual context differ, show side-by-side values with specific incompatibility reasons, without calculating a combined change or implying a sole cause. Other observed environmental changes remain visible for interpretation.

Export, import, and privacy

  • Provide a portable, versioned, machine-readable export in v1 with raw measurements, summaries, coverage, provenance, health findings, and advice. A human-readable report may be additional. Preserve units and identities so another reader can evaluate compatibility.
  • Redact hostname, username, home paths, serial numbers, and stable hardware identifiers by default. Retain non-identifying specifications needed to interpret results. A full-fidelity export is available only through an explicit, clearly labelled choice with a preview of identifying fields.
  • If redaction would remove comparison-critical information, explain the conflict and require a different export choice. Never silently disclose identifiers or silently discard critical comparison data.
  • Support v1 import after schema and integrity checks. Mark imported records and preserve their original identities; do not make import look like a locally measured run.

The user identified no missing requirement in the final check. No new decision ticket or fog item emerged. Runtime-specific storage mechanisms remain for Choose Odin’s runtime and Linux compatibility contract, while the TUI presentation is evaluated in Evaluate the Bongbetic TUI workflow with InkUI. The discovery questions used for this decision are attached as Odin local-results questionnaire.

## Resolution The user confirmed the recommendations across five live rounds and confirmed the completed contract. This defines product behavior for local results; it does not select a database engine or claim that persistence has been implemented. ### Run record and provenance - Keep a full run record for every execution, including completed, incomplete, cancelled, invalid, and safety-stopped runs. Preserve partial samples and cleanup status with their actual outcomes; never present them as completed measurements or scores. - Preserve raw trials and native statistics, units, timestamps and time order, sample counts and ranges, validity and coverage outcomes, summary calculations, and score inputs. Retain the exact reason for each unavailable or unfinished workload. Preserve the scoring rule and calibration release used for any published score. A missing required measurement still withholds the full headline under [Define Odin’s median score and comparability contract](https://git.bongbetic.com/xavierk/odin/issues/10). - Preserve enough context to interpret each result: workload definitions and versions, helper and parser versions, profile, protocol, scoring and schema versions, calibration release, execution mode, software mode, kernel and relevant configuration, driver, selected device and filesystem mapping, permissions, and measured thermal or power context. Record observed values and unavailable context distinctly; do not infer a stable condition from absent evidence. - Keep the source, observation time, scope, severity, confidence, stale/unknown status, and supporting evidence for health findings and optimization advice. Preserve the separation between health evidence and performance scoring agreed in [Define health findings and the optimization advice rubric](https://git.bongbetic.com/xavierk/odin/issues/11). - Version the record schema and calculation identity. Preserve the original measurement and calculation record as recorded; a later reader may derive a new view or migrate a copy, but must not silently rewrite historical evidence or scores. User notes live separately from the original record. ### Local history and recovery - Save by default in a per-user local data directory, separate from the path selected for storage benchmarking. Let users choose or open the result location. Keep runs until explicit deletion; show storage use and offer deliberate cleanup. No automatic age or size eviction in v1. - A failed or interrupted save must leave previously committed history intact. Recover or flag any incomplete record and state whether the current run was saved. The implementation mechanism is open, but this observable atomicity and recovery behavior is required. - The v1 terminal history must list and filter runs, show full evidence, compare two runs, export, and delete a selected run after confirmation. Users may add or edit a note such as an upgrade marker without changing the original evidence. - Compare only directly compatible run identities under the existing scoring contract. When scoring rules, calibration, workload protocol, browser mode, software mode, storage mode, or physical/virtual context differ, show side-by-side values with specific incompatibility reasons, without calculating a combined change or implying a sole cause. Other observed environmental changes remain visible for interpretation. ### Export, import, and privacy - Provide a portable, versioned, machine-readable export in v1 with raw measurements, summaries, coverage, provenance, health findings, and advice. A human-readable report may be additional. Preserve units and identities so another reader can evaluate compatibility. - Redact hostname, username, home paths, serial numbers, and stable hardware identifiers by default. Retain non-identifying specifications needed to interpret results. A full-fidelity export is available only through an explicit, clearly labelled choice with a preview of identifying fields. - If redaction would remove comparison-critical information, explain the conflict and require a different export choice. Never silently disclose identifiers or silently discard critical comparison data. - Support v1 import after schema and integrity checks. Mark imported records and preserve their original identities; do not make import look like a locally measured run. The user identified no missing requirement in the final check. No new decision ticket or fog item emerged. Runtime-specific storage mechanisms remain for [Choose Odin’s runtime and Linux compatibility contract](https://git.bongbetic.com/xavierk/odin/issues/9), while the TUI presentation is evaluated in [Evaluate the Bongbetic TUI workflow with InkUI](https://git.bongbetic.com/xavierk/odin/issues/13). The discovery questions used for this decision are attached as [Odin local-results questionnaire](https://git.bongbetic.com/attachments/c8754a21-db5b-4f9f-8552-de24e50a3d2b).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/odin#12