Define health findings and the optimization advice rubric #11

Closed
opened 2026-09-25 17:48:52 +00:00 by xavierk · 1 comment
Owner

Part of Find the way to Odin’s build-ready specification.

Question

Which observable conditions justify each severity, bottleneck diagnosis, driver finding, optimization recommendation, and NVMe/HDD replacement recommendation? Set evidence prerequisites, thresholds or vendor-specific interpretation, uncertainty, stale/missing telemetry, kernel configuration context, VM restrictions and what Odin must never infer. Define non-destructive defaults, optional privileges, safe load/write bounds, refusal/abort behavior, and advice-only versus any explicitly invoked action.

Resolve with the human; use health warnings that cannot be masked by a favorable performance result.

Part of [Find the way to Odin’s build-ready specification](https://git.bongbetic.com/xavierk/odin/issues/1). <!-- wayfinder-map: 1 --> ## Question Which observable conditions justify each severity, bottleneck diagnosis, driver finding, optimization recommendation, and NVMe/HDD replacement recommendation? Set evidence prerequisites, thresholds or vendor-specific interpretation, uncertainty, stale/missing telemetry, kernel configuration context, VM restrictions and what Odin must never infer. Define non-destructive defaults, optional privileges, safe load/write bounds, refusal/abort behavior, and advice-only versus any explicitly invoked action. Resolve with the human; use health warnings that cannot be masked by a favorable performance result.
xavierk added the wayfinder:grilling label 2026-09-25 17:48:52 +00:00
xavierk added a new dependency 2026-09-25 17:49:55 +00:00
xavierk added a new dependency 2026-09-25 17:49:55 +00:00
xavierk added a new dependency 2026-09-25 17:49:56 +00:00
xavierk added a new dependency 2026-09-25 17:49:57 +00:00
xavierk added a new dependency 2026-09-25 17:49:57 +00:00
xavierk added a new dependency 2026-09-25 17:49:59 +00:00
xavierk added a new dependency 2026-09-25 17:50:03 +00:00
xavierk added a new dependency 2026-09-25 17:50:07 +00:00
xavierk added a new dependency 2026-09-25 17:50:10 +00:00
xavierk self-assigned this 2026-09-27 05:57:00 +00:00
Author
Owner

Resolution

The human accepted the recommendations in three live decision rounds and confirmed the completed policy. This is Odin’s health-finding and optimization-advice contract, not a claim of implemented or physically validated diagnostics.

Severity, confidence, and coverage

Finding severity expresses response urgency independently of performance and confidence:

  • Informational: an observation or historical condition without evidence requiring current investigation; provide context or monitoring advice where appropriate.
  • Warning: evidence requiring investigation, such as new corrected hardware errors, unexplained new interface/device errors, or a condition needing maintenance planning.
  • Critical: a condition requiring prompt action, including credible device failure or unsafe operation. Stop decisions follow the safety rules below, rather than relying only on the severity label.

Finding confidence is directly observed, corroborated, or tentative. Show the supporting evidence and apply confidence to the actual claim: an error may be directly observed while its cause remains tentative. Do not turn these labels into invented numerical probabilities. New corrected hardware errors warrant a warning and investigation; historical counts alone do not establish ongoing failure.

Missing, inaccessible, uninterpretable, or stale evidence leaves the relevant current condition unknown, never healthy. Unknown is a coverage state, not a severity. Distinguish current observations, changes observed during the run, and historical evidence. Preserve observation time, collection scope, source/tool interpretation, and any limitations. A historical event can justify follow-up without proving that the fault is active now. Describe successful limited checks as no concerning evidence observed within their tested scope, not a certificate of overall health.

Drive findings and replacement advice

Credible device-reported failure, applicable reliability degradation, a read-only failure state, or repeated device-attributable integrity errors justify back up promptly and arrange replacement/service. State the evidence and attribution limits; a read-only filesystem or ambiguous I/O error alone does not establish a failing physical drive.

Endurance consumption alone warrants monitoring and replacement planning, not a claim of certain or imminent failure. Temperature alarms warrant stopping relevant heat-producing workloads and investigating cooling, not automatic drive replacement. Interface errors warrant connection/path investigation; do not equate them with failed media. Low measured storage performance alone is a performance limitation, not evidence of hardware failure.

Use documented protocol alarms and applicable vendor/kernel limits, recording their source and applicability. Do not invent universal temperature limits, SMART raw-value thresholds, or remaining-life formulas. Unknown transport safety or unknown interpretation produces an explicit coverage gap. Elevated permissions do not bypass transport or interpretation restrictions. Safe read-only discovery must not guess device-changing commands or start self-tests.

Memory, GPU, and system diagnosis

A passing memory check covers only actually tested bytes and completed patterns. Missing EDAC/RAS evidence or a zero observed delta does not certify RAM. Treat detected errors according to the evidence and retain the distinction between observed corruption, corrected events, and historical counts.

GPU discovery, driver loading, successful API execution, verified output, and presentation are different claims. Report the stage that succeeded or failed, distinguishing missing permissions/API support, software or virtual rendering, execution failure, device loss, and incorrect output. A slow result or unavailable API alone does not establish a broken driver.

VM findings describe the guest-visible system unless physical-device attribution is established. Preserve kernel configuration, power/thermal conditions, and other measurement context when explaining limitations; their mere presence is not proof of a bottleneck.

Safety, privileges, and recovery

Start unprivileged. Explain missing evidence and offer an explicitly selected, narrowly scoped privileged diagnostic check when appropriate. Declining retains the run with the corresponding evidence unavailable. Privilege is not authorization to tune, repair, install drivers, or use an unsafe transport.

Retain all agreed resource, write, profile, and deadline bounds from Choose Odin’s workload suite and run profiles, including explicit storage-write consent and separately selected extended modules. This decision does not relax those budgets or add automatic self-tests.

Stop affected workloads on active thermal alarms, device loss, integrity errors, or breached resource/write limits. Stop all later scheduled workloads affecting the same component as well. Existing critical conditions prevent relevant load from starting. Memory corruption stops the entire run. When safety or measurement integrity cannot be isolated to one component, stop the entire run. Missing health telemetry alone permits otherwise safely bounded workloads with the coverage gap shown.

Do not automatically retry a safety stop. Preserve evidence and partial results, terminate workers where possible, and report cleanup failures or unresolved I/O. Do not claim that timeout, process termination, or attempted cleanup guarantees driver recovery or completion of kernel I/O. Cancellation and terminal/process implementation belong to the runtime contract and must satisfy these safety outcomes.

Optimization advice and score integrity

Require workload-specific corroborating evidence before asserting a bottleneck. A low score, high utilization, pressure reading, or configuration choice alone does not establish a sole cause. Describe plausible causes as hypotheses when attribution remains uncertain. Every recommendation must connect to observed evidence, state relevant tradeoffs, and provide a verification step. Do not promise a speedup without measurement. Tuning, driver changes, and repairs remain advice-only.

Preserve Define Odin’s median score and comparability contract: health findings remain prominent outside numerical aggregation; they never numerically penalize speed. Detected memory errors or failed output verification withhold the headline. Safety stops or unfinished required workloads also withhold it. Missing health telemetry alone does not suppress an otherwise complete valid performance score. Slow valid measurements remain valid. Severity and score eligibility are distinct rules: classifying a newly corrected memory error as a warning does not override the existing memory-error eligibility rule.

Handoff and evidence

This decision uses the existing investigations: Establish safe CPU, memory, and kernel measurements, Establish GPU driver evidence and performance workloads, and Establish storage measurements and trustworthy health advice. No new benchmark, stress test, privileged probe, or system modification was performed.

The remaining work is already owned by existing decisions: the runtime ticket specifies process/privilege mechanisms and compatibility; the local-results ticket specifies evidence persistence; the interaction ticket evaluates how these distinctions are presented; and Set distro, VM, and physical-hardware validation gates must cover qualified alarm/parser interpretations, scope/attribution, safety stops, privilege denial, unavailable evidence, and cleanup failures. Unsupported interpretations remain unknown until qualified. No new fog or separate decision ticket was exposed by this resolution.

## Resolution The human accepted the recommendations in three live decision rounds and confirmed the completed policy. This is Odin’s health-finding and optimization-advice contract, not a claim of implemented or physically validated diagnostics. ### Severity, confidence, and coverage Finding severity expresses response urgency independently of performance and confidence: - **Informational:** an observation or historical condition without evidence requiring current investigation; provide context or monitoring advice where appropriate. - **Warning:** evidence requiring investigation, such as new corrected hardware errors, unexplained new interface/device errors, or a condition needing maintenance planning. - **Critical:** a condition requiring prompt action, including credible device failure or unsafe operation. Stop decisions follow the safety rules below, rather than relying only on the severity label. Finding confidence is **directly observed**, **corroborated**, or **tentative**. Show the supporting evidence and apply confidence to the actual claim: an error may be directly observed while its cause remains tentative. Do not turn these labels into invented numerical probabilities. New corrected hardware errors warrant a warning and investigation; historical counts alone do not establish ongoing failure. Missing, inaccessible, uninterpretable, or stale evidence leaves the relevant current condition **unknown**, never healthy. Unknown is a coverage state, not a severity. Distinguish current observations, changes observed during the run, and historical evidence. Preserve observation time, collection scope, source/tool interpretation, and any limitations. A historical event can justify follow-up without proving that the fault is active now. Describe successful limited checks as no concerning evidence observed within their tested scope, not a certificate of overall health. ### Drive findings and replacement advice Credible device-reported failure, applicable reliability degradation, a read-only failure state, or repeated device-attributable integrity errors justify **back up promptly and arrange replacement/service**. State the evidence and attribution limits; a read-only filesystem or ambiguous I/O error alone does not establish a failing physical drive. Endurance consumption alone warrants monitoring and replacement planning, not a claim of certain or imminent failure. Temperature alarms warrant stopping relevant heat-producing workloads and investigating cooling, not automatic drive replacement. Interface errors warrant connection/path investigation; do not equate them with failed media. Low measured storage performance alone is a performance limitation, not evidence of hardware failure. Use documented protocol alarms and applicable vendor/kernel limits, recording their source and applicability. Do not invent universal temperature limits, SMART raw-value thresholds, or remaining-life formulas. Unknown transport safety or unknown interpretation produces an explicit coverage gap. Elevated permissions do not bypass transport or interpretation restrictions. Safe read-only discovery must not guess device-changing commands or start self-tests. ### Memory, GPU, and system diagnosis A passing memory check covers only actually tested bytes and completed patterns. Missing EDAC/RAS evidence or a zero observed delta does not certify RAM. Treat detected errors according to the evidence and retain the distinction between observed corruption, corrected events, and historical counts. GPU discovery, driver loading, successful API execution, verified output, and presentation are different claims. Report the stage that succeeded or failed, distinguishing missing permissions/API support, software or virtual rendering, execution failure, device loss, and incorrect output. A slow result or unavailable API alone does not establish a broken driver. VM findings describe the guest-visible system unless physical-device attribution is established. Preserve kernel configuration, power/thermal conditions, and other measurement context when explaining limitations; their mere presence is not proof of a bottleneck. ### Safety, privileges, and recovery Start unprivileged. Explain missing evidence and offer an explicitly selected, narrowly scoped privileged diagnostic check when appropriate. Declining retains the run with the corresponding evidence unavailable. Privilege is not authorization to tune, repair, install drivers, or use an unsafe transport. Retain all agreed resource, write, profile, and deadline bounds from [Choose Odin’s workload suite and run profiles](https://git.bongbetic.com/xavierk/odin/issues/8#issuecomment-6705), including explicit storage-write consent and separately selected extended modules. This decision does not relax those budgets or add automatic self-tests. Stop affected workloads on active thermal alarms, device loss, integrity errors, or breached resource/write limits. Stop all later scheduled workloads affecting the same component as well. Existing critical conditions prevent relevant load from starting. Memory corruption stops the entire run. When safety or measurement integrity cannot be isolated to one component, stop the entire run. Missing health telemetry alone permits otherwise safely bounded workloads with the coverage gap shown. Do not automatically retry a safety stop. Preserve evidence and partial results, terminate workers where possible, and report cleanup failures or unresolved I/O. Do not claim that timeout, process termination, or attempted cleanup guarantees driver recovery or completion of kernel I/O. Cancellation and terminal/process implementation belong to the runtime contract and must satisfy these safety outcomes. ### Optimization advice and score integrity Require workload-specific corroborating evidence before asserting a bottleneck. A low score, high utilization, pressure reading, or configuration choice alone does not establish a sole cause. Describe plausible causes as hypotheses when attribution remains uncertain. Every recommendation must connect to observed evidence, state relevant tradeoffs, and provide a verification step. Do not promise a speedup without measurement. Tuning, driver changes, and repairs remain advice-only. Preserve [Define Odin’s median score and comparability contract](https://git.bongbetic.com/xavierk/odin/issues/10#issuecomment-6723): health findings remain prominent outside numerical aggregation; they never numerically penalize speed. Detected memory errors or failed output verification withhold the headline. Safety stops or unfinished required workloads also withhold it. Missing health telemetry alone does not suppress an otherwise complete valid performance score. Slow valid measurements remain valid. Severity and score eligibility are distinct rules: classifying a newly corrected memory error as a warning does not override the existing memory-error eligibility rule. ### Handoff and evidence This decision uses the existing investigations: [Establish safe CPU, memory, and kernel measurements](https://git.bongbetic.com/xavierk/odin/issues/2), [Establish GPU driver evidence and performance workloads](https://git.bongbetic.com/xavierk/odin/issues/3), and [Establish storage measurements and trustworthy health advice](https://git.bongbetic.com/xavierk/odin/issues/4). No new benchmark, stress test, privileged probe, or system modification was performed. The remaining work is already owned by existing decisions: the runtime ticket specifies process/privilege mechanisms and compatibility; the local-results ticket specifies evidence persistence; the interaction ticket evaluates how these distinctions are presented; and [Set distro, VM, and physical-hardware validation gates](https://git.bongbetic.com/xavierk/odin/issues/14) must cover qualified alarm/parser interpretations, scope/attribution, safety stops, privilege denial, unavailable evidence, and cleanup failures. Unsupported interpretations remain unknown until qualified. No new fog or separate decision ticket was exposed by this resolution.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/odin#11