Define monitoring signals and honest warm-up estimates #67

Closed
opened 2026-09-11 04:08:37 +00:00 by xavierk · 1 comment
Owner

Parent map: Fenris TUI polish and hourly history

Question

What precise state, precedence, text, and timing contract makes monitoring and data readiness truthful at a glance?

HITL grilling + domain-modeling. Approved direction: blinking green Monitoring, explicit Collecting, steady amber Paused, red Error/Stale; text always visible. Preserve ADR 0002 projection rules and ADR 0003 boot/runtime and deliberate-disable semantics.

Resolve mapping from actual timer/collection/store evidence to status, including healthy time between collection runs, UI refresh versus collector liveness, unknown/unavailable status, external service stop versus deliberate pause, failures with fresh data, and conflicting facts. Define blink cadence and reduced-motion/non-colour fallback. Keep boot enablement, last outcome, freshness and collection activity distinct; no false healthy signal.

Separately resolve conditional remaining-hours estimates for projection warm-up, including qualifying UTC days, low coverage, pauses, stale data, missing baseline/unsupported counters, and cases with no defensible ETA. State remaining qualifying evidence and assumptions, never promise a fixed countdown. First graph-data availability must not wait for projection confidence; coordinate with the hourly-data contract without inventing its answer. Preserve TUI/CLI parity where already binding and explicit quit-versus-pause semantics.

Answer must include a decision table, exact wording, state precedence and example transitions; history/counter/projection absence needs explanatory text, not a blank or misleading zero.

Parent map: [Fenris TUI polish and hourly history](https://git.bongbetic.com/xavierk/Fenris/issues/65) ## Question What precise state, precedence, text, and timing contract makes monitoring and data readiness truthful at a glance? HITL grilling + domain-modeling. Approved direction: blinking green Monitoring, explicit Collecting, steady amber Paused, red Error/Stale; text always visible. Preserve ADR 0002 projection rules and ADR 0003 boot/runtime and deliberate-disable semantics. Resolve mapping from actual timer/collection/store evidence to status, including healthy time between collection runs, UI refresh versus collector liveness, unknown/unavailable status, external service stop versus deliberate pause, failures with fresh data, and conflicting facts. Define blink cadence and reduced-motion/non-colour fallback. Keep boot enablement, last outcome, freshness and collection activity distinct; no false healthy signal. Separately resolve conditional remaining-hours estimates for projection warm-up, including qualifying UTC days, low coverage, pauses, stale data, missing baseline/unsupported counters, and cases with no defensible ETA. State remaining qualifying evidence and assumptions, never promise a fixed countdown. First graph-data availability must not wait for projection confidence; coordinate with the hourly-data contract without inventing its answer. Preserve TUI/CLI parity where already binding and explicit quit-versus-pause semantics. Answer must include a decision table, exact wording, state precedence and example transitions; history/counter/projection absence needs explanatory text, not a blank or misleading zero.
xavierk added this to the Wayfinder: Fenris TUI polish and hourly history milestone 2026-09-11 04:08:37 +00:00
xavierk added the wayfinder:grilling label 2026-09-11 04:08:37 +00:00
xavierk added a new dependency 2026-09-11 04:08:51 +00:00
xavierk added a new dependency 2026-09-11 04:08:51 +00:00
xavierk added a new dependency 2026-09-11 05:12:04 +00:00
xavierk self-assigned this 2026-09-13 18:34:52 +00:00
Author
Owner

Resolution — approved monitoring-signal and warm-up contract

Human accepted every Round 1 and Round 2 recommendation in live grilling ("recommendation accepted" ×2). Planning only; wording below is spec-ready for the companion spec. One mechanical gap fixed at resolution: Round 2's accepted glyph set omitted a glyph for Interrupted; ⊘ was added (single-width, consistent with the accepted set).

Status lattice — eight states

Text label always visible; colour never the sole carrier (glyph + text). No animation in any state except Monitoring's dot.

State Colour Glyph Motion Wording / explanation line
Monitoring green ● blink 750 ms on / 750 ms off, status dot only, never text ● Monitoring
Collecting green ◐ steady ◐ Collecting — run in flight (≤ 90 s)
Paused amber ‖ steady ‖ Paused — monitoring paused — paused time excluded from your usage habit
Waiting amber ○ steady ○ Waiting — last sample X ago / awaiting first sample (empty store) / awaiting another sample (one sample, no compatible pair, per #66)
Interrupted red ⊘ steady ⊘ Interrupted — collection stopped outside Fenris — monitoring period still open
Error red ✖ steady ✖ Error — last run failed (exit N) (+ last good sample X ago when data is still fresh) or observation store unreadable — see journal
Stale red ◌ steady ◌ Stale — last sample X days ago (timer active, no failure recorded)
Unknown grey ? steady ? Unknown — service state unavailable

Reduced motion (Textual SCREEN_REDUCE_MOTION or user preference): static ●, no pulse — which is exactly the CLI's permanent form.

Precedence

Base order: Error > Interrupted > Paused > Stale > Waiting > Monitoring > Unknown.

  • Paused outranks Stale: paused 49 days → amber Paused with last sample 49 days ago as a fact line; staleness is the expected consequence of the user's own choice, not an alarm.
  • Error outranks Interrupted: the stop rides in the explanation line.
  • Unknown only when the service query fails and no store-derived fact (open period, freshness, deliberate-pause row) places the state higher.
  • Collecting is an overlay, not a rung: shown while the oneshot is in flight, overriding every base state except a store fault (a run cannot publish against an unreadable store; red Error stays). While overlaying Paused/Interrupted or a retry after failure, the explanation line names the underlying state (e.g. run in flight — paused). It reverts within ≤ 90 s to whatever the outcome earns.

Distinct facts — never folded into the status word

  • Freshness: Last sample: X ago always displayed (own axis).
  • Last outcome: ok / failed (exit N) / none.
  • Boot enablement: Start at boot: on/off; enabled≠active is explained by the state text itself.
  • Collection activity: the Collecting overlay plus the poll below.

Blink means exactly "monitoring enabled and data fresh" — never data refresh.

Detection and timing

  • A 5 s lightweight systemctl show poll (timer + service units) drives the status strip; the full store refresh stays at 5-min cadence. Status transitions are re-evaluated every poll tick from the cached newest-sample timestamp + current clock, so aging boundaries (Waiting→Stale at 48 h) cross within ~5 s without a store read; only a new sample needs the refresh. Poll failure → Unknown.
  • Collecting detection: service ActiveState = activating (oneshot in flight).
  • Boundaries adopt the existing constants unchanged: fresh ≤ 690 s, stale ≥ 48 h, run timeout 90 s (ADR 0003), cadence 5 min.

Canonical example transitions

Scenario Shows
Failed run, data 4 min old ✖ Error — last run failed (exit 3) · last good sample 4 min ago
Next run succeeds ● Monitoring (self-clears ≈ 5 min)
Failed run, retry in flight ◐ Collecting — run in flight — retry → Error or Monitoring on outcome
Suspend 3 h, wake, timer fires ○ Waiting → ◐ Collecting → ● Monitoring
External stop, data fresh ⊘ Interrupted — collection stopped outside Fenris
External stop, 3 days later ⊘ Interrupted — + last sample 3 days ago (names both facts; never silently degrades to Stale-without-stop)
Pause, 49 days ‖ Paused (amber) — last sample 49 days ago
Fresh install, zero samples ○ Waiting — awaiting first sample
First sample, no pair yet ○ Waiting — awaiting another sample (#66 wording)
Store unreadable ✖ Error — observation store unreadable — see journal; nothing store-derived renders (ADR 0005)

Warm-up and conditional estimates

Confidence block during warm-up (counts per #71: current controller segment; qualifying = represented monitored UTC date with ≥ 50 % coverage; gate = 14 represented / 12 qualifying):

Building evidence — 9 of 14 days observed · 7 qualifying
First lifespan estimate after 12 qualifying days

No countdown by hours: days are the grain, and a day can still fail qualification — a promised ETA would be fabricated confidence.

Withheld-estimate reason lines (one per cause, never a blank or misleading zero):

  • No endurance baseline — set a rated TBW to see an estimate
  • Drive does not report write counters
  • warm-up progress line (above)
  • stale: the estimate is shown frozen at the last evidence endpoint with estimate not updating — last sample X ago (#71 frozen-endpoint rule)
  • paused days: excluded from the day count; the Paused status carries that fact — no double accounting in the block.

First graph-data availability is independent of estimate availability (#66): the graph renders from first usable evidence while the estimate stays gated.

Once warm-up clears, the existing frozen-spec lifespan line renders unchanged with its Limited/Supported label; this contract adds only the progress block, the reason lines and the frozen note — no new steady-state format is invented here.

CLI parity

fenris status adopts the identical status vocabulary, precedence, glyphs (statically rendered) and reason lines. Collecting appears in CLI output only if a run is in flight at query time.

Preserved unchanged

ADR 0002 projection math and confidence categories; ADR 0003 lifecycle, sanctioned toggle and cadence; ADR 0005 failure contract; #66 hourly-history and first-data contract; #71 UTC accounting and warm-up gate; quit-versus-pause semantics; no fabricated history; the TUI stays read-only.

## Resolution — approved monitoring-signal and warm-up contract Human accepted every Round 1 and Round 2 recommendation in live grilling ("recommendation accepted" ×2). Planning only; wording below is spec-ready for the companion spec. One mechanical gap fixed at resolution: Round 2's accepted glyph set omitted a glyph for Interrupted; **⊘** was added (single-width, consistent with the accepted set). ### Status lattice — eight states Text label always visible; colour never the sole carrier (glyph + text). No animation in any state except Monitoring's dot. | State | Colour | Glyph | Motion | Wording / explanation line | |---|---|---|---|---| | Monitoring | green | ● | blink 750 ms on / 750 ms off, status dot only, never text | `● Monitoring` | | Collecting | green | ◐ | steady | `◐ Collecting` — run in flight (≤ 90 s) | | Paused | amber | ‖ | steady | `‖ Paused` — `monitoring paused — paused time excluded from your usage habit` | | Waiting | amber | ○ | steady | `○ Waiting` — `last sample X ago` / `awaiting first sample` (empty store) / `awaiting another sample` (one sample, no compatible pair, per #66) | | Interrupted | red | ⊘ | steady | `⊘ Interrupted` — `collection stopped outside Fenris — monitoring period still open` | | Error | red | ✖ | steady | `✖ Error` — `last run failed (exit N)` (+ `last good sample X ago` when data is still fresh) or `observation store unreadable — see journal` | | Stale | red | ◌ | steady | `◌ Stale` — `last sample X days ago` (timer active, no failure recorded) | | Unknown | grey | ? | steady | `? Unknown` — `service state unavailable` | Reduced motion (Textual `SCREEN_REDUCE_MOTION` or user preference): static ●, no pulse — which is exactly the CLI's permanent form. ### Precedence Base order: **Error > Interrupted > Paused > Stale > Waiting > Monitoring > Unknown**. - **Paused outranks Stale**: paused 49 days → amber Paused with `last sample 49 days ago` as a fact line; staleness is the expected consequence of the user's own choice, not an alarm. - **Error outranks Interrupted**: the stop rides in the explanation line. - **Unknown** only when the service query fails *and* no store-derived fact (open period, freshness, deliberate-pause row) places the state higher. - **Collecting is an overlay**, not a rung: shown while the oneshot is in flight, overriding every base state **except a store fault** (a run cannot publish against an unreadable store; red Error stays). While overlaying Paused/Interrupted or a retry after failure, the explanation line names the underlying state (e.g. `run in flight — paused`). It reverts within ≤ 90 s to whatever the outcome earns. ### Distinct facts — never folded into the status word - Freshness: `Last sample: X ago` always displayed (own axis). - Last outcome: ok / failed (exit N) / none. - Boot enablement: `Start at boot: on/off`; enabled≠active is explained by the state text itself. - Collection activity: the Collecting overlay plus the poll below. Blink means exactly "monitoring enabled and data fresh" — never data refresh. ### Detection and timing - A **5 s lightweight `systemctl show` poll** (timer + service units) drives the status strip; the full store refresh stays at 5-min cadence. Status transitions are re-evaluated every poll tick from the cached newest-sample timestamp + current clock, so aging boundaries (Waiting→Stale at 48 h) cross within ~5 s without a store read; only a *new* sample needs the refresh. Poll failure → Unknown. - Collecting detection: service `ActiveState` = activating (oneshot in flight). - Boundaries adopt the existing constants unchanged: fresh ≤ 690 s, stale ≥ 48 h, run timeout 90 s (ADR 0003), cadence 5 min. ### Canonical example transitions | Scenario | Shows | |---|---| | Failed run, data 4 min old | `✖ Error` — `last run failed (exit 3) · last good sample 4 min ago` | | Next run succeeds | `● Monitoring` (self-clears ≈ 5 min) | | Failed run, retry in flight | `◐ Collecting` — `run in flight — retry` → Error or Monitoring on outcome | | Suspend 3 h, wake, timer fires | `○ Waiting` → `◐ Collecting` → `● Monitoring` | | External stop, data fresh | `⊘ Interrupted` — `collection stopped outside Fenris` | | External stop, 3 days later | `⊘ Interrupted` — + `last sample 3 days ago` (names both facts; never silently degrades to Stale-without-stop) | | Pause, 49 days | `‖ Paused` (amber) — `last sample 49 days ago` | | Fresh install, zero samples | `○ Waiting` — `awaiting first sample` | | First sample, no pair yet | `○ Waiting` — `awaiting another sample` (#66 wording) | | Store unreadable | `✖ Error` — `observation store unreadable — see journal`; nothing store-derived renders (ADR 0005) | ### Warm-up and conditional estimates Confidence block during warm-up (counts per #71: current controller segment; qualifying = represented monitored UTC date with ≥ 50 % coverage; gate = 14 represented / 12 qualifying): ``` Building evidence — 9 of 14 days observed · 7 qualifying First lifespan estimate after 12 qualifying days ``` No countdown by hours: days are the grain, and a day can still fail qualification — a promised ETA would be fabricated confidence. Withheld-estimate reason lines (one per cause, never a blank or misleading zero): - `No endurance baseline — set a rated TBW to see an estimate` - `Drive does not report write counters` - warm-up progress line (above) - stale: the estimate is shown **frozen at the last evidence endpoint** with `estimate not updating — last sample X ago` (#71 frozen-endpoint rule) - paused days: excluded from the day count; the Paused status carries that fact — no double accounting in the block. First graph-data availability is independent of estimate availability (#66): the graph renders from first usable evidence while the estimate stays gated. Once warm-up clears, the **existing frozen-spec lifespan line renders unchanged** with its Limited/Supported label; this contract adds only the progress block, the reason lines and the frozen note — no new steady-state format is invented here. ### CLI parity `fenris status` adopts the identical status vocabulary, precedence, glyphs (statically rendered) and reason lines. Collecting appears in CLI output only if a run is in flight at query time. ### Preserved unchanged ADR 0002 projection math and confidence categories; ADR 0003 lifecycle, sanctioned toggle and cadence; ADR 0005 failure contract; #66 hourly-history and first-data contract; #71 UTC accounting and warm-up gate; quit-versus-pause semantics; no fabricated history; the TUI stays read-only.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Reference: xavierk/Fenris#67