docs(spec): fenris-redesign.md — implementation-ready specification from ADRs 0001-0006 with two-way traceability matrix
This commit is contained in:
@@ -0,0 +1,654 @@
|
|||||||
|
# Fenris redesign specification
|
||||||
|
|
||||||
|
**Status: implementation-ready.** Assembled by [Write the Fenris redesign specification and close the map](https://git.bongbetic.com/xavierk/Fenris/issues/19), executing the assembly decision [Assemble the implementation-ready specification](https://git.bongbetic.com/xavierk/Fenris/issues/17) (all seven recommendations accepted) on the Wayfinder map [Chart Fenris's persistent TUI monitoring redesign](https://git.bongbetic.com/xavierk/Fenris/issues/1).
|
||||||
|
|
||||||
|
**Canonical roles.** [ADRs 0001–0006](../adr/) are the immutable rationale records — the *why*. [Acceptance criteria](acceptance-criteria.md) are the single register of testable statements — the *definition of done*. This document normatively restates every **operative contract** — the *what* — so an implementer never needs Wayfinder-ticket access: schema column sets, constants, rule tables, unit and CLI definitions, and the Panes TUI layout. Nothing here overrides an ADR or restates a criterion as a criterion.
|
||||||
|
|
||||||
|
## How to read this document
|
||||||
|
|
||||||
|
- **Binding language.** *Must*, *exactly*, and *never* are normative. Terminology follows the glossary in [`CONTEXT.md`](../../CONTEXT.md): *observation history*, *usage-adjusted theoretical lifespan*, *projection confidence*, *monitoring period*, *observation store*, *hour observation*, *day aggregate*, *controller segment*, *degraded identity*, *endurance baseline*, *verified override*, *unverified override*, *sustained regime*, *habit change*, *scenario range*, *coverage*, *collection run*, *deliberate disable*, *store fault*.
|
||||||
|
- **Ordering.** Sections follow data flow: system context → collector acquisition → observation store → controller identity & segmentation → hour/day derivation → projection & confidence → Panes TUI → service lifecycle & sanctioned toggle → failure & recovery → installation. Each section opens with its ADR links and criterion-ID block.
|
||||||
|
- **Implementation boundary.** This specification plans the redesign; it does not implement it. The complete handoff is this document + the [criteria register](acceptance-criteria.md) + [ADRs 0001–0006](../adr/) + the glossary. Given/When/Then test specs are derived by the implementer at implementation time.
|
||||||
|
|
||||||
|
### Normative constants index
|
||||||
|
|
||||||
|
Every constant is defined once, in the section named below; other sections cite, never redefine. All are named constants in code, not configuration.
|
||||||
|
|
||||||
|
| Constant | Value | Defined in |
|
||||||
|
|---|---|---|
|
||||||
|
| Collection cadence (default) | 5 min (`OnUnitInactiveSec`) | §8.2 |
|
||||||
|
| First-boot delay | 2 min (`OnBootSec`) | §8.2 |
|
||||||
|
| Timer accuracy window | 30 s (`AccuracySec`) | §8.2 |
|
||||||
|
| Collection-run timeout | 90 s (`TimeoutStartSec`) | §8.2 |
|
||||||
|
| Fresh threshold | newest sample within 2 × cadence + `AccuracySec` + 60 s | §8.9 |
|
||||||
|
| Missed → stale boundary | 48 h | §8.9, §6.7 |
|
||||||
|
| Powered-off hour threshold | power-on-hours delta < 90 % of the hour's wall-clock span | §5.1 |
|
||||||
|
| Active hour threshold | DUW delta ≥ 256 MiB in the hour | §5.1 |
|
||||||
|
| Raw-sample retention | 14 days | §3.4 |
|
||||||
|
| Warming gate | 14 distinct UTC day aggregates, ≤ 2 below 50 % coverage | §6.6 |
|
||||||
|
| Supported coverage floor | 80 % | §6.7 |
|
||||||
|
| Horizon agreement | 7/28/90-day rates within a factor of 2 | §6.7 |
|
||||||
|
| Burst guard | no single day ≥ 50 % of trailing 28-day bytes | §6.7 |
|
||||||
|
| Young-regime cap | regime < 7 days old → Limited | §6.4 |
|
||||||
|
| Habit-change trigger | trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean, 3 consecutive days | §6.4 |
|
||||||
|
| Regime span cap (default) | full observation history capped at 90 days | §6.4 |
|
||||||
|
| Scenario horizons | 7 / 28 / 90 days | §6.5 |
|
||||||
|
| Implied-baseline eligibility | ≥ 2 Percentage-Used increments within the current controller segment | §6.3 |
|
||||||
|
| Rated-TBW conversion | `E_rated = entered_TBW × 10¹²` bytes | §6.3 |
|
||||||
|
| Implied-baseline validity window | 1 ≤ p ≤ 254 | §6.3 |
|
||||||
|
| Wear-disagreement note | vendor wear vs. observed write rate by more than a factor of 2 | §6.1 |
|
||||||
|
| Capacity validation tolerance | ± 1 % | §6.2 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. System context
|
||||||
|
|
||||||
|
**ADRs:** [0001](../adr/0001-observation-store-sqlite.md), [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md), [0006](../adr/0006-collector-acquisition-path.md). **Criteria:** CI-3, LC-1, LC-5, ST-1.
|
||||||
|
|
||||||
|
### 1.1 Scope
|
||||||
|
|
||||||
|
Fenris observes one configured NVMe drive's real-world use and translates the observation history into a usage-adjusted theoretical lifespan. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer, persistent compact observation storage, and categorical projection confidence reflecting the length, completeness, and stability of real usage history.
|
||||||
|
|
||||||
|
Standing constraints, binding on every section:
|
||||||
|
|
||||||
|
- Linux with systemd and polkit only; no other init system is supported.
|
||||||
|
- Exactly one configured NVMe drive — the device named by `/etc/fenris/fenris.conf` (§8.3).
|
||||||
|
- Fully local: no telemetry, no network fetching, no automatic vendor-data retrieval.
|
||||||
|
- The HTML dashboard and HTTP server are gone; nothing of the daemonization, PID files, or `/run` state survives.
|
||||||
|
- CLI `status` and `sample` are retained (§8.8).
|
||||||
|
|
||||||
|
### 1.2 Components and privilege boundaries
|
||||||
|
|
||||||
|
| Component | Privilege | Path | Role |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `fenris-collect.service` | root oneshot unit | `/usr/libexec/fenris/fenris-collect` | The only code path that interrogates the device and writes the observation store. |
|
||||||
|
| `fenris-collect.timer` | system timer | — | Schedules collection runs; `WantedBy=timers.target`. |
|
||||||
|
| `fenris-monitor` | root helper | `/usr/libexec/fenris/fenris-monitor` | Fixed privileged operations: `enable`/`disable` (optional `--now`), the collect trigger, monitoring-period bookkeeping, and baseline persistence. The only binary polkit authorizes. |
|
||||||
|
| `fenris` | unprivileged | `/usr/local/bin/fenris` | Human entry point: no arguments opens the TUI; subcommands are the CLI (§8.8). Never a unit. |
|
||||||
|
| Observation store | root-written, group-read | `/var/lib/fenris/observations.db` | Single SQLite database in WAL mode (§3). The TUI and `status` open it read-only. |
|
||||||
|
| Configuration | world-readable | `/etc/fenris/fenris.conf` | Exactly one key: the device selector (§8.3). |
|
||||||
|
|
||||||
|
The TUI and CLI are ordinary unprivileged processes. Elevation is exclusively polkit, exclusively for `fenris-monitor` (§8.5). There is no `/run/fenris` coordination surface and no export layer: systemd serializes collection runs, the observation store holds state, failures go to the journal.
|
||||||
|
|
||||||
|
### 1.3 Data flow
|
||||||
|
|
||||||
|
1. The timer fires; `fenris-collect.service` runs `fenris-collect`.
|
||||||
|
2. The collector acquires counters and thermal evidence from `smartctl -a -j` and controller identity from sysfs (§2), normalizes identity exactly once (§2.3), and either fails the whole run or writes one complete sample.
|
||||||
|
3. The collector derives and validates hour observations and day aggregates, advances controller segmentation and period bookkeeping, prunes raw samples, and commits (§3–§5, §9.1).
|
||||||
|
4. Readers — the TUI and `fenris status` — open the store read-only and **recompute the projection on every read** (§6); nothing derived is ever stored (§3.7).
|
||||||
|
|
||||||
|
Control flow is separate: the human drives the TUI/CLI; privileged operations route through `fenris-monitor` under polkit to `systemctl`; period rows record *intent* (only the sanctioned path), while the collector records *observed fact* (§8.5–§8.6, §9.8).
|
||||||
|
|
||||||
|
### 1.4 Cross-cutting prohibitions
|
||||||
|
|
||||||
|
These are operative contracts; each is restated in its home section and gated by the criteria block [CI-3](acceptance-criteria.md):
|
||||||
|
|
||||||
|
1. No code path outside `fenris-collect` interrogates the device (§2.1, §8.7).
|
||||||
|
2. Polkit authorizes exactly one binary, `fenris-monitor`, under `com.bongbetic.fenris.monitor` `auth_admin` (§8.5).
|
||||||
|
3. No `/run/fenris` coordination surface or export layer exists (§1.2).
|
||||||
|
4. No absent hour is ever interpolated, estimated, or fabricated (§5.3, §9.3).
|
||||||
|
5. No alerting, notification, or escalation machinery exists anywhere (§9.6–§9.7).
|
||||||
|
6. `/etc/fenris/fenris.conf` holds exactly one key — the device selector (§8.3).
|
||||||
|
7. No synthetic or capacity-derived baseline is ever created, including for legacy history (§6.1, §3.5).
|
||||||
|
8. Readers never partially interpret a newer-schema store (§3.6, §9.5).
|
||||||
|
9. Projections are never stored; always recomputed on read (§3.7, §6.10).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Collector acquisition
|
||||||
|
|
||||||
|
**ADR:** [0006](../adr/0006-collector-acquisition-path.md). **Criteria:** AC-1–AC-5; miss absorption per [0005](../adr/0005-failure-detection-and-recovery.md) §5.
|
||||||
|
|
||||||
|
### 2.1 Channels — the hard pin
|
||||||
|
|
||||||
|
Every collection run acquires exactly two ways:
|
||||||
|
|
||||||
|
- **Counters and thermal evidence** — solely from `smartctl -a -j <device>`: `data_units_written`, `data_units_read`, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning` — consumed as-is (smartmontools already trims the strings it copies).
|
||||||
|
- **Controller identity** — solely from sysfs (`/sys/class/nvme/<ctrl>/`): `subnqn`, `sn`, `mn`, `fr`, `transport`.
|
||||||
|
|
||||||
|
No other acquisition path exists anywhere in the codebase. There is no fallback: libnvme bindings and the `nvme` CLI JSON interface are excluded (ADR 0006, *Considered options*).
|
||||||
|
|
||||||
|
### 2.2 All-or-nothing runs
|
||||||
|
|
||||||
|
Any acquisition failure — missing `smartctl` binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the **whole** collection run. A partial sample (identity without counters, or counters without identity) is never written: a transient read failure must never push a healthy drive down the degraded-identity path (§4). The miss surfaces through freshness grading (§8.9) and the flat retry cadence (§9.6), never as degraded identity.
|
||||||
|
|
||||||
|
### 2.3 Identity normalization — once, at write time
|
||||||
|
|
||||||
|
One collector-side function normalizes every identity field, applied exactly once at write time:
|
||||||
|
|
||||||
|
- strip trailing spaces and newlines;
|
||||||
|
- no case folding;
|
||||||
|
- empty-after-strip is stored blank.
|
||||||
|
|
||||||
|
Padded and unpadded renderings of the same field therefore yield byte-identical stored values — a collector implementation change can never split a drive's own history. A future acquisition-path change must deliver byte-identical normalized identity values, or the change itself forces a controller-segment boundary.
|
||||||
|
|
||||||
|
### 2.4 Segment metadata sourcing
|
||||||
|
|
||||||
|
`transport` comes from the NVMe class sysfs directory. `vid`/`ssvid` come from the PCI node (`/sys/class/nvme/<ctrl>/device/{vendor,subsystem_vendor}`) when present and are stored null otherwise. Both are segment **metadata only** (§4.2), never key components.
|
||||||
|
|
||||||
|
### 2.5 Prerequisites
|
||||||
|
|
||||||
|
`make install` verifies `smartctl` is present and fails cleanly otherwise (§10.1). The acquisition path adds no Python dependency and no OS package beyond smartmontools; the dependency lockfile (§10.5) is untouched by this section.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Observation store
|
||||||
|
|
||||||
|
**ADR:** [0001](../adr/0001-observation-store-sqlite.md) as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) and [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14). **Criteria:** ST-1–ST-12; FL-5.
|
||||||
|
|
||||||
|
### 3.1 Substrate and access
|
||||||
|
|
||||||
|
- One SQLite database in **WAL mode** at `/var/lib/fenris/observations.db`. An unprivileged reader querying during a collector write sees a consistent snapshot.
|
||||||
|
- The database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI and `status` open it **read-only**. No `/run` snapshot, no export layer.
|
||||||
|
- `/var/lib/fenris` is created by the installer with root-written group-read permissions; the database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet (§7.6, §10.1).
|
||||||
|
- Migration, schema changes, prune, and import are each single transactions — a killed timer run can never leave partial state.
|
||||||
|
|
||||||
|
### 3.2 Entities and column sets
|
||||||
|
|
||||||
|
The schema carries exactly six entities:
|
||||||
|
|
||||||
|
**`samples`** — recent raw samples (14-day retention, §3.4): timestamp (UTC); the normalized controller-identity fields captured at acquisition (§2.3); raw integer `data_units_written`, `data_units_read`; `percentage_used`; `available_spare`; `media_errors`; `power_on_hours`; `power_cycles`; `unsafe_shutdowns`; temperature; `critical_warning`.
|
||||||
|
|
||||||
|
**`hour_observations`** — one row per UTC hour: the usage-habit split `seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown` (summing to 3600, §5.1); DUW/DUR deltas; temperature min/avg/max; sample count; coverage flag. Classification thresholds belong to the projection model (§5.1), not the store.
|
||||||
|
|
||||||
|
**`day_aggregates`** — one row per UTC day, the habit-evidence grain: each day row carries, at minimum, the day's activity-split sums, write deltas, and coverage share — the inputs the evidence gates of §6.6 consume — derived monotonically from its hour rows.
|
||||||
|
|
||||||
|
**`monitoring_periods`** — `started_at`; `ended_at` (NULL = open); `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not (§5.2, §8.6).
|
||||||
|
|
||||||
|
**`controller_segments`** — spans of unchanged controller identity and monotonic counters; write deltas are never computed across a segment boundary. Columns: the identity key (§4.1) and the frozen metadata snapshot of §4.2, plus the segment's span bounds.
|
||||||
|
|
||||||
|
**`endurance_baseline`** — one active row, replaced on edit (§6.2): the rated-TBW value in bytes (`E_rated = entered_TBW × 10¹²`); mandatory provenance — source URL, document revision, entry date, model string, nominal capacity; frozen validation facts — detected model, detected capacity bytes, `validated_by` (`machine`/`user`), `validated_at`.
|
||||||
|
|
||||||
|
### 3.3 Time model
|
||||||
|
|
||||||
|
Hours and days are UTC-bounded. Day derivation from hour rows is monotonic; DST-ambiguous 23- or 25-hour days never exist in the store.
|
||||||
|
|
||||||
|
### 3.4 Retention
|
||||||
|
|
||||||
|
Raw samples are pruned opportunistically by the collector to **14 days**. Hour observations and day aggregates are retained indefinitely.
|
||||||
|
|
||||||
|
### 3.5 Legacy migration
|
||||||
|
|
||||||
|
The migration procedure, invoked from the entry points below, is **idempotent and interruption-safe**:
|
||||||
|
|
||||||
|
1. If the store already carries the legacy-import marker, do nothing.
|
||||||
|
2. `history.jsonl` is the sole authority: import raw samples and derive hour observations and day aggregates from them.
|
||||||
|
3. `hourly.jsonl` is never trusted as input: mismatches against derived data are diffed and logged.
|
||||||
|
4. Open one implicit `monitoring_periods` row at the first legacy sample, closed `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts — a sample present means powered on; a DUW delta means writes occurred.
|
||||||
|
5. The import is a single transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
|
||||||
|
6. Only after commit are legacy files renamed `*.migrated` — never deleted.
|
||||||
|
7. Malformed legacy lines are quarantined with a logged count, never silently dropped.
|
||||||
|
|
||||||
|
No synthetic or capacity-derived baseline is ever created for legacy history (§1.4–7). Entry points: the installer's import detection at `./data/history.jsonl` (or an explicit path) (§10.1); `fenris import <path>` for later finds (§8.8); and the collector's first new-version run, which performs this same procedure (ADR 0001 §6).
|
||||||
|
|
||||||
|
### 3.6 Schema versioning
|
||||||
|
|
||||||
|
`PRAGMA user_version` plus ordered migration steps in code, each in its own transaction. The collector refuses to run against an unknown **newer** version; readers refuse symmetrically with the exact wording of §9.5 and never partially interpret. *Reconciliation note:* ADR 0004 §6 describes upgrade-time migrations as "governed by a `schema_version` table" — the operative mechanism is this section's `user_version` (ADR 0001 §8, criterion ST-12); there is one version authority, not two.
|
||||||
|
|
||||||
|
### 3.7 Nothing derived is stored
|
||||||
|
|
||||||
|
Projections are not stored; there is no separate latest-status table and no stored health flag. The freshest sample timestamp is the store's own staleness signal (§8.9). The baseline lives in the database (§6.2); `/etc/fenris/` holds only operational configuration (§8.3).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Controller identity and segmentation
|
||||||
|
|
||||||
|
**Decisions:** [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11), [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14), [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15). **ADRs:** [0001](../adr/0001-observation-store-sqlite.md) §3 (as amended), [0002](../adr/0002-projection-model-sustained-regime.md) §§8–9 (as amended). **Criteria:** ID-1–ID-4, PR-9, PR-15, PR-16.
|
||||||
|
|
||||||
|
### 4.1 Identity key ladder
|
||||||
|
|
||||||
|
The controller-segment identity key is the **normalized, kernel-exposed subsystem NQN** (`subnqn`), with fallbacks, in order:
|
||||||
|
|
||||||
|
1. kernel-exposed subsystem NQN;
|
||||||
|
2. the kernel composite;
|
||||||
|
3. `model|serial`.
|
||||||
|
|
||||||
|
`fr` (firmware revision) is metadata only — it may go stale after a mid-segment firmware update. Identity change and DUW decrease are **independent axes** (§4.3).
|
||||||
|
|
||||||
|
### 4.2 Frozen metadata snapshot
|
||||||
|
|
||||||
|
Each segment freezes, at open, a fully nullable metadata snapshot — immutable thereafter: normalized `subnqn`, `sn`, `mn`, `fr`, plus `vid`, `ssvid`, `transport`, and the `identity_degraded` flag. All columns are nullable so incompleteness stays explicit: legacy-imported segments carry `mn` with NULLs (§4.4); degraded segments carry whatever was observed. These are human diagnostics, never key components. `cntlid` is excluded — it distinguishes controllers within one subsystem, out of scope for a single-drive monitor.
|
||||||
|
|
||||||
|
### 4.3 Segmentation axes
|
||||||
|
|
||||||
|
- **DUW decrease, unchanged identity** — a segment boundary within the same drive. Prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.8).
|
||||||
|
- **Identity-key change** — quarantines prior history from projection entirely: it describes a different drive (§6.8).
|
||||||
|
- **Degraded identity** — a segment whose identity key is **blank** (every rung of the ladder empty). `identity_degraded` is set at segment open exactly when the key is blank; keys from the kernel-composite or `model|serial` rungs are not degraded. Blank-key semantics extend identity-change rules verbatim: any visible change of the recorded key — including to or from blank — is a controller-identity change and quarantines; equal blank keys continue the segment, segmented by DUW monotonicity alone. Even a degraded→healthy transition quarantines, so the projection window only ever spans segments sharing one key (§6.8).
|
||||||
|
- Ephemeral markers (model "Linux", non-pcie transport) are segment metadata, never confidence facts.
|
||||||
|
|
||||||
|
The confidence consequence of degraded identity — capped at Limited with its fixed contributing fact — is §6.7's rule.
|
||||||
|
|
||||||
|
### 4.4 Legacy identity
|
||||||
|
|
||||||
|
Legacy history imports under a labeled, model-scoped **legacy identity** (mn-only segments), so it never blends with the post-redesign identity of the same physical drive.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Hour and day derivation
|
||||||
|
|
||||||
|
**ADRs:** [0002](../adr/0002-projection-model-sustained-regime.md) §§4–6; [0001](../adr/0001-observation-store-sqlite.md) §3; [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) §2 (power-on-hours evidence). **Criteria:** PR-4–PR-6, ST-4, FL-3.
|
||||||
|
|
||||||
|
### 5.1 Hour classification
|
||||||
|
|
||||||
|
Each UTC hour is classified by named constants, in this order of evidence:
|
||||||
|
|
||||||
|
- **Powered-off** — the hour's power-on-hours delta is below **90 %** of its wall-clock span.
|
||||||
|
- **Active** — DUW delta ≥ **256 MiB** in the hour.
|
||||||
|
- **Idle** — powered on, sampled, below the active threshold.
|
||||||
|
- **Unknown** — everything else: unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable by design), or inconsistent counters.
|
||||||
|
|
||||||
|
There is no configuration surface for these thresholds; they are documented constants in one projection module.
|
||||||
|
|
||||||
|
### 5.2 Denominator and disabled time
|
||||||
|
|
||||||
|
The projection denominator is **wall-clock seconds inside monitoring periods**, including powered-off and unknown time. Disabled periods — wall-clock outside monitoring periods — are excluded from numerator and denominator. **Disabled time is not an hour state.**
|
||||||
|
|
||||||
|
### 5.3 Gaps and coverage — never backfill
|
||||||
|
|
||||||
|
No absent hour is ever interpolated, estimated, or fabricated. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage. Power-on-hours classification (§5.1) is the only inference admitted. **Coverage** is the share of wall-clock seconds inside monitoring periods whose classification is known rather than unknown — a first-class displayed fact (§6.10, §7.3).
|
||||||
|
|
||||||
|
### 5.4 Day aggregates
|
||||||
|
|
||||||
|
One row per UTC day, derived monotonically from hour rows (§3.2–§3.3) — the grain at which usage-habit evidence is judged (§6.6).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Projection and confidence
|
||||||
|
|
||||||
|
**ADRs:** [0002](../adr/0002-projection-model-sustained-regime.md) as amended by [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15); baseline per [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12). **Criteria:** PR-1–PR-17, CI-4.
|
||||||
|
|
||||||
|
### 6.1 One projection; baseline precedence
|
||||||
|
|
||||||
|
Exactly **one** usage-adjusted theoretical lifespan is computed, against the endurance baseline chosen by precedence:
|
||||||
|
|
||||||
|
1. **Verified override** — a rated-TBW override with complete provenance whose applicability to the detected drive was confirmed by machine match or explicit user attestation;
|
||||||
|
2. **Unverified override** — a rated-TBW override knowingly stored with incomplete provenance; always presented as user-supplied, never as verified;
|
||||||
|
3. **Implied baseline** — derived from vendor wear (§6.3), eligible only per §6.3's gate;
|
||||||
|
4. otherwise the projection is **Unavailable**.
|
||||||
|
|
||||||
|
Percentage Used is context, never a second projection: it renders as a vendor-wear context line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The legacy PU-slope regression and `capacity × 600` synthesis are gone; no synthetic or capacity-derived baseline is ever created (§1.4–7).
|
||||||
|
|
||||||
|
### 6.2 Endurance baseline: provenance and validation
|
||||||
|
|
||||||
|
The baseline lives in the observation store's `endurance_baseline` table (§3.2) and is edited via the CLI (§8.8) — `/etc/fenris/` holds no baseline.
|
||||||
|
|
||||||
|
- **Mandatory provenance:** source URL, document revision, entry date, model string, nominal capacity.
|
||||||
|
- **One active row**, replaced on edit.
|
||||||
|
- **Verification is derived at read** — complete provenance and a drive match (machine or attested) — never a stored boolean.
|
||||||
|
- **Unverified tier:** incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier.
|
||||||
|
- **Entry-time validation** (unprivileged, live sysfs read of the configured device): normalized model containment, with an interactive confirm recorded as `validated_by = user`; nominal capacity within ± 1 %.
|
||||||
|
- **Read-time applicability:** a model match against the current controller segment (§4). A mismatch is **retained — never auto-deleted** — and leaves the projection Unavailable.
|
||||||
|
- Persistence goes through the polkit-guarded `fenris-monitor` verb after CLI-side validation (§8.5).
|
||||||
|
|
||||||
|
### 6.3 Arithmetic
|
||||||
|
|
||||||
|
```text
|
||||||
|
rate = regime DUW delta bytes / in-period wall-clock seconds
|
||||||
|
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
|
||||||
|
E_rated = entered_TBW × 10¹² bytes
|
||||||
|
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
|
||||||
|
```
|
||||||
|
|
||||||
|
- `E_rated` is exact: rated TBW converts to bytes by × 10¹².
|
||||||
|
- `E_implied` is computed **only** for `1 ≤ p ≤ 254`; Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable. The implied baseline is labeled *implied from vendor wear estimate* and shown with few significant digits.
|
||||||
|
- **Implied-baseline eligibility:** the implied tier is used only after ≥ 2 Percentage-Used increments within the current controller segment; until then the projection is Unavailable with the fixed phrase *"vendor wear estimate too coarse to imply endurance"*.
|
||||||
|
|
||||||
|
### 6.4 Sustained regime and habit change
|
||||||
|
|
||||||
|
The headline rate is the **sustained-regime** rate: regime DUW bytes ÷ in-period wall-clock seconds. The default regime is the full observation history capped at **90 days**.
|
||||||
|
|
||||||
|
A **habit change** is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for **3 consecutive days**. The new regime starts at the **first day of divergence**, is adopted automatically, and is labeled *"usage habit changed N days ago"*; the scenario range keeps the longer horizons visible. A regime younger than **7 days** caps projection confidence at Limited evidence.
|
||||||
|
|
||||||
|
### 6.5 Scenario range
|
||||||
|
|
||||||
|
The 7-, 28-, and 90-day rates are computed **independently of the regime** and shown as the scenario range. Only horizons the history actually covers appear — no placeholders. The scenario range is the only spread shown anywhere (§6.9).
|
||||||
|
|
||||||
|
### 6.6 Minimum evidence
|
||||||
|
|
||||||
|
Warming up until **14 distinct UTC day aggregates** of which at most **2** fall below 50 % coverage. The projection still renders while warming up, labeled with its facts (e.g. *"warming up: N of 14 qualifying days"*). Every Unavailable condition renders **no lifespan number**.
|
||||||
|
|
||||||
|
### 6.7 Confidence rule table
|
||||||
|
|
||||||
|
Confidence renders as **state plus contributing facts, never a percentage**. Three states:
|
||||||
|
|
||||||
|
- **Unavailable** — no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
|
||||||
|
- **Supported** — verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80 % **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50 % of trailing 28-day bytes **and** regime ≥ 7 days old **and** the current controller segment's identity key is not degraded.
|
||||||
|
- **Limited** — every other case with a baseline and a positive rate; the failing facts are shown.
|
||||||
|
|
||||||
|
**Staleness:** a newest day aggregate older than **48 hours** drops confidence one level (Supported → Limited) and is shown as a contributing fact.
|
||||||
|
|
||||||
|
**Degraded identity:** a controller segment whose identity key is blank (§4.3) caps confidence at **Limited evidence**, with the contributing fact *"controller identity unavailable — replacement detection relies on write-counter continuity only"* rendered in every state. Supported is unreachable while the current segment is degraded. The cap combines idempotently with the staleness drop (both land at Limited).
|
||||||
|
|
||||||
|
### 6.8 Segment-break effects
|
||||||
|
|
||||||
|
- **DUW decrease, unchanged identity:** prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.6).
|
||||||
|
- **Controller-identity change** — including any to-or-from-blank key change (§4.3): prior history is quarantined from projection entirely.
|
||||||
|
- Since even degraded→healthy transitions quarantine, the projection window only ever spans segments sharing one key; no cross-segment propagation rule is needed.
|
||||||
|
|
||||||
|
### 6.9 Zero rate and uncertainty
|
||||||
|
|
||||||
|
Zero rate renders *"no finite projection from this history"* — never infinity, never zero. No statistical confidence interval appears anywhere; the scenario range is the only spread.
|
||||||
|
|
||||||
|
### 6.10 The projection contract
|
||||||
|
|
||||||
|
The projection function hands the TUI and `status` exactly: the confidence state; the contributing facts — including the degraded-identity fact when the current segment's key is blank; the headline remaining time when one exists; the scenario range; the Percentage-Used context line; the disclosure text (§6.11). Recomputed on read, never stored.
|
||||||
|
|
||||||
|
### 6.11 User-facing language
|
||||||
|
|
||||||
|
Adopted from the endurance research as fixed by ADR 0002 §12; rendered identically by TUI and `status`.
|
||||||
|
|
||||||
|
**Headline wording** (equivalent phrasing required):
|
||||||
|
|
||||||
|
> Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
|
||||||
|
|
||||||
|
**Fixed phrases** (exact): *no finite projection from this history* (zero rate); *vendor wear estimate too coarse to imply endurance* (§6.3); *usage habit changed N days ago* (§6.4); *controller identity unavailable — replacement detection relies on write-counter continuity only* (§6.7); *observation store unreadable* (§9.4); *observation store written by a newer Fenris — upgrade Fenris* (§9.5); *no observations yet* with an enable hint (§8.9); *configuration error: ⟨reason⟩* (§8.3).
|
||||||
|
|
||||||
|
**Confidence rendering:** state plus contributing facts, in the research's evidence style, e.g.
|
||||||
|
|
||||||
|
> Supported evidence · verified manufacturer TBW · 42 calendar days · 96 % interval coverage · 6 weekly cycles · recent and 28-day rates agree
|
||||||
|
|
||||||
|
Never "82 % confidence" or "95 % accurate".
|
||||||
|
|
||||||
|
**The six disclosures** (verbatim, always available — TUI disclosures view and `status`):
|
||||||
|
|
||||||
|
1. This is an endurance projection, not a predicted hardware-failure date.
|
||||||
|
2. Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated.
|
||||||
|
3. Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold.
|
||||||
|
4. DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes.
|
||||||
|
5. Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
|
||||||
|
6. Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Panes TUI
|
||||||
|
|
||||||
|
**Decisions:** [Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3) (Variant A adopted), [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6) (Textual). **ADRs:** [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) §§8, 10; [0004](../adr/0004-install-upgrade-removal-lifecycle.md) §10. **Criteria:** TUI-1–TUI-4, CI-1, CI-2, CI-4. The [prototype](https://git.bongbetic.com/xavierk/Fenris/src/branch/prototype/tui-information-architecture/prototype/tui-ia) is visual reference only; this section is normative.
|
||||||
|
|
||||||
|
### 7.1 Framework and floor
|
||||||
|
|
||||||
|
The TUI is built on **Textual**. It runs on Python 3.9+, gated at install time (§10.1) — never a runtime crash. The tty-passthrough mechanism below was validated live under Textual on a real terminal (prototype decision).
|
||||||
|
|
||||||
|
### 7.2 Layout — one dense keyboard-first screen
|
||||||
|
|
||||||
|
Variant A **Panes**: everything on one screen, no page navigation. The screen is a grid of four regions:
|
||||||
|
|
||||||
|
1. **Headline band** — full width, top: the lifespan headline (or its no-projection wording) with its regime line (*"if current habits continue · sustained regime: N days at R GB/day"*); the confidence state with contributing facts; the scenario range.
|
||||||
|
2. **Usage-history pane** — left, wider column: the write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus their legend; the habit-split bar with active/idle/powered-off/unknown shares.
|
||||||
|
3. **Drive-health and settings pane** — right, narrower column: health facts (model, temperature, spare, media errors, unsafe shutdowns, power-on hours, power cycles, capacity); the vendor-wear context line (Percentage Used · total written of rated — *context, not a second projection*); a read-only settings view (device selector, cadence with drop-in pointer, raw retention, endurance baseline value with its provenance label). Edits happen via CLI / drop-ins, not in the TUI.
|
||||||
|
4. **Service strip** — full width, bottom: the four separate service facts (§7.3), the monitoring-period line, the action legend.
|
||||||
|
|
||||||
|
Exact proportions, glyphs, and borders follow the prototype's validated arrangement as visual reference; the region arrangement, contents, and bindings above are normative.
|
||||||
|
|
||||||
|
### 7.3 Content contracts per region
|
||||||
|
|
||||||
|
- **Confidence is evidence:** state plus contributing facts, never a percentage (§6.7); the headline band renders the §6.10 contract in full, including the disclosure affordance (`d`).
|
||||||
|
- **Four separate service facts, always:** boot enablement (enabled/disabled) · runtime activity (timer active/inactive) · last collect outcome (ok/FAILED, age, reason) · freshness (fresh/missed/stale with newest-sample age, §8.9). They are never merged into one "service status".
|
||||||
|
- **Monitoring-period line:** open-since / closed with end cause; deliberate-disable count where nonzero.
|
||||||
|
- The scenario range shows only covered horizons (§6.5); the vendor-wear context line carries the >2× disagreement note when it applies (§6.1).
|
||||||
|
|
||||||
|
### 7.4 Keybindings and asymmetry
|
||||||
|
|
||||||
|
Production bindings:
|
||||||
|
|
||||||
|
| Key | Action |
|
||||||
|
|---|---|
|
||||||
|
| `p` | Pause — **asks for confirmation** (y pause · n cancel), stating that paused time is excluded from the usage habit while powered-off time would still count. |
|
||||||
|
| `r` | Resume — **no confirmation** (benign; friction invites raw-systemctl escapes). |
|
||||||
|
| `c` | Collect now — synchronous outcome (§8.7), no confirmation. |
|
||||||
|
| `d` | Disclosures — the six disclosures of §6.11. |
|
||||||
|
| `q` | Quit. |
|
||||||
|
|
||||||
|
No bare start/stop exists anywhere; no page navigation keys exist (variant switching was prototype-only). Framework defaults apply for focus and scrolling otherwise.
|
||||||
|
|
||||||
|
### 7.5 Privileged actions and tty passthrough
|
||||||
|
|
||||||
|
Pause, resume, collect-now, and baseline operations run through `fenris-monitor` as a **terminal-attached subprocess**: the TUI suspends, the platform polkit agent prompts on the real terminal, and control returns cleanly with the outcome reflected in the service facts. Where no polkit agent exists the operation fails cleanly with the printed root equivalent (§8.5).
|
||||||
|
|
||||||
|
### 7.6 State rendering obligations
|
||||||
|
|
||||||
|
- From a synthetic observation store, the TUI renders **every** realizable combination of confidence state × freshness grade × baseline tier exactly as the §6.7 rule table and §8.9 constants dictate — headline number only when the rules allow it, contributing facts always, never a percentage (criterion CI-1).
|
||||||
|
- Empty store: *"no observations yet"* with an enable hint; the first-run prompt is an opt-in that enables the timer and opens the first period in one step (§10.1, dormant install).
|
||||||
|
- A `configuration error: ⟨reason⟩` fact renders when the device selector is invalid (§8.3); a store fault suppresses everything store-dependent (§9.4); a newer schema renders its fixed phrase (§9.5).
|
||||||
|
- Every TUI action has a CLI twin with identical outcomes and wording (§8.8, CI-2).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Service lifecycle and sanctioned toggle
|
||||||
|
|
||||||
|
**ADR:** [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12). **Criteria:** LC-1–LC-10, CI-2, CI-3.
|
||||||
|
|
||||||
|
### 8.1 Units
|
||||||
|
|
||||||
|
Exactly two system units exist:
|
||||||
|
|
||||||
|
- `fenris-collect.timer` — `WantedBy=timers.target`.
|
||||||
|
- `fenris-collect.service` — `Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code.
|
||||||
|
|
||||||
|
The TUI and CLI are ordinary unprivileged processes and never units.
|
||||||
|
|
||||||
|
### 8.2 Cadence
|
||||||
|
|
||||||
|
Shipped defaults: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no` (no suspend catch-up — absent hours classify through power-on-hours evidence, §5.1), `TimeoutStartSec=90s` so a hung device interrogation fails visibly as a bounded failed run retried next interval. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); **no interval key exists in configuration**.
|
||||||
|
|
||||||
|
### 8.3 Configuration
|
||||||
|
|
||||||
|
`/etc/fenris/fenris.conf` holds exactly one key: the **device selector**, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run — there is no reload path. An invalid selector is a bounded failed run (journal + failed unit result, retried next interval); `status` and the TUI also read the world-readable file directly and surface `configuration error: ⟨reason⟩`.
|
||||||
|
|
||||||
|
### 8.4 Entry points
|
||||||
|
|
||||||
|
Two privileged binaries — `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable`/`disable` with optional `--now`, the collect trigger, monitoring-period bookkeeping, and `baseline set`/`baseline clear` persistence for the CLI-validated baseline; the only binary the polkit policy authorizes). One unprivileged `fenris` wrapper (§1.2). Root invokes the helpers directly; unprivileged users go through polkit.
|
||||||
|
|
||||||
|
### 8.5 Sanctioned toggle and polkit
|
||||||
|
|
||||||
|
- Pause = `fenris-monitor disable --now`; Resume = `enable --now`. Both perform the systemctl operation **and** the monitoring-period bookkeeping in one step. The human-facing twins `fenris monitor pause` / `fenris monitor resume` map to these and always act immediately; pause asks for confirmation in both TUI and CLI, resume does not (§7.4).
|
||||||
|
- Polkit action `com.bongbetic.fenris.monitor` (`auth_admin`) covers the toggle **and** the collect trigger **and** baseline persistence — authorizing exactly the one binary `fenris-monitor`.
|
||||||
|
- Where no polkit agent exists the operation fails cleanly and prints the root equivalent.
|
||||||
|
- This is the **only** sanctioned control path: a raw `systemctl stop`/`disable` never records `user_disabled` — only the sanctioned path records intent (§8.6).
|
||||||
|
|
||||||
|
### 8.6 Period-row idempotent matrix
|
||||||
|
|
||||||
|
| Situation | Effect on `monitoring_periods` |
|
||||||
|
|---|---|
|
||||||
|
| First-ever enable | Opens a period at the enable moment (hours before the first successful sample are unknown-but-inside — correct when the device errors). |
|
||||||
|
| Resume with an open period (a raw `systemctl stop` intervened) | No row changes; the gap remains inside as unknown seconds. |
|
||||||
|
| Resume with no open period | Opens a new row at the resume moment. |
|
||||||
|
| Pause with an open period | Closes it `user_disabled` at the pause moment. |
|
||||||
|
| Pause otherwise | No-op. |
|
||||||
|
| Raw `systemctl stop`/`disable` outside the helper | An unexplained gap, never `user_disabled`. |
|
||||||
|
|
||||||
|
### 8.7 On-demand collection
|
||||||
|
|
||||||
|
`fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits; the outcome (freshness line or journal hint) is reported synchronously. No confirmation is required. No code path outside `fenris-collect` touches the device; the TUI never samples in-process.
|
||||||
|
|
||||||
|
### 8.8 CLI surface
|
||||||
|
|
||||||
|
| Command | Behavior |
|
||||||
|
|---|---|
|
||||||
|
| `fenris` (no arguments) | Opens the TUI (§7). |
|
||||||
|
| `fenris status` | Read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness. Never auto-samples, never prompts. |
|
||||||
|
| `fenris sample` | On-demand collection via the helper path (§8.7). |
|
||||||
|
| `fenris monitor pause` / `resume` | The sanctioned toggle (§8.5), pause asking confirmation. |
|
||||||
|
| `fenris baseline set` / `clear` | CLI-side validation (§6.2), then polkit-guarded persistence. |
|
||||||
|
| `fenris import ⟨path⟩` | The idempotent single-transaction legacy import (§3.5). |
|
||||||
|
| `--device` | Rejected with a pointer to the configuration file. |
|
||||||
|
| `start`, `stop`, `run` | Rejected with one-line migration pointers — never aliased (an alias would silently change meaning). |
|
||||||
|
|
||||||
|
`fenris.sh` is retired: not shipped, removed from the repository; the README maps its five menu options to their successors. Headless administration has full parity: every TUI action has a CLI twin (pause, resume, collect-now, baseline set/clear, the status fact set) with identical outcomes and wording.
|
||||||
|
|
||||||
|
### 8.9 Freshness grading
|
||||||
|
|
||||||
|
Constants defined once, consumed by TUI and CLI alike; the grade derives from the **newest sample timestamp**, never a stored flag:
|
||||||
|
|
||||||
|
- **fresh** — newest sample within 2 × cadence + `AccuracySec` + 60 s (11.5 min at default cadence);
|
||||||
|
- **missed** — between that and 48 h (a contributing fact);
|
||||||
|
- **stale** — ≥ 48 h, matching the §6.7 evidence gate;
|
||||||
|
- **empty store** — *"no observations yet"* with an enable hint.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Failure and recovery
|
||||||
|
|
||||||
|
**ADR:** [0005](../adr/0005-failure-detection-and-recovery.md). **Criteria:** FL-1–FL-8.
|
||||||
|
|
||||||
|
The posture: **visible degradation, never fabrication.**
|
||||||
|
|
||||||
|
### 9.1 Write-boundary validation
|
||||||
|
|
||||||
|
The collector validates every row it would write against the store's domain invariants: hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count. A violating run **writes nothing**, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Store invariant: everything persisted is well-formed.
|
||||||
|
|
||||||
|
### 9.2 Reader defense
|
||||||
|
|
||||||
|
Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact. Under a single trusted writer they should never see one.
|
||||||
|
|
||||||
|
### 9.3 No backfill, ever
|
||||||
|
|
||||||
|
Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. Power-on-hours classification (§5.1) is the only inference admitted.
|
||||||
|
|
||||||
|
### 9.4 Store faults — degrade, never recreate over
|
||||||
|
|
||||||
|
An unreadable or corrupt database is a store fault: readers surface *"observation store unreadable"* with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and **never recreates or overwrites** an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside; the next run starts a fresh store; if the legacy import never completed, the still-present `history.jsonl` is re-imported (§3.5). No built-in destructive command exists.
|
||||||
|
|
||||||
|
### 9.5 Newer schema — symmetric refusal
|
||||||
|
|
||||||
|
The TUI and `status` detect a `user_version` newer than they understand and display *"observation store written by a newer Fenris — upgrade Fenris"* without partial interpretation, matching the collector's refusal (§3.6) and the forward-only upgrade rule (§10.2).
|
||||||
|
|
||||||
|
### 9.6 Repeated collector failures — flat cadence, no escalation
|
||||||
|
|
||||||
|
The timer's retry is the recovery path; the freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff, no notification machinery; a persistent failure reads as stale exactly like any other gap.
|
||||||
|
|
||||||
|
### 9.7 Drive-reported anomalies — facts, not alerts
|
||||||
|
|
||||||
|
`critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status` (§7.2); no alerting or notification surface exists. The projection is unaffected: endurance math consumes writes, not warnings.
|
||||||
|
|
||||||
|
### 9.8 Orphaned samples — the collector re-anchors observed fact
|
||||||
|
|
||||||
|
When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the **run moment**, never backdated. This records observed fact, not intent — only the sanctioned path records a `user_disabled` close (§8.5). Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Installation
|
||||||
|
|
||||||
|
**ADR:** [0004](../adr/0004-install-upgrade-removal-lifecycle.md). **Criteria:** IN-1–IN-10.
|
||||||
|
|
||||||
|
### 10.1 Install
|
||||||
|
|
||||||
|
`sudo make install`:
|
||||||
|
|
||||||
|
1. Builds a wheel from the checkout and installs it, with pinned dependencies (§10.5), into the dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only — after install, nothing references it.
|
||||||
|
2. Verifies `python3 ≥ 3.9` and `smartctl` presence, failing cleanly otherwise (never a runtime crash).
|
||||||
|
3. Creates `/var/lib/fenris` with root-written group-read permissions and the `fenris` read group; the database file is created lazily by the first write (§3.1).
|
||||||
|
4. Places units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/` — recording **every** placed file in an explicit manifest consumed by upgrade and uninstall (§10.6).
|
||||||
|
5. **Never enables or starts units.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle — `fenris monitor resume` or the first-run TUI prompt — enabling the timer and opening the first period in one step.
|
||||||
|
6. Detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs the idempotent single-transaction import (§3.5), and reports imported counts.
|
||||||
|
|
||||||
|
### 10.2 Upgrade
|
||||||
|
|
||||||
|
`sudo make upgrade` installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed **and** it is active — safe with `Persistent=no`), leaves timer state untouched, and **never kills an in-flight collection run**: a running oneshot finishes on its mapped interpreter; at worst one old-code run completes to the store. It then applies forward-only observation-store schema migrations (§3.6). `/var/lib/fenris` is never rebuilt.
|
||||||
|
|
||||||
|
### 10.3 Rollback
|
||||||
|
|
||||||
|
Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade does not exist.
|
||||||
|
|
||||||
|
### 10.4 Uninstall and purge
|
||||||
|
|
||||||
|
- `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. Journal entries age out naturally.
|
||||||
|
- `make purge` additionally removes configuration and store.
|
||||||
|
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
|
||||||
|
|
||||||
|
### 10.5 Dependencies
|
||||||
|
|
||||||
|
Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing. The acquisition path adds no Python dependency and no OS package beyond smartmontools (§2.5).
|
||||||
|
|
||||||
|
### 10.6 Placement manifest
|
||||||
|
|
||||||
|
Installed artifacts sit only at their fixed locations — units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/`, configuration at `/etc/fenris`, observation store under `/var/lib/fenris`, venv at `/opt/fenris`, wrapper at `/usr/local/bin/fenris` — and every placed file is recorded in the manifest (criterion IN-10).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Appendix A: Traceability matrix
|
||||||
|
|
||||||
|
Built as assembly's first step (assembly decision, recommendation 5). Two-way: every ADR section maps to at least one criterion ID; every criterion cites its ADR or ticket.
|
||||||
|
|
||||||
|
### A.1 ADR section → criteria
|
||||||
|
|
||||||
|
| ADR section | Criteria |
|
||||||
|
|---|---|
|
||||||
|
| 0001 §1 Substrate | ST-1, ST-2 |
|
||||||
|
| 0001 §2 Access | ST-1, CI-3 (/run) |
|
||||||
|
| 0001 §3 Entities (incl. #12/#14 amendments) | ST-3, ST-5, PR-13, ID-2, CI-3 (no stored projections) |
|
||||||
|
| 0001 §4 Day boundary | ST-4 |
|
||||||
|
| 0001 §5 Retention | ST-5 |
|
||||||
|
| 0001 §6 Migration | ST-6, ST-7, ST-8, ST-9, ST-10, ST-11 |
|
||||||
|
| 0001 §7 Projection inputs | ST-3, PR-14, CI-3 (one-key config) |
|
||||||
|
| 0001 §8 Versioning | ST-12, FL-5 |
|
||||||
|
| 0001 §9 Collector health | LC-10 |
|
||||||
|
| 0002 §1 One projection | PR-1 |
|
||||||
|
| 0002 §2 Rate/formulas/regime/scenario | PR-2, PR-17 |
|
||||||
|
| 0002 §3 Habit change | PR-3 |
|
||||||
|
| 0002 §4 Hour classification | PR-4 |
|
||||||
|
| 0002 §5 Denominator | PR-5 |
|
||||||
|
| 0002 §6 Minimum evidence | PR-6 |
|
||||||
|
| 0002 §7 Staleness | PR-7 |
|
||||||
|
| 0002 §8 Confidence table (incl. #15 amendment) | PR-8, PR-15, CI-1, CI-4 |
|
||||||
|
| 0002 §9 Segment breaks (incl. #15 amendment) | PR-9, PR-16 |
|
||||||
|
| 0002 §10 Implied eligibility | PR-10 |
|
||||||
|
| 0002 §11 Uncertainty | PR-11 |
|
||||||
|
| 0002 §12 Language | CI-4 |
|
||||||
|
| 0002 §13 Contract | PR-12 |
|
||||||
|
| 0003 §1 Units | LC-1, CI-3 (/run) |
|
||||||
|
| 0003 §2 Cadence | LC-2, LC-3 |
|
||||||
|
| 0003 §3 Configuration | LC-4, CI-3 |
|
||||||
|
| 0003 §4 Entry points | LC-5, IN-10 |
|
||||||
|
| 0003 §5 Sanctioned toggle | LC-6, CI-2, CI-3 |
|
||||||
|
| 0003 §6 Period rows | LC-7 |
|
||||||
|
| 0003 §7 On-demand collection | LC-8, CI-2, CI-3 |
|
||||||
|
| 0003 §8 TUI controls | TUI-1, TUI-2, TUI-4 |
|
||||||
|
| 0003 §9 CLI compatibility | LC-9, CI-2 |
|
||||||
|
| 0003 §10 Freshness constants | LC-10, CI-1, CI-2 |
|
||||||
|
| 0004 §1 Delivery | IN-1 |
|
||||||
|
| 0004 §2 Layout and manifest | IN-2, IN-10 |
|
||||||
|
| 0004 §3 Privilege | IN-3, CI-3 |
|
||||||
|
| 0004 §4 Dormant install | IN-3 |
|
||||||
|
| 0004 §5 Legacy import | IN-4 |
|
||||||
|
| 0004 §6 Upgrade | IN-5 |
|
||||||
|
| 0004 §7 Rollback | IN-6 |
|
||||||
|
| 0004 §8 Removal | IN-7 |
|
||||||
|
| 0004 §9 Dependencies | IN-8, AC-5 |
|
||||||
|
| 0004 §10 Scaffolding and floor | IN-9, TUI-3 |
|
||||||
|
| 0005 §1 Malformed observations | FL-1, FL-2 |
|
||||||
|
| 0005 §2 Missed observations | FL-3, CI-3 |
|
||||||
|
| 0005 §3 Store faults | FL-4 |
|
||||||
|
| 0005 §4 Newer schema | FL-5 |
|
||||||
|
| 0005 §5 Repeated failures | FL-6 |
|
||||||
|
| 0005 §6 Drive-reported anomalies | FL-7, CI-3 |
|
||||||
|
| 0005 §7 Orphaned samples | FL-8 |
|
||||||
|
| 0006 §1 Pin | AC-1 |
|
||||||
|
| 0006 §2 Hard pin, no fallback | AC-3 |
|
||||||
|
| 0006 §3 Normalization | AC-2 |
|
||||||
|
| 0006 §4 Segment metadata sourcing | AC-4 |
|
||||||
|
| 0006 §5 Prerequisites | AC-5 |
|
||||||
|
|
||||||
|
### A.2 Ticket decisions → criteria
|
||||||
|
|
||||||
|
| Ticket | Criteria |
|
||||||
|
|---|---|
|
||||||
|
| [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6) | TUI-3 |
|
||||||
|
| [Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3) | TUI-1, TUI-2, TUI-4 |
|
||||||
|
| [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11) | ID-1, ID-3 |
|
||||||
|
| [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) | PR-13, PR-14, ST-3 |
|
||||||
|
| [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14) | ID-2 |
|
||||||
|
| [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15) | PR-15, PR-16, ID-4 |
|
||||||
|
| [Define cross-cutting acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/issues/13) | the register itself |
|
||||||
|
| [Assemble the implementation-ready specification](https://git.bongbetic.com/xavierk/Fenris/issues/17) / [Write the Fenris redesign specification and close the map](https://git.bongbetic.com/xavierk/Fenris/issues/19) | this document and this appendix |
|
||||||
|
|
||||||
|
### A.3 Criterion → source
|
||||||
|
|
||||||
|
Every criterion carries its citation inline in the [register](acceptance-criteria.md): CI-1–CI-4 (ADR 0002 §§6–8, 0003 §§1–10, 0001 §§2–3/8, 0005 §§2/4–6); ST-1–ST-12 (ADR 0001, with ST-3 amended by tickets #12/#14); LC-1–LC-10 (ADR 0003); PR-1–PR-12 (ADR 0002), PR-13–PR-14 (ticket #12), PR-15–PR-16 (ticket #15, ADR 0002 §§8–9 as amended), PR-17 (ADR 0002 §2); ID-1–ID-4 (tickets #11/#14/#15, ADR 0001 §3 as amended); TUI-1–TUI-4 (tickets #3/#6, ADR 0003 §§8/10, ADR 0004 §10); FL-1–FL-8 (ADR 0005); IN-1–IN-10 (ADR 0004, with IN-10 also citing ADR 0003 §4); AC-1–AC-5 (ADR 0006).
|
||||||
|
|
||||||
|
### A.4 Assembly result
|
||||||
|
|
||||||
|
- **Every ADR 0001–0006 section maps to at least one criterion** — the A.1 table is complete; no orphan sections.
|
||||||
|
- **Every criterion cites its ADR or ticket** — verified in the register; no orphan criteria.
|
||||||
|
- **Decided-but-uncitered gaps found and filled inline in the register during assembly:** CI-3 bullet (no `/run` coordination surface; ADR 0003 §1, ADR 0001 §2), PR-17 (projection arithmetic; ADR 0002 §2), TUI-4 (normative Panes layout and bindings; ticket #3), IN-10 (fixed artifact placement; ADR 0004 §2, ADR 0003 §4).
|
||||||
|
- **No genuinely undecided behavior remained** — no blocking ticket was raised.
|
||||||
|
- **One reconciliation:** ADR 0004 §6's "`schema_version` table" wording resolves to ADR 0001 §8's `PRAGMA user_version` as the single version authority (§3.6); criterion ST-12 already fixed the mechanism.
|
||||||
Reference in New Issue
Block a user