Compare commits

...
Author SHA1 Message Date
xavierk 7f006c7df7 feat(status): read-only CLI status command (issue #27)
Implement fenris status as the read-only CLI twin of the TUI,
composing from the observation store and allow-listed systemctl
properties per spec section 8.8.

New module src/fenris/status.py:
- Freshness grading with shared constants (section 8.9, LC-10)
- Configuration error from direct config reads (section 8.3, LC-4)
- Store fault / newer-schema exact phrases (section 9.4-9.5, FL-4/FL-5)
- Drive anomalies as ordinary facts (section 9.7, FL-7)
- Four separate service facts (section 7.3, LC-9, CI-2)
- Projection recomputed on read, never stored (section 6.10)
- Retired command rejection with migration pointers (section 8.8)
- Six disclosures via --disclosures flag (section 6.11, CI-4)

Updated fenris.py:
- Replaced old cmd_status with new status module integration
- Added retired command handlers (start/stop/run)
- Added global --device flag rejection

Tests: 43 new, 182 total passing, zero regressions.
Closes #27.
2026-09-01 23:49:23 +05:30
xavierk b99ebfe9dd fix(projection): regime dynamics, warming gate, habit change detection, horizon coverage
Fixes #26.

Changes:
- Fix warming gate: check total_days < 14 OR days_below_coverage > 2
- Rewrite _detect_habit_change: correct consecutive-day scanning
- Fix _compute_horizon_rate: require history spans full horizon

Tests: 29 new tests covering PR-2,3,6,7,9,15,16. 139 total passing.
2026-09-01 23:32:00 +05:30
xavierk 7802a72606 feat: implement projection core pure function (closes #25) 2026-09-01 23:16:33 +05:30
xavierk d005412c0d feat: implement legacy history migration (closes #24) 2026-09-01 23:09:05 +05:30
xavierk 6a1841c447 feat: segment observation history by controller identity (closes #23) 2026-09-01 23:03:42 +05:30
xavierk 7d219c4697 feat(hour/day derivation): hour classification, monitoring periods, day aggregates, pruning
Hour classification (PR-4):
- Powered-off: POH delta < 90% of wall-clock span
- Active: DUW delta >= 256 MiB
- Idle: powered on + sampled + below active threshold
- Unknown: unsampled without POH evidence
- Four splits sum to exactly wall_clock_seconds
- Disabled time is never an hour state

Monitoring periods (FL-8):
- ensure_period_open: opens period at run moment if none exists
- close_period: closes with end cause
- is_inside_period: checks timestamp against period bounds
- Never backdated; wall-clock outside periods excluded from denominator

Day aggregates (ST-4, PR-5):
- Derived monotonically from hour rows
- UTC-bounded; no 23/25-hour days
- Coverage: known seconds / period wall-clock
- Gap hours inside periods contribute unknown seconds
- Hours outside periods excluded entirely
- No absent hour interpolated/estimated/fabricated (FL-3)

Raw sample pruning (ST-5):
- Prunes samples older than 14 days
- Hour observations and day aggregates retained indefinitely

Closes #22
2026-09-01 22:58:03 +05:30
xavierk 2217b00ff6 feat: collector tracer bullet (#21)\n\nSmartctl acquisition with validation\nSysfs controller identity acquisition\nIdentity normalization (strip, no case fold, blank handling)\nSQLite store in WAL mode with six entities\nSchema versioning via PRAGMA user_version\nInvariant validation (negative bytes, etc.)\n17 passing tests across two test files\n\nCloses #21 2026-09-01 22:33:26 +05:30
xavierk 566d1c81b1 docs(spec): fenris-redesign.md — implementation-ready specification from ADRs 0001-0006 with two-way traceability matrix 2026-08-31 23:43:22 +05:30
xavierk ff51be2d8a docs(spec): fill assembly traceability gaps — PR-17 projection arithmetic, TUI-4 Panes layout, IN-10 artifact placement, CI-3 /run prohibition 2026-08-31 23:43:22 +05:30
xavierk b2ef1308b0 docs(adr): 0006 collector acquisition path — smartctl counters + sysfs identity, hard pin, no partial samples; SLOT-B filled (AC-1–5) 2026-08-31 23:02:25 +05:30
xavierk 580fab26be docs(adr): amend 0002 — blank identity key caps confidence at Limited, blank-key change semantics; SLOT-A filled (PR-15/16, ID-4); glossary term degraded identity 2026-08-31 22:30:40 +05:30
xavierk bf83ac5481 docs(spec): acceptance criteria for the redesign — subsystem gates, A/P/M evidence classes, ADR traceability; state-matrix and parity gates, prohibition set 2026-08-31 22:13:20 +05:30
xavierk 305bdd4779 docs(adr): amend 0001 — controller-segment metadata snapshot: normalized identity diagnostics, vid/ssvid/transport, frozen at open, nullable 2026-08-31 21:57:54 +05:30
xavierk 4d332bef53 docs(adr): amend 0001 + 0003 — baseline provenance and validation; glossary terms for verified and unverified override 2026-08-31 21:02:09 +05:30
xavierk d45d931a0d docs(adr): 0005 failure and recovery — refuse bad writes, never backfill, degrade store faults; glossary term for store fault 2026-08-31 20:14:23 +05:30
xavierk ffa6f22777 docs(adr): 0004 installation lifecycle — Makefile venv install, dormant install, sanctioned teardown 2026-08-31 18:58:05 +05:30
xavierk de1e8c753b docs(agents): issue-tracker and domain guidance for agents; ignore .pi session state 2026-08-31 17:25:10 +05:30
xavierk 26a6703152 docs(adr): 0003 service lifecycle — timer-driven collector, sanctioned control helper; glossary terms for collection run, deliberate disable 2026-08-31 16:14:52 +05:30
xavierk 43467f8957 docs(adr): 0002 projection model — sustained regime, categorical confidence; glossary terms for regime, habit change, scenario range, coverage 2026-08-31 15:38:41 +05:30
xavierk a63da74e44 docs(adr): 0001 observation store in SQLite; glossary terms for store entities 2026-08-31 14:48:14 +05:30
36 changed files with 7444 additions and 40 deletions
+1
View File
@@ -1,4 +1,5 @@
__pycache__/ __pycache__/
.pi/
*.pyc *.pyc
.commandcode/ .commandcode/
data/fenris.pid data/fenris.pid
+9
View File
@@ -0,0 +1,9 @@
## Agent skills
### Issue tracker
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
### Domain docs
This is a single-context repository. See `docs/agents/domain.md`.
+85
View File
@@ -0,0 +1,85 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Store fault**:
The condition where the observation store is present but cannot be read or trusted — unreadable, corrupt, or written by a newer Fenris — degrading every view that depends on it rather than crashing or guessing.
_Avoid_: Database error, corruption, broken data
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Degraded identity**:
The condition where a controller segment's identity key is blank because no identifier rung produced a value; replacement detection then relies on write-counter continuity alone, and projection confidence is capped.
_Avoid_: Identity error, unknown device, virtual drive
**Endurance baseline**:
The write-endurance value a projection consumes, chosen by precedence: a verified override when one exists, otherwise an unverified override, otherwise a coarse implied baseline derived from vendor wear — each labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Verified override**:
A rated-TBW override with complete provenance whose applicability to the detected drive was confirmed by machine match or explicit user attestation; the strongest endurance baseline.
_Avoid_: Confirmed TBW, trusted value
**Unverified override**:
A rated-TBW override knowingly stored with incomplete provenance; always presented as user-supplied, never as verified.
_Avoid_: Forced entry, fallback baseline
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
+46
View File
@@ -0,0 +1,46 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
Amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12): the `endurance_baseline` field set and validation contract — derived verification, entry-time unprivileged sysfs validation, and read-time controller-segment applicability.
Amended by [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14): the `controller_segments` metadata snapshot — normalized identity diagnostics plus `vid`/`ssvid`/`transport` and a degraded flag, frozen at segment open, all nullable.
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment. Each row carries a metadata snapshot frozen when the segment opens and immutable thereafter ([Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14)): normalized `subnqn`, `sn`, `mn`, `fr`, plus `vid`, `ssvid`, `transport`, and an `identity_degraded` flag — human diagnostics, never key components (`cntlid` excluded: it distinguishes controllers within one subsystem, out of scope for a single-drive monitor). `fr` may go stale after a mid-segment firmware update; counter discontinuities belong to the DUW-monotonic axis. Every metadata column is nullable — legacy-imported segments carry `mn` with NULLs, degraded segments whatever was observed — so incompleteness stays explicit.
- `endurance_baseline` — one active row, replaced on edit ([Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12)): the rated-TBW value in bytes (`E_rated = entered_TBW × 10¹²`) plus mandatory provenance — source URL, document revision, entry date, model string, nominal capacity — and frozen validation facts (detected model, detected capacity bytes, `validated_by` `machine`/`user`, `validated_at`). Verification is derived at read — complete provenance and a drive match (machine or attested), never a stored boolean; incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier. Entry validation is an unprivileged live sysfs read of the configured device (normalized model containment with an interactive confirm recorded as `validated_by = user`; capacity within ±1%); at projection time applicability is a model match against the current controller segment, and a mismatch is retained — never auto-deleted — leaving the projection Unavailable.
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -0,0 +1,52 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
Amended by [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15): a blank (degraded) identity key caps confidence at Limited evidence, and identity-change semantics extend verbatim to blank keys.
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old **and** the current controller segment's identity key is not degraded.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- **Degraded identity** ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15)): a controller segment whose identity key is blank — every rung of the key ladder empty — is identity-degraded. Supported is unreachable while the current segment is degraded, because a blank key cannot detect a replacement; the fact "controller identity unavailable — replacement detection relies on write-counter continuity only" renders with every state, and the cap combines idempotently with the staleness drop (both land at Limited). Ephemeral markers (model "Linux", non-pcie transport) are segment metadata, never confidence facts.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive. Degraded keys get no special casing ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15)): a blank-key segment is marked identity-degraded and segmented by DUW monotonicity alone, any visible change of the recorded key — including to or from blank — is a controller-identity change and quarantines, and equal blank keys continue the segment. Since even a degraded→healthy transition quarantines, the projection window only ever spans segments sharing one key, so a degraded current segment needs no cross-segment propagation rule.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts — including the degraded-identity fact when the current segment's key is blank — the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -0,0 +1,33 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
Amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12): the helper gains a `baseline` verb that persists the CLI-validated endurance-baseline row — same fixed-operation, polkit-mediated pattern.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger, monitoring-period bookkeeping, and `baseline set`/`baseline clear` persistence for the CLI-validated endurance baseline; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
@@ -0,0 +1,31 @@
# 4. Installation lifecycle: Makefile-delivered venv, dormant install, sanctioned teardown
## Status
Accepted — resolves [Define installation, upgrade, and removal behavior](https://git.bongbetic.com/xavierk/Fenris/issues/9) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris runs today from the source checkout (`fenris.py`, `fenris.sh`, `data/`): code, state, and control all live relative to wherever the checkout sits. The redesign fixes system artifacts — helpers in `/usr/libexec/fenris` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), the observation store at `/var/lib/fenris/observations.db` ([ADR 0001](0001-observation-store-sqlite.md)), configuration at `/etc/fenris/fenris.conf` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), Textual on Python 3.9+ ([framework decision](https://git.bongbetic.com/xavierk/Fenris/issues/6)) — but nothing says how those artifacts are delivered, upgraded, or removed, or what happens when the checkout moves or disappears.
## Decision
1. **Delivery.** `sudo make install` builds a wheel from the checkout and installs it, with pinned dependencies, into a dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only: after install, nothing references it.
2. **Layout and manifest.** Units in `/etc/systemd/system` (`fenris-collect.{timer,service}`, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)); helpers in `/usr/libexec/fenris`; polkit policy in `/usr/share/polkit-1/actions/`; configuration and observation store in their ADR-fixed locations. The installer records every file it places in an explicit manifest consumed by upgrade and uninstall.
3. **Privilege.** One root installer (`sudo make install`); at runtime, elevation is exclusively polkit (`auth_admin`, `fenris-monitor` only, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)). The installer never enables or starts units.
4. **Dormant install.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle (`fenris monitor resume [--now]`, or the first-run TUI prompt), which enables the timer and opens the first period in one step.
5. **Legacy import.** The installer detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs [ADR 0001](0001-observation-store-sqlite.md)'s idempotent single-transaction import, and reports imported counts — existing observations never depend on checkout survival. `fenris import <path>` remains available for later finds.
6. **Upgrade.** `sudo make upgrade` builds and installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed and it is active — safe with `Persistent=no`), leaves timer state untouched, and never kills an in-flight collection run: a running oneshot finishes on its mapped interpreter, so at worst one old-code run completes to the store and the next run uses the new code. It then applies forward-only observation-store schema migrations governed by a `schema_version` table. `/var/lib/fenris` is never rebuilt.
7. **Rollback.** Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade is explicitly unsupported.
8. **Removal.** `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. `make purge` additionally removes configuration and store. Journal entries age out naturally.
9. **Dependencies.** Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing.
10. **Scaffolding and floor.** The installer creates `/var/lib/fenris` with [ADR 0001](0001-observation-store-sqlite.md)'s root-written group-read permissions and verifies `python3 ≥ 3.9`, failing cleanly otherwise — the Textual contingency becomes an install-time gate rather than a runtime crash. The database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet.
## Consequences
- Installed Fenris survives checkout deletion; the checkout is only where builds happen.
- Teardown preserves monitoring-period semantics: deliberate removal excludes the uninstalled span from the usage habit instead of leaving it as unknown-inside.
- Installs are reproducible; dependency drift cannot ride in on an upgrade.
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
- Rollback support is exactly one generation deep, no further.
- The README documents install, upgrade, uninstall/purge, and legacy import alongside [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s menu-successor mapping.
@@ -0,0 +1,27 @@
# 5. Failure and recovery: visible degradation, never fabrication
## Status
Accepted — resolves [Define failure and recovery behavior](https://git.bongbetic.com/xavierk/Fenris/issues/10) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The observation store ([ADR 0001](0001-observation-store-sqlite.md)) and the service lifecycle ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) settled single-writer transactions, bounded single failed runs, freshness grading (fresh / missed / stale), and absent-hour classification through power-on-hours evidence. Left open by the [failure ticket](https://git.bongbetic.com/xavierk/Fenris/issues/10): behavior per failure class — malformed observations inside the store, missed observations, store faults (unreadable, corrupt, or newer-schema database), and repeated collector failures — and how the habit record re-anchors after the store itself is lost.
## Decision
1. **Malformed observations — refuse at the write boundary.** The collector validates every row it would write against the store's domain invariants (hour seconds sum to 3600, non-negative DUW delta within a controller segment, coverage consistent with sample count). A violating run writes nothing for that run, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact, but under a single trusted writer they should never see one. Store invariant: everything persisted is well-formed.
2. **Missed observations — never backfill.** Fenris never interpolates, estimates, or fabricates an hour. Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s power-on-hours classification is the only inference admitted.
3. **Store faults — degrade, never recreate over.** An unreadable or corrupt database is a store fault: readers surface a "observation store unreadable" fact with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside, the next run starts a fresh store, and if the legacy import never completed, the still-present `history.jsonl` is re-imported. No built-in destructive command exists.
4. **Newer schema — readers refuse symmetrically.** The TUI and `status` detect a `user_version` newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal in [ADR 0001](0001-observation-store-sqlite.md) and the forward-only upgrade rule of [ADR 0004](0004-install-upgrade-removal-lifecycle.md).
5. **Repeated collector failures — flat cadence, no escalation.** The timer's retry is the recovery path; the settled freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff and no notification machinery; a persistent failure reads as stale exactly like any other gap.
6. **Drive-reported anomalies — facts, not alerts.** `critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status`; no alerting or notification surface exists. Fenris observes and projects; it does not alarm. The projection is unaffected: endurance math consumes writes, not warnings.
7. **Orphaned samples — the collector re-anchors observed fact.** When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent: only the sanctioned path of [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) records a `user_disabled` close. Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
## Consequences
- Validation lives at one boundary — the collector — so the store's contract is "everything in it is well-formed" and readers only defend against the impossible.
- No synthetic data can ever enter the habit record; confidence categories can be trusted to reflect real evidence.
- Store-fault recovery can lose history; the mitigation is the one-file backup story of [ADR 0001](0001-observation-store-sqlite.md), kept human-sanctioned so loss is never silent.
- Period bookkeeping splits by epistemics: the helper records intent, the collector records observed fact.
- Fenris stays fully local and silent: no notification, escalation, or alerting machinery anywhere.
@@ -0,0 +1,28 @@
# 6. Collector acquisition path: smartctl counters, sysfs identity
## Status
Accepted — resolves [Choose the collector's NVMe acquisition path](https://git.bongbetic.com/xavierk/Fenris/issues/16) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The collector ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) must acquire SMART/Health counters, thermal evidence, and controller identity each run. The [controller-identity research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/controller-identity/docs/research/controller-identity.md) fixed the identity key to the normalized, kernel-exposed subsystem NQN and warned that normalization must be specified once and applied at write time — or a collector implementation change can split a drive's own history. [ADR 0004](0004-install-upgrade-removal-lifecycle.md) pins exact Python dependencies in a dedicated venv, and the [segment-metadata decision](https://git.bongbetic.com/xavierk/Fenris/issues/14) froze nullable `vid`/`ssvid`/`transport` alongside the identity fields. Three first-party paths were candidates: the official libnvme Python bindings (SWIG; sysfs-backed attribute getters delivering normalized values), `nvme` CLI JSON output, and the incumbent `smartctl -j` plus sysfs reads.
## Decision
1. **Pin.** Every collection run acquires counters and thermal evidence solely from `smartctl -a -j <device>` and controller identity (`subnqn`, `sn`, `mn`, `fr`, `transport`) solely from sysfs (`/sys/class/nvme/<ctrl>/`). No other acquisition path exists anywhere in the codebase.
2. **Hard pin, no fallback.** Any acquisition failure — missing binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the whole collection run; [ADR 0005](0005-failure-detection-and-recovery.md)'s flat retry and freshness grading absorb the miss. A partial sample (identity without counters, or counters without identity) is never written: a transient read failure must not push a healthy drive down the degraded-identity path.
3. **Normalization once, at write time.** One collector-side function normalizes every identity field: trailing spaces and newlines stripped, no case folding, empty-after-strip stored blank. `smartctl` counter and thermal fields are consumed as-is (smartmontools already trims the strings it copies). Padded and unpadded renderings of the same field therefore yield byte-identical stored values.
4. **Segment metadata sourcing.** `transport` comes from the NVMe class sysfs directory; `vid`/`ssvid` from the PCI node (`/sys/class/nvme/<ctrl>/device/{vendor,subsystem_vendor}`) when present, null otherwise — metadata only, never key components.
5. **Prerequisites.** `make install` verifies `smartctl` is present and fails cleanly otherwise. The acquisition path adds no Python dependency and no OS package beyond smartmontools; the [ADR 0004](0004-install-upgrade-removal-lifecycle.md) lockfile is untouched.
## Considered options
- **libnvme Python bindings** — the purest API and natively-normalized getters, but the SWIG module is not on PyPI: entering the venv requires the distro's `python3-libnvme` through `--system-site-packages` or a from-source build, coupling the exact-lockfile venv to the system Python and the distro's shipping choices. Rejected on dependency weight for one privileged five-minute oneshot.
- **`nvme` CLI JSON** — one binary covers counters and identity, but it adds an OS package for what smartmontools already provides, emits untrimmed strings, and reports `subnqn` from Identify data rather than the kernel: when a controller reports an empty NQN the kernel synthesizes one for sysfs while `id-ctrl` JSON omits the field, so the identity ladder would drop a rung depending on the drive. Rejected on packaging and identity-key consistency.
## Consequences
- The venv stays pure-Python; the two acquisition channels per run (subprocess JSON plus sysfs reads) hide behind one acquisition function, gated by acceptance criteria AC-1–AC-5.
- Identity is read from exactly the source the identity key names; libnvme's getters wrap the same sysfs attributes, so the values agree byte-for-byte where both exist.
- Switching acquisition path later is history-sensitive: a future path must deliver byte-identical normalized identity values, or the change itself forces a controller-segment boundary.
+26
View File
@@ -0,0 +1,26 @@
# Domain Docs
How engineering skills should consume this repository’s domain documentation.
## Layout
This is a single-context repository:
```text
/
├── CONTEXT.md
├── docs/adr/
└── ...
```
## Before exploring
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
## Use the glossary’s vocabulary
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
## Flag ADR conflicts
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
+85
View File
@@ -0,0 +1,85 @@
# Issue tracker: Gitea
Issues for this repository live in Gitea at:
https://git.bongbetic.com/xavierk/Fenris/issues
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
## General operations
- List: `tea issues list`
- Read: `tea issues <index> --comments`
- Create: `tea issues create --title "<title>" --description "<body>"`
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
- Assign: `tea issues edit <index> --add-assignees "<username>"`
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
- Comment: `tea comments add <index> --description "<comment>"`
- Close: `tea issues close <index>`
- Reopen: `tea issues reopen <index>`
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
## When a skill says “publish to the issue tracker”
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
## When a skill says “fetch the relevant ticket”
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
## Wayfinding operations
Wayfinder maps and decision tickets are Gitea issues.
### Map and ticket grouping
- A map has the label `wayfinder:map`.
- Create one milestone named `Wayfinder: <map title>` for the effort.
- Assign the map and all its tickets to that milestone.
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
### Blocking
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
```bash
tea api -X POST \
repos/{owner}/{repo}/issues/<blocked>/dependencies \
-F index=<blocker> \
-f owner=xavierk \
-f repo=Fenris
```
List blockers:
```bash
tea api repos/{owner}/{repo}/issues/<index>/dependencies
```
Remove the relationship with the same payload and `-X DELETE`.
### Frontier
List open issues in the map’s milestone. Exclude:
- the issue labelled `wayfinder:map`
- assigned tickets, because assignment is the claim
- tickets whose dependency query returns any open issue
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
### Claim
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
### Resolve
1. Add the answer as a resolution comment.
2. Close the ticket.
3. Re-fetch the map immediately before editing it.
4. Append a linked one-line context pointer to `Decisions so far`.
5. Create newly visible tickets, then wire dependencies in a second pass.
+126
View File
@@ -0,0 +1,126 @@
# Acceptance criteria: the Fenris redesign
Status: Accepted — resolves [Define cross-cutting acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/issues/13) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). These criteria are the accepted definition of done for the finished redesign; the implementation-ready specification assembles them with ADRs 0001–0006 at handoff.
## Framework
- **Canonical term**: *acceptance criterion* — one testable behavioral statement. "Behavioral gate" is avoided as a synonym. Wording follows the repository glossary (`CONTEXT.md`).
- **Evidence classes** — every criterion carries exactly one:
- **A** — automated test (unit/integration, fixture-driven).
- **P** — scripted system probe on a host with systemd, polkit, and the configured NVMe device.
- **M** — manual checklist, reserved for interactions a fixture cannot capture (live polkit agent prompts, TUI keyboard feel).
- Whatever can be automated must be; **M** only where automation cannot reach.
- **Traceability-only**: every number and behavior cites the ADR or ticket that fixed it. Nothing undecided enters here; new demands become new tickets, never criteria.
- **Organization**: criteria are grouped by subsystem, with cross-cutting invariants spanning them. Coverage spans all decided areas — observation store and migration, collector lifecycle and privileges, projection and confidence, controller identity, the Panes TUI, failure paths, and installation lifecycle.
- **Placeholders**: none remain. SLOT-B was filled by [Choose the collector's NVMe acquisition path](https://git.bongbetic.com/xavierk/Fenris/issues/16) as AC-1–AC-5; SLOT-A was filled by [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15) as PR-15, PR-16, and ID-4.
- **Test-plan boundary**: Given/When/Then test specs are derived by the implementer at implementation time. This effort produces criteria only.
## Cross-cutting invariants
- **CI-1** (A; ADR 0002 §§6–8, ADR 0003 §10) *Exhaustive state matrix*: from a synthetic observation store, the TUI and `fenris status` render every realizable combination of confidence state (Unavailable, Limited, Supported) × freshness grade (fresh, missed, stale, empty store) × endurance-baseline tier (verified override, unverified override, implied, none) exactly as the ADR 0002 rule table and ADR 0003 freshness constants dictate — headline number present only when the rules allow it, contributing facts always, never a percentage.
- **CI-2** (P/A; ADR 0003 §§4–9) *TUI/CLI parity*: every TUI action has a CLI twin — pause (`fenris monitor pause`), resume (`fenris monitor resume`), collect-now (`fenris sample`), baseline set/clear, and the status fact set — with identical outcomes and wording.
- **CI-3** *Prohibition set* (A = test, P = probe; each cites its clause):
- No code path outside `fenris-collect` interrogates the device (ADR 0003 §7).
- No `/run/fenris` coordination surface or export layer exists anywhere; state lives in the observation store and coordination in systemd (ADR 0003 §1, ADR 0001 §2).
- Polkit authorizes exactly one binary, `fenris-monitor`, under `com.bongbetic.fenris.monitor` `auth_admin` (ADR 0003 §5, ADR 0004 §3).
- No absent hour is ever interpolated, estimated, or fabricated (ADR 0005 §2).
- No alerting, notification, or escalation machinery exists anywhere (ADR 0005 §§5–6).
- `/etc/fenris/fenris.conf` holds exactly one key — the device selector (ADR 0003 §3).
- No synthetic or capacity-derived baseline is ever created, including for legacy history (ADR 0002, Consequences).
- Readers never partially interpret a newer-schema store (ADR 0001 §8, ADR 0005 §4).
- Projections are never stored; always recomputed on read (ADR 0001 §3, ADR 0002 §13).
- **CI-4** (A; ADR 0002 §§11–12) *User-facing language*: TUI and status render the endurance research's required wording and six disclosures as adopted; zero-rate and unavailable cases use their exact phrasing; the scenario range is the only spread shown anywhere.
## Observation store and legacy migration (ADR 0001)
- **ST-1** (P) One SQLite database in WAL mode at `/var/lib/fenris/observations.db`, root-owned and group-readable through the `fenris` read group; the TUI opens it read-only.
- **ST-2** (A) An unprivileged reader querying during a collector write sees a consistent snapshot.
- **ST-3** (A) The schema carries `samples`, `hour_observations`, `day_aggregates`, `monitoring_periods`, `controller_segments`, and `endurance_baseline` with the ADR 0001 column sets as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) and [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14).
- **ST-4** (A) Hours and days are UTC-bounded; day derivation from hours is monotonic; no 23- or 25-hour days exist.
- **ST-5** (A) Raw samples are pruned opportunistically to 14 days; hour observations and day aggregates are retained indefinitely.
- **ST-6** (P) Legacy import is one transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
- **ST-7** (A) Import is idempotent: a second run no-ops on the legacy-import marker.
- **ST-8** (A) Legacy files are renamed `*.migrated` only after commit and never deleted.
- **ST-9** (A) Malformed legacy lines are quarantined with a logged count, never silently dropped.
- **ST-10** (A) `hourly.jsonl` is never trusted: mismatches against derived data are diffed and logged.
- **ST-11** (A) Migration opens one implicit monitoring period at the first legacy sample, closed `end_cause = migrated` at the migration moment; pre-migration hours carry an unknown activity split except directly evidenced facts.
- **ST-12** (A) Schema versioning: `PRAGMA user_version` with ordered, per-step transactional migrations; the collector refuses an unknown newer version.
## Collector lifecycle and privilege boundaries (ADR 0003)
- **LC-1** (P) Exactly two system units exist — `fenris-collect.timer` (`timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`); the TUI and CLI are ordinary unprivileged processes and never units.
- **LC-2** (P) Timer defaults ship as `OnBootSec=2min`, `OnUnitInactiveSec=5min`, `AccuracySec=30s`, `Persistent=no`, `TimeoutStartSec=90s`; cadence changes are documented drop-ins and no interval key exists in configuration.
- **LC-3** (P) A hung device interrogation fails visibly within `TimeoutStartSec=90s` as a bounded failed run retried next interval.
- **LC-4** (A/P) `/etc/fenris/fenris.conf` holds exactly the device selector (stable `/dev/disk/by-id/…` path; raw nodes warned), re-read every run; an invalid selector is a bounded failed run surfaced as `configuration error: <reason>` in `status` and the TUI.
- **LC-5** (P) Two privileged binaries ship at `/usr/libexec/fenris/fenris-collect` and `/usr/libexec/fenris/fenris-monitor`; the unprivileged `fenris` wrapper opens the TUI with no arguments.
- **LC-6** (P+M) Pause = `fenris-monitor disable --now` asks for confirmation; Resume = `enable --now` does not; both perform the systemctl operation and period bookkeeping in one step under polkit `com.bongbetic.fenris.monitor` (`auth_admin`), failing cleanly with the printed root equivalent where no polkit agent exists. (M covers the live agent prompt.)
- **LC-7** (A) The period-row idempotent matrix of ADR 0003 §6 holds exactly: first-ever enable opens; resume with an open period changes nothing; resume without one opens anew; pause with an open period closes `user_disabled`; pause otherwise no-ops; a raw systemctl stop/disable never records `user_disabled`.
- **LC-8** (P) `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, block until exit, and report the outcome (freshness line or journal hint) synchronously; the TUI never samples in-process.
- **LC-9** (A/P) CLI compatibility: `status` is a read-only composition (projection facts, enabled/active, last collect outcome, `journalctl` hint on failure or staleness) that never auto-samples and never prompts; `sample` is retained via the helper; `--device` is rejected with a pointer to the configuration file; `start`, `stop`, and `run` are rejected with one-line migration pointers; `fenris.sh` is not shipped and is removed from the repository; the README maps its five menu options to successors.
- **LC-10** (A) Freshness constants are defined once and shared by TUI and CLI: fresh = newest sample within 2× cadence + `AccuracySec` + 60 s; missed between that and 48 h; stale ≥ 48 h; an empty store reads "no observations yet" with an enable hint; freshness derives from the newest sample timestamp, never a stored health flag.
## Projection contract and confidence (ADR 0002)
- **PR-1** (A) Exactly one projection, from the precedence-chosen baseline (verified override → unverified override → implied → unavailable); Percentage Used renders as a vendor-wear context line, with a note when it disagrees with the observed write rate by more than a factor of 2; the PU-slope regression and `capacity × 600` synthesis are gone.
- **PR-2** (A) The headline rate is the sustained-regime rate (regime DUW bytes ÷ in-period wall-clock seconds), default regime = full history capped at 90 days; the 7/28/90-day scenario range is computed independently and shows only covered horizons, with no placeholders.
- **PR-3** (A) Habit change: trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean for 3 consecutive days starts a new regime at the first divergence day, adopted automatically and labeled "usage habit changed N days ago"; a regime younger than 7 days caps confidence at Limited evidence.
- **PR-4** (A) Hour classification uses the named constants: powered-off below 90% of power-on-hours span; active at ≥ 256 MiB DUW; idle below it while powered on and sampled; unknown otherwise; disabled time is wall-clock outside monitoring periods, never an hour state.
- **PR-5** (A) The denominator is wall-clock seconds inside monitoring periods including powered-off and unknown time; disabled periods are excluded from numerator and denominator; unexplained gaps keep the aggregate counter delta, remain as unknown seconds, and reduce coverage.
- **PR-6** (A) Warming up until 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage; the projection still renders with its facts while warming; every Unavailable condition renders no lifespan number.
- **PR-7** (A) A newest day aggregate older than 48 h drops confidence one level and is shown as a contributing fact.
- **PR-8** (A) The confidence rule table of ADR 0002 §8 holds verbatim, rendering state plus contributing facts and never a percentage.
- **PR-9** (A) Segment breaks: a DUW decrease with unchanged identity keeps prior day aggregates as habit evidence with the projection Unavailable until re-warm; a controller-identity change quarantines prior history from projection entirely.
- **PR-10** (A) The implied baseline is eligible only after ≥ 2 Percentage-Used increments within the current controller segment; until then, Unavailable with "vendor wear estimate too coarse to imply endurance".
- **PR-11** (A) Zero rate renders "no finite projection from this history" — never infinity or zero; no statistical confidence interval appears anywhere.
- **PR-12** (A) The projection contract hands the TUI exactly: confidence state, contributing facts, headline remaining time when one exists, scenario range, Percentage-Used context line, disclosure text — recomputed on read, never stored.
- **PR-13** (A) Baseline provenance and validation per [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12): mandatory provenance (URL, revision, entry date, model, nominal capacity); one active row replaced on edit; verification derived at read (machine match or recorded attestation), never a stored boolean; incomplete provenance stores only behind explicit acknowledgment as the unverified tier; entry-time unprivileged sysfs validation (normalized model containment; capacity within ±1%; interactive confirm recorded as `validated_by = user`); read-time applicability is a model match against the current controller segment, with a mismatch retained — never auto-deleted — leaving the projection Unavailable.
- **PR-14** (P) `baseline set` / `baseline clear` persist through the polkit-guarded `fenris-monitor` verb after CLI-side validation.
- **PR-15** (A) An identity-degraded controller segment (blank identity key — every rung of the key ladder empty) caps projection confidence at Limited evidence, with the contributing fact "controller identity unavailable — replacement detection relies on write-counter continuity only" rendered in every state; the cap combines idempotently with the 48-hour staleness drop, and ephemeral markers (model "Linux", non-pcie transport) never render as confidence facts ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15); ADR 0002 §§8–9 as amended).
- **PR-16** (A) Identity-change semantics extend to blank keys verbatim: any visible change of the recorded identity key — including to or from a blank key — quarantines prior history from projection as a controller-identity change, while equal blank keys continue the segment segmented by DUW monotonicity alone (ADR 0002 §9 as amended).
- **PR-17** (A) Projection arithmetic is exactly `E_rated = entered_TBW × 10¹²` bytes, `E_implied = 100 · W_t / p` computed only for `1 ≤ p ≤ 254` (Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable), and `projected = max(E_baseline − W_t, 0) / rate` for `rate > 0` (ADR 0002 §2).
## Controller identity ([Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11), [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14); ADR 0001 §3 as amended)
- **ID-1** (A) The controller-segment identity key is the normalized kernel-exposed subsystem NQN, with the kernel composite then model|serial as fallbacks; FR is metadata only; identity change and DUW decrease act as independent axes.
- **ID-2** (A) Segments freeze a fully nullable metadata snapshot at open — normalized `subnqn`/`sn`/`mn`/`fr` plus `vid`/`ssvid`/`transport` and `identity_degraded` — immutable thereafter, with `cntlid` excluded.
- **ID-3** (A) Legacy history imports under a labeled model-scoped legacy identity (mn-only segments).
- **ID-4** (A) `identity_degraded` is set at segment open exactly when the identity key is blank; keys from the kernel-composite or `model|serial` rungs are not degraded ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15)).
## Panes TUI ([Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3), [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6); ADR 0003 §§8, 10; ADR 0004 §10)
- **TUI-1** (A) Variant A "Panes": one dense keyboard-first screen; confidence rendered as evidence (state + contributing facts); boot enablement, runtime activity, last collect outcome, and freshness displayed as four separate facts.
- **TUI-2** (M) Pause/resume asymmetry and polkit tty passthrough work in a live terminal: pause confirms, resume does not, and the platform agent prompts without breaking the TUI.
- **TUI-3** (P) Textual runs on Python 3.9+, gated at install time, never a runtime crash.
- **TUI-4** (A) The Panes screen layout is normative: a full-width headline band (lifespan headline or its no-projection wording, confidence state with contributing facts, scenario range); a usage-history pane on the left (write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus legend, habit-split bar with active/idle/powered-off/unknown shares); a drive-health and settings pane on the right (health facts, vendor-wear context line, read-only settings with the endurance baseline and its provenance label); a full-width service strip at the bottom (the four separate service facts, the monitoring-period line, the action legend). Production bindings are `p` pause (asks), `r` resume (does not), `c` collect now, `d` disclosures, `q` quit ([Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3)); the prototype branch is visual reference only.
## Failure and recovery (ADR 0005)
- **FL-1** (A) The collector validates every row against the store invariants (hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count); a violating run writes nothing, logs the refused row, and fails visibly.
- **FL-2** (A) Readers defensively exclude and count malformed rows as a contributing fact.
- **FL-3** (A) No backfill ever: gaps remain unknown seconds; degradation flows only through coverage, freshness facts, and confidence categories.
- **FL-4** (P/A) A store fault surfaces "observation store unreadable" with a journal hint and suppresses everything else store-dependent; the collector treats it as a bounded failed run and never recreates or overwrites the file; recovery is the documented human-sanctioned move-aside (with `history.jsonl` re-import if the legacy import never completed); no built-in destructive command exists.
- **FL-5** (A) A newer-schema store renders "observation store written by a newer Fenris — upgrade Fenris" in TUI and status, with no partial interpretation.
- **FL-6** (A/P) Repeated collector failures retry at flat cadence with no backoff or notification; persistence reads as stale exactly like any other gap.
- **FL-7** (A) `critical_warning`, media errors, and unsafe shutdowns render as ordinary facts in TUI and status and never affect the projection.
- **FL-8** (A) A collection run finding no open monitoring period opens one at the run moment, never backdated.
## Installation, upgrade, and removal (ADR 0004)
- **IN-1** (P) `sudo make install` builds a wheel from the checkout and installs pinned dependencies into the dedicated venv at `/opt/fenris`, with a `/usr/local/bin/fenris` wrapper; after install nothing references the checkout.
- **IN-2** (P) The installer records every placed file in an explicit manifest consumed by upgrade and uninstall.
- **IN-3** (P) The installer never enables or starts units: a fresh install is dormant (units disabled, nothing running, no monitoring period); the only opt-in is the sanctioned toggle — `fenris monitor resume [--now]` or the first-run TUI prompt — enabling the timer and opening the first period in one step.
- **IN-4** (P) Install-time legacy import detects `./data/history.jsonl` (or an explicit path), runs the idempotent single-transaction import, and reports imported counts; `fenris import <path>` remains available.
- **IN-5** (P) `sudo make upgrade` installs into the same venv, syncs units and polkit against the manifest (`daemon-reload`; timer restarted only if unit contents changed and it is active), never kills an in-flight collection run, then applies forward-only schema migrations; `/var/lib/fenris` is never rebuilt.
- **IN-6** (P) Before migrations, `observations.db` is snapshotted to a one-generation `.bak`; rollback is reinstall-previous plus restore; automatic schema downgrade does not exist.
- **IN-7** (P) `make uninstall` performs the sanctioned disable first (open period closes `user_disabled`), then removes venv, helpers, units, polkit policy, and wrapper while keeping `/etc/fenris` and the observation store; `make purge` additionally removes configuration and store.
- **IN-8** (P) Dependencies are exact pins in a committed lockfile installed by both install and upgrade; refreshing pins is an explicit `make update-deps` step, never an install side effect.
- **IN-9** (P) The installer verifies `python3 ≥ 3.9` and fails cleanly otherwise; `/var/lib/fenris` is created with root-written group-read permissions; the database file is created lazily by the first write.
- **IN-10** (P) Installed artifacts sit only at their fixed locations — units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/`, configuration at `/etc/fenris`, observation store under `/var/lib/fenris` — and every placed file is recorded in the manifest (ADR 0004 §2; ADR 0003 §4).
## Collector acquisition path (ADR 0006)
- **AC-1** (P) Each collection run acquires counters and thermal evidence solely from `smartctl -a -j <device>` and controller identity (`subnqn`, `sn`, `mn`, `fr`, `transport`) solely from sysfs; no other acquisition path exists anywhere in the codebase.
- **AC-2** (A) Identity normalization is applied exactly once, at write time — trailing spaces and newlines stripped, no case folding, empty-after-strip stored blank — so padded and unpadded renderings of the same field yield byte-identical stored values.
- **AC-3** (A) Any acquisition failure — missing binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the whole collection run; no partial sample (identity without counters, or counters without identity) is ever written; the miss surfaces through ADR 0005 freshness, never as degraded identity.
- **AC-4** (P) `vid`/`ssvid` are read from the PCI sysfs node when present and stored null otherwise; they are segment metadata only, never key components.
- **AC-5** (P) `make install` verifies `smartctl` and fails cleanly otherwise; the acquisition path adds no Python dependency and no OS package beyond smartmontools (ADR 0004 §9).
+654
View File
@@ -0,0 +1,654 @@
# Fenris redesign specification
**Status: implementation-ready.** Assembled by [Write the Fenris redesign specification and close the map](https://git.bongbetic.com/xavierk/Fenris/issues/19), executing the assembly decision [Assemble the implementation-ready specification](https://git.bongbetic.com/xavierk/Fenris/issues/17) (all seven recommendations accepted) on the Wayfinder map [Chart Fenris's persistent TUI monitoring redesign](https://git.bongbetic.com/xavierk/Fenris/issues/1).
**Canonical roles.** [ADRs 0001–0006](../adr/) are the immutable rationale records — the *why*. [Acceptance criteria](acceptance-criteria.md) are the single register of testable statements — the *definition of done*. This document normatively restates every **operative contract** — the *what* — so an implementer never needs Wayfinder-ticket access: schema column sets, constants, rule tables, unit and CLI definitions, and the Panes TUI layout. Nothing here overrides an ADR or restates a criterion as a criterion.
## How to read this document
- **Binding language.** *Must*, *exactly*, and *never* are normative. Terminology follows the glossary in [`CONTEXT.md`](../../CONTEXT.md): *observation history*, *usage-adjusted theoretical lifespan*, *projection confidence*, *monitoring period*, *observation store*, *hour observation*, *day aggregate*, *controller segment*, *degraded identity*, *endurance baseline*, *verified override*, *unverified override*, *sustained regime*, *habit change*, *scenario range*, *coverage*, *collection run*, *deliberate disable*, *store fault*.
- **Ordering.** Sections follow data flow: system context → collector acquisition → observation store → controller identity & segmentation → hour/day derivation → projection & confidence → Panes TUI → service lifecycle & sanctioned toggle → failure & recovery → installation. Each section opens with its ADR links and criterion-ID block.
- **Implementation boundary.** This specification plans the redesign; it does not implement it. The complete handoff is this document + the [criteria register](acceptance-criteria.md) + [ADRs 0001–0006](../adr/) + the glossary. Given/When/Then test specs are derived by the implementer at implementation time.
### Normative constants index
Every constant is defined once, in the section named below; other sections cite, never redefine. All are named constants in code, not configuration.
| Constant | Value | Defined in |
|---|---|---|
| Collection cadence (default) | 5 min (`OnUnitInactiveSec`) | §8.2 |
| First-boot delay | 2 min (`OnBootSec`) | §8.2 |
| Timer accuracy window | 30 s (`AccuracySec`) | §8.2 |
| Collection-run timeout | 90 s (`TimeoutStartSec`) | §8.2 |
| Fresh threshold | newest sample within 2 × cadence + `AccuracySec` + 60 s | §8.9 |
| Missed → stale boundary | 48 h | §8.9, §6.7 |
| Powered-off hour threshold | power-on-hours delta < 90 % of the hour's wall-clock span | §5.1 |
| Active hour threshold | DUW delta ≥ 256 MiB in the hour | §5.1 |
| Raw-sample retention | 14 days | §3.4 |
| Warming gate | 14 distinct UTC day aggregates, ≤ 2 below 50 % coverage | §6.6 |
| Supported coverage floor | 80 % | §6.7 |
| Horizon agreement | 7/28/90-day rates within a factor of 2 | §6.7 |
| Burst guard | no single day ≥ 50 % of trailing 28-day bytes | §6.7 |
| Young-regime cap | regime < 7 days old → Limited | §6.4 |
| Habit-change trigger | trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean, 3 consecutive days | §6.4 |
| Regime span cap (default) | full observation history capped at 90 days | §6.4 |
| Scenario horizons | 7 / 28 / 90 days | §6.5 |
| Implied-baseline eligibility | ≥ 2 Percentage-Used increments within the current controller segment | §6.3 |
| Rated-TBW conversion | `E_rated = entered_TBW × 10¹²` bytes | §6.3 |
| Implied-baseline validity window | 1 ≤ p ≤ 254 | §6.3 |
| Wear-disagreement note | vendor wear vs. observed write rate by more than a factor of 2 | §6.1 |
| Capacity validation tolerance | ± 1 % | §6.2 |
---
## 1. System context
**ADRs:** [0001](../adr/0001-observation-store-sqlite.md), [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md), [0006](../adr/0006-collector-acquisition-path.md). **Criteria:** CI-3, LC-1, LC-5, ST-1.
### 1.1 Scope
Fenris observes one configured NVMe drive's real-world use and translates the observation history into a usage-adjusted theoretical lifespan. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer, persistent compact observation storage, and categorical projection confidence reflecting the length, completeness, and stability of real usage history.
Standing constraints, binding on every section:
- Linux with systemd and polkit only; no other init system is supported.
- Exactly one configured NVMe drive — the device named by `/etc/fenris/fenris.conf` (§8.3).
- Fully local: no telemetry, no network fetching, no automatic vendor-data retrieval.
- The HTML dashboard and HTTP server are gone; nothing of the daemonization, PID files, or `/run` state survives.
- CLI `status` and `sample` are retained (§8.8).
### 1.2 Components and privilege boundaries
| Component | Privilege | Path | Role |
|---|---|---|---|
| `fenris-collect.service` | root oneshot unit | `/usr/libexec/fenris/fenris-collect` | The only code path that interrogates the device and writes the observation store. |
| `fenris-collect.timer` | system timer | — | Schedules collection runs; `WantedBy=timers.target`. |
| `fenris-monitor` | root helper | `/usr/libexec/fenris/fenris-monitor` | Fixed privileged operations: `enable`/`disable` (optional `--now`), the collect trigger, monitoring-period bookkeeping, and baseline persistence. The only binary polkit authorizes. |
| `fenris` | unprivileged | `/usr/local/bin/fenris` | Human entry point: no arguments opens the TUI; subcommands are the CLI (§8.8). Never a unit. |
| Observation store | root-written, group-read | `/var/lib/fenris/observations.db` | Single SQLite database in WAL mode (§3). The TUI and `status` open it read-only. |
| Configuration | world-readable | `/etc/fenris/fenris.conf` | Exactly one key: the device selector (§8.3). |
The TUI and CLI are ordinary unprivileged processes. Elevation is exclusively polkit, exclusively for `fenris-monitor` (§8.5). There is no `/run/fenris` coordination surface and no export layer: systemd serializes collection runs, the observation store holds state, failures go to the journal.
### 1.3 Data flow
1. The timer fires; `fenris-collect.service` runs `fenris-collect`.
2. The collector acquires counters and thermal evidence from `smartctl -a -j` and controller identity from sysfs (§2), normalizes identity exactly once (§2.3), and either fails the whole run or writes one complete sample.
3. The collector derives and validates hour observations and day aggregates, advances controller segmentation and period bookkeeping, prunes raw samples, and commits (§3–§5, §9.1).
4. Readers — the TUI and `fenris status` — open the store read-only and **recompute the projection on every read** (§6); nothing derived is ever stored (§3.7).
Control flow is separate: the human drives the TUI/CLI; privileged operations route through `fenris-monitor` under polkit to `systemctl`; period rows record *intent* (only the sanctioned path), while the collector records *observed fact* (§8.5–§8.6, §9.8).
### 1.4 Cross-cutting prohibitions
These are operative contracts; each is restated in its home section and gated by the criteria block [CI-3](acceptance-criteria.md):
1. No code path outside `fenris-collect` interrogates the device (§2.1, §8.7).
2. Polkit authorizes exactly one binary, `fenris-monitor`, under `com.bongbetic.fenris.monitor` `auth_admin` (§8.5).
3. No `/run/fenris` coordination surface or export layer exists (§1.2).
4. No absent hour is ever interpolated, estimated, or fabricated (§5.3, §9.3).
5. No alerting, notification, or escalation machinery exists anywhere (§9.6–§9.7).
6. `/etc/fenris/fenris.conf` holds exactly one key — the device selector (§8.3).
7. No synthetic or capacity-derived baseline is ever created, including for legacy history (§6.1, §3.5).
8. Readers never partially interpret a newer-schema store (§3.6, §9.5).
9. Projections are never stored; always recomputed on read (§3.7, §6.10).
---
## 2. Collector acquisition
**ADR:** [0006](../adr/0006-collector-acquisition-path.md). **Criteria:** AC-1–AC-5; miss absorption per [0005](../adr/0005-failure-detection-and-recovery.md) §5.
### 2.1 Channels — the hard pin
Every collection run acquires exactly two ways:
- **Counters and thermal evidence** — solely from `smartctl -a -j <device>`: `data_units_written`, `data_units_read`, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning` — consumed as-is (smartmontools already trims the strings it copies).
- **Controller identity** — solely from sysfs (`/sys/class/nvme/<ctrl>/`): `subnqn`, `sn`, `mn`, `fr`, `transport`.
No other acquisition path exists anywhere in the codebase. There is no fallback: libnvme bindings and the `nvme` CLI JSON interface are excluded (ADR 0006, *Considered options*).
### 2.2 All-or-nothing runs
Any acquisition failure — missing `smartctl` binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the **whole** collection run. A partial sample (identity without counters, or counters without identity) is never written: a transient read failure must never push a healthy drive down the degraded-identity path (§4). The miss surfaces through freshness grading (§8.9) and the flat retry cadence (§9.6), never as degraded identity.
### 2.3 Identity normalization — once, at write time
One collector-side function normalizes every identity field, applied exactly once at write time:
- strip trailing spaces and newlines;
- no case folding;
- empty-after-strip is stored blank.
Padded and unpadded renderings of the same field therefore yield byte-identical stored values — a collector implementation change can never split a drive's own history. A future acquisition-path change must deliver byte-identical normalized identity values, or the change itself forces a controller-segment boundary.
### 2.4 Segment metadata sourcing
`transport` comes from the NVMe class sysfs directory. `vid`/`ssvid` come from the PCI node (`/sys/class/nvme/<ctrl>/device/{vendor,subsystem_vendor}`) when present and are stored null otherwise. Both are segment **metadata only** (§4.2), never key components.
### 2.5 Prerequisites
`make install` verifies `smartctl` is present and fails cleanly otherwise (§10.1). The acquisition path adds no Python dependency and no OS package beyond smartmontools; the dependency lockfile (§10.5) is untouched by this section.
---
## 3. Observation store
**ADR:** [0001](../adr/0001-observation-store-sqlite.md) as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) and [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14). **Criteria:** ST-1–ST-12; FL-5.
### 3.1 Substrate and access
- One SQLite database in **WAL mode** at `/var/lib/fenris/observations.db`. An unprivileged reader querying during a collector write sees a consistent snapshot.
- The database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI and `status` open it **read-only**. No `/run` snapshot, no export layer.
- `/var/lib/fenris` is created by the installer with root-written group-read permissions; the database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet (§7.6, §10.1).
- Migration, schema changes, prune, and import are each single transactions — a killed timer run can never leave partial state.
### 3.2 Entities and column sets
The schema carries exactly six entities:
**`samples`** — recent raw samples (14-day retention, §3.4): timestamp (UTC); the normalized controller-identity fields captured at acquisition (§2.3); raw integer `data_units_written`, `data_units_read`; `percentage_used`; `available_spare`; `media_errors`; `power_on_hours`; `power_cycles`; `unsafe_shutdowns`; temperature; `critical_warning`.
**`hour_observations`** — one row per UTC hour: the usage-habit split `seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown` (summing to 3600, §5.1); DUW/DUR deltas; temperature min/avg/max; sample count; coverage flag. Classification thresholds belong to the projection model (§5.1), not the store.
**`day_aggregates`** — one row per UTC day, the habit-evidence grain: each day row carries, at minimum, the day's activity-split sums, write deltas, and coverage share — the inputs the evidence gates of §6.6 consume — derived monotonically from its hour rows.
**`monitoring_periods`** — `started_at`; `ended_at` (NULL = open); `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not (§5.2, §8.6).
**`controller_segments`** — spans of unchanged controller identity and monotonic counters; write deltas are never computed across a segment boundary. Columns: the identity key (§4.1) and the frozen metadata snapshot of §4.2, plus the segment's span bounds.
**`endurance_baseline`** — one active row, replaced on edit (§6.2): the rated-TBW value in bytes (`E_rated = entered_TBW × 10¹²`); mandatory provenance — source URL, document revision, entry date, model string, nominal capacity; frozen validation facts — detected model, detected capacity bytes, `validated_by` (`machine`/`user`), `validated_at`.
### 3.3 Time model
Hours and days are UTC-bounded. Day derivation from hour rows is monotonic; DST-ambiguous 23- or 25-hour days never exist in the store.
### 3.4 Retention
Raw samples are pruned opportunistically by the collector to **14 days**. Hour observations and day aggregates are retained indefinitely.
### 3.5 Legacy migration
The migration procedure, invoked from the entry points below, is **idempotent and interruption-safe**:
1. If the store already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples and derive hour observations and day aggregates from them.
3. `hourly.jsonl` is never trusted as input: mismatches against derived data are diffed and logged.
4. Open one implicit `monitoring_periods` row at the first legacy sample, closed `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts — a sample present means powered on; a DUW delta means writes occurred.
5. The import is a single transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
6. Only after commit are legacy files renamed `*.migrated` — never deleted.
7. Malformed legacy lines are quarantined with a logged count, never silently dropped.
No synthetic or capacity-derived baseline is ever created for legacy history (§1.4–7). Entry points: the installer's import detection at `./data/history.jsonl` (or an explicit path) (§10.1); `fenris import <path>` for later finds (§8.8); and the collector's first new-version run, which performs this same procedure (ADR 0001 §6).
### 3.6 Schema versioning
`PRAGMA user_version` plus ordered migration steps in code, each in its own transaction. The collector refuses to run against an unknown **newer** version; readers refuse symmetrically with the exact wording of §9.5 and never partially interpret. *Reconciliation note:* ADR 0004 §6 describes upgrade-time migrations as "governed by a `schema_version` table" — the operative mechanism is this section's `user_version` (ADR 0001 §8, criterion ST-12); there is one version authority, not two.
### 3.7 Nothing derived is stored
Projections are not stored; there is no separate latest-status table and no stored health flag. The freshest sample timestamp is the store's own staleness signal (§8.9). The baseline lives in the database (§6.2); `/etc/fenris/` holds only operational configuration (§8.3).
---
## 4. Controller identity and segmentation
**Decisions:** [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11), [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14), [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15). **ADRs:** [0001](../adr/0001-observation-store-sqlite.md) §3 (as amended), [0002](../adr/0002-projection-model-sustained-regime.md) §§8–9 (as amended). **Criteria:** ID-1–ID-4, PR-9, PR-15, PR-16.
### 4.1 Identity key ladder
The controller-segment identity key is the **normalized, kernel-exposed subsystem NQN** (`subnqn`), with fallbacks, in order:
1. kernel-exposed subsystem NQN;
2. the kernel composite;
3. `model|serial`.
`fr` (firmware revision) is metadata only — it may go stale after a mid-segment firmware update. Identity change and DUW decrease are **independent axes** (§4.3).
### 4.2 Frozen metadata snapshot
Each segment freezes, at open, a fully nullable metadata snapshot — immutable thereafter: normalized `subnqn`, `sn`, `mn`, `fr`, plus `vid`, `ssvid`, `transport`, and the `identity_degraded` flag. All columns are nullable so incompleteness stays explicit: legacy-imported segments carry `mn` with NULLs (§4.4); degraded segments carry whatever was observed. These are human diagnostics, never key components. `cntlid` is excluded — it distinguishes controllers within one subsystem, out of scope for a single-drive monitor.
### 4.3 Segmentation axes
- **DUW decrease, unchanged identity** — a segment boundary within the same drive. Prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.8).
- **Identity-key change** — quarantines prior history from projection entirely: it describes a different drive (§6.8).
- **Degraded identity** — a segment whose identity key is **blank** (every rung of the ladder empty). `identity_degraded` is set at segment open exactly when the key is blank; keys from the kernel-composite or `model|serial` rungs are not degraded. Blank-key semantics extend identity-change rules verbatim: any visible change of the recorded key — including to or from blank — is a controller-identity change and quarantines; equal blank keys continue the segment, segmented by DUW monotonicity alone. Even a degraded→healthy transition quarantines, so the projection window only ever spans segments sharing one key (§6.8).
- Ephemeral markers (model "Linux", non-pcie transport) are segment metadata, never confidence facts.
The confidence consequence of degraded identity — capped at Limited with its fixed contributing fact — is §6.7's rule.
### 4.4 Legacy identity
Legacy history imports under a labeled, model-scoped **legacy identity** (mn-only segments), so it never blends with the post-redesign identity of the same physical drive.
---
## 5. Hour and day derivation
**ADRs:** [0002](../adr/0002-projection-model-sustained-regime.md) §§4–6; [0001](../adr/0001-observation-store-sqlite.md) §3; [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) §2 (power-on-hours evidence). **Criteria:** PR-4–PR-6, ST-4, FL-3.
### 5.1 Hour classification
Each UTC hour is classified by named constants, in this order of evidence:
- **Powered-off** — the hour's power-on-hours delta is below **90 %** of its wall-clock span.
- **Active** — DUW delta ≥ **256 MiB** in the hour.
- **Idle** — powered on, sampled, below the active threshold.
- **Unknown** — everything else: unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable by design), or inconsistent counters.
There is no configuration surface for these thresholds; they are documented constants in one projection module.
### 5.2 Denominator and disabled time
The projection denominator is **wall-clock seconds inside monitoring periods**, including powered-off and unknown time. Disabled periods — wall-clock outside monitoring periods — are excluded from numerator and denominator. **Disabled time is not an hour state.**
### 5.3 Gaps and coverage — never backfill
No absent hour is ever interpolated, estimated, or fabricated. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage. Power-on-hours classification (§5.1) is the only inference admitted. **Coverage** is the share of wall-clock seconds inside monitoring periods whose classification is known rather than unknown — a first-class displayed fact (§6.10, §7.3).
### 5.4 Day aggregates
One row per UTC day, derived monotonically from hour rows (§3.2–§3.3) — the grain at which usage-habit evidence is judged (§6.6).
---
## 6. Projection and confidence
**ADRs:** [0002](../adr/0002-projection-model-sustained-regime.md) as amended by [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15); baseline per [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12). **Criteria:** PR-1–PR-17, CI-4.
### 6.1 One projection; baseline precedence
Exactly **one** usage-adjusted theoretical lifespan is computed, against the endurance baseline chosen by precedence:
1. **Verified override** — a rated-TBW override with complete provenance whose applicability to the detected drive was confirmed by machine match or explicit user attestation;
2. **Unverified override** — a rated-TBW override knowingly stored with incomplete provenance; always presented as user-supplied, never as verified;
3. **Implied baseline** — derived from vendor wear (§6.3), eligible only per §6.3's gate;
4. otherwise the projection is **Unavailable**.
Percentage Used is context, never a second projection: it renders as a vendor-wear context line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The legacy PU-slope regression and `capacity × 600` synthesis are gone; no synthetic or capacity-derived baseline is ever created (§1.4–7).
### 6.2 Endurance baseline: provenance and validation
The baseline lives in the observation store's `endurance_baseline` table (§3.2) and is edited via the CLI (§8.8) — `/etc/fenris/` holds no baseline.
- **Mandatory provenance:** source URL, document revision, entry date, model string, nominal capacity.
- **One active row**, replaced on edit.
- **Verification is derived at read** — complete provenance and a drive match (machine or attested) — never a stored boolean.
- **Unverified tier:** incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier.
- **Entry-time validation** (unprivileged, live sysfs read of the configured device): normalized model containment, with an interactive confirm recorded as `validated_by = user`; nominal capacity within ± 1 %.
- **Read-time applicability:** a model match against the current controller segment (§4). A mismatch is **retained — never auto-deleted** — and leaves the projection Unavailable.
- Persistence goes through the polkit-guarded `fenris-monitor` verb after CLI-side validation (§8.5).
### 6.3 Arithmetic
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
- `E_rated` is exact: rated TBW converts to bytes by × 10¹².
- `E_implied` is computed **only** for `1 ≤ p ≤ 254`; Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable. The implied baseline is labeled *implied from vendor wear estimate* and shown with few significant digits.
- **Implied-baseline eligibility:** the implied tier is used only after ≥ 2 Percentage-Used increments within the current controller segment; until then the projection is Unavailable with the fixed phrase *"vendor wear estimate too coarse to imply endurance"*.
### 6.4 Sustained regime and habit change
The headline rate is the **sustained-regime** rate: regime DUW bytes ÷ in-period wall-clock seconds. The default regime is the full observation history capped at **90 days**.
A **habit change** is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for **3 consecutive days**. The new regime starts at the **first day of divergence**, is adopted automatically, and is labeled *"usage habit changed N days ago"*; the scenario range keeps the longer horizons visible. A regime younger than **7 days** caps projection confidence at Limited evidence.
### 6.5 Scenario range
The 7-, 28-, and 90-day rates are computed **independently of the regime** and shown as the scenario range. Only horizons the history actually covers appear — no placeholders. The scenario range is the only spread shown anywhere (§6.9).
### 6.6 Minimum evidence
Warming up until **14 distinct UTC day aggregates** of which at most **2** fall below 50 % coverage. The projection still renders while warming up, labeled with its facts (e.g. *"warming up: N of 14 qualifying days"*). Every Unavailable condition renders **no lifespan number**.
### 6.7 Confidence rule table
Confidence renders as **state plus contributing facts, never a percentage**. Three states:
- **Unavailable** — no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported** — verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80 % **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50 % of trailing 28-day bytes **and** regime ≥ 7 days old **and** the current controller segment's identity key is not degraded.
- **Limited** — every other case with a baseline and a positive rate; the failing facts are shown.
**Staleness:** a newest day aggregate older than **48 hours** drops confidence one level (Supported → Limited) and is shown as a contributing fact.
**Degraded identity:** a controller segment whose identity key is blank (§4.3) caps confidence at **Limited evidence**, with the contributing fact *"controller identity unavailable — replacement detection relies on write-counter continuity only"* rendered in every state. Supported is unreachable while the current segment is degraded. The cap combines idempotently with the staleness drop (both land at Limited).
### 6.8 Segment-break effects
- **DUW decrease, unchanged identity:** prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.6).
- **Controller-identity change** — including any to-or-from-blank key change (§4.3): prior history is quarantined from projection entirely.
- Since even degraded→healthy transitions quarantine, the projection window only ever spans segments sharing one key; no cross-segment propagation rule is needed.
### 6.9 Zero rate and uncertainty
Zero rate renders *"no finite projection from this history"* — never infinity, never zero. No statistical confidence interval appears anywhere; the scenario range is the only spread.
### 6.10 The projection contract
The projection function hands the TUI and `status` exactly: the confidence state; the contributing facts — including the degraded-identity fact when the current segment's key is blank; the headline remaining time when one exists; the scenario range; the Percentage-Used context line; the disclosure text (§6.11). Recomputed on read, never stored.
### 6.11 User-facing language
Adopted from the endurance research as fixed by ADR 0002 §12; rendered identically by TUI and `status`.
**Headline wording** (equivalent phrasing required):
> Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
**Fixed phrases** (exact): *no finite projection from this history* (zero rate); *vendor wear estimate too coarse to imply endurance* (§6.3); *usage habit changed N days ago* (§6.4); *controller identity unavailable — replacement detection relies on write-counter continuity only* (§6.7); *observation store unreadable* (§9.4); *observation store written by a newer Fenris — upgrade Fenris* (§9.5); *no observations yet* with an enable hint (§8.9); *configuration error: ⟨reason⟩* (§8.3).
**Confidence rendering:** state plus contributing facts, in the research's evidence style, e.g.
> Supported evidence · verified manufacturer TBW · 42 calendar days · 96 % interval coverage · 6 weekly cycles · recent and 28-day rates agree
Never "82 % confidence" or "95 % accurate".
**The six disclosures** (verbatim, always available — TUI disclosures view and `status`):
1. This is an endurance projection, not a predicted hardware-failure date.
2. Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated.
3. Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold.
4. DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes.
5. Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
6. Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
---
## 7. Panes TUI
**Decisions:** [Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3) (Variant A adopted), [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6) (Textual). **ADRs:** [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) §§8, 10; [0004](../adr/0004-install-upgrade-removal-lifecycle.md) §10. **Criteria:** TUI-1–TUI-4, CI-1, CI-2, CI-4. The [prototype](https://git.bongbetic.com/xavierk/Fenris/src/branch/prototype/tui-information-architecture/prototype/tui-ia) is visual reference only; this section is normative.
### 7.1 Framework and floor
The TUI is built on **Textual**. It runs on Python 3.9+, gated at install time (§10.1) — never a runtime crash. The tty-passthrough mechanism below was validated live under Textual on a real terminal (prototype decision).
### 7.2 Layout — one dense keyboard-first screen
Variant A **Panes**: everything on one screen, no page navigation. The screen is a grid of four regions:
1. **Headline band** — full width, top: the lifespan headline (or its no-projection wording) with its regime line (*"if current habits continue · sustained regime: N days at R GB/day"*); the confidence state with contributing facts; the scenario range.
2. **Usage-history pane** — left, wider column: the write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus their legend; the habit-split bar with active/idle/powered-off/unknown shares.
3. **Drive-health and settings pane** — right, narrower column: health facts (model, temperature, spare, media errors, unsafe shutdowns, power-on hours, power cycles, capacity); the vendor-wear context line (Percentage Used · total written of rated — *context, not a second projection*); a read-only settings view (device selector, cadence with drop-in pointer, raw retention, endurance baseline value with its provenance label). Edits happen via CLI / drop-ins, not in the TUI.
4. **Service strip** — full width, bottom: the four separate service facts (§7.3), the monitoring-period line, the action legend.
Exact proportions, glyphs, and borders follow the prototype's validated arrangement as visual reference; the region arrangement, contents, and bindings above are normative.
### 7.3 Content contracts per region
- **Confidence is evidence:** state plus contributing facts, never a percentage (§6.7); the headline band renders the §6.10 contract in full, including the disclosure affordance (`d`).
- **Four separate service facts, always:** boot enablement (enabled/disabled) · runtime activity (timer active/inactive) · last collect outcome (ok/FAILED, age, reason) · freshness (fresh/missed/stale with newest-sample age, §8.9). They are never merged into one "service status".
- **Monitoring-period line:** open-since / closed with end cause; deliberate-disable count where nonzero.
- The scenario range shows only covered horizons (§6.5); the vendor-wear context line carries the >2× disagreement note when it applies (§6.1).
### 7.4 Keybindings and asymmetry
Production bindings:
| Key | Action |
|---|---|
| `p` | Pause — **asks for confirmation** (y pause · n cancel), stating that paused time is excluded from the usage habit while powered-off time would still count. |
| `r` | Resume — **no confirmation** (benign; friction invites raw-systemctl escapes). |
| `c` | Collect now — synchronous outcome (§8.7), no confirmation. |
| `d` | Disclosures — the six disclosures of §6.11. |
| `q` | Quit. |
No bare start/stop exists anywhere; no page navigation keys exist (variant switching was prototype-only). Framework defaults apply for focus and scrolling otherwise.
### 7.5 Privileged actions and tty passthrough
Pause, resume, collect-now, and baseline operations run through `fenris-monitor` as a **terminal-attached subprocess**: the TUI suspends, the platform polkit agent prompts on the real terminal, and control returns cleanly with the outcome reflected in the service facts. Where no polkit agent exists the operation fails cleanly with the printed root equivalent (§8.5).
### 7.6 State rendering obligations
- From a synthetic observation store, the TUI renders **every** realizable combination of confidence state × freshness grade × baseline tier exactly as the §6.7 rule table and §8.9 constants dictate — headline number only when the rules allow it, contributing facts always, never a percentage (criterion CI-1).
- Empty store: *"no observations yet"* with an enable hint; the first-run prompt is an opt-in that enables the timer and opens the first period in one step (§10.1, dormant install).
- A `configuration error: ⟨reason⟩` fact renders when the device selector is invalid (§8.3); a store fault suppresses everything store-dependent (§9.4); a newer schema renders its fixed phrase (§9.5).
- Every TUI action has a CLI twin with identical outcomes and wording (§8.8, CI-2).
---
## 8. Service lifecycle and sanctioned toggle
**ADR:** [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12). **Criteria:** LC-1–LC-10, CI-2, CI-3.
### 8.1 Units
Exactly two system units exist:
- `fenris-collect.timer` — `WantedBy=timers.target`.
- `fenris-collect.service` — `Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code.
The TUI and CLI are ordinary unprivileged processes and never units.
### 8.2 Cadence
Shipped defaults: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no` (no suspend catch-up — absent hours classify through power-on-hours evidence, §5.1), `TimeoutStartSec=90s` so a hung device interrogation fails visibly as a bounded failed run retried next interval. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); **no interval key exists in configuration**.
### 8.3 Configuration
`/etc/fenris/fenris.conf` holds exactly one key: the **device selector**, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run — there is no reload path. An invalid selector is a bounded failed run (journal + failed unit result, retried next interval); `status` and the TUI also read the world-readable file directly and surface `configuration error: ⟨reason⟩`.
### 8.4 Entry points
Two privileged binaries — `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable`/`disable` with optional `--now`, the collect trigger, monitoring-period bookkeeping, and `baseline set`/`baseline clear` persistence for the CLI-validated baseline; the only binary the polkit policy authorizes). One unprivileged `fenris` wrapper (§1.2). Root invokes the helpers directly; unprivileged users go through polkit.
### 8.5 Sanctioned toggle and polkit
- Pause = `fenris-monitor disable --now`; Resume = `enable --now`. Both perform the systemctl operation **and** the monitoring-period bookkeeping in one step. The human-facing twins `fenris monitor pause` / `fenris monitor resume` map to these and always act immediately; pause asks for confirmation in both TUI and CLI, resume does not (§7.4).
- Polkit action `com.bongbetic.fenris.monitor` (`auth_admin`) covers the toggle **and** the collect trigger **and** baseline persistence — authorizing exactly the one binary `fenris-monitor`.
- Where no polkit agent exists the operation fails cleanly and prints the root equivalent.
- This is the **only** sanctioned control path: a raw `systemctl stop`/`disable` never records `user_disabled` — only the sanctioned path records intent (§8.6).
### 8.6 Period-row idempotent matrix
| Situation | Effect on `monitoring_periods` |
|---|---|
| First-ever enable | Opens a period at the enable moment (hours before the first successful sample are unknown-but-inside — correct when the device errors). |
| Resume with an open period (a raw `systemctl stop` intervened) | No row changes; the gap remains inside as unknown seconds. |
| Resume with no open period | Opens a new row at the resume moment. |
| Pause with an open period | Closes it `user_disabled` at the pause moment. |
| Pause otherwise | No-op. |
| Raw `systemctl stop`/`disable` outside the helper | An unexplained gap, never `user_disabled`. |
### 8.7 On-demand collection
`fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits; the outcome (freshness line or journal hint) is reported synchronously. No confirmation is required. No code path outside `fenris-collect` touches the device; the TUI never samples in-process.
### 8.8 CLI surface
| Command | Behavior |
|---|---|
| `fenris` (no arguments) | Opens the TUI (§7). |
| `fenris status` | Read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness. Never auto-samples, never prompts. |
| `fenris sample` | On-demand collection via the helper path (§8.7). |
| `fenris monitor pause` / `resume` | The sanctioned toggle (§8.5), pause asking confirmation. |
| `fenris baseline set` / `clear` | CLI-side validation (§6.2), then polkit-guarded persistence. |
| `fenris import ⟨path⟩` | The idempotent single-transaction legacy import (§3.5). |
| `--device` | Rejected with a pointer to the configuration file. |
| `start`, `stop`, `run` | Rejected with one-line migration pointers — never aliased (an alias would silently change meaning). |
`fenris.sh` is retired: not shipped, removed from the repository; the README maps its five menu options to their successors. Headless administration has full parity: every TUI action has a CLI twin (pause, resume, collect-now, baseline set/clear, the status fact set) with identical outcomes and wording.
### 8.9 Freshness grading
Constants defined once, consumed by TUI and CLI alike; the grade derives from the **newest sample timestamp**, never a stored flag:
- **fresh** — newest sample within 2 × cadence + `AccuracySec` + 60 s (11.5 min at default cadence);
- **missed** — between that and 48 h (a contributing fact);
- **stale** — ≥ 48 h, matching the §6.7 evidence gate;
- **empty store** — *"no observations yet"* with an enable hint.
---
## 9. Failure and recovery
**ADR:** [0005](../adr/0005-failure-detection-and-recovery.md). **Criteria:** FL-1–FL-8.
The posture: **visible degradation, never fabrication.**
### 9.1 Write-boundary validation
The collector validates every row it would write against the store's domain invariants: hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count. A violating run **writes nothing**, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Store invariant: everything persisted is well-formed.
### 9.2 Reader defense
Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact. Under a single trusted writer they should never see one.
### 9.3 No backfill, ever
Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. Power-on-hours classification (§5.1) is the only inference admitted.
### 9.4 Store faults — degrade, never recreate over
An unreadable or corrupt database is a store fault: readers surface *"observation store unreadable"* with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and **never recreates or overwrites** an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside; the next run starts a fresh store; if the legacy import never completed, the still-present `history.jsonl` is re-imported (§3.5). No built-in destructive command exists.
### 9.5 Newer schema — symmetric refusal
The TUI and `status` detect a `user_version` newer than they understand and display *"observation store written by a newer Fenris — upgrade Fenris"* without partial interpretation, matching the collector's refusal (§3.6) and the forward-only upgrade rule (§10.2).
### 9.6 Repeated collector failures — flat cadence, no escalation
The timer's retry is the recovery path; the freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff, no notification machinery; a persistent failure reads as stale exactly like any other gap.
### 9.7 Drive-reported anomalies — facts, not alerts
`critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status` (§7.2); no alerting or notification surface exists. The projection is unaffected: endurance math consumes writes, not warnings.
### 9.8 Orphaned samples — the collector re-anchors observed fact
When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the **run moment**, never backdated. This records observed fact, not intent — only the sanctioned path records a `user_disabled` close (§8.5). Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
---
## 10. Installation
**ADR:** [0004](../adr/0004-install-upgrade-removal-lifecycle.md). **Criteria:** IN-1–IN-10.
### 10.1 Install
`sudo make install`:
1. Builds a wheel from the checkout and installs it, with pinned dependencies (§10.5), into the dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only — after install, nothing references it.
2. Verifies `python3 ≥ 3.9` and `smartctl` presence, failing cleanly otherwise (never a runtime crash).
3. Creates `/var/lib/fenris` with root-written group-read permissions and the `fenris` read group; the database file is created lazily by the first write (§3.1).
4. Places units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/` — recording **every** placed file in an explicit manifest consumed by upgrade and uninstall (§10.6).
5. **Never enables or starts units.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle — `fenris monitor resume` or the first-run TUI prompt — enabling the timer and opening the first period in one step.
6. Detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs the idempotent single-transaction import (§3.5), and reports imported counts.
### 10.2 Upgrade
`sudo make upgrade` installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed **and** it is active — safe with `Persistent=no`), leaves timer state untouched, and **never kills an in-flight collection run**: a running oneshot finishes on its mapped interpreter; at worst one old-code run completes to the store. It then applies forward-only observation-store schema migrations (§3.6). `/var/lib/fenris` is never rebuilt.
### 10.3 Rollback
Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade does not exist.
### 10.4 Uninstall and purge
- `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. Journal entries age out naturally.
- `make purge` additionally removes configuration and store.
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
### 10.5 Dependencies
Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing. The acquisition path adds no Python dependency and no OS package beyond smartmontools (§2.5).
### 10.6 Placement manifest
Installed artifacts sit only at their fixed locations — units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/`, configuration at `/etc/fenris`, observation store under `/var/lib/fenris`, venv at `/opt/fenris`, wrapper at `/usr/local/bin/fenris` — and every placed file is recorded in the manifest (criterion IN-10).
---
## Appendix A: Traceability matrix
Built as assembly's first step (assembly decision, recommendation 5). Two-way: every ADR section maps to at least one criterion ID; every criterion cites its ADR or ticket.
### A.1 ADR section → criteria
| ADR section | Criteria |
|---|---|
| 0001 §1 Substrate | ST-1, ST-2 |
| 0001 §2 Access | ST-1, CI-3 (/run) |
| 0001 §3 Entities (incl. #12/#14 amendments) | ST-3, ST-5, PR-13, ID-2, CI-3 (no stored projections) |
| 0001 §4 Day boundary | ST-4 |
| 0001 §5 Retention | ST-5 |
| 0001 §6 Migration | ST-6, ST-7, ST-8, ST-9, ST-10, ST-11 |
| 0001 §7 Projection inputs | ST-3, PR-14, CI-3 (one-key config) |
| 0001 §8 Versioning | ST-12, FL-5 |
| 0001 §9 Collector health | LC-10 |
| 0002 §1 One projection | PR-1 |
| 0002 §2 Rate/formulas/regime/scenario | PR-2, PR-17 |
| 0002 §3 Habit change | PR-3 |
| 0002 §4 Hour classification | PR-4 |
| 0002 §5 Denominator | PR-5 |
| 0002 §6 Minimum evidence | PR-6 |
| 0002 §7 Staleness | PR-7 |
| 0002 §8 Confidence table (incl. #15 amendment) | PR-8, PR-15, CI-1, CI-4 |
| 0002 §9 Segment breaks (incl. #15 amendment) | PR-9, PR-16 |
| 0002 §10 Implied eligibility | PR-10 |
| 0002 §11 Uncertainty | PR-11 |
| 0002 §12 Language | CI-4 |
| 0002 §13 Contract | PR-12 |
| 0003 §1 Units | LC-1, CI-3 (/run) |
| 0003 §2 Cadence | LC-2, LC-3 |
| 0003 §3 Configuration | LC-4, CI-3 |
| 0003 §4 Entry points | LC-5, IN-10 |
| 0003 §5 Sanctioned toggle | LC-6, CI-2, CI-3 |
| 0003 §6 Period rows | LC-7 |
| 0003 §7 On-demand collection | LC-8, CI-2, CI-3 |
| 0003 §8 TUI controls | TUI-1, TUI-2, TUI-4 |
| 0003 §9 CLI compatibility | LC-9, CI-2 |
| 0003 §10 Freshness constants | LC-10, CI-1, CI-2 |
| 0004 §1 Delivery | IN-1 |
| 0004 §2 Layout and manifest | IN-2, IN-10 |
| 0004 §3 Privilege | IN-3, CI-3 |
| 0004 §4 Dormant install | IN-3 |
| 0004 §5 Legacy import | IN-4 |
| 0004 §6 Upgrade | IN-5 |
| 0004 §7 Rollback | IN-6 |
| 0004 §8 Removal | IN-7 |
| 0004 §9 Dependencies | IN-8, AC-5 |
| 0004 §10 Scaffolding and floor | IN-9, TUI-3 |
| 0005 §1 Malformed observations | FL-1, FL-2 |
| 0005 §2 Missed observations | FL-3, CI-3 |
| 0005 §3 Store faults | FL-4 |
| 0005 §4 Newer schema | FL-5 |
| 0005 §5 Repeated failures | FL-6 |
| 0005 §6 Drive-reported anomalies | FL-7, CI-3 |
| 0005 §7 Orphaned samples | FL-8 |
| 0006 §1 Pin | AC-1 |
| 0006 §2 Hard pin, no fallback | AC-3 |
| 0006 §3 Normalization | AC-2 |
| 0006 §4 Segment metadata sourcing | AC-4 |
| 0006 §5 Prerequisites | AC-5 |
### A.2 Ticket decisions → criteria
| Ticket | Criteria |
|---|---|
| [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6) | TUI-3 |
| [Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3) | TUI-1, TUI-2, TUI-4 |
| [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11) | ID-1, ID-3 |
| [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) | PR-13, PR-14, ST-3 |
| [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14) | ID-2 |
| [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15) | PR-15, PR-16, ID-4 |
| [Define cross-cutting acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/issues/13) | the register itself |
| [Assemble the implementation-ready specification](https://git.bongbetic.com/xavierk/Fenris/issues/17) / [Write the Fenris redesign specification and close the map](https://git.bongbetic.com/xavierk/Fenris/issues/19) | this document and this appendix |
### A.3 Criterion → source
Every criterion carries its citation inline in the [register](acceptance-criteria.md): CI-1–CI-4 (ADR 0002 §§6–8, 0003 §§1–10, 0001 §§2–3/8, 0005 §§2/4–6); ST-1–ST-12 (ADR 0001, with ST-3 amended by tickets #12/#14); LC-1–LC-10 (ADR 0003); PR-1–PR-12 (ADR 0002), PR-13–PR-14 (ticket #12), PR-15–PR-16 (ticket #15, ADR 0002 §§8–9 as amended), PR-17 (ADR 0002 §2); ID-1–ID-4 (tickets #11/#14/#15, ADR 0001 §3 as amended); TUI-1–TUI-4 (tickets #3/#6, ADR 0003 §§8/10, ADR 0004 §10); FL-1–FL-8 (ADR 0005); IN-1–IN-10 (ADR 0004, with IN-10 also citing ADR 0003 §4); AC-1–AC-5 (ADR 0006).
### A.4 Assembly result
- **Every ADR 0001–0006 section maps to at least one criterion** — the A.1 table is complete; no orphan sections.
- **Every criterion cites its ADR or ticket** — verified in the register; no orphan criteria.
- **Decided-but-uncitered gaps found and filled inline in the register during assembly:** CI-3 bullet (no `/run` coordination surface; ADR 0003 §1, ADR 0001 §2), PR-17 (projection arithmetic; ADR 0002 §2), TUI-4 (normative Panes layout and bindings; ticket #3), IN-10 (fixed artifact placement; ADR 0004 §2, ADR 0003 §4).
- **No genuinely undecided behavior remained** — no blocking ticket was raised.
- **One reconciliation:** ADR 0004 §6's "`schema_version` table" wording resolves to ADR 0001 §8's `PRAGMA user_version` as the single version authority (§3.6); criterion ST-12 already fixed the mechanism.
+49 -40
View File
@@ -21,7 +21,11 @@ import mimetypes
from datetime import datetime, timezone, timedelta from datetime import datetime, timezone, timedelta
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__)) # Add src/ to path for package imports
_SCRIPT_DIR = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, os.path.join(_SCRIPT_DIR, "src"))
SCRIPT_DIR = _SCRIPT_DIR
DATA_DIR = os.path.join(SCRIPT_DIR, "data") DATA_DIR = os.path.join(SCRIPT_DIR, "data")
DATA_FILE = os.path.join(DATA_DIR, "history.jsonl") DATA_FILE = os.path.join(DATA_DIR, "history.jsonl")
HOURLY_FILE = os.path.join(DATA_DIR, "hourly.jsonl") HOURLY_FILE = os.path.join(DATA_DIR, "hourly.jsonl")
@@ -973,38 +977,18 @@ def cmd_stop(args):
def cmd_status(args): def cmd_status(args):
running = False try:
if os.path.exists(PID_FILE): from fenris.status import render_status, check_retired_flag
with open(PID_FILE) as f: from pathlib import Path
pid = int(f.read().strip())
running = pid_alive(pid) show_disclosures = getattr(args, "disclosures", False)
print(f"Fenris daemon: {'RUNNING (pid ' + str(pid) + ')' if running else 'not running (stale pid file)'}") store_path = Path("/var/lib/fenris/observations.db")
else: output = render_status(store_path=store_path, show_disclosures=show_disclosures)
print("Fenris daemon: not running") print(output)
rows = load_history() except ImportError:
if rows: # Fallback if fenris package not importable
summ = compute_summary(rows) print("Error: cannot import fenris.status module. Is the package installed?")
latest = rows[-1] sys.exit(1)
print(f"Samples collected: {len(rows)}")
print(f"Last sample: {latest['ts']}")
print(f"Wear (percentage_used): {latest.get('percentage_used')}%")
print(f"Total written: {latest.get('bytes_written', 0) / 1e9:.1f} GB")
print(f"Written (24h rolling): {summ['window24h']['gb']:.2f} GB over {summ['window24h']['coverage_hours']:.1f}h")
print(f"Write rate: {summ['gb_per_hour']:.2f} GB/h ({summ['gb_per_day']:.1f} GB/day)")
if summ["seconds_remaining"]:
print(f"Projected life remaining: {summ['breakdown']['human']} (≈{summ['breakdown']['days']:.0f} days / {summ['breakdown']['hours']:.0f} hours / {summ['breakdown']['years']:.2f} years)")
print(f"Endurance: {summ['endurance_tb']:.1f} TB total, {summ['remaining_tb']:.1f} TB remaining" + (" (estimated)" if summ["endurance_estimated"] else ""))
if summ["wear_model_days"]:
print(f"Wear-model cross-check: ~{summ['wear_model_days']:.0f} days at current wear rate")
if summ["preliminary"]:
print("Note: preliminary — less than 24h coverage")
else:
print("Projected life remaining: — (no writes in window or no endurance data)")
hourly = load_hourly()
if hourly:
print(f"Hourly buckets: {len(hourly)} (last {hourly[-1]['hour']}: {hourly[-1]['bytes_written']/1e9:.2f} GB)")
else:
print("No samples collected yet.")
def cmd_run(args): def cmd_run(args):
@@ -1023,8 +1007,22 @@ def cmd_sample_once(args):
sys.exit(1) sys.exit(1)
def cmd_retired(args):
"""Handle retired commands with migration pointers (§8.8)."""
from fenris.status import check_retired_command
cmd = sys.argv[1] if len(sys.argv) > 1 else ""
ptr = check_retired_command(cmd)
if ptr:
print(ptr)
else:
print("Unknown command. Use 'fenris status' or 'fenris sample'.")
sys.exit(1)
def main(): def main():
p = argparse.ArgumentParser(description="Fenris — NVMe wear monitor & dashboard (by Bongbetic)") p = argparse.ArgumentParser(description="Fenris — NVMe wear monitor & dashboard (by Bongbetic)")
p.add_argument("--device", default=None,
help="(retired — device is configured in /etc/fenris/fenris.conf)")
sub = p.add_subparsers(dest="cmd", required=True) sub = p.add_subparsers(dest="cmd", required=True)
def add_common(sp): def add_common(sp):
@@ -1032,19 +1030,30 @@ def main():
sp.add_argument("--interval", type=int, default=300, help="seconds between samples (default 300)") sp.add_argument("--interval", type=int, default=300, help="seconds between samples (default 300)")
sp.add_argument("--port", type=int, default=8420, help="dashboard HTTP port (default 8420)") sp.add_argument("--port", type=int, default=8420, help="dashboard HTTP port (default 8420)")
sp = sub.add_parser("start", help="start monitoring in background") sp = sub.add_parser("start", help="(retired — use 'fenris monitor resume')")
add_common(sp); sp.set_defaults(func=cmd_start) add_common(sp); sp.set_defaults(func=cmd_retired)
sp = sub.add_parser("stop", help="stop background monitoring") sp = sub.add_parser("stop", help="(retired — use 'fenris monitor pause')")
sp.set_defaults(func=cmd_stop) sp.set_defaults(func=cmd_retired)
sp = sub.add_parser("status", help="show daemon + latest wear stats") sp = sub.add_parser("status", help="show read-only status (§8.8)")
sp.add_argument("-d", "--disclosures", action="store_true",
help="show the six disclosures (§6.11)")
sp.set_defaults(func=cmd_status) sp.set_defaults(func=cmd_status)
sp = sub.add_parser("run", help="run in foreground (used internally by 'start')") sp = sub.add_parser("run", help="(retired — use 'fenris monitor resume')")
add_common(sp); sp.set_defaults(func=cmd_run) add_common(sp); sp.set_defaults(func=cmd_retired)
sp = sub.add_parser("sample", help="take one sample immediately and print it") sp = sub.add_parser("sample", help="take one sample immediately and print it")
sp.add_argument("--device", default=detect_device()) sp.add_argument("--device", default=detect_device())
sp.set_defaults(func=cmd_sample_once) sp.set_defaults(func=cmd_sample_once)
args = p.parse_args() args = p.parse_args()
# Reject retired --device flag (§8.8) — only if explicitly passed
if getattr(args, "device", None) is not None:
from fenris.status import check_retired_flag
ptr = check_retired_flag("--device")
if ptr:
print(ptr)
sys.exit(1)
args.func(args) args.func(args)
+22
View File
@@ -0,0 +1,22 @@
[project]
name = "fenris"
version = "0.3.0"
description = "NVMe wear monitor with persistent TUI"
requires-python = ">=3.9"
dependencies = [
"textual>=0.40.0",
]
[project.optional-dependencies]
dev = [
"pytest>=7.0.0",
"pytest-cov>=4.0.0",
]
[tool.pytest.ini_options]
testpaths = ["tests"]
python_files = ["test_*.py"]
python_functions = ["test_*"]
markers = [
"slow: marks tests as slow",
]
+2
View File
@@ -0,0 +1,2 @@
"""Fenris: NVMe wear monitor with persistent TUI."""
__version__ = "0.3.0"
+325
View File
@@ -0,0 +1,325 @@
"""Collector: acquires counters and identity, writes to observation store.
This module implements the thinnest complete write path:
- Acquire counters and thermal evidence from smartctl -a -j
- Acquire controller identity from sysfs
- Normalize identity exactly once at write time
- Validate every row against store invariants
- Commit one well-formed sample
No code path outside the collector interrogates the device.
"""
import json
import sqlite3
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, Optional, Tuple
from .store import init_store, get_store_path
class AcquisitionError(Exception):
"""Raised when acquisition fails - whole run is refused."""
pass
class InvariantViolationError(Exception):
"""Raised when a row would violate store invariants - writes nothing."""
pass
def acquire_from_smartctl(smartctl_data: Dict[str, Any]) -> Dict[str, Any]:
"""Acquire counters and thermal evidence from smartctl -a -j data.
Validates that all required fields are present.
Raises AcquisitionError on any failure.
"""
required_fields = [
"nvme_smart_health_information_log",
"user_capacity",
"model_name",
"serial_number",
"firmware_version",
]
for field in required_fields:
if field not in smartctl_data:
raise AcquisitionError(f"Missing required field in smartctl data: {field}")
log = smartctl_data["nvme_smart_health_information_log"]
required_log_fields = [
"data_units_written",
"data_units_read",
"percentage_used",
"power_on_hours",
"temperature",
]
for field in required_log_fields:
if field not in log:
raise AcquisitionError(f"Missing required field in SMART log: {field}")
return {
"model": smartctl_data["model_name"],
"serial": smartctl_data["serial_number"],
"firmware_rev": smartctl_data["firmware_version"],
"capacity_bytes": smartctl_data["user_capacity"]["bytes"],
"percentage_used": log["percentage_used"],
"available_spare": log.get("available_spare"),
"media_errors": log.get("media_errors", 0),
"power_on_hours": log["power_on_hours"],
"power_cycles": log.get("power_cycles"),
"unsafe_shutdowns": log.get("unsafe_shutdowns"),
"temperature_c": log["temperature"],
"data_units_written": log["data_units_written"],
"data_units_read": log["data_units_read"],
"bytes_written": log["data_units_written"] * 512000,
"bytes_read": log["data_units_read"] * 512000,
"critical_warning": log.get("critical_warning", 0),
}
def acquire_from_sysfs(sysfs_path: Path) -> Dict[str, Any]:
"""Acquire controller identity from sysfs.
Reads identity from:
- /sys/class/nvme/<ctrl>/subsysnqn (primary)
- /sys/class/nvme/<ctrl>/model
- /sys/class/nvme/<ctrl>/serial
- /sys/class/nvme/<ctrl>/firmware_rev
- /sys/class/nvme/<ctrl>/transport/ (optional)
Raises AcquisitionError on any failure.
"""
identity_files = {
"subnqn": "subsysnqn",
"mn": "model",
"sn": "serial",
"fr": "firmware_rev",
}
identity = {}
for key, filename in identity_files.items():
filepath = sysfs_path / filename
if not filepath.exists():
raise AcquisitionError(f"Missing sysfs file: {filepath}")
try:
value = filepath.read_text().strip()
identity[key] = value if value else ""
except Exception as e:
raise AcquisitionError(f"Failed to read {filepath}: {e}")
# Transport info (optional)
transport_dir = sysfs_path / "transport"
if transport_dir.exists():
try:
transport_file = transport_dir / "trstring"
if transport_file.exists():
identity["transport"] = transport_file.read_text().strip()
else:
identity["transport"] = None
except Exception:
identity["transport"] = None
else:
identity["transport"] = None
# vid/ssvid from PCI node (optional, metadata only - never key components)
# PCI device directory is the sysfs_path itself (the controller dir is a symlink to PCI)
pci_device = sysfs_path
for attr, key in [("vendor", "vid"), ("subsystem_vendor", "ssvid")]:
filepath = pci_device / attr
if filepath.exists():
try:
value = filepath.read_text().strip()
identity[key] = value if value else None
except Exception:
identity[key] = None
else:
identity[key] = None
return identity
def normalize_identity(identity: Dict[str, Any]) -> str:
"""Normalize identity exactly once at write time.
Rules:
- Strip trailing spaces and newlines
- No case folding
- Empty-after-strip stored blank
Returns normalized identity key.
"""
# Primary key: normalized kernel-exposed subsystem NQN
key = identity.get("subnqn", "")
if key:
key = key.rstrip()
return key
# Fallback 1: kernel composite (not implemented yet)
# Fallback 2: model|serial
mn = identity.get("mn", "").rstrip()
sn = identity.get("sn", "").rstrip()
if mn or sn:
return f"{mn}|{sn}"
# All keys blank - degraded identity
return ""
def compute_identity_degraded(identity: Dict[str, Any]) -> bool:
"""Check if identity is degraded (all key rungs empty)."""
key = normalize_identity(identity)
return key == ""
def validate_sample_invariants(sample: Dict[str, Any], conn: sqlite3.Connection) -> None:
"""Validate sample against store invariants.
Raises InvariantViolationError if any invariant is violated.
"""
# TODO: Implement more complex invariants as needed
# For now, just check basic constraints
if sample.get("bytes_written", 0) < 0:
raise InvariantViolationError("Negative bytes_written")
if sample.get("bytes_read", 0) < 0:
raise InvariantViolationError("Negative bytes_read")
def write_sample(
sample: Dict[str, Any],
identity: Dict[str, Any],
conn: sqlite3.Connection,
clock,
) -> Dict[str, Any]:
"""Write one sample to the observation store.
Identity normalization happens exactly once here.
Returns segment info for the caller.
"""
from .segment import find_current_segment, should_open_new_segment, open_segment
# Normalize identity exactly once at write time
identity_key = normalize_identity(identity)
identity_degraded = compute_identity_degraded(identity)
# Find current segment
current_segment = find_current_segment(conn)
# Determine if we need a new segment
should_open, reason = should_open_new_segment(
current_segment, identity_key, sample["bytes_written"], conn
)
# Open new segment if needed
segment_opened = False
if should_open:
now = clock.utcnow()
open_segment(conn, now, identity, identity_key, identity_degraded)
segment_opened = True
# Insert sample
cursor = conn.execute(
"""
INSERT INTO samples (
ts, device, subnqn, sn, mn, fr, capacity_bytes,
percentage_used, available_spare, media_errors, power_on_hours,
power_cycles, unsafe_shutdowns, temperature_c,
data_units_written, data_units_read, bytes_written, bytes_read,
critical_warning
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
sample["ts"],
sample["device"],
identity.get("subnqn", ""),
identity.get("sn", ""),
identity.get("mn", ""),
identity.get("fr", ""),
sample["capacity_bytes"],
sample["percentage_used"],
sample["available_spare"],
sample["media_errors"],
sample["power_on_hours"],
sample["power_cycles"],
sample["unsafe_shutdowns"],
sample["temperature_c"],
sample["data_units_written"],
sample["data_units_read"],
sample["bytes_written"],
sample["bytes_read"],
sample["critical_warning"],
),
)
conn.commit()
return {
"segment_opened": segment_opened,
"segment_reason": reason,
"identity_key": identity_key,
"identity_degraded": identity_degraded,
}
def run_collection(
smartctl_data: Dict[str, Any],
sysfs_path: Path,
config: Dict[str, Any],
clock,
) -> Dict[str, Any]:
"""Run one collection run.
This is the main entry point for the collector.
Returns the run outcome.
"""
try:
# Acquire counters and thermal evidence
counters = acquire_from_smartctl(smartctl_data)
# Acquire controller identity
identity = acquire_from_sysfs(sysfs_path)
# Build sample with injected clock
sample = {
"ts": clock.utcnow().isoformat(),
"device": config["device"],
**counters,
**identity,
}
# Initialize store if needed
store_path = get_store_path(config)
conn = init_store(store_path)
# Run legacy import if needed (idempotent)
from .legacy import import_legacy_history
history_path = Path(config.get("data_dir", ".")) / "history.jsonl"
if history_path.exists():
import_legacy_history(conn, history_path, clock=clock)
try:
# Validate invariants
validate_sample_invariants(sample, conn)
# Write sample
write_sample(sample, identity, conn, clock)
return {
"ok": True,
"sample_count": 1,
"store_path": str(store_path),
}
finally:
conn.close()
except (AcquisitionError, InvariantViolationError) as e:
return {
"ok": False,
"error": str(e),
"error_type": type(e).__name__,
}
+165
View File
@@ -0,0 +1,165 @@
"""Day aggregate derivation per spec §5.4, §3.3.
One row per UTC day, derived monotonically from hour rows — the grain at
which usage-habit evidence is judged. No absent hour is ever interpolated,
estimated, or fabricated (§5.3, FL-3).
Coverage: the share of wall-clock seconds inside monitoring periods whose
usage-habit classification is known rather than unknown (§5.3).
"""
import sqlite3
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
@dataclass(frozen=True)
class DayAggregate:
"""One UTC day's aggregated stats."""
day: str # ISO 8601 UTC date, e.g. "2026-09-01"
seconds_active: int
seconds_idle: int
seconds_powered_off: int
seconds_unknown: int
bytes_written_delta: int
bytes_read_delta: int
sample_count: int
coverage: float
def _period_wall_clock_for_day(conn: sqlite3.Connection, day: str) -> int:
"""Total wall-clock seconds inside monitoring periods for a UTC day.
Clamps each period to the day boundary [dayT00:00, dayT24:00).
"""
day_start = datetime.fromisoformat(f"{day}T00:00:00+00:00")
day_end = day_start + timedelta(days=1)
day_start_str = day_start.isoformat()
day_end_str = day_end.isoformat()
cursor = conn.execute(
"SELECT started_at, ended_at FROM monitoring_periods "
"WHERE (ended_at IS NULL OR ended_at > ?) AND started_at < ? "
"ORDER BY started_at",
(day_start_str, day_end_str),
)
total = 0
for row in cursor.fetchall():
period_start = row[0]
period_end = row[1]
effective_start = max(period_start, day_start_str)
if period_end is not None:
effective_end = min(period_end, day_end_str)
else:
effective_end = day_end_str
if effective_start < effective_end:
s = datetime.fromisoformat(effective_start)
e = datetime.fromisoformat(effective_end)
total += int((e - s).total_seconds())
return total
def _hour_overlaps_period(conn: sqlite3.Connection, hour_iso: str) -> bool:
"""Check if an hour's wall-clock span overlaps any monitoring period."""
hour_start = datetime.fromisoformat(hour_iso)
hour_end = hour_start + timedelta(hours=1)
hs = hour_start.isoformat()
he = hour_end.isoformat()
cursor = conn.execute(
"SELECT 1 FROM monitoring_periods "
"WHERE started_at < ? AND (ended_at IS NULL OR ended_at > ?) "
"LIMIT 1",
(he, hs),
)
return cursor.fetchone() is not None
def derive_day(conn: sqlite3.Connection, day: str) -> DayAggregate | None:
"""Derive a single day aggregate from its hour rows + monitoring periods.
Only hours overlapping a monitoring period contribute to the aggregate.
Gap hours inside periods contribute unknown seconds. Hours outside all
monitoring periods are excluded entirely (§5.2).
Returns None if no hours exist for the day.
"""
cursor = conn.execute(
"SELECT hour, active_seconds, idle_seconds, powered_off_seconds, unknown_seconds, "
" bytes_written_delta, bytes_read_delta, sample_count "
"FROM hour_observations "
"WHERE hour LIKE ? "
"ORDER BY hour",
(day + "T%",),
)
rows = cursor.fetchall()
if not rows:
return None
total_active = 0
total_idle = 0
total_powered_off = 0
total_unknown_from_hours = 0
total_bw = 0
total_br = 0
total_samples = 0
total_hour_wall_clock = 0
for row in rows:
# Only count hours overlapping a monitoring period
if not _hour_overlaps_period(conn, row[0]):
continue
total_active += row[1]
total_idle += row[2]
total_powered_off += row[3]
total_unknown_from_hours += row[4]
total_bw += row[5]
total_br += row[6]
total_samples += row[7]
total_hour_wall_clock += row[1] + row[2] + row[3] + row[4]
# Wall-clock seconds inside monitoring periods for this day
period_wc = _period_wall_clock_for_day(conn, day)
# Gap seconds = period wall-clock - sum of existing hour wall-clock
gap_seconds = max(0, period_wc - total_hour_wall_clock)
total_unknown = total_unknown_from_hours + gap_seconds
# Coverage: known seconds / period wall-clock (§5.2, §5.3)
known_seconds = total_active + total_idle + total_powered_off
coverage = known_seconds / period_wc if period_wc > 0 else 0.0
return DayAggregate(
day=day,
seconds_active=total_active,
seconds_idle=total_idle,
seconds_powered_off=total_powered_off,
seconds_unknown=total_unknown,
bytes_written_delta=total_bw,
bytes_read_delta=total_br,
sample_count=total_samples,
coverage=coverage,
)
def derive_all_days(conn: sqlite3.Connection) -> list[DayAggregate]:
"""Derive day aggregates for all days that have hour rows.
Returns days sorted by date.
"""
cursor = conn.execute(
"SELECT DISTINCT substr(hour, 1, 10) as day FROM hour_observations ORDER BY day"
)
days = [row[0] for row in cursor.fetchall()]
results = []
for day in days:
agg = derive_day(conn, day)
if agg is not None:
results.append(agg)
return results
+94
View File
@@ -0,0 +1,94 @@
"""Hour classification per spec §5.1.
Each UTC hour is classified by named constants, in this order of evidence:
- Powered-off: power-on-hours delta < 90% of wall-clock span
- Active: DUW delta >= 256 MiB in the hour
- Idle: powered on, sampled, below active threshold
- Unknown: everything else (unsampled without POH evidence)
Four splits sum to exactly wall_clock_seconds. Disabled time is never an
hour state — it is wall-clock outside monitoring periods (§5.2).
"""
from dataclasses import dataclass
# Spec §5.1: Active hour threshold — 256 MiB DUW delta
ACTIVE_THRESHOLD_BYTES = 256 * 1024 * 1024 # 256 MiB
# Spec §5.1: Powered-off threshold — 90% of wall-clock span
POWERED_OFF_THRESHOLD_PERCENT = 0.90
@dataclass(frozen=True)
class HourSplit:
"""Usage-habit split for one UTC hour. Fields sum to wall_clock_seconds."""
seconds_active: int
seconds_idle: int
seconds_powered_off: int
seconds_unknown: int
def classify_hour(
wall_clock_seconds: int,
poh_delta: int,
duw_delta: int,
dur_delta: int,
sampled_seconds: int | None = None,
) -> HourSplit:
"""Classify a UTC hour into the four usage-habit states.
Args:
wall_clock_seconds: Total seconds in this hour boundary (3600 for a
full hour, less for partial-hours at period edges).
poh_delta: Power-on-hours delta since previous sample (in seconds).
duw_delta: Data-units-written delta since previous sample (in bytes).
dur_delta: Data-units-read delta since previous sample (in bytes).
sampled_seconds: Seconds within this hour covered by a sample.
None or 0 means no sample fell in this hour.
Returns:
HourSplit whose four fields sum to wall_clock_seconds.
"""
if sampled_seconds is None:
sampled_seconds = 0
# Clamp sampled_seconds to wall_clock_seconds
sampled_seconds = min(sampled_seconds, wall_clock_seconds)
# --- Decision order per spec §5.1 ---
# 1. Powered-off: POH delta < 90% of wall-clock span
powered_off_threshold = wall_clock_seconds * POWERED_OFF_THRESHOLD_PERCENT
if poh_delta < powered_off_threshold:
return HourSplit(
seconds_active=0,
seconds_idle=0,
seconds_powered_off=wall_clock_seconds,
seconds_unknown=0,
)
# 2. Active: DUW delta >= 256 MiB
if duw_delta >= ACTIVE_THRESHOLD_BYTES:
return HourSplit(
seconds_active=wall_clock_seconds,
seconds_idle=0,
seconds_powered_off=0,
seconds_unknown=0,
)
# 3. Idle: powered on, sampled, below active threshold
# Unsampled portion within the hour is unknown
if sampled_seconds > 0:
return HourSplit(
seconds_active=0,
seconds_idle=sampled_seconds,
seconds_powered_off=0,
seconds_unknown=wall_clock_seconds - sampled_seconds,
)
# 4. Unknown: unsampled without POH evidence
return HourSplit(
seconds_active=0,
seconds_idle=0,
seconds_powered_off=0,
seconds_unknown=wall_clock_seconds,
)
+440
View File
@@ -0,0 +1,440 @@
"""Legacy migration: import history.jsonl into the observation store.
Spec §3.5, ADR 0001 §6. Idempotent and interruption-safe.
Entry points:
- Installer import detection at ./data/history.jsonl (§10.1)
- fenris import <path> (§8.8)
- Collector's first new-version run (ADR 0001 §6)
The import is a single transaction — a scripted kill mid-import leaves the
store fully pre- or fully post-migration. history.jsonl is the sole authority;
hourly.jsonl is diffed and logged but never trusted. Malformed lines are
quarantined with a logged count, never silently dropped. Legacy files are
renamed *.migrated only after commit and never deleted.
"""
import json
import logging
import sqlite3
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
from .hour_classify import classify_hour
from .segment import open_segment, normalize_identity
from .monitoring_periods import close_period
logger = logging.getLogger(__name__)
# Legacy import marker — stored in a metadata table
LEGACY_IMPORT_MARKER = "legacy_imported"
def _ensure_metadata_table(conn: sqlite3.Connection) -> None:
"""Create the metadata table if it doesn't exist."""
conn.execute("""
CREATE TABLE IF NOT EXISTS store_metadata (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
)
""")
def is_legacy_imported(conn: sqlite3.Connection) -> bool:
"""Check if legacy history has already been imported.
Spec §3.5.1: If the store already carries the legacy-import marker, do nothing.
"""
_ensure_metadata_table(conn)
cursor = conn.execute(
"SELECT value FROM store_metadata WHERE key = ?",
(LEGACY_IMPORT_MARKER,),
)
row = cursor.fetchone()
return row is not None and row[0] == "true"
def _parse_history_line(line: str, line_num: int) -> Optional[Dict[str, Any]]:
"""Parse a single line from history.jsonl.
Returns None for malformed lines (quarantined, not silently dropped).
"""
line = line.strip()
if not line:
return None
try:
record = json.loads(line)
except json.JSONDecodeError as e:
logger.warning("Malformed JSON at line %d: %s", line_num, e)
return None
# Validate required fields
required_fields = [
"timestamp", "model", "serial", "firmware_version",
"data_units_written", "data_units_read", "percentage_used",
"power_on_hours", "temperature",
]
for field in required_fields:
if field not in record:
logger.warning("Missing field '%s' at line %d", field, line_num)
return None
return record
def _record_to_sample(record: Dict[str, Any]) -> Dict[str, Any]:
"""Convert a legacy history.jsonl record to a sample dict."""
# Legacy records use different field names
duw = record["data_units_written"]
dur = record["data_units_read"]
return {
"ts": record["timestamp"],
"device": record.get("device", "/dev/nvme0"),
"subnqn": record.get("subsystem_nqn", ""),
"sn": record["serial"],
"mn": record["model"],
"fr": record["firmware_version"],
"capacity_bytes": record.get("capacity_bytes", 0),
"percentage_used": record["percentage_used"],
"available_spare": record.get("available_spare"),
"media_errors": record.get("media_errors", 0),
"power_on_hours": record["power_on_hours"],
"power_cycles": record.get("power_cycles"),
"unsafe_shutdowns": record.get("unsafe_shutdowns"),
"temperature_c": record["temperature"],
"data_units_written": duw,
"data_units_read": dur,
"bytes_written": duw * 512000,
"bytes_read": dur * 512000,
"critical_warning": record.get("critical_warning", 0),
}
def _derive_hour_observation(
samples: List[Dict[str, Any]],
hour_start: datetime,
) -> Dict[str, Any]:
"""Derive a single hour observation from samples in that hour.
Pre-migration hours carry an unknown activity split except directly
evidenced facts — a sample present means powered on; a DUW delta means
writes occurred (§3.5.4).
"""
hour_end = hour_start.replace(hour=hour_start.hour + 1) if hour_start.hour < 23 else hour_start.replace(hour=0, day=hour_start.day + 1)
# Filter samples in this hour
hour_samples = []
for s in samples:
ts = datetime.fromisoformat(s["ts"])
if hour_start <= ts < hour_end:
hour_samples.append(s)
if not hour_samples:
return None
# Sort by timestamp
hour_samples.sort(key=lambda x: x["ts"])
# Compute deltas from first to last sample in the hour
first = hour_samples[0]
last = hour_samples[-1]
duw_delta = last["bytes_written"] - first["bytes_written"]
dur_delta = last["bytes_read"] - first["bytes_read"]
poh_delta = (last["power_on_hours"] - first["power_on_hours"]) * 3600
# Temperature stats
temps = [s["temperature_c"] for s in hour_samples]
# Classify the hour
split = classify_hour(
wall_clock_seconds=3600,
poh_delta=poh_delta,
duw_delta=duw_delta,
dur_delta=dur_delta,
sampled_seconds=3600, # Legacy samples cover the full hour
)
return {
"hour": hour_start.strftime("%Y-%m-%dT%H:00:00Z"),
"active_seconds": split.seconds_active,
"idle_seconds": split.seconds_idle,
"powered_off_seconds": split.seconds_powered_off,
"unknown_seconds": split.seconds_unknown,
"bytes_written_delta": duw_delta,
"bytes_read_delta": dur_delta,
"temperature_min": min(temps),
"temperature_avg": sum(temps) / len(temps),
"temperature_max": max(temps),
"sample_count": len(hour_samples),
"coverage": 1.0 if split.seconds_unknown == 0 else (3600 - split.seconds_unknown) / 3600,
}
def _diff_hourly_jsonl(
hourly_path: Path,
derived_hours: Dict[str, Dict[str, Any]],
) -> None:
"""Diff hourly.jsonl against derived data and log mismatches.
Spec §3.5.3: hourly.jsonl is never trusted; mismatches are diffed and logged.
"""
if not hourly_path.exists():
logger.info("No hourly.jsonl found for diffing")
return
try:
with open(hourly_path, "r") as f:
for line_num, line in enumerate(f, 1):
line = line.strip()
if not line:
continue
try:
record = json.loads(line)
except json.JSONDecodeError:
logger.warning("Malformed hourly.jsonl at line %d", line_num)
continue
hour_key = record.get("hour")
if hour_key not in derived_hours:
logger.info("Hourly.jsonl has hour %s not in derived data", hour_key)
continue
derived = derived_hours[hour_key]
mismatches = []
for field in ["bytes_written_delta", "bytes_read_delta", "sample_count"]:
if field in record and record[field] != derived.get(field):
mismatches.append(
f"{field}: hourly={record[field]} derived={derived.get(field)}"
)
if mismatches:
logger.info(
"Hourly.jsonl mismatch for %s: %s",
hour_key,
"; ".join(mismatches),
)
except Exception as e:
logger.warning("Failed to diff hourly.jsonl: %s", e)
def import_legacy_history(
conn: sqlite3.Connection,
history_path: Path,
hourly_path: Optional[Path] = None,
clock=None,
) -> Dict[str, Any]:
"""Import legacy history.jsonl into the observation store.
This is the main entry point for legacy migration. It is:
- Idempotent: second run no-ops on the legacy-import marker
- Interruption-safe: single transaction
- Never creates synthetic baselines
Args:
conn: Connection to the observation store
history_path: Path to history.jsonl
hourly_path: Optional path to hourly.jsonl for diffing
clock: Injected clock (for testing)
Returns:
Dict with migration outcome
"""
# Check idempotency
if is_legacy_imported(conn):
return {"ok": True, "skipped": True, "reason": "already_imported"}
# Read and parse history.jsonl
samples = []
malformed_count = 0
if not history_path.exists():
return {"ok": False, "error": f"History file not found: {history_path}"}
with open(history_path, "r") as f:
for line_num, line in enumerate(f, 1):
record = _parse_history_line(line, line_num)
if record is None:
malformed_count += 1
continue
samples.append(_record_to_sample(record))
if malformed_count > 0:
logger.warning("Quarantined %d malformed lines from history.jsonl", malformed_count)
if not samples:
return {"ok": False, "error": "No valid samples found in history.jsonl"}
# Sort samples by timestamp
samples.sort(key=lambda x: x["ts"])
# Derive hour observations
hour_observations = {}
for sample in samples:
ts = datetime.fromisoformat(sample["ts"])
hour_start = ts.replace(minute=0, second=0, microsecond=0)
hour_key = hour_start.strftime("%Y-%m-%dT%H:00:00Z")
if hour_key not in hour_observations:
hour_observations[hour_key] = {
"hour": hour_start,
"samples": [],
}
hour_observations[hour_key]["samples"].append(sample)
derived_hours = {}
for hour_key, hour_data in hour_observations.items():
obs = _derive_hour_observation(hour_data["samples"], hour_data["hour"])
if obs is not None:
derived_hours[hour_key] = obs
# Diff against hourly.jsonl if provided
if hourly_path:
_diff_hourly_jsonl(hourly_path, derived_hours)
# Single transaction for the entire import
try:
# Begin transaction
conn.execute("BEGIN IMMEDIATE")
# 1. Insert raw samples
for sample in samples:
conn.execute(
"""
INSERT INTO samples (
ts, device, subnqn, sn, mn, fr, capacity_bytes,
percentage_used, available_spare, media_errors, power_on_hours,
power_cycles, unsafe_shutdowns, temperature_c,
data_units_written, data_units_read, bytes_written, bytes_read,
critical_warning
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
sample["ts"],
sample["device"],
sample["subnqn"],
sample["sn"],
sample["mn"],
sample["fr"],
sample["capacity_bytes"],
sample["percentage_used"],
sample["available_spare"],
sample["media_errors"],
sample["power_on_hours"],
sample["power_cycles"],
sample["unsafe_shutdowns"],
sample["temperature_c"],
sample["data_units_written"],
sample["data_units_read"],
sample["bytes_written"],
sample["bytes_read"],
sample["critical_warning"],
),
)
# 2. Insert derived hour observations
for hour_key, obs in sorted(derived_hours.items()):
conn.execute(
"""
INSERT OR REPLACE INTO hour_observations (
hour, active_seconds, idle_seconds, powered_off_seconds,
unknown_seconds, bytes_written_delta, bytes_read_delta,
temperature_min, temperature_avg, temperature_max,
sample_count, coverage
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
obs["hour"],
obs["active_seconds"],
obs["idle_seconds"],
obs["powered_off_seconds"],
obs["unknown_seconds"],
obs["bytes_written_delta"],
obs["bytes_read_delta"],
obs["temperature_min"],
obs["temperature_avg"],
obs["temperature_max"],
obs["sample_count"],
obs["coverage"],
),
)
# 3. Open implicit monitoring period at first legacy sample
first_sample_ts = datetime.fromisoformat(samples[0]["ts"])
conn.execute(
"INSERT INTO monitoring_periods (started_at) VALUES (?)",
(first_sample_ts.isoformat(),),
)
# 4. Close period with end_cause = migrated at migration moment
if clock:
migration_time = clock.utcnow()
else:
migration_time = datetime.now(timezone.utc)
conn.execute(
"UPDATE monitoring_periods SET ended_at = ?, end_cause = ? WHERE ended_at IS NULL",
(migration_time.isoformat(), "migrated"),
)
# 5. Open legacy controller segment (mn-only)
# Legacy identity is model-scoped only (§4.4)
first_sample = samples[0]
legacy_identity = {
"mn": first_sample["mn"],
"sn": "", # Legacy segments are mn-only
"subnqn": "",
"fr": "",
}
legacy_identity_key = f"legacy|{first_sample['mn']}"
open_segment(
conn,
migration_time,
legacy_identity,
legacy_identity_key,
identity_degraded=False,
)
# 6. Set legacy import marker
_ensure_metadata_table(conn)
conn.execute(
"INSERT OR REPLACE INTO store_metadata (key, value) VALUES (?, ?)",
(LEGACY_IMPORT_MARKER, "true"),
)
# Commit
conn.commit()
except Exception as e:
conn.rollback()
raise RuntimeError(f"Migration failed: {e}") from e
# 7. Rename legacy files to *.migrated (only after commit)
try:
migrated_path = history_path.with_suffix(history_path.suffix + ".migrated")
history_path.rename(migrated_path)
logger.info("Renamed %s to %s", history_path, migrated_path)
if hourly_path and hourly_path.exists():
hourly_migrated = hourly_path.with_suffix(hourly_path.suffix + ".migrated")
hourly_path.rename(hourly_migrated)
logger.info("Renamed %s to %s", hourly_path, hourly_migrated)
except Exception as e:
# Non-fatal: files weren't renamed but migration succeeded
logger.warning("Failed to rename legacy files: %s", e)
return {
"ok": True,
"skipped": False,
"samples_imported": len(samples),
"hours_imported": len(derived_hours),
"malformed_lines": malformed_count,
"first_sample": samples[0]["ts"],
"last_sample": samples[-1]["ts"],
"legacy_identity_key": legacy_identity_key,
}
+124
View File
@@ -0,0 +1,124 @@
"""Monitoring period bookkeeping per spec §5.2, §8.6, §9.8.
A monitoring period is a span during which Fenris monitoring is enabled.
Powered-off time stays inside a period; deliberately disabled time does not.
Key contracts:
- Run finding no open period opens one at the run moment, never backdated (§9.8)
- Wall-clock outside periods excluded from numerator and denominator (§5.2)
- End causes: user_disabled, migrated, unknown_gap
"""
import sqlite3
from datetime import datetime
def ensure_period_open(conn: sqlite3.Connection, run_time: datetime) -> None:
"""Ensure a monitoring period is open. If none exists, open one at run_time.
Spec §9.8: A collection run finding no open monitoring period opens one
at the run moment, never backdated.
"""
if get_open_period(conn) is not None:
return # Already open — no-op
ts = run_time.isoformat()
conn.execute(
"INSERT INTO monitoring_periods (started_at) VALUES (?)",
(ts,),
)
conn.commit()
def close_period(
conn: sqlite3.Connection,
closed_at: datetime,
end_cause: str,
) -> None:
"""Close the current open monitoring period.
Spec §8.6: Pause with an open period closes it user_disabled.
If no period is open, this is a no-op (pause otherwise).
"""
open_period = get_open_period(conn)
if open_period is None:
return # No-op
ts = closed_at.isoformat()
conn.execute(
"UPDATE monitoring_periods SET ended_at = ?, end_cause = ? WHERE id = ?",
(ts, end_cause, open_period["id"]),
)
conn.commit()
def get_open_period(conn: sqlite3.Connection) -> dict | None:
"""Return the currently open monitoring period, or None."""
cursor = conn.execute(
"SELECT id, started_at, ended_at, end_cause "
"FROM monitoring_periods WHERE ended_at IS NULL LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0],
"started_at": row[1],
"ended_at": row[2],
"end_cause": row[3],
}
def is_inside_period(conn: sqlite3.Connection, ts: datetime) -> bool:
"""Check if a timestamp falls inside any monitoring period.
Spec §5.2: Wall-clock outside periods is excluded from numerator/denominator.
"""
ts_str = ts.isoformat()
cursor = conn.execute(
"SELECT 1 FROM monitoring_periods "
"WHERE started_at <= ? AND (ended_at IS NULL OR ended_at > ?) "
"LIMIT 1",
(ts_str, ts_str),
)
return cursor.fetchone() is not None
def wall_clock_in_periods(
conn: sqlite3.Connection,
start: datetime,
end: datetime,
) -> int:
"""Compute total wall-clock seconds between start and end that fall inside
any monitoring period.
Used for denominator computation (§5.2).
"""
start_str = start.isoformat()
end_str = end.isoformat()
cursor = conn.execute(
"SELECT started_at, ended_at FROM monitoring_periods "
"WHERE ended_at IS NULL OR ended_at > ? "
"ORDER BY started_at",
(start_str,),
)
total = 0
for row in cursor.fetchall():
period_start = row[0]
period_end = row[1] # None if open
# Clip period to [start, end]
effective_start = max(period_start, start_str)
if period_end is not None:
effective_end = min(period_end, end_str)
else:
effective_end = end_str
if effective_start < effective_end:
# Parse for arithmetic
s = datetime.fromisoformat(effective_start)
e = datetime.fromisoformat(effective_end)
total += int((e - s).total_seconds())
return total
+580
View File
@@ -0,0 +1,580 @@
"""Projection core: the pure-function read path (spec §6).
Recomputes the complete projection contract on every read, never stores
anything derived. Takes a read-only observation store connection and an
injected clock; returns a ProjectionResult with confidence state,
contributing facts, headline remaining time (when one exists), scenario
range, Percentage-Used context line, and disclosure text.
Baseline precedence (§6.1):
verified override → unverified override → implied → unavailable
Confidence rule table (§6.7):
Supported — all conjuncts satisfied
Limited — baseline + positive rate, failing facts shown
Unavailable — no applicable baseline / zero rate / identity change
Arithmetic (§6.3):
rate = regime DUW bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
Criteria: PR-1–PR-17, CI-4.
"""
import sqlite3
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
from enum import Enum
from typing import Any, Dict, List, Optional, Tuple
# ---------------------------------------------------------------------------
# Constants (spec §6)
# ---------------------------------------------------------------------------
HORIZON_DAYS = (7, 28, 90)
TBW_TO_BYTES = 10 ** 12
IMPLIED_P_MIN = 1
IMPLIED_P_MAX = 254
IMPLIED_MIN_PU_INCREMENTS = 2
WARMING_MIN_DAYS = 14
WARMING_MAX_LOW_COVERAGE = 2
WARMING_COVERAGE_FLOOR = 0.50
SUPPORTED_COVERAGE_FLOOR = 0.80
HORIZON_AGREEMENT_FACTOR = 2
BURST_GUARD_FRACTION = 0.50
BURST_GUARD_LOOKBACK = 28
YOUNG_REGIME_DAYS = 7
HABIT_CHANGE_SHORT_WINDOW = 7
HABIT_CHANGE_LONG_WINDOW = 28
HABIT_CHANGE_UPPER_FACTOR = 2
HABIT_CHANGE_LOWER_FACTOR = 0.5
HABIT_CHANGE_CONSECUTIVE_DAYS = 3
STALENESS_HOURS = 48
WEAR_DISAGREEMENT_FACTOR = 2
class ConfidenceState(Enum):
UNSUPPORTED = "Unavailable"
LIMITED = "Limited"
SUPPORTED = "Supported"
class BaselineTier(Enum):
VERIFIED = "verified_override"
UNVERIFIED = "unverified_override"
IMPLIED = "implied"
NONE = "none"
@dataclass(frozen=True)
class ScenarioRange:
rates: Dict[int, float]
min_days: int
max_days: int
@dataclass(frozen=True)
class ProjectionResult:
confidence_state: ConfidenceState
contributing_facts: List[str]
headline_remaining_seconds: Optional[float]
scenario_range: Optional[ScenarioRange]
pu_context_line: str
disclosure_text: List[str]
baseline_tier: BaselineTier
baseline_label: str
regime_days: Optional[int]
habit_change_fact: Optional[str]
warming_fact: Optional[str]
staleness_fact: Optional[str]
degraded_identity_fact: Optional[str]
zero_rate_fact: Optional[str]
# ---------------------------------------------------------------------------
# Store queries
# ---------------------------------------------------------------------------
def _get_baseline(conn: sqlite3.Connection) -> Optional[Dict[str, Any]]:
cursor = conn.execute(
"SELECT id, tbw_terabytes, source_url, document_revision, entry_date, "
" model_string, nominal_capacity_bytes, validated_by, verified "
"FROM endurance_baseline LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0], "tbw_terabytes": row[1], "source_url": row[2],
"document_revision": row[3], "entry_date": row[4],
"model_string": row[5], "nominal_capacity_bytes": row[6],
"validated_by": row[7], "verified": bool(row[8]),
}
def _get_current_segment(conn: sqlite3.Connection) -> Optional[Dict[str, Any]]:
cursor = conn.execute(
"SELECT id, opened_at, identity_key, identity_degraded, mn "
"FROM controller_segments ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0], "opened_at": row[1], "identity_key": row[2],
"identity_degraded": bool(row[3]), "mn": row[4],
}
def _get_days_in_segment(conn, segment_opened_at):
cursor = conn.execute(
"SELECT day, bytes_written_delta, coverage, sample_count "
"FROM day_aggregates WHERE day >= ? ORDER BY day",
(segment_opened_at[:10],),
)
return [{"day": r[0], "bytes_written": r[1], "coverage": r[2], "sample_count": r[3]}
for r in cursor.fetchall()]
def _get_all_days(conn):
cursor = conn.execute(
"SELECT day, bytes_written_delta, coverage, sample_count "
"FROM day_aggregates ORDER BY day"
)
return [{"day": r[0], "bytes_written": r[1], "coverage": r[2], "sample_count": r[3]}
for r in cursor.fetchall()]
def _get_latest_pu(conn):
cursor = conn.execute("SELECT percentage_used FROM samples ORDER BY id DESC LIMIT 1")
row = cursor.fetchone()
return row[0] if row else None
def _get_pu_increments_in_segment(conn, segment_opened_at):
cursor = conn.execute(
"SELECT COUNT(DISTINCT percentage_used) FROM samples WHERE ts >= ?",
(segment_opened_at,),
)
row = cursor.fetchone()
return max(0, (row[0] if row else 0) - 1)
def _wall_clock_in_range(conn, start, end):
start_str = start.isoformat()
end_str = end.isoformat()
cursor = conn.execute(
"SELECT started_at, ended_at FROM monitoring_periods "
"WHERE (ended_at IS NULL OR ended_at > ?) AND started_at < ? "
"ORDER BY started_at", (start_str, end_str),
)
total = 0
for row in cursor.fetchall():
eff_start = max(row[0], start_str)
eff_end = min(row[1], end_str) if row[1] is not None else end_str
if eff_start < eff_end:
total += int((datetime.fromisoformat(eff_end) - datetime.fromisoformat(eff_start)).total_seconds())
return total
# ---------------------------------------------------------------------------
# Baseline resolution (§6.1, §6.2)
# ---------------------------------------------------------------------------
def _resolve_baseline(conn, current_segment):
baseline = _get_baseline(conn)
facts = []
if baseline is None:
return BaselineTier.NONE, None, "no baseline", facts
mandatory = [baseline["source_url"], baseline["document_revision"],
baseline["entry_date"], baseline["model_string"],
baseline["nominal_capacity_bytes"]]
provenance_complete = all(f is not None and f != "" for f in mandatory)
model_matches = True
if current_segment is not None and baseline["model_string"] is not None:
seg_mn = (current_segment.get("mn") or "").lower()
bl_model = (baseline["model_string"] or "").lower()
model_matches = bl_model in seg_mn or seg_mn in bl_model
if provenance_complete and model_matches and baseline["verified"]:
label = "verified manufacturer TBW (%.1f TB)" % baseline["tbw_terabytes"]
return BaselineTier.VERIFIED, baseline, label, facts
if not model_matches:
facts.append(
"baseline model '%s' does not match current drive '%s'"
" — baseline retained but not applicable"
% (baseline.get("model_string", ""),
current_segment.get("mn", "") if current_segment else "")
)
return BaselineTier.NONE, baseline, "baseline model mismatch", facts
if not provenance_complete:
label = "unverified TBW (%.1f TB) — user-supplied" % baseline["tbw_terabytes"]
return BaselineTier.UNVERIFIED, baseline, label, facts
label = "verified manufacturer TBW (%.1f TB)" % baseline["tbw_terabytes"]
return BaselineTier.VERIFIED, baseline, label, facts
# ---------------------------------------------------------------------------
# Rate computation
# ---------------------------------------------------------------------------
def _compute_regime_rate(days, conn, regime_start_day, clock_now):
regime_bytes = sum(d["bytes_written"] for d in days if d["day"] >= regime_start_day)
regime_start_dt = datetime.fromisoformat(regime_start_day + "T00:00:00+00:00")
regime_wc = _wall_clock_in_range(conn, regime_start_dt, clock_now)
if regime_wc <= 0:
return None, regime_bytes, 0
return regime_bytes / regime_wc, regime_bytes, regime_wc
def _compute_horizon_rate(days, conn, horizon_days, clock_now):
cutoff = (clock_now - timedelta(days=horizon_days)).strftime("%Y-%m-%d")
# History must span the full horizon — no placeholders
if not days or days[0]["day"] > cutoff:
return None
h_bytes = sum(d["bytes_written"] for d in days if d["day"] >= cutoff)
covered = sum(1 for d in days if d["day"] >= cutoff)
if covered == 0:
return None
h_start = datetime.fromisoformat(cutoff + "T00:00:00+00:00")
h_wc = _wall_clock_in_range(conn, h_start, clock_now)
if h_wc <= 0:
return None
return h_bytes / h_wc
# ---------------------------------------------------------------------------
# Habit change detection (§6.4)
# ---------------------------------------------------------------------------
def _detect_habit_change(days):
"""Detect habit change per spec §6.4.
Trailing 7-day mean >= 2x (or <= 0.5x) the preceding 28-day mean
for 3 consecutive days. Returns (change_day, days_since) or None.
The first divergence day is the earliest day in the consecutive run.
"""
need = HABIT_CHANGE_SHORT_WINDOW + HABIT_CHANGE_LONG_WINDOW
if len(days) < need:
return None
def _ratio_at(end_idx):
"""Compute 7-day / preceding-28-day mean ratio ending at end_idx."""
if end_idx < HABIT_CHANGE_SHORT_WINDOW - 1:
return None
se = end_idx + 1
ss = se - HABIT_CHANGE_SHORT_WINDOW
s_bytes = sum(d["bytes_written"] for d in days[ss:se])
s_mean = s_bytes / HABIT_CHANGE_SHORT_WINDOW
le = ss
ls = le - HABIT_CHANGE_LONG_WINDOW
if ls < 0:
return None
l_bytes = sum(d["bytes_written"] for d in days[ls:le])
l_mean = l_bytes / HABIT_CHANGE_LONG_WINDOW
if l_mean == 0:
return None
return s_mean / l_mean
# Scan backwards from the most recent day
for i in range(len(days) - 1, HABIT_CHANGE_LONG_WINDOW + HABIT_CHANGE_SHORT_WINDOW - 2, -1):
ratio = _ratio_at(i)
if ratio is None:
continue
is_upper = ratio >= HABIT_CHANGE_UPPER_FACTOR
is_lower = ratio <= HABIT_CHANGE_LOWER_FACTOR
if not (is_upper or is_lower):
continue
# Count consecutive days going backwards from i
consecutive = 1
for j in range(i - 1, HABIT_CHANGE_LONG_WINDOW + HABIT_CHANGE_SHORT_WINDOW - 3, -1):
r = _ratio_at(j)
if r is None:
break
if (is_upper and r >= HABIT_CHANGE_UPPER_FACTOR) or \
(is_lower and r <= HABIT_CHANGE_LOWER_FACTOR):
consecutive += 1
else:
break
if consecutive >= HABIT_CHANGE_CONSECUTIVE_DAYS:
change_idx = i - consecutive + 1
change_day = days[change_idx]["day"]
days_since = (datetime.fromisoformat(days[-1]["day"]) - datetime.fromisoformat(change_day)).days
return change_day, days_since
return None
# ---------------------------------------------------------------------------
# Confidence rule table (§6.7)
# ---------------------------------------------------------------------------
def _evaluate_confidence(tier, rate, regime_days, days, current_segment,
clock_now, warming_days, warming_low_coverage,
habit_change, staleness_hours, scenario_range):
facts = []
if tier == BaselineTier.NONE:
facts.append("no applicable endurance baseline")
return ConfidenceState.UNSUPPORTED, facts
if rate is None or rate <= 0:
facts.append("no finite projection from this history")
return ConfidenceState.UNSUPPORTED, facts
supported_facts = []
failing = False
# 1. Verified baseline
if tier != BaselineTier.VERIFIED:
failing = True
else:
supported_facts.append("verified manufacturer TBW")
# 2. >= 14 qualifying days
qualifying = sum(1 for d in days if d["coverage"] >= WARMING_COVERAGE_FLOOR and d["sample_count"] > 0)
if qualifying < WARMING_MIN_DAYS:
failing = True
else:
supported_facts.append("%d calendar days" % qualifying)
# 3. Coverage >= 80%
total_wc = len(days) * 86400
total_known = sum(int(d["coverage"] * 86400) for d in days)
avg_cov = total_known / total_wc if total_wc > 0 else 0.0
if avg_cov < SUPPORTED_COVERAGE_FLOOR:
failing = True
else:
supported_facts.append("%d%% interval coverage" % int(avg_cov * 100))
# 4. Fresh (< 48h)
if staleness_hours is not None and staleness_hours > STALENESS_HOURS:
failing = True
elif staleness_hours is not None:
supported_facts.append("recent data")
# 5. Horizon agreement
if scenario_range is not None and len(scenario_range.rates) >= 2:
rl = list(scenario_range.rates.values())
if min(rl) > 0 and max(rl) / min(rl) > HORIZON_AGREEMENT_FACTOR:
failing = True
else:
supported_facts.append("%d weekly cycles" % len(scenario_range.rates))
else:
failing = True
# 6. Burst guard
if not failing and len(days) >= BURST_GUARD_LOOKBACK:
t28 = sum(d["bytes_written"] for d in days[-BURST_GUARD_LOOKBACK:])
for d in days[-BURST_GUARD_LOOKBACK:]:
if t28 > 0 and d["bytes_written"] >= BURST_GUARD_FRACTION * t28:
failing = True
break
if not failing:
supported_facts.append("no burst days")
# 7. Regime >= 7 days
if regime_days < YOUNG_REGIME_DAYS:
failing = True
# 8. Degraded identity
if current_segment and current_segment.get("identity_degraded"):
failing = True
facts.append("controller identity unavailable — replacement detection relies on write-counter continuity only")
if not failing:
return ConfidenceState.SUPPORTED, supported_facts
# Limited
limited_facts = list(supported_facts)
if staleness_hours is not None and staleness_hours > STALENESS_HOURS:
limited_facts.append("newest data %dh old (≥48h)" % staleness_hours)
if habit_change is not None:
limited_facts.append("usage habit changed %d days ago" % habit_change[1])
if regime_days < YOUNG_REGIME_DAYS:
limited_facts.append("regime only %d days old (≥7 required)" % regime_days)
if current_segment and current_segment.get("identity_degraded"):
degraded_fact = "controller identity unavailable — replacement detection relies on write-counter continuity only"
if degraded_fact not in limited_facts:
limited_facts.append(degraded_fact)
return ConfidenceState.LIMITED, limited_facts
# ---------------------------------------------------------------------------
# Disclosure text (§6.11)
# ---------------------------------------------------------------------------
DISCLOSURES = [
"This is an endurance projection, not a predicted hardware-failure date.",
("Percentage Used is vendor-specific; 100 means estimated endurance consumed "
"but may not mean failure, it can exceed 100, and 255 is saturated."),
("Rated TBW can be a warranty/endurance threshold with separate time and "
"eligibility terms, not a failure threshold."),
("DUW is upward-rounded host writes excluding metadata and selected commands, "
"not exact physical NAND writes."),
("Projection quality depends on baseline provenance, history duration and "
"completeness, recentness, stability, and representative usage cycles; "
"future workload and firmware behavior remain outside the observed evidence."),
("Gaps can preserve an aggregate counter delta without preserving hourly "
"timing; unexplained and deliberately disabled periods must be distinguished."),
]
# ---------------------------------------------------------------------------
# PU context line (§6.1)
# ---------------------------------------------------------------------------
def _build_pu_context_line(conn, rate, days, clock_now):
pu = _get_latest_pu(conn)
if pu is None:
return "Percentage Used: unknown"
if rate is None or rate <= 0 or not days:
return "Percentage Used: %d%%" % pu
total_bytes = sum(d["bytes_written"] for d in days)
if total_bytes <= 0:
return "Percentage Used: %d%%" % pu
total_days_count = len(days)
if total_days_count == 0:
return "Percentage Used: %d%%" % pu
pu_daily = total_bytes / total_days_count
obs_daily = rate * 86400
if pu_daily > 0:
ratio = obs_daily / pu_daily
if ratio > WEAR_DISAGREEMENT_FACTOR or ratio < 1.0 / WEAR_DISAGREEMENT_FACTOR:
return ("Percentage Used: %d%% — vendor wear estimate disagrees "
"with observed write rate (>2× difference)") % pu
return "Percentage Used: %d%%" % pu
# ---------------------------------------------------------------------------
# Main projection function
# ---------------------------------------------------------------------------
def compute_projection(conn, clock_now):
facts = []
habit_change_fact = None
warming_fact = None
staleness_fact = None
degraded_identity_fact = None
zero_rate_fact = None
current_segment = _get_current_segment(conn)
tier, baseline, baseline_label, baseline_facts = _resolve_baseline(conn, current_segment)
facts.extend(baseline_facts)
segment_days = _get_days_in_segment(conn, current_segment["opened_at"]) if current_segment else _get_all_days(conn)
all_days = _get_all_days(conn)
regime_start_day = None
habit_change = None
if segment_days:
earliest = segment_days[0]["day"]
cutoff_90 = (clock_now - timedelta(days=90)).strftime("%Y-%m-%d")
regime_start_day = max(earliest, cutoff_90)
habit_change = _detect_habit_change(segment_days)
if habit_change is not None:
regime_start_day = habit_change[0]
habit_change_fact = "usage habit changed %d days ago" % habit_change[1]
facts.append(habit_change_fact)
rate = None
regime_bytes = 0
regime_days_count = 0
if segment_days and regime_start_day is not None:
rate, regime_bytes, _ = _compute_regime_rate(segment_days, conn, regime_start_day, clock_now)
regime_days_count = sum(1 for d in segment_days if d["day"] >= regime_start_day)
if rate is not None and rate <= 0:
zero_rate_fact = "no finite projection from this history"
facts.append(zero_rate_fact)
scenario = None
horizon_rates = {}
for h in HORIZON_DAYS:
hr = _compute_horizon_rate(all_days, conn, h, clock_now)
if hr is not None:
horizon_rates[h] = hr
if horizon_rates:
scenario = ScenarioRange(rates=horizon_rates, min_days=min(horizon_rates), max_days=max(horizon_rates))
total_days_count = len(segment_days)
days_below_coverage = sum(1 for d in segment_days
if d["coverage"] < WARMING_COVERAGE_FLOOR or d["sample_count"] == 0)
qualifying = total_days_count - days_below_coverage
if total_days_count < WARMING_MIN_DAYS or days_below_coverage > WARMING_MAX_LOW_COVERAGE:
warming_fact = "warming up: %d of %d qualifying days" % (qualifying, WARMING_MIN_DAYS)
facts.append(warming_fact)
staleness_hours = None
if segment_days:
newest_dt = datetime.fromisoformat(segment_days[-1]["day"] + "T12:00:00+00:00")
staleness_hours = int((clock_now - newest_dt).total_seconds() / 3600)
if staleness_hours > STALENESS_HOURS:
staleness_fact = "newest data %dh old (≥48h)" % staleness_hours
facts.append(staleness_fact)
if current_segment and current_segment.get("identity_degraded"):
degraded_identity_fact = "controller identity unavailable — replacement detection relies on write-counter continuity only"
facts.append(degraded_identity_fact)
state, conf_facts = _evaluate_confidence(
tier, rate, regime_days_count, segment_days, current_segment, clock_now,
qualifying, 0, habit_change, staleness_hours, scenario,
)
all_facts = list(facts)
for cf in conf_facts:
if cf not in all_facts:
all_facts.append(cf)
headline_seconds = None
if state != ConfidenceState.UNSUPPORTED and rate is not None and rate > 0 and baseline is not None:
if tier in (BaselineTier.VERIFIED, BaselineTier.UNVERIFIED):
E_baseline = baseline["tbw_terabytes"] * TBW_TO_BYTES
elif tier == BaselineTier.IMPLIED:
p = _get_latest_pu(conn)
if p is not None and IMPLIED_P_MIN <= p <= IMPLIED_P_MAX:
E_baseline = 100 * regime_bytes / p
else:
E_baseline = None
else:
E_baseline = None
if E_baseline is not None:
headline_seconds = max(E_baseline - regime_bytes, 0) / rate
pu_line = _build_pu_context_line(conn, rate, segment_days, clock_now)
return ProjectionResult(
confidence_state=state,
contributing_facts=all_facts,
headline_remaining_seconds=headline_seconds,
scenario_range=scenario,
pu_context_line=pu_line,
disclosure_text=list(DISCLOSURES),
baseline_tier=tier,
baseline_label=baseline_label,
regime_days=regime_days_count,
habit_change_fact=habit_change_fact,
warming_fact=warming_fact,
staleness_fact=staleness_fact,
degraded_identity_fact=degraded_identity_fact,
zero_rate_fact=zero_rate_fact,
)
+31
View File
@@ -0,0 +1,31 @@
"""Raw sample pruning per spec §3.4, ST-5.
Raw samples are pruned opportunistically to 14 days.
Hour observations and day aggregates are retained indefinitely.
"""
import sqlite3
from datetime import datetime, timedelta, timezone
# Spec §3.4: Raw-sample retention
RAW_SAMPLE_RETENTION_DAYS = 14
def prune_old_samples(
conn: sqlite3.Connection,
now: datetime,
retention_days: int = RAW_SAMPLE_RETENTION_DAYS,
) -> int:
"""Remove raw samples older than retention_days.
Args:
conn: Connection to the observation store.
now: Current UTC time.
retention_days: Number of days to retain (default 14).
Returns:
Number of samples removed.
"""
cutoff = (now - timedelta(days=retention_days)).isoformat()
cursor = conn.execute("DELETE FROM samples WHERE ts < ?", (cutoff,))
conn.commit()
return cursor.rowcount
+147
View File
@@ -0,0 +1,147 @@
"""Controller segment management.
Handles identity-based segmentation of observation history:
- Find current (most recent) segment
- Determine if a new segment should open
- Open new segments with frozen metadata snapshot
Segmentation axes (independent):
- Identity key change → quarantines prior history
- DUW decrease with unchanged identity → new segment, prior history stays as habit evidence
Blank-key semantics (PR-16):
- To/from blank is an identity change → quarantines
- Equal blanks continue the segment, segmented by DUW monotonicity alone
"""
import sqlite3
from datetime import datetime
from typing import Any, Dict, Optional, Tuple
from .collector import normalize_identity
def find_current_segment(conn: sqlite3.Connection) -> Optional[Dict[str, Any]]:
"""Find the most recent (open) controller segment.
Returns the segment dict or None if no segments exist.
"""
cursor = conn.execute(
"SELECT id, opened_at, identity_key, identity_degraded, "
"subnqn, sn, mn, fr, vid, ssvid, transport "
"FROM controller_segments ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0],
"opened_at": row[1],
"identity_key": row[2],
"identity_degraded": bool(row[3]),
"subnqn": row[4],
"sn": row[5],
"mn": row[6],
"fr": row[7],
"vid": row[8],
"ssvid": row[9],
"transport": row[10],
}
def get_last_duw(conn: sqlite3.Connection, segment_id: int) -> Optional[int]:
"""Get the bytes_written from the most recent sample in a segment.
Returns None if no samples exist in the segment.
"""
# Samples don't have a segment_id FK yet, so we need to find the
# latest sample before the segment's opened_at, or the latest sample
# if this is the first segment.
#
# For now, we'll use a simpler approach: get the latest sample's bytes_written.
# TODO: Add segment_id FK to samples table in next schema migration
cursor = conn.execute(
"SELECT bytes_written FROM samples ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
return row[0] if row else None
def should_open_new_segment(
current_segment: Optional[Dict[str, Any]],
new_identity_key: str,
new_bytes_written: int,
conn: sqlite3.Connection,
) -> Tuple[bool, Optional[str]]:
"""Determine if a new segment should open.
Returns (should_open, reason).
reason is None if no new segment, or a string describing why.
"""
# No current segment → must open first segment
if current_segment is None:
return True, "first_segment"
old_key = current_segment["identity_key"] or ""
# Identity key change (including to/from blank)
if old_key != new_identity_key:
return True, "identity_change"
# DUW decrease (counter reset or controller replacement with same identity)
last_duw = get_last_duw(conn, current_segment["id"])
if last_duw is not None and new_bytes_written < last_duw:
return True, "duw_decrease"
# Same identity, DUW non-decreasing → continue segment
return False, None
def open_segment(
conn: sqlite3.Connection,
now: datetime,
identity: Dict[str, Any],
identity_key: str,
identity_degraded: bool,
) -> Dict[str, Any]:
"""Open a new controller segment with frozen metadata snapshot.
The metadata is immutable once frozen.
"""
cursor = conn.execute(
"""
INSERT INTO controller_segments (
opened_at, identity_key, identity_degraded,
subnqn, sn, mn, fr, vid, ssvid, transport
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
now.isoformat(),
identity_key if identity_key else None,
identity_degraded,
identity.get("subnqn") or None,
identity.get("sn") or None,
identity.get("mn") or None,
identity.get("fr") or None,
identity.get("vid") or None,
identity.get("ssvid") or None,
identity.get("transport") or None,
),
)
segment_id = cursor.lastrowid
return {
"id": segment_id,
"opened_at": now.isoformat(),
"identity_key": identity_key if identity_key else None,
"identity_degraded": identity_degraded,
"subnqn": identity.get("subnqn") or None,
"sn": identity.get("sn") or None,
"mn": identity.get("mn") or None,
"fr": identity.get("fr") or None,
"vid": identity.get("vid") or None,
"ssvid": identity.get("ssvid") or None,
"transport": identity.get("transport") or None,
}
+599
View File
@@ -0,0 +1,599 @@
"""Read-only CLI status command: the CLI twin of the TUI (spec §8.8, LC-9, CI-2).
Composes from the observation store (read-only) and allow-listed systemctl
properties: projection facts, four separate service facts (boot enablement,
runtime activity, last collect outcome, freshness), and a journalctl hint on
failure or staleness. Never auto-samples, never prompts.
Freshness constants are defined once here and shared with the TUI (§8.9):
fresh — newest sample within 2 × cadence + AccuracySec + 60 s
missed — between fresh and 48 h
stale — ≥ 48 h
empty — no observations yet
Criteria: LC-9, CI-2, CI-4, FL-4, FL-5, FL-7.
"""
import sqlite3
import subprocess
import sys
from datetime import datetime, timedelta, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
from .projection import compute_projection, ConfidenceState, DISCLOSURES
from .store import SCHEMA_VERSION
# ---------------------------------------------------------------------------
# Freshness constants (§8.9, §8.2)
# ---------------------------------------------------------------------------
CADENCE_DEFAULT_S = 300 # 5 min
ACCURACY_SEC = 30
FRESH_THRESHOLD_S = 2 * CADENCE_DEFAULT_S + ACCURACY_SEC + 60 # 690 s
STALENESS_THRESHOLD_S = 48 * 3600 # 48 h
# ---------------------------------------------------------------------------
# Configuration reading (§8.3)
# ---------------------------------------------------------------------------
CONFIG_PATH = Path("/etc/fenris/fenris.conf")
def read_config() -> Dict[str, Any]:
"""Read the world-readable configuration file.
Returns a dict with at least 'device'.
Raises ConfigError with a reason string on any failure.
"""
if not CONFIG_PATH.exists():
raise ConfigError("configuration file not found at %s" % CONFIG_PATH)
try:
text = CONFIG_PATH.read_text()
except OSError as e:
raise ConfigError("cannot read %s: %s" % (CONFIG_PATH, e))
device = None
for line in text.splitlines():
line = line.strip()
if not line or line.startswith("#"):
continue
if "=" in line:
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip('"').strip("'")
if key == "device":
device = value
break
if not device:
raise ConfigError("no device selector in %s" % CONFIG_PATH)
return {"device": device}
class ConfigError(Exception):
"""Configuration is invalid — surfaced in status as a fact (§8.3)."""
pass
# ---------------------------------------------------------------------------
# Store opening (read-only, §3, §9.4, §9.5)
# ---------------------------------------------------------------------------
def open_store_readonly(store_path: Path) -> sqlite3.Connection:
"""Open the observation store read-only.
Raises StoreFault if unreadable, NewerSchema if user_version > SCHEMA_VERSION.
"""
if not store_path.exists():
raise StoreFault("observation store not found at %s" % store_path)
try:
conn = sqlite3.connect("file:%s?mode=ro" % store_path, uri=True)
conn.row_factory = sqlite3.Row
except sqlite3.Error as e:
raise StoreFault("observation store unreadable: %s" % e)
try:
cursor = conn.execute("PRAGMA user_version")
version = cursor.fetchone()[0]
except sqlite3.Error as e:
conn.close()
raise StoreFault("observation store unreadable: %s" % e)
if version > SCHEMA_VERSION:
conn.close()
raise NewerSchema(version)
return conn
class StoreFault(Exception):
"""Store is present but unreadable or corrupt (§9.4)."""
pass
class NewerSchema(Exception):
"""Store has a newer user_version (§9.5)."""
def __init__(self, version: int):
self.version = version
super().__init__("schema version %d" % version)
# ---------------------------------------------------------------------------
# Service state queries (§8.8 — allow-listed systemctl properties)
# ---------------------------------------------------------------------------
def _systemctl_show(unit: str, *properties: str) -> Dict[str, str]:
"""Query systemctl show for specific properties. Returns empty dict on failure."""
try:
result = subprocess.run(
["systemctl", "show", unit, "--property=" + ",".join(properties)],
capture_output=True, text=True, timeout=5,
)
if result.returncode != 0:
return {}
out = {}
for line in result.stdout.splitlines():
if "=" in line:
key, _, value = line.partition("=")
out[key.strip()] = value.strip()
return out
except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
return {}
def _journalctl_hint(unit: str, lines: int = 5) -> Optional[str]:
"""Get the last N journal lines for a unit. Returns None on failure."""
try:
result = subprocess.run(
["journalctl", "-u", unit, "--no-pager", "-n", str(lines), "--output=short-iso"],
capture_output=True, text=True, timeout=5,
)
if result.returncode != 0 or not result.stdout.strip():
return None
return result.stdout.strip()
except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
return None
def query_service_state() -> Dict[str, Any]:
"""Query systemctl for the four separate service facts (§7.3, LC-9).
Returns dict with keys:
boot_enabled: bool
timer_active: bool
last_collect_ok: Optional[bool]
last_collect_age_s: Optional[int]
last_collect_reason: Optional[str]
"""
timer_props = _systemctl_show(
"fenris-collect.timer",
"UnitFileState", "ActiveState", "LastTriggerUSec",
)
service_props = _systemctl_show(
"fenris-collect.service",
"ActiveState", "ExecMainStatus", "ExecMainExitTimestamp",
)
boot_enabled_str = timer_props.get("UnitFileState", "")
boot_enabled = boot_enabled_str == "enabled"
active_state = timer_props.get("ActiveState", "inactive")
timer_active = active_state == "active"
last_collect_ok = None
last_collect_age_s = None
last_collect_reason = None
last_trigger = timer_props.get("LastTriggerUSec", "")
if last_trigger and last_trigger != "n/a":
try:
trigger_dt = datetime.fromisoformat(last_trigger.replace("Z", "+00:00"))
now = datetime.now(timezone.utc)
last_collect_age_s = int((now - trigger_dt).total_seconds())
except (ValueError, TypeError):
pass
exec_status = service_props.get("ExecMainStatus", "")
if exec_status:
try:
exit_code = int(exec_status)
last_collect_ok = exit_code == 0
if exit_code != 0:
last_collect_reason = "exit code %d" % exit_code
except (ValueError, TypeError):
pass
return {
"boot_enabled": boot_enabled,
"timer_active": timer_active,
"last_collect_ok": last_collect_ok,
"last_collect_age_s": last_collect_age_s,
"last_collect_reason": last_collect_reason,
}
# ---------------------------------------------------------------------------
# Freshness grading (§8.9)
# ---------------------------------------------------------------------------
def grade_freshness(newest_sample_ts: Optional[str], clock_now: datetime) -> str:
"""Grade freshness from the newest sample timestamp (never a stored flag).
Returns 'fresh', 'missed', 'stale', or 'empty'.
"""
if newest_sample_ts is None:
return "empty"
try:
ts = datetime.fromisoformat(newest_sample_ts)
if ts.tzinfo is None:
ts = ts.replace(tzinfo=timezone.utc)
else:
ts = ts.astimezone(timezone.utc)
except (ValueError, TypeError):
return "empty"
age_s = (clock_now - ts).total_seconds()
if age_s <= FRESH_THRESHOLD_S:
return "fresh"
elif age_s < STALENESS_THRESHOLD_S:
return "missed"
else:
return "stale"
def freshness_age_human(age_s: Optional[int]) -> str:
"""Human-readable age string for freshness fact."""
if age_s is None:
return "unknown age"
if age_s < 60:
return "%ds ago" % age_s
if age_s < 3600:
return "%dm ago" % (age_s // 60)
if age_s < 86400:
return "%dh %dm ago" % (age_s // 3600, (age_s % 3600) // 60)
return "%dd ago" % (age_s // 86400)
# ---------------------------------------------------------------------------
# Drive anomalies (§9.7 — FL-7)
# ---------------------------------------------------------------------------
def _query_drive_facts(conn: sqlite3.Connection) -> List[str]:
"""Query drive-reported anomalies from the latest sample (§9.7, FL-7).
critical_warning, media errors, and unsafe shutdowns render as ordinary
facts and never affect the projection.
"""
cursor = conn.execute(
"SELECT critical_warning, media_errors, unsafe_shutdowns, "
"temperature_c, available_spare "
"FROM samples ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return []
facts = []
cw = row[0]
if cw and cw != 0:
facts.append("critical warning: %s" % hex(cw) if isinstance(cw, int) else str(cw))
me = row[1]
if me and me > 0:
facts.append("media errors: %d" % me)
us = row[2]
if us and us > 0:
facts.append("unsafe shutdowns: %d" % us)
return facts
# ---------------------------------------------------------------------------
# Retired command rejection (§8.8)
# ---------------------------------------------------------------------------
RETIRED_COMMANDS = {
"start": "use 'fenris monitor resume' to enable monitoring",
"stop": "use 'fenris monitor pause' to disable monitoring",
"run": "use 'fenris monitor resume' to enable monitoring; the timer runs in the background",
}
MIGRATION_POINTERS = {
"--device": "device is configured in /etc/fenris/fenris.conf",
}
def check_retired_command(cmd: str) -> Optional[str]:
"""Check if a command is retired and return the migration pointer, or None."""
return RETIRED_COMMANDS.get(cmd)
def check_retired_flag(flag: str) -> Optional[str]:
"""Check if a flag is retired and return the migration pointer, or None."""
return MIGRATION_POINTERS.get(flag)
# ---------------------------------------------------------------------------
# Formatting
# ---------------------------------------------------------------------------
def _format_projection(proj, freshness: str, service: Dict[str, Any],
drive_facts: List[str], config_error: Optional[str],
store_fault: Optional[str], newer_schema: Optional[str],
journal_hint: Optional[str]) -> str:
"""Format the complete status output."""
lines = []
# --- Store/system fault overrides (§9.4, §9.5) ---
if store_fault:
lines.append("observation store unreadable")
if journal_hint:
lines.append("")
lines.append("Recent journal entries:")
lines.append(journal_hint)
return "\n".join(lines)
if newer_schema:
lines.append("observation store written by a newer Fenris — upgrade Fenris")
return "\n".join(lines)
# --- Configuration error (§8.3) ---
if config_error:
lines.append("configuration error: %s" % config_error)
lines.append("")
# --- Empty store (§8.9) ---
if freshness == "empty":
lines.append("no observations yet")
lines.append("")
lines.append("Enable monitoring: fenris monitor resume")
_append_service_facts(lines, service)
return "\n".join(lines)
# --- Projection headline ---
headline = _format_headline(proj)
lines.append(headline)
lines.append("")
# --- Confidence state + contributing facts (§6.7, §6.11) ---
state_label = proj.confidence_state.value
if proj.contributing_facts:
facts_str = " · ".join(proj.contributing_facts)
lines.append("%s evidence · %s" % (state_label, facts_str))
else:
lines.append("%s evidence" % state_label)
lines.append("")
# --- Scenario range (§6.5) ---
if proj.scenario_range and proj.scenario_range.rates:
parts = []
for horizon in sorted(proj.scenario_range.rates.keys()):
rate_gb_day = proj.scenario_range.rates[horizon] * 86400 / 1e9
parts.append("%dd: %.2f GB/day" % (horizon, rate_gb_day))
lines.append("scenario range: %s" % " · ".join(parts))
lines.append("")
# --- PU context line (§6.1) ---
lines.append(proj.pu_context_line)
lines.append("")
# --- Drive anomalies (§9.7, FL-7) ---
if drive_facts:
for fact in drive_facts:
lines.append(fact)
lines.append("")
# --- Four separate service facts (§7.3, LC-9) ---
_append_service_facts(lines, service)
# --- Journal hint on failure or staleness (§8.8) ---
if journal_hint:
if freshness in ("missed", "stale"):
lines.append("")
lines.append("Recent journal entries:")
lines.append(journal_hint)
return "\n".join(lines)
def _format_headline(proj) -> str:
"""Format the lifespan headline or its no-projection wording (§6.11)."""
if proj.headline_remaining_seconds is None:
if proj.zero_rate_fact:
return "no finite projection from this history"
if proj.warming_fact:
return proj.warming_fact
return "no projection available"
secs = proj.headline_remaining_seconds
if secs <= 0:
return "endurance exhausted"
# Human-readable time
years = int(secs // 31557600)
rem = secs % 31557600
days = int(rem // 86400)
rem %= 86400
hours = int(rem // 3600)
parts = []
if years:
parts.append("%d yr" % years)
if days or years:
parts.append("%d d" % days)
parts.append("%d h" % hours)
remaining_human = " ".join(parts)
# Regime line
regime_parts = []
if proj.regime_days:
regime_parts.append("sustained regime: %d days" % proj.regime_days)
headline = "%s remaining" % remaining_human
if regime_parts:
headline += " · %s" % " · ".join(regime_parts)
return headline
def _append_service_facts(lines: List[str], service: Dict[str, Any]) -> None:
"""Append the four separate service facts (§7.3, LC-9)."""
boot = "enabled" if service.get("boot_enabled") else "disabled"
activity = "active" if service.get("timer_active") else "inactive"
if service.get("last_collect_ok") is True:
collect = "ok"
elif service.get("last_collect_ok") is False:
collect = "FAILED"
if service.get("last_collect_reason"):
collect += " (%s)" % service["last_collect_reason"]
else:
collect = "unknown"
collect_age = ""
if service.get("last_collect_age_s") is not None:
collect_age = " %s" % freshness_age_human(service["last_collect_age_s"])
freshness_str = service.get("freshness", "unknown")
freshness_age = ""
if service.get("freshness_age_s") is not None:
freshness_age = " (%s)" % freshness_age_human(service["freshness_age_s"])
lines.append("boot: %s · timer: %s · last collect: %s%s · freshness: %s%s"
% (boot, activity, collect, collect_age, freshness_str, freshness_age))
def format_disclosures() -> str:
"""Format the six disclosures (§6.11, CI-4)."""
lines = []
lines.append("Disclosures")
lines.append("")
for i, disc in enumerate(DISCLOSURES, 1):
lines.append("%d. %s" % (i, disc))
return "\n".join(lines)
# ---------------------------------------------------------------------------
# Main status entry point
# ---------------------------------------------------------------------------
def get_status(store_path: Optional[Path] = None, clock_now: Optional[datetime] = None,
query_services: bool = True, query_journal: bool = True) -> str:
"""Render the complete read-only status (§8.8, LC-9).
This is the single entry point for 'fenris status'. It never auto-samples,
never prompts, and never writes to the store.
"""
if clock_now is None:
clock_now = datetime.now(timezone.utc)
# --- Configuration (§8.3) ---
config_error = None
device = None
try:
config = read_config()
device = config["device"]
except ConfigError as e:
config_error = str(e)
# --- Service state ---
service = {}
if query_services:
service = query_service_state()
# --- Store open ---
store_fault = None
newer_schema = None
conn = None
if store_path is None:
store_path = Path("/var/lib/fenris/observations.db")
try:
conn = open_store_readonly(store_path)
except StoreFault as e:
store_fault = str(e)
except NewerSchema as e:
newer_schema = str(e)
# --- Store fault / newer schema short-circuit ---
if store_fault or newer_schema:
journal_hint = None
if query_journal:
journal_hint = _journalctl_hint("fenris-collect.service")
service["freshness"] = "unknown"
service["freshness_age_s"] = None
return _format_projection(
None, "unknown", service, [], config_error, store_fault, newer_schema, journal_hint
)
# --- Freshness grading (§8.9) ---
try:
cursor = conn.execute("SELECT ts FROM samples ORDER BY id DESC LIMIT 1")
row = cursor.fetchone()
newest_ts = row[0] if row else None
except sqlite3.Error:
newest_ts = None
freshness = grade_freshness(newest_ts, clock_now)
# Freshness age for the service fact
freshness_age_s = None
if newest_ts:
try:
ts = datetime.fromisoformat(newest_ts)
if ts.tzinfo is None:
ts = ts.replace(tzinfo=timezone.utc)
freshness_age_s = int((clock_now - ts).total_seconds())
except (ValueError, TypeError):
pass
service["freshness"] = freshness
service["freshness_age_s"] = freshness_age_s
# --- Drive anomalies (§9.7, FL-7) ---
drive_facts = []
try:
drive_facts = _query_drive_facts(conn)
except sqlite3.Error:
pass
# --- Projection (§6 — recomputed on read, never stored) ---
try:
proj = compute_projection(conn, clock_now)
except Exception:
proj = None
# --- Journal hint on failure or staleness (§8.8) ---
journal_hint = None
if query_journal and freshness in ("missed", "stale"):
journal_hint = _journalctl_hint("fenris-collect.service")
# --- Compose output ---
result = _format_projection(
proj, freshness, service, drive_facts, config_error,
None, None, journal_hint,
)
conn.close()
return result
def render_status(store_path: Optional[Path] = None, clock_now: Optional[datetime] = None,
query_services: bool = True, query_journal: bool = True,
show_disclosures: bool = False) -> str:
"""High-level status renderer: status + optional disclosures.
Used by the CLI entry point.
"""
parts = [get_status(store_path, clock_now, query_services, query_journal)]
if show_disclosures:
parts.append("")
parts.append(format_disclosures())
return "\n\n".join(parts)
+196
View File
@@ -0,0 +1,196 @@
"""Observation store: SQLite database for persisting observation history.
This module handles:
- Store initialization with WAL mode
- Schema versioning with PRAGMA user_version
- The six entities: samples, hour_observations, day_aggregates,
monitoring_periods, controller_segments, endurance_baseline
"""
import sqlite3
from pathlib import Path
from typing import Optional
# Schema version - increment on each migration
SCHEMA_VERSION = 1
def get_store_path(config: dict) -> Path:
"""Get the store path from config."""
return Path(config["store_path"])
def init_store(store_path: Path) -> sqlite3.Connection:
"""Initialize the observation store if not present.
Creates the database with WAL mode and all six entities.
Returns a connection to the store.
"""
conn = sqlite3.connect(str(store_path))
# Enable WAL mode for concurrent reads during writes
conn.execute("PRAGMA journal_mode=WAL")
# Check if this is a new database
cursor = conn.execute("PRAGMA user_version")
current_version = cursor.fetchone()[0]
if current_version == 0:
# New database - create schema
_create_schema(conn)
conn.execute(f"PRAGMA user_version={SCHEMA_VERSION}")
conn.commit()
elif current_version > SCHEMA_VERSION:
# Unknown newer version - refuse
conn.close()
raise ValueError(
f"Observation store written by a newer Fenris (version {current_version}) "
f"— upgrade Fenris"
)
elif current_version < SCHEMA_VERSION:
# Older version - apply migrations
_apply_migrations(conn, current_version)
conn.execute(f"PRAGMA user_version={SCHEMA_VERSION}")
conn.commit()
return conn
def _create_schema(conn: sqlite3.Connection):
"""Create the initial schema with all six entities."""
# Samples: raw collection runs (14-day retention)
conn.execute("""
CREATE TABLE IF NOT EXISTS samples (
id INTEGER PRIMARY KEY AUTOINCREMENT,
ts TEXT NOT NULL, -- ISO 8601 UTC timestamp
device TEXT NOT NULL,
-- Normalized controller-identity fields captured at acquisition
subnqn TEXT,
sn TEXT,
mn TEXT,
fr TEXT,
capacity_bytes INTEGER,
percentage_used INTEGER,
available_spare INTEGER,
media_errors INTEGER,
power_on_hours INTEGER,
power_cycles INTEGER,
unsafe_shutdowns INTEGER,
temperature_c INTEGER,
data_units_written INTEGER,
data_units_read INTEGER,
bytes_written INTEGER,
bytes_read INTEGER,
critical_warning INTEGER
)
""")
# Hour observations: UTC-hour usage-habit split
conn.execute("""
CREATE TABLE IF NOT EXISTS hour_observations (
id INTEGER PRIMARY KEY AUTOINCREMENT,
hour TEXT NOT NULL UNIQUE, -- ISO 8601 UTC hour (e.g., "2026-09-01T12:00:00Z")
active_seconds INTEGER DEFAULT 0,
idle_seconds INTEGER DEFAULT 0,
powered_off_seconds INTEGER DEFAULT 0,
unknown_seconds INTEGER DEFAULT 0,
bytes_written_delta INTEGER DEFAULT 0,
bytes_read_delta INTEGER DEFAULT 0,
temperature_min INTEGER,
temperature_avg REAL,
temperature_max INTEGER,
sample_count INTEGER DEFAULT 0,
coverage REAL DEFAULT 0.0
)
""")
# Day aggregates: derived from hour observations
conn.execute("""
CREATE TABLE IF NOT EXISTS day_aggregates (
id INTEGER PRIMARY KEY AUTOINCREMENT,
day TEXT NOT NULL UNIQUE, -- ISO 8601 UTC day (e.g., "2026-09-01")
active_seconds INTEGER DEFAULT 0,
idle_seconds INTEGER DEFAULT 0,
powered_off_seconds INTEGER DEFAULT 0,
unknown_seconds INTEGER DEFAULT 0,
bytes_written_delta INTEGER DEFAULT 0,
bytes_read_delta INTEGER DEFAULT 0,
sample_count INTEGER DEFAULT 0,
coverage REAL DEFAULT 0.0
)
""")
# Monitoring periods: tracking when monitoring was enabled/disabled
conn.execute("""
CREATE TABLE IF NOT EXISTS monitoring_periods (
id INTEGER PRIMARY KEY AUTOINCREMENT,
started_at TEXT NOT NULL, -- ISO 8601 UTC timestamp
ended_at TEXT, -- NULL if currently active
end_cause TEXT CHECK(end_cause IN ('user_disabled', 'migrated', 'unknown_gap'))
)
""")
# Controller segments: identity key plus metadata snapshot
conn.execute("""
CREATE TABLE IF NOT EXISTS controller_segments (
id INTEGER PRIMARY KEY AUTOINCREMENT,
opened_at TEXT NOT NULL, -- ISO 8601 UTC timestamp
identity_key TEXT, -- Normalized identity key (NULL if degraded)
identity_degraded BOOLEAN DEFAULT 0,
subnqn TEXT,
sn TEXT,
mn TEXT,
fr TEXT,
vid TEXT,
ssvid TEXT,
transport TEXT
)
""")
# Endurance baseline: one active row, replaced on edit
conn.execute("""
CREATE TABLE IF NOT EXISTS endurance_baseline (
id INTEGER PRIMARY KEY AUTOINCREMENT,
tbw_terabytes REAL NOT NULL,
source_url TEXT,
document_revision TEXT,
entry_date TEXT,
model_string TEXT,
nominal_capacity_bytes INTEGER,
validated_by TEXT, -- 'user' or 'machine_match'
verified BOOLEAN DEFAULT 0,
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL
)
""")
# Metadata table for store state (e.g., legacy import marker)
conn.execute("""
CREATE TABLE IF NOT EXISTS store_metadata (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
)
""")
def _apply_migrations(conn: sqlite3.Connection, current_version: int):
"""Apply forward-only migrations from current_version to SCHEMA_VERSION."""
# Future migrations will go here
# For now, just upgrade to current version
pass
def is_store_faulty(store_path: Path) -> bool:
"""Check if the store is present but cannot be read or trusted."""
if not store_path.exists():
return False
try:
conn = sqlite3.connect(f"file:{store_path}?mode=ro", uri=True)
conn.execute("PRAGMA user_version")
conn.close()
return False
except sqlite3.Error:
return True
+359
View File
@@ -0,0 +1,359 @@
"""Collector tracer bullet test.
Tests the thinnest complete write path through the system:
- Input: smartctl-JSON fixture, sysfs fixture tree, config fixture, injected clock
- Output: resulting store contents, run outcome
Seam: write side of the observation store database file.
"""
import json
import os
import sqlite3
import tempfile
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, Generator
import pytest
# Add src to path for imports
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.collector import run_collection, AcquisitionError, InvariantViolationError
from fenris.store import init_store, get_store_path
# Fixtures
@pytest.fixture
def smartctl_fixture() -> Dict[str, Any]:
"""Minimal smartctl -a -j output with required fields."""
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0,
"temperature": 35,
"available_spare": 100,
"available_spare_threshold": 10,
"percentage_used": 5,
"data_units_written": 12345678,
"data_units_read": 9876543,
"power_on_hours": 8765,
"power_cycles": 1234,
"unsafe_shutdowns": 5,
"media_errors": 0,
"num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
@pytest.fixture
def sysfs_fixture_tree(tmp_path: Path) -> Path:
"""Create a minimal sysfs fixture tree with controller identity."""
ctrl_dir = tmp_path / "sys" / "class" / "nvme" / "nvme0"
ctrl_dir.mkdir(parents=True)
# Controller identity files
(ctrl_dir / "subsysnqn").write_text("nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc\n")
(ctrl_dir / "model").write_text("Samsung SSD 970 EVO Plus 1TB\n")
(ctrl_dir / "serial").write_text("S4EWNX0N123456\n")
(ctrl_dir / "firmware_rev").write_text("2B2QEXM7\n")
# Transport info (optional, but we'll include it)
transport_dir = ctrl_dir / "transport"
transport_dir.mkdir()
(transport_dir / "address").write_text("0000:03:00.0")
(transport_dir / "trstring").write_text("pcie")
return tmp_path
@pytest.fixture
def config_fixture(tmp_path: Path) -> Dict[str, Any]:
"""Configuration fixture naming the device."""
return {
"device": "/dev/nvme0",
"store_path": str(tmp_path / "observations.db"),
}
@pytest.fixture
def clock_fixture():
"""Injected clock returning fixed time."""
class FakeClock:
def __init__(self):
self.now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
def utcnow(self):
return self.now
return FakeClock()
# Test: Collector writes one well-formed sample
def test_collector_writes_one_sample(
smartctl_fixture: Dict[str, Any],
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given a smartctl fixture and sysfs fixture tree,
when the collector runs,
then one well-formed sample is written to the observation store."""
result = run_collection(
smartctl_data=smartctl_fixture,
sysfs_path=sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config=config_fixture,
clock=clock_fixture,
)
# Verify run succeeded
assert result["ok"] is True, f"Collection failed: {result.get('error')}"
# Verify store contents
conn = sqlite3.connect(config_fixture["store_path"])
cursor = conn.execute("SELECT COUNT(*) FROM samples")
count = cursor.fetchone()[0]
assert count == 1
cursor = conn.execute("SELECT * FROM samples")
row = cursor.fetchone()
assert row is not None
# Verify row contents match fixtures
# Row structure: id, ts, device, subnqn, sn, mn, fr, capacity_bytes,
# percentage_used, available_spare, media_errors, power_on_hours, power_cycles,
# unsafe_shutdowns, temperature_c, data_units_written, data_units_read,
# bytes_written, bytes_read, critical_warning
assert row[2] == "/dev/nvme0" # device
assert row[3] == "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc" # subnqn
assert row[4] == "S4EWNX0N123456" # sn
assert row[5] == "Samsung SSD 970 EVO Plus 1TB" # mn
assert row[6] == "2B2QEXM7" # fr
assert row[7] == 1024000000000 # capacity_bytes
assert row[8] == 5 # percentage_used
assert row[15] == 12345678 # data_units_written
assert row[17] == 12345678 * 512000 # bytes_written
conn.close()
# Test: Identity normalization applied exactly once at write time
def test_identity_normalization(
smartctl_fixture: Dict[str, Any],
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given sysfs identity with trailing spaces/newlines,
when the collector writes,
then identity is normalized exactly once at write time."""
# Create identity with trailing whitespace
identity = {
"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc \n",
"mn": "Samsung SSD 970 EVO Plus 1TB\n",
"sn": "S4EWNX0N123456\n",
"fr": "2B2QEXM7",
"transport": "pcie",
}
from fenris.collector import normalize_identity
# Normalize once
key1 = normalize_identity(identity)
# Normalize again - should be identical
key2 = normalize_identity(identity)
assert key1 == key2
assert key1 == "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc"
assert "\n" not in key1
assert key1 == key1.rstrip() # No trailing whitespace
# Test: Any acquisition failure fails the whole run
def test_acquisition_failure_fails_run(
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given a smartctl fixture with missing fields,
when the collector runs,
then the whole run fails and writes nothing."""
# Missing required field
bad_smartctl = {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3]},
# Missing nvme_smart_health_information_log
}
result = run_collection(
smartctl_data=bad_smartctl,
sysfs_path=sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config=config_fixture,
clock=clock_fixture,
)
# Verify run failed
assert result["ok"] is False
assert "Missing required field" in result["error"]
# Verify nothing was written
if os.path.exists(config_fixture["store_path"]):
conn = sqlite3.connect(config_fixture["store_path"])
cursor = conn.execute("SELECT COUNT(*) FROM samples")
count = cursor.fetchone()[0]
assert count == 0
conn.close()
else:
# Store wasn't even created - also valid
pass
# Test: Store initializes with six entities
def test_store_initialization(config_fixture: Dict[str, Any]):
"""Given no existing store,
when the collector runs,
then the store is initialized with six entities."""
store_path = Path(config_fixture["store_path"])
# Store shouldn't exist yet
assert not store_path.exists()
# Initialize store
conn = init_store(store_path)
# Verify all six entities exist
cursor = conn.execute("SELECT name FROM sqlite_master WHERE type='table'")
tables = {row[0] for row in cursor.fetchall()}
expected_tables = {
"samples",
"hour_observations",
"day_aggregates",
"monitoring_periods",
"controller_segments",
"endurance_baseline",
}
# sqlite_sequence is a system table created by AUTOINCREMENT
expected_tables.add("sqlite_sequence")
expected_tables.add("store_metadata")
assert expected_tables == tables
conn.close()
# Test: Schema versioning with PRAGMA user_version
def test_schema_versioning(config_fixture: Dict[str, Any]):
"""Given a store with unknown newer version,
when the collector runs,
then it refuses to proceed."""
store_path = Path(config_fixture["store_path"])
# Create a store with newer version
conn = sqlite3.connect(str(store_path))
conn.execute("PRAGMA journal_mode=WAL")
conn.execute("PRAGMA user_version=999") # Unknown newer version
conn.close()
# Try to initialize - should fail
with pytest.raises(ValueError, match="newer Fenris"):
init_store(store_path)
def test_schema_version_current(config_fixture: Dict[str, Any]):
"""Given a store with current version,
when the collector runs,
then it proceeds without migration."""
from fenris.store import SCHEMA_VERSION
store_path = Path(config_fixture["store_path"])
# Initialize store
conn1 = init_store(store_path)
conn1.close()
# Open again - should succeed
conn2 = init_store(store_path)
# Verify version is current
cursor = conn2.execute("PRAGMA user_version")
version = cursor.fetchone()[0]
assert version == SCHEMA_VERSION
conn2.close()
def test_schema_version_older(config_fixture: Dict[str, Any]):
"""Given a store with older version,
when the collector runs,
then it applies migrations and proceeds."""
from fenris.store import SCHEMA_VERSION
store_path = Path(config_fixture["store_path"])
# Create a store with older version
conn = sqlite3.connect(str(store_path))
conn.execute("PRAGMA journal_mode=WAL")
conn.execute("PRAGMA user_version=0") # Older version
conn.close()
# Initialize store - should apply migrations
conn = init_store(store_path)
# Verify version is current
cursor = conn.execute("PRAGMA user_version")
version = cursor.fetchone()[0]
assert version == SCHEMA_VERSION
conn.close()
# Test: Invariant-violating run writes nothing
def test_invariant_violation_writes_nothing(
smartctl_fixture: Dict[str, Any],
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given a smartctl fixture that would violate store invariants,
when the collector runs,
then it writes nothing and fails visibly."""
# Create a fixture that would cause invariant violation
# (negative bytes_written - we'll mock this)
bad_smartctl = smartctl_fixture.copy()
bad_smartctl["nvme_smart_health_information_log"] = {
**smartctl_fixture["nvme_smart_health_information_log"],
"data_units_written": -1, # This will cause negative bytes_written
}
result = run_collection(
smartctl_data=bad_smartctl,
sysfs_path=sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config=config_fixture,
clock=clock_fixture,
)
# Verify run failed
assert result["ok"] is False
assert "InvariantViolation" in result["error_type"]
# Verify nothing was written
if os.path.exists(config_fixture["store_path"]):
conn = sqlite3.connect(config_fixture["store_path"])
cursor = conn.execute("SELECT COUNT(*) FROM samples")
count = cursor.fetchone()[0]
assert count == 0
conn.close()
+234
View File
@@ -0,0 +1,234 @@
"""Day aggregate tests.
From spec §5.4, §3.3, ST-4, ST-5:
- One row per UTC day, derived monotonically from hour rows
- No 23/25-hour days (UTC-bounded, DST never applies)
- Coverage: share of wall-clock seconds inside monitoring periods
whose classification is known
- No absent hour ever interpolated/estimated/fabricated (FL-3)
"""
import sqlite3
import sys
from datetime import datetime, timezone, timedelta
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store
from fenris.monitoring_periods import ensure_period_open, close_period
from fenris.day_aggregate import derive_day, derive_all_days, DayAggregate
from fenris.hour_classify import HourSplit, ACTIVE_THRESHOLD_BYTES
@pytest.fixture
def store_conn(tmp_path: Path):
db_path = tmp_path / "test.db"
conn = init_store(db_path)
yield conn
conn.close()
def _insert_hour(conn, hour_iso, split, bytes_written_delta=0, bytes_read_delta=0,
sample_count=1, coverage=1.0):
conn.execute(
"INSERT INTO hour_observations "
"(hour, active_seconds, idle_seconds, powered_off_seconds, unknown_seconds, "
" bytes_written_delta, bytes_read_delta, sample_count, coverage) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
(
hour_iso,
split.seconds_active,
split.seconds_idle,
split.seconds_powered_off,
split.seconds_unknown,
bytes_written_delta,
bytes_read_delta,
sample_count,
coverage,
),
)
conn.commit()
class TestDeriveDay:
"""Spec §5.4: One row per UTC day, derived monotonically from hour rows."""
def test_single_hour_day(self, store_conn):
t_start = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
hour = "2026-09-01T12:00:00+00:00"
split = HourSplit(seconds_active=3600, seconds_idle=0, seconds_powered_off=0, seconds_unknown=0)
_insert_hour(store_conn, hour, split, bytes_written_delta=1024*1024*100)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.day == "2026-09-01"
assert day.seconds_active == 3600
assert day.seconds_idle == 0
assert day.seconds_powered_off == 0
assert day.bytes_written_delta == 1024*1024*100
assert day.sample_count == 1
def test_multiple_hours(self, store_conn):
t_start = datetime(2026, 9, 1, 10, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
hours = [
("2026-09-01T10:00:00+00:00", HourSplit(1800, 1800, 0, 0), 50*1024*1024),
("2026-09-01T11:00:00+00:00", HourSplit(3600, 0, 0, 0), 200*1024*1024),
("2026-09-01T12:00:00+00:00", HourSplit(0, 0, 3600, 0), 0),
]
for h, s, bw in hours:
_insert_hour(store_conn, h, s, bytes_written_delta=bw)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.seconds_active == 1800 + 3600 + 0
assert day.seconds_idle == 1800 + 0 + 0
assert day.seconds_powered_off == 0 + 0 + 3600
assert day.bytes_written_delta == 50*1024*1024 + 200*1024*1024 + 0
assert day.sample_count == 3
def test_no_hours_returns_none(self, store_conn):
day = derive_day(store_conn, "2026-09-01")
assert day is None
class TestUtcBounded:
"""Spec §3.3, ST-4: Hours and days are UTC-bounded. No 23/25-hour days."""
def test_day_keys_are_utc_date_strings(self, store_conn):
t_start = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
hour = "2026-09-01T12:00:00+00:00"
split = HourSplit(0, 0, 3600, 0)
_insert_hour(store_conn, hour, split)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.day == "2026-09-01"
def test_cross_midnight_hours_produce_two_days(self, store_conn):
# Period spans both days
t_start = datetime(2026, 9, 1, 23, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 2, 1, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
hours = [
("2026-09-01T23:00:00+00:00", HourSplit(3600, 0, 0, 0), 100),
("2026-09-02T00:00:00+00:00", HourSplit(0, 0, 3600, 0), 0),
]
for h, s, bw in hours:
_insert_hour(store_conn, h, s, bytes_written_delta=bw)
days = derive_all_days(store_conn)
day_keys = [d.day for d in days]
assert "2026-09-01" in day_keys
assert "2026-09-02" in day_keys
assert len(days) == 2
class TestCoverage:
"""Spec §5.2, §5.3, PR-5: Coverage is known share of wall-clock seconds
inside monitoring periods."""
def test_full_coverage_all_known(self, store_conn):
"""All hours in period classified -> coverage = 1.0."""
t_start = datetime(2026, 9, 1, 0, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 3, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
hours = [
("2026-09-01T00:00:00+00:00", HourSplit(3600, 0, 0, 0), 100),
("2026-09-01T01:00:00+00:00", HourSplit(0, 3600, 0, 0), 50),
("2026-09-01T02:00:00+00:00", HourSplit(0, 0, 3600, 0), 0),
]
for h, s, bw in hours:
_insert_hour(store_conn, h, s, bytes_written_delta=bw)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.coverage == pytest.approx(1.0)
def test_gap_reduces_coverage(self, store_conn):
"""Missing hour -> unknown seconds reduce coverage."""
t_start = datetime(2026, 9, 1, 0, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 3, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
_insert_hour(store_conn, "2026-09-01T00:00:00+00:00",
HourSplit(3600, 0, 0, 0), bytes_written_delta=100)
_insert_hour(store_conn, "2026-09-01T02:00:00+00:00",
HourSplit(0, 3600, 0, 0), bytes_written_delta=50)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.coverage == pytest.approx(7200 / 10800)
assert day.seconds_unknown == 3600
def test_outside_period_excluded(self, store_conn):
"""Hours outside any monitoring period excluded from denominator."""
t_start = datetime(2026, 9, 1, 1, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 2, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
# Hour 01 is inside the period
_insert_hour(store_conn, "2026-09-01T01:00:00+00:00",
HourSplit(3600, 0, 0, 0), bytes_written_delta=100)
# Hour 02 is outside (period ends at 02:00)
_insert_hour(store_conn, "2026-09-01T02:00:00+00:00",
HourSplit(0, 3600, 0, 0), bytes_written_delta=50)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.coverage == pytest.approx(1.0)
def test_unknown_in_period_reduces_coverage(self, store_conn):
"""Unknown seconds inside period count in denominator but not numerator."""
t_start = datetime(2026, 9, 1, 0, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 1, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
_insert_hour(store_conn, "2026-09-01T00:00:00+00:00",
HourSplit(1800, 0, 0, 1800), bytes_written_delta=100)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.coverage == pytest.approx(0.5)
assert day.seconds_unknown == 1800
class TestNoFabrication:
"""Spec §5.3, FL-3: No absent hour is ever interpolated/estimated/fabricated."""
def test_missing_hours_stay_unknown(self, store_conn):
"""Gap hours are never filled in — they remain as unknown seconds."""
t_start = datetime(2026, 9, 1, 0, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 3, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
# Only hour 02 — hours 00 and 01 are gaps
_insert_hour(store_conn, "2026-09-01T02:00:00+00:00",
HourSplit(3600, 0, 0, 0), bytes_written_delta=100)
day = derive_day(store_conn, "2026-09-01")
assert day is not None
assert day.seconds_unknown == 2 * 3600
assert day.seconds_active == 3600
assert day.sample_count == 1
+218
View File
@@ -0,0 +1,218 @@
"""Hour classification tests.
From spec §5.1 — each UTC hour is classified by named constants:
- Powered-off: power-on-hours delta < 90% of wall-clock span
- Active: DUW delta >= 256 MiB
- Idle: powered on + sampled + below active threshold
- Unknown: everything else
Four splits sum to exactly 3600s. Disabled time is never an hour state.
"""
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.hour_classify import classify_hour, HourSplit, ACTIVE_THRESHOLD_BYTES
HOUR_SECONDS = 3600
class TestPoweredOff:
"""Spec §5.1: Powered-off when power-on-hours delta < 90% of wall-clock span."""
def test_below_90_percent_poh_is_powered_off(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=0,
duw_delta=0,
dur_delta=0,
)
assert split.seconds_powered_off == HOUR_SECONDS
assert split.seconds_active == 0
assert split.seconds_idle == 0
assert split.seconds_unknown == 0
def test_exactly_90_percent_poh_is_not_powered_off(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=3240,
duw_delta=0,
dur_delta=0,
)
assert split.seconds_powered_off == 0
assert split.seconds_unknown == HOUR_SECONDS
def test_just_below_90_percent_poh_is_powered_off(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=3239,
duw_delta=0,
dur_delta=0,
)
assert split.seconds_powered_off == HOUR_SECONDS
def test_100_percent_poh_is_not_powered_off(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=HOUR_SECONDS,
duw_delta=0,
dur_delta=0,
)
assert split.seconds_powered_off == 0
class TestActive:
"""Spec §5.1: Active when DUW delta >= 256 MiB."""
def test_above_256_mib_is_active(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=HOUR_SECONDS,
duw_delta=ACTIVE_THRESHOLD_BYTES,
dur_delta=1000,
sampled_seconds=HOUR_SECONDS,
)
assert split.seconds_active == HOUR_SECONDS
def test_just_below_256_mib_is_idle(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=HOUR_SECONDS,
duw_delta=ACTIVE_THRESHOLD_BYTES - 1,
dur_delta=1000,
sampled_seconds=HOUR_SECONDS,
)
assert split.seconds_idle == HOUR_SECONDS
assert split.seconds_active == 0
def test_powered_off_takes_priority_over_active_writes(self):
"""Spec §5.1 order: powered-off checked first."""
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=0,
duw_delta=ACTIVE_THRESHOLD_BYTES,
dur_delta=1000,
)
assert split.seconds_powered_off == HOUR_SECONDS
assert split.seconds_active == 0
class TestIdle:
"""Spec §5.1: Idle when powered on, sampled, below active threshold."""
def test_idle_with_writes_below_threshold(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=HOUR_SECONDS,
duw_delta=1024 * 1024, # 1 MiB
dur_delta=500,
sampled_seconds=HOUR_SECONDS,
)
assert split.seconds_idle == HOUR_SECONDS
assert split.seconds_active == 0
def test_idle_no_writes(self):
"""Powered on, sampled, zero writes -> idle."""
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=HOUR_SECONDS,
duw_delta=0,
dur_delta=0,
sampled_seconds=HOUR_SECONDS,
)
assert split.seconds_idle == HOUR_SECONDS
class TestUnknown:
"""Spec §5.1: Unknown when unsampled without POH evidence."""
def test_no_sample_no_writes_is_unknown(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=HOUR_SECONDS,
duw_delta=0,
dur_delta=0,
)
assert split.seconds_unknown == HOUR_SECONDS
def test_partial_unknown(self):
"""Partial hour -> split unknown for unsampled portion."""
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=HOUR_SECONDS,
duw_delta=0,
dur_delta=0,
sampled_seconds=1800,
)
assert split.seconds_unknown == 1800
assert split.seconds_idle == 1800
def test_powered_off_with_no_sample_still_powered_off(self):
"""POH < 90% is powered off regardless of sample status."""
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=0,
duw_delta=0,
dur_delta=0,
sampled_seconds=0,
)
assert split.seconds_powered_off == HOUR_SECONDS
assert split.seconds_unknown == 0
class TestSumTo3600:
"""Spec §5.1: Four splits sum to exactly wall_clock_seconds."""
def test_all_scenarios_sum_to_wall_clock(self):
scenarios = [
{"wall_clock_seconds": 3600, "poh_delta": 0, "duw_delta": 0, "dur_delta": 0},
{"wall_clock_seconds": 3600, "poh_delta": 3600, "duw_delta": 0, "dur_delta": 0},
{"wall_clock_seconds": 3600, "poh_delta": 3600, "duw_delta": ACTIVE_THRESHOLD_BYTES, "dur_delta": 1000, "sampled_seconds": 3600},
{"wall_clock_seconds": 3600, "poh_delta": 3600, "duw_delta": 1000, "dur_delta": 500, "sampled_seconds": 3600},
{"wall_clock_seconds": 3600, "poh_delta": 3239, "duw_delta": 0, "dur_delta": 0},
{"wall_clock_seconds": 1800, "poh_delta": 0, "duw_delta": 0, "dur_delta": 0},
{"wall_clock_seconds": 3600, "poh_delta": 3600, "duw_delta": 0, "dur_delta": 0, "sampled_seconds": 1800},
]
for s in scenarios:
split = classify_hour(**s)
total = (
split.seconds_active
+ split.seconds_idle
+ split.seconds_powered_off
+ split.seconds_unknown
)
assert total == s["wall_clock_seconds"], (
f"Scenario {s}: splits sum to {total}, expected {s['wall_clock_seconds']}"
)
def test_partial_wall_clock(self):
split = classify_hour(
wall_clock_seconds=1800,
poh_delta=0,
duw_delta=0,
dur_delta=0,
)
total = (
split.seconds_active
+ split.seconds_idle
+ split.seconds_powered_off
+ split.seconds_unknown
)
assert total == 1800
assert split.seconds_powered_off == 1800
class TestDisabledTimeNotAnHourState:
"""Spec §5.2: Disabled time is not an hour state — excluded from numerator/denominator."""
def test_classify_hour_has_no_disabled_state(self):
split = classify_hour(
wall_clock_seconds=HOUR_SECONDS,
poh_delta=0,
duw_delta=0,
dur_delta=0,
)
assert not hasattr(split, "seconds_disabled")
+96
View File
@@ -0,0 +1,96 @@
"""Test identity normalization per specification.
From spec §2.3:
- strip trailing spaces and newlines
- no case folding
- empty-after-strip stored blank
Padded and unpadded renderings of the same field yield byte-identical stored values.
"""
import pytest
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.collector import normalize_identity
def test_strip_trailing_spaces():
"""Trailing spaces are stripped."""
identity = {"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678 "}
result = normalize_identity(identity)
assert result == "nqn.2014-08.org.nvmexpress:uuid:12345678"
def test_strip_trailing_newlines():
"""Trailing newlines are stripped."""
identity = {"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678\n"}
result = normalize_identity(identity)
assert result == "nqn.2014-08.org.nvmexpress:uuid:12345678"
def test_strip_trailing_spaces_and_newlines():
"""Trailing spaces and newlines are stripped."""
identity = {"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678 \n\n"}
result = normalize_identity(identity)
assert result == "nqn.2014-08.org.nvmexpress:uuid:12345678"
def test_no_case_folding():
"""Case is preserved - no case folding."""
identity = {"subnqn": "NQN.2014-08.ORG.NVMEXPRESS:UUID:12345678"}
result = normalize_identity(identity)
assert result == "NQN.2014-08.ORG.NVMEXPRESS:UUID:12345678"
def test_empty_after_strip_stored_blank():
"""Empty after strip is stored as blank string."""
identity = {"subnqn": " \n\n "}
result = normalize_identity(identity)
assert result == ""
def test_fallback_to_model_serial():
"""When subnqn is empty, falls back to model|serial."""
identity = {
"subnqn": "",
"mn": "Samsung SSD 970 EVO Plus 1TB",
"sn": "S4EWNX0N123456",
}
result = normalize_identity(identity)
assert result == "Samsung SSD 970 EVO Plus 1TB|S4EWNX0N123456"
def test_fallback_model_serial_normalized():
"""Model and serial are also normalized."""
identity = {
"subnqn": "",
"mn": "Samsung SSD 970 EVO Plus 1TB\n",
"sn": "S4EWNX0N123456 ",
}
result = normalize_identity(identity)
assert result == "Samsung SSD 970 EVO Plus 1TB|S4EWNX0N123456"
def test_all_keys_blank_returns_blank():
"""When all keys are blank, returns blank (degraded identity)."""
identity = {
"subnqn": "",
"mn": "",
"sn": "",
}
result = normalize_identity(identity)
assert result == ""
def test_byte_identical_for_padded_unpadded():
"""Padded and unpadded renderings yield byte-identical values."""
padded = {"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678 \n"}
unpadded = {"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678"}
result_padded = normalize_identity(padded)
result_unpadded = normalize_identity(unpadded)
assert result_padded == result_unpadded
assert result_padded == "nqn.2014-08.org.nvmexpress:uuid:12345678"
+404
View File
@@ -0,0 +1,404 @@
"""Legacy migration tests.
Tests the idempotent, interruption-safe import of history.jsonl into the
observation store.
"""
import json
import os
import sqlite3
import tempfile
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict
import pytest
# Add src to path for imports
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.legacy import import_legacy_history, is_legacy_imported, _parse_history_line
from fenris.store import init_store
# Fixtures
@pytest.fixture
def history_fixture() -> str:
"""Minimal history.jsonl content with two samples."""
samples = [
{
"timestamp": "2026-08-30T10:00:00Z",
"device": "/dev/nvme0",
"model": "Samsung SSD 970 EVO Plus 1TB",
"serial": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
"capacity_bytes": 1024000000000,
"data_units_written": 1000000,
"data_units_read": 500000,
"percentage_used": 5,
"power_on_hours": 8765,
"temperature": 35,
"available_spare": 100,
"media_errors": 0,
"power_cycles": 1234,
"unsafe_shutdowns": 5,
"critical_warning": 0,
},
{
"timestamp": "2026-08-30T11:00:00Z",
"device": "/dev/nvme0",
"model": "Samsung SSD 970 EVO Plus 1TB",
"serial": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
"capacity_bytes": 1024000000000,
"data_units_written": 1001000,
"data_units_read": 501000,
"percentage_used": 5,
"power_on_hours": 8766,
"temperature": 36,
"available_spare": 100,
"media_errors": 0,
"power_cycles": 1234,
"unsafe_shutdowns": 5,
"critical_warning": 0,
},
]
return "\n".join(json.dumps(s) for s in samples)
@pytest.fixture
def hourly_fixture() -> str:
"""Minimal hourly.jsonl content for diffing."""
hours = [
{
"hour": "2026-08-30T10:00:00Z",
"bytes_written_delta": 512000000,
"bytes_read_delta": 256000000,
"sample_count": 1,
},
{
"hour": "2026-08-30T11:00:00Z",
"bytes_written_delta": 513000000,
"bytes_read_delta": 257000000,
"sample_count": 1,
},
]
return "\n".join(json.dumps(h) for h in hours)
@pytest.fixture
def config_fixture(tmp_path: Path) -> Dict[str, Any]:
"""Configuration fixture."""
return {
"device": "/dev/nvme0",
"store_path": str(tmp_path / "observations.db"),
"data_dir": str(tmp_path),
}
@pytest.fixture
def clock_fixture():
"""Injected clock returning fixed time."""
class FakeClock:
def __init__(self):
self.now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
def utcnow(self):
return self.now
return FakeClock()
# Test: Idempotency - second run no-ops
def test_import_idempotent(
history_fixture: str,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given a store with legacy import marker,
when import is called again,
then it no-ops."""
# Create history file
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_fixture)
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# First import
result1 = import_legacy_history(conn, history_path, clock=clock_fixture)
assert result1["ok"] is True
assert result1["skipped"] is False
# Second import (should no-op)
history_path2 = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path2.write_text(history_fixture)
result2 = import_legacy_history(conn, history_path2, clock=clock_fixture)
assert result2["ok"] is True
assert result2["skipped"] is True
conn.close()
# Test: Interruption safety - single transaction
def test_import_single_transaction(
history_fixture: str,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given history.jsonl,
when import runs,
then the entire import is a single transaction."""
# Create history file
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_fixture)
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# Import
result = import_legacy_history(conn, history_path, clock=clock_fixture)
assert result["ok"] is True
# Verify all data was imported atomically
cursor = conn.execute("SELECT COUNT(*) FROM samples")
assert cursor.fetchone()[0] == 2
cursor = conn.execute("SELECT COUNT(*) FROM hour_observations")
assert cursor.fetchone()[0] == 2
cursor = conn.execute("SELECT COUNT(*) FROM monitoring_periods")
assert cursor.fetchone()[0] == 1
conn.close()
# Test: Legacy files renamed to *.migrated after commit
def test_legacy_files_renamed(
history_fixture: str,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given history.jsonl and hourly.jsonl,
when import commits,
then files are renamed to *.migrated."""
# Create files
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_fixture)
hourly_path = Path(config_fixture["data_dir"]) / "hourly.jsonl"
hourly_path.write_text("{}")
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# Import
result = import_legacy_history(conn, history_path, hourly_path=hourly_path, clock=clock_fixture)
assert result["ok"] is True
# Verify files renamed
assert not history_path.exists()
assert history_path.with_suffix(history_path.suffix + ".migrated").exists()
assert not hourly_path.exists()
assert hourly_path.with_suffix(hourly_path.suffix + ".migrated").exists()
conn.close()
# Test: Malformed lines quarantined with logged count
def test_malformed_lines_quarantined(
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given history.jsonl with malformed lines,
when import runs,
then malformed lines are quarantined with logged count."""
# Create history with malformed lines
history_content = "\n".join([
'{"timestamp": "2026-08-30T10:00:00Z", "model": "Test", "serial": "123", "firmware_version": "1.0", "data_units_written": 1000, "data_units_read": 500, "percentage_used": 5, "power_on_hours": 100, "temperature": 35}',
'NOT JSON',
'{"timestamp": "2026-08-30T11:00:00Z", "model": "Test", "serial": "123", "firmware_version": "1.0", "data_units_written": 1001, "data_units_read": 501, "percentage_used": 5, "power_on_hours": 101, "temperature": 36}',
])
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_content)
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# Import
result = import_legacy_history(conn, history_path, clock=clock_fixture)
assert result["ok"] is True
assert result["malformed_lines"] == 1
assert result["samples_imported"] == 2
conn.close()
# Test: hourly.jsonl diffed and logged but never trusted
def test_hourly_jsonl_diffed(
history_fixture: str,
hourly_fixture: str,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given history.jsonl and hourly.jsonl with mismatches,
when import runs,
then mismatches are diffed and logged."""
# Create files
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_fixture)
hourly_path = Path(config_fixture["data_dir"]) / "hourly.jsonl"
hourly_path.write_text(hourly_fixture)
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# Import (should not fail even with mismatches)
result = import_legacy_history(conn, history_path, hourly_path=hourly_path, clock=clock_fixture)
assert result["ok"] is True
# Verify data was imported from history.jsonl, not hourly.jsonl
cursor = conn.execute("SELECT bytes_written_delta FROM hour_observations ORDER BY hour")
deltas = [row[0] for row in cursor.fetchall()]
# Should match history.jsonl derived values, not hourly.jsonl
assert len(deltas) == 2
conn.close()
# Test: Legacy identity is mn-only
def test_legacy_identity_mn_only(
history_fixture: str,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given history.jsonl,
when import runs,
then legacy segment has mn-only identity."""
# Create history file
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_fixture)
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# Import
result = import_legacy_history(conn, history_path, clock=clock_fixture)
assert result["ok"] is True
assert result["legacy_identity_key"] == "legacy|Samsung SSD 970 EVO Plus 1TB"
# Verify segment has mn-only identity
cursor = conn.execute("SELECT identity_key, mn, sn, subnqn FROM controller_segments")
row = cursor.fetchone()
assert row[0] == "legacy|Samsung SSD 970 EVO Plus 1TB"
assert row[1] == "Samsung SSD 970 EVO Plus 1TB"
assert row[2] is None # sn is None for legacy
assert row[3] is None # subnqn is None for legacy
conn.close()
# Test: No synthetic baseline created
def test_no_synthetic_baseline(
history_fixture: str,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given history.jsonl,
when import runs,
then no endurance baseline is created."""
# Create history file
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_fixture)
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# Import
result = import_legacy_history(conn, history_path, clock=clock_fixture)
assert result["ok"] is True
# Verify no baseline created
cursor = conn.execute("SELECT COUNT(*) FROM endurance_baseline")
assert cursor.fetchone()[0] == 0
conn.close()
# Test: Monitoring period opened and closed
def test_monitoring_period_opened_closed(
history_fixture: str,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given history.jsonl,
when import runs,
then one monitoring period is opened at first sample and closed at migration."""
# Create history file
history_path = Path(config_fixture["data_dir"]) / "history.jsonl"
history_path.write_text(history_fixture)
# Initialize store
store_path = Path(config_fixture["store_path"])
conn = init_store(store_path)
# Import
result = import_legacy_history(conn, history_path, clock=clock_fixture)
assert result["ok"] is True
# Verify monitoring period
cursor = conn.execute("SELECT started_at, ended_at, end_cause FROM monitoring_periods")
row = cursor.fetchone()
assert row[0] == "2026-08-30T10:00:00+00:00" # First sample time
assert row[1] == clock_fixture.now.isoformat() # Migration time
assert row[2] == "migrated"
conn.close()
# Test: Parse history line
def test_parse_history_line_valid():
"""Given a valid history line,
when parsed,
then returns the record."""
line = '{"timestamp": "2026-08-30T10:00:00Z", "model": "Test", "serial": "123", "firmware_version": "1.0", "data_units_written": 1000, "data_units_read": 500, "percentage_used": 5, "power_on_hours": 100, "temperature": 35}'
result = _parse_history_line(line, 1)
assert result is not None
assert result["model"] == "Test"
def test_parse_history_line_malformed():
"""Given a malformed JSON line,
when parsed,
then returns None."""
result = _parse_history_line("NOT JSON", 1)
assert result is None
def test_parse_history_line_missing_field():
"""Given a line with missing required field,
when parsed,
then returns None."""
line = '{"timestamp": "2026-08-30T10:00:00Z", "model": "Test"}'
result = _parse_history_line(line, 1)
assert result is None
+147
View File
@@ -0,0 +1,147 @@
"""Monitoring period tests.
From spec §5.2, §8.6, §9.8:
- Run finding no open period opens one at the run moment, never backdated
- Wall-clock outside periods excluded from numerator/denominator
- Powered-off time stays inside a period; disabled time does not
- End causes: user_disabled, migrated, unknown_gap
"""
import sqlite3
import sys
from datetime import datetime, timezone, timedelta
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store
from fenris.monitoring_periods import (
ensure_period_open,
close_period,
get_open_period,
is_inside_period,
)
@pytest.fixture
def store_conn(tmp_path: Path):
"""Initialize an observation store and return a connection."""
db_path = tmp_path / "test.db"
conn = init_store(db_path)
yield conn
conn.close()
class TestEnsurePeriodOpen:
"""Spec §9.8: Run finding no open period opens one at the run moment."""
def test_opens_period_when_none_exists(self, store_conn):
"""First collection run opens a period at the run moment."""
run_time = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, run_time)
period = get_open_period(store_conn)
assert period is not None
assert period["started_at"] == "2026-09-01T12:00:00+00:00"
assert period["ended_at"] is None
assert period["end_cause"] is None
def test_no_opener_when_already_open(self, store_conn):
"""If a period is already open, no new period is created."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t1)
ensure_period_open(store_conn, t2)
period = get_open_period(store_conn)
assert period is not None
assert period["started_at"] == "2026-09-01T12:00:00+00:00"
def test_never_backdated(self, store_conn):
"""Period starts at the run moment, not the beginning of the hour."""
run_time = datetime(2026, 9, 1, 12, 3, 45, tzinfo=timezone.utc)
ensure_period_open(store_conn, run_time)
period = get_open_period(store_conn)
assert period["started_at"] == "2026-09-01T12:03:45+00:00"
def test_new_period_after_close(self, store_conn):
"""After closing a period, next run opens a new one."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
t3 = datetime(2026, 9, 1, 14, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t1)
close_period(store_conn, t2, "user_disabled")
ensure_period_open(store_conn, t3)
period = get_open_period(store_conn)
assert period is not None
assert period["started_at"] == "2026-09-01T14:00:00+00:00"
class TestClosePeriod:
"""Spec §8.6: Pause closes with user_disabled."""
def test_close_with_user_disabled(self, store_conn):
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t1)
close_period(store_conn, t2, "user_disabled")
period = get_open_period(store_conn)
assert period is None
cursor = store_conn.execute(
"SELECT ended_at, end_cause FROM monitoring_periods WHERE id = 1"
)
row = cursor.fetchone()
assert row[0] == "2026-09-01T13:00:00+00:00"
assert row[1] == "user_disabled"
def test_close_nonexistent_is_noop(self, store_conn):
"""Closing when no period is open is a no-op (spec §8.6 pause otherwise)."""
t = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
close_period(store_conn, t, "user_disabled")
cursor = store_conn.execute("SELECT COUNT(*) FROM monitoring_periods")
assert cursor.fetchone()[0] == 0
class TestIsInsidePeriod:
"""Spec §5.2: Wall-clock outside periods excluded from numerator/denominator."""
def test_inside_open_period(self, store_conn):
t_start = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t_inside = datetime(2026, 9, 1, 12, 30, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
assert is_inside_period(store_conn, t_inside) is True
def test_outside_closed_period(self, store_conn):
t_start = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t_end = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
t_outside = datetime(2026, 9, 1, 14, 0, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t_start)
close_period(store_conn, t_end, "user_disabled")
assert is_inside_period(store_conn, t_outside) is False
def test_outside_no_periods(self, store_conn):
t = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
assert is_inside_period(store_conn, t) is False
def test_inside_second_period(self, store_conn):
"""Two periods with a gap; time in second period is inside."""
t1_start = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t1_end = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
t2_start = datetime(2026, 9, 1, 14, 0, 0, tzinfo=timezone.utc)
t_inside = datetime(2026, 9, 1, 14, 30, 0, tzinfo=timezone.utc)
ensure_period_open(store_conn, t1_start)
close_period(store_conn, t1_end, "user_disabled")
ensure_period_open(store_conn, t2_start)
assert is_inside_period(store_conn, t_inside) is True
+901
View File
@@ -0,0 +1,901 @@
"""Projection core tests.
Covers acceptance criteria:
- PR-1: Exactly one projection from precedence-chosen baseline; PU context only
- PR-8: Confidence rule table holds verbatim; state + facts, never percentage
- PR-10: Implied baseline eligible only after >=2 PU increments
- PR-11: Zero rate renders fixed phrase; scenario range only spread
- PR-12: Contract hands over exactly: state, facts, headline, scenario, PU, disclosures
- PR-13: Baseline provenance and validation per register
- PR-17: Arithmetic exactly E_rated = TBW * 10^12, E_implied = 100*W/p, projected = max(E-W,0)/rate
"""
import sqlite3
from datetime import datetime, timedelta, timezone
from pathlib import Path
import pytest
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store
from fenris.monitoring_periods import ensure_period_open, close_period
from fenris.projection import (
compute_projection, ConfidenceState, BaselineTier, ScenarioRange,
TBW_TO_BYTES, HORIZON_DAYS, WARMING_MIN_DAYS, STALENESS_HOURS,
YOUNG_REGIME_DAYS, DISCLOSURES,
)
@pytest.fixture
def store(tmp_path):
conn = init_store(tmp_path / "test.db")
yield conn
conn.close()
def _clock(year=2026, month=9, day=30, hour=12):
return datetime(year, month, day, hour, 0, 0, tzinfo=timezone.utc)
def _insert_baseline(conn, tbw_tb=1.0, verified=True, model="Samsung SSD 970 EVO Plus 1TB",
source_url="https://example.com/spec", doc_rev="v1.0",
entry_date="2026-01-01", nominal_cap=1024000000000):
conn.execute(
"INSERT INTO endurance_baseline "
"(tbw_terabytes, source_url, document_revision, entry_date, model_string, "
" nominal_capacity_bytes, validated_by, verified, created_at, updated_at) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(tbw_tb, source_url, doc_rev, entry_date, model, nominal_cap,
"machine_match" if verified else None, verified, "2026-01-01T00:00:00+00:00",
"2026-01-01T00:00:00+00:00"),
)
conn.commit()
def _insert_segment(conn, opened_at="2026-09-01T00:00:00+00:00",
identity_key="nqn.test", degraded=False,
mn="Samsung SSD 970 EVO Plus 1TB"):
conn.execute(
"INSERT INTO controller_segments "
"(opened_at, identity_key, identity_degraded, subnqn, sn, mn, fr, vid, ssvid, transport) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(opened_at, identity_key, degraded, "nqn.test", "SN123", mn, "FW1", "0x144d", "0x144d", "pcie"),
)
conn.commit()
def _insert_day(conn, day, bw=1024*1024*100, coverage=0.95, samples=24):
conn.execute(
"INSERT INTO day_aggregates (day, active_seconds, idle_seconds, powered_off_seconds, "
"unknown_seconds, bytes_written_delta, bytes_read_delta, sample_count, coverage) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
(day, 3600, 0, 0, 0, bw, 0, samples, coverage),
)
conn.commit()
def _insert_sample(conn, ts, pu=5):
conn.execute(
"INSERT INTO samples (ts, device, data_units_written, data_units_read, "
"percentage_used, bytes_written, bytes_read, power_on_hours) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?)",
(ts, "/dev/nvme0n1", 1000000, 500000, pu, 512000000000, 256000000000, 8765),
)
conn.commit()
def _open_period(conn, start="2026-09-01T00:00:00+00:00"):
ensure_period_open(conn, datetime.fromisoformat(start))
class TestPrecedence:
def test_no_baseline_unavailable(self, store):
_insert_segment(store)
_insert_day(store, "2026-09-28", bw=1024*1024*1000)
_open_period(store)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert result.baseline_tier == BaselineTier.NONE
assert result.headline_remaining_seconds is None
def test_verified_baseline_chosen(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(14):
d = (datetime(2026, 9, 15) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.baseline_tier == BaselineTier.VERIFIED
assert "verified manufacturer TBW" in result.baseline_label
def test_pu_is_context_not_second_projection(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=10)
result = compute_projection(store, _clock())
assert result.pu_context_line.startswith("Percentage Used:")
assert "%" in result.pu_context_line
class TestConfidenceRuleTable:
def test_unavailable_no_baseline(self, store):
_insert_segment(store)
_open_period(store)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("no applicable endurance baseline" in f for f in result.contributing_facts)
def test_unavailable_zero_rate(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=0)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("no finite projection" in f for f in result.contributing_facts)
def test_limited_young_regime(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(5):
d = (datetime(2026, 9, 25) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.LIMITED
assert any("regime only" in f and "days old" in f for f in result.contributing_facts)
def test_limited_degraded_identity(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store, degraded=True)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.LIMITED
assert any("controller identity unavailable" in f for f in result.contributing_facts)
def test_state_plus_facts_never_percentage(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.confidence_state in ConfidenceState
for f in result.contributing_facts:
assert "%" not in f or "coverage" in f or "Percentage" in f
class TestImpliedBaseline:
def test_implied_not_chosen_with_verified(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.baseline_tier == BaselineTier.VERIFIED
class TestZeroRate:
def test_zero_rate_fixed_phrase(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=0)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("no finite projection from this history" in f for f in result.contributing_facts)
assert result.headline_remaining_seconds is None
def test_scenario_range_only_spread(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
if result.scenario_range is not None:
assert isinstance(result.scenario_range, ScenarioRange)
assert hasattr(result.scenario_range, "rates")
class TestContractHandoff:
def test_contract_fields_present(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert isinstance(result.confidence_state, ConfidenceState)
assert isinstance(result.contributing_facts, list)
assert isinstance(result.pu_context_line, str)
assert isinstance(result.disclosure_text, list)
assert isinstance(result.baseline_tier, BaselineTier)
assert isinstance(result.baseline_label, str)
def test_recomputed_on_read(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
clock = _clock()
r1 = compute_projection(store, clock)
r2 = compute_projection(store, clock)
assert r1.confidence_state == r2.confidence_state
assert r1.headline_remaining_seconds == r2.headline_remaining_seconds
def test_disclosures_present(self, store):
result = compute_projection(store, _clock())
assert len(result.disclosure_text) == 6
for d in DISCLOSURES:
assert d in result.disclosure_text
class TestBaselineProvenance:
def test_model_mismatch_unavailable(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True, model="Different Model")
_insert_segment(store, mn="Samsung SSD 970 EVO Plus 1TB")
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("does not match" in f for f in result.contributing_facts)
assert result.baseline_tier == BaselineTier.NONE
class TestArithmetic:
def test_rated_tbw_conversion(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store, start="2026-09-01T00:00:00+00:00")
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
if result.headline_remaining_seconds is not None:
E_rated = 1.0 * TBW_TO_BYTES
regime_bytes = 30 * 1024 * 1024 * 100
# Actual wall-clock: Sep 1 00:00 -> Sep 30 12:00 = 29.5 days
period_start = datetime(2026, 9, 1, 0, 0, 0, tzinfo=timezone.utc)
period_end = _clock()
actual_wc = int((period_end - period_start).total_seconds())
rate = regime_bytes / actual_wc
expected = max(E_rated - regime_bytes, 0) / rate
assert abs(result.headline_remaining_seconds - expected) < 1.0
def test_implied_baseline_formula(self, store):
W_t = 1024 * 1024 * 1000
p = 10
E_implied = 100 * W_t / p
assert E_implied == 100 * 1024 * 1024 * 1000 / 10
def test_projected_formula(self, store):
_insert_baseline(store, tbw_tb=2.0, verified=True)
_insert_segment(store)
_open_period(store, start="2026-09-01T00:00:00+00:00")
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
if result.headline_remaining_seconds is not None:
E_rated = 2.0 * TBW_TO_BYTES
regime_bytes = 30 * 1024 * 1024 * 100
period_start = datetime(2026, 9, 1, 0, 0, 0, tzinfo=timezone.utc)
period_end = _clock()
actual_wc = int((period_end - period_start).total_seconds())
rate = regime_bytes / actual_wc
expected = max(E_rated - regime_bytes, 0) / rate
assert abs(result.headline_remaining_seconds - expected) < 1.0
def test_wearing_rate_proportional(self, store):
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
r_slow = compute_projection(store, _clock())
store.execute("DELETE FROM day_aggregates")
store.commit()
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=2*1024*1024*100)
r_fast = compute_projection(store, _clock())
if r_slow.headline_remaining_seconds is not None and r_fast.headline_remaining_seconds is not None:
assert r_fast.headline_remaining_seconds < r_slow.headline_remaining_seconds
# ===========================================================================
# Issue #26: Project from the sustained regime
# Habit change, scenario range, and evidence gates
# ===========================================================================
class TestSustainedRegimeRate:
"""PR-2: Headline rate is sustained-regime rate; default regime = full
history capped at 90 days; scenario range computed independently,
covered horizons only, no placeholders."""
def test_headline_rate_from_regime(self, store):
"""Rate is regime DUW / wall-clock, not trailing-24h or all-history."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
# 30 days of 100 MiB/day
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
# Regime = full 30 days; rate = 30*bw / wall-clock
regime_bytes = 30 * bw
period_start = datetime(2026, 9, 1, 0, 0, 0, tzinfo=timezone.utc)
wc = int((_clock() - period_start).total_seconds())
expected_rate = regime_bytes / wc
if result.headline_remaining_seconds is not None:
E = 10.0 * TBW_TO_BYTES
expected_seconds = max(E - regime_bytes, 0) / expected_rate
assert abs(result.headline_remaining_seconds - expected_seconds) < 1.0
def test_regime_capped_at_90_days(self, store):
"""Default regime is full history capped at 90 days."""
_insert_baseline(store, tbw_tb=100.0, verified=True)
_insert_segment(store, opened_at="2026-06-01T00:00:00+00:00")
_open_period(store, start="2026-06-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
# 120 days of data (Jun 1 - Sep 28)
for i in range(120):
d = (datetime(2026, 6, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-28T12:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=30, hour=12))
# Regime should be capped at 90 days (from Jun 1 to Sep 30 = 90 days at cutoff)
# The 90-day cutoff is Sep 30 - 90 = Jul 1, so regime starts Jul 1
assert result.regime_days is not None
assert result.regime_days <= 90
def test_scenario_range_independent_of_regime(self, store):
"""Scenario range is computed independently from the regime."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
if result.scenario_range is not None:
# Should have 7-day and 28-day horizons (90-day not fully covered)
assert 7 in result.scenario_range.rates
assert 28 in result.scenario_range.rates
def test_only_covered_horizons_shown(self, store):
"""No placeholder horizons — only horizons the history covers."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-20T00:00:00+00:00")
_open_period(store, start="2026-09-20T00:00:00+00:00")
bw = 100 * 1024 * 1024
# Only 10 days of data
for i in range(10):
d = (datetime(2026, 9, 20) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
if result.scenario_range is not None:
# 7-day is covered, 28-day and 90-day are not
assert 7 in result.scenario_range.rates
assert 28 not in result.scenario_range.rates
assert 90 not in result.scenario_range.rates
class TestHabitChange:
"""PR-3: Habit change triggers at 2x/0.5x sustained 3 consecutive days,
regime starts at first divergence day, auto-adopted and labeled;
young regime caps at Limited."""
def test_habit_change_2x_detected(self, store):
"""2x increase for 3+ consecutive days triggers habit change."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-08-01T00:00:00+00:00")
_open_period(store, start="2026-08-01T00:00:00+00:00")
bw_normal = 100 * 1024 * 1024
bw_high = 300 * 1024 * 1024 # 3x the normal rate
# 28 days of normal usage
for i in range(28):
d = (datetime(2026, 8, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_normal)
# 10 days of high usage (3x > 2x threshold)
for i in range(10):
d = (datetime(2026, 8, 29) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_high)
_insert_sample(store, "2026-09-08T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=8, hour=12))
assert result.habit_change_fact is not None
assert "usage habit changed" in result.habit_change_fact
assert "days ago" in result.habit_change_fact
def test_habit_change_05x_detected(self, store):
"""0.5x decrease for 3+ consecutive days triggers habit change."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-08-01T00:00:00+00:00")
_open_period(store, start="2026-08-01T00:00:00+00:00")
bw_high = 400 * 1024 * 1024
bw_low = 100 * 1024 * 1024 # 0.25x < 0.5x threshold
# 28 days of high usage
for i in range(28):
d = (datetime(2026, 8, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_high)
# 10 days of low usage
for i in range(10):
d = (datetime(2026, 8, 29) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_low)
_insert_sample(store, "2026-09-08T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=8, hour=12))
assert result.habit_change_fact is not None
assert "usage habit changed" in result.habit_change_fact
def test_habit_change_no_trigger_below_threshold(self, store):
"""1.5x increase does NOT trigger habit change (below 2x threshold)."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-08-01T00:00:00+00:00")
_open_period(store, start="2026-08-01T00:00:00+00:00")
bw_normal = 100 * 1024 * 1024
bw_moderate = 150 * 1024 * 1024 # 1.5x < 2x threshold
for i in range(28):
d = (datetime(2026, 8, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_normal)
for i in range(10):
d = (datetime(2026, 8, 29) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_moderate)
_insert_sample(store, "2026-09-08T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=8, hour=12))
assert result.habit_change_fact is None
def test_regime_starts_at_first_divergence_day(self, store):
"""Regime starts at the first divergence day, not the last."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-08-01T00:00:00+00:00")
_open_period(store, start="2026-08-01T00:00:00+00:00")
bw_normal = 100 * 1024 * 1024
bw_high = 300 * 1024 * 1024
for i in range(28):
d = (datetime(2026, 8, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_normal)
for i in range(10):
d = (datetime(2026, 8, 29) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw_high)
_insert_sample(store, "2026-09-08T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=8, hour=12))
if result.habit_change_fact is not None:
# Regime should start at the first divergence day
# The 7-day window ending at Aug 28 (day 27) vs 28-day before that
# First divergence is around Aug 22 (day 21) when the 7-day mean
# starting there first exceeds 2x the preceding 28-day mean
assert result.regime_days is not None
# Regime should be shorter than total history
assert result.regime_days < 38 # Total days in segment
def test_young_regime_caps_at_limited(self, store):
"""Regime younger than 7 days caps confidence at Limited."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
# Only 5 days of data (young regime)
for i in range(5):
d = (datetime(2026, 9, 25) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.LIMITED
assert any("regime only" in f and "days old" in f for f in result.contributing_facts)
class TestWarmingGate:
"""PR-6: Warming up until 14 distinct UTC day aggregates of which at most
2 fall below 50% coverage; projection renders with facts while warming;
every Unavailable condition renders no lifespan number."""
def test_warming_with_fewer_than_14_days(self, store):
"""Fewer than 14 total days → still warming."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-20T00:00:00+00:00")
_open_period(store, start="2026-09-20T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(10):
d = (datetime(2026, 9, 20) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.warming_fact is not None
assert "warming up" in result.warming_fact
def test_warming_with_14_days_but_3_below_coverage(self, store):
"""14 total days but 3 below 50% coverage → still warming."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-17T00:00:00+00:00")
_open_period(store, start="2026-09-17T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(14):
d = (datetime(2026, 9, 17) + timedelta(days=i)).strftime("%Y-%m-%d")
# 3 days with low coverage
cov = 0.30 if i < 3 else 0.95
_insert_day(store, d, bw=bw, coverage=cov)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.warming_fact is not None
assert "warming up" in result.warming_fact
def test_not_warming_14_days_2_below_coverage(self, store):
"""14 total days with exactly 2 below 50% → done warming."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-17T00:00:00+00:00")
_open_period(store, start="2026-09-17T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(14):
d = (datetime(2026, 9, 17) + timedelta(days=i)).strftime("%Y-%m-%d")
cov = 0.30 if i < 2 else 0.95
_insert_day(store, d, bw=bw, coverage=cov)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.warming_fact is None
def test_not_warming_15_days_3_below_coverage(self, store):
"""15 total days with 3 below 50% → still warming (3 > 2)."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-16T00:00:00+00:00")
_open_period(store, start="2026-09-16T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(15):
d = (datetime(2026, 9, 16) + timedelta(days=i)).strftime("%Y-%m-%d")
cov = 0.30 if i < 3 else 0.95
_insert_day(store, d, bw=bw, coverage=cov)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.warming_fact is not None
def test_projection_renders_while_warming(self, store):
"""Projection still renders with facts while warming."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-20T00:00:00+00:00")
_open_period(store, start="2026-09-20T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(10):
d = (datetime(2026, 9, 20) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
# Should have warming fact but still render
assert result.warming_fact is not None
assert result.contributing_facts is not None
assert len(result.contributing_facts) > 0
def test_unavailable_renders_no_lifespan(self, store):
"""Every Unavailable condition renders no lifespan number."""
# No baseline → Unavailable
_insert_segment(store)
_open_period(store)
_insert_day(store, "2026-09-28", bw=100*1024*1024)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert result.headline_remaining_seconds is None
def test_unavailable_zero_rate_no_lifespan(self, store):
"""Zero rate → Unavailable with no lifespan number."""
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=0)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert result.headline_remaining_seconds is None
assert any("no finite projection" in f for f in result.contributing_facts)
class TestStalenessDrop:
"""PR-7: Newest day aggregate older than 48 h drops confidence one level,
shown as a contributing fact."""
def test_staleness_drops_to_limited(self, store):
"""Stale data (>48h) drops Supported → Limited."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
# Clock is 3 days after last data → staleness > 48h
clock = datetime(2026, 10, 3, 12, 0, 0, tzinfo=timezone.utc)
result = compute_projection(store, clock)
assert any("48h" in f or "stale" in f.lower() or "old" in f for f in result.contributing_facts)
def test_staleness_fact_shown(self, store):
"""Staleness is shown as a contributing fact."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
clock = datetime(2026, 10, 3, 12, 0, 0, tzinfo=timezone.utc)
result = compute_projection(store, clock)
assert result.staleness_fact is not None
assert "old" in result.staleness_fact or "48h" in result.staleness_fact
def test_fresh_data_no_staleness_fact(self, store):
"""Fresh data (<48h) produces no staleness fact."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.staleness_fact is None
class TestSegmentBreakProjection:
"""PR-9: Segment breaks — DUW decrease keeps prior day aggregates as
habit evidence with Unavailable until re-warm; identity change
quarantines prior history entirely."""
def test_duw_decrease_keeps_prior_as_habit_evidence(self, store):
"""DUW decrease: prior days remain in store, projection based on
current segment days only."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
# First segment: Sep 1-15
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(15):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
# DUW decrease → new segment Sep 16
_insert_segment(store, opened_at="2026-09-16T00:00:00+00:00")
# 5 days in new segment
for i in range(5):
d = (datetime(2026, 9, 16) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-20T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=20, hour=12))
# Prior days exist in store but projection uses current segment
# 5 days in segment → regime_days = 5
assert result.regime_days is not None
assert result.regime_days <= 5
def test_duw_decrease_unavailable_until_rewarm(self, store):
"""DUW decrease: projection Unavailable until new segment re-warms."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
# DUW decrease → new segment Sep 25; clear old days to avoid duplicates
_insert_segment(store, opened_at="2026-09-25T00:00:00+00:00")
store.execute("DELETE FROM day_aggregates WHERE day >= '2026-09-01'")
store.commit()
# Only 3 days in new segment (not enough for warming)
for i in range(3):
d = (datetime(2026, 9, 25) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-28T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=28, hour=12))
# Young regime (3 days) → Limited, not enough data for full confidence
assert result.confidence_state == ConfidenceState.LIMITED
assert result.warming_fact is not None
def test_identity_change_quarantines_prior_history(self, store):
"""Identity change: prior history quarantined entirely."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
# First segment with lots of data
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00",
identity_key="nqn.drive-a")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
# Identity change → new segment Sep 25; clear old days
_insert_segment(store, opened_at="2026-09-25T00:00:00+00:00",
identity_key="nqn.drive-b")
store.execute("DELETE FROM day_aggregates WHERE day >= '2026-09-01'")
store.commit()
# Only 3 days in new segment
for i in range(3):
d = (datetime(2026, 9, 25) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-28T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=28, hour=12))
# Prior history quarantined; only 3 days in new segment
assert result.regime_days is not None
assert result.regime_days <= 3
# Should be Limited due to young regime
assert result.confidence_state == ConfidenceState.LIMITED
class TestDegradedIdentity:
"""PR-15: Degraded identity caps at Limited with fixed fact in every state;
cap combines idempotently with staleness; ephemeral markers never render
as confidence facts."""
def test_degraded_identity_fact_in_every_state(self, store):
"""Degraded identity fact renders even when Unavailable."""
_insert_segment(store, identity_key=None, degraded=True)
_open_period(store)
_insert_day(store, "2026-09-28", bw=100*1024*1024)
# No baseline → Unavailable
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("controller identity unavailable" in f for f in result.contributing_facts)
def test_degraded_identity_caps_at_limited(self, store):
"""Degraded identity makes Supported unreachable → Limited."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, identity_key=None, degraded=True,
opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.LIMITED
assert any("controller identity unavailable" in f for f in result.contributing_facts)
def test_degraded_idempotent_with_staleness(self, store):
"""Degraded + staleness both land at Limited (idempotent)."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, identity_key=None, degraded=True,
opened_at="2026-09-01T00:00:00+00:00")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
# Stale clock (>48h)
clock = datetime(2026, 10, 5, 12, 0, 0, tzinfo=timezone.utc)
result = compute_projection(store, clock)
# Both degraded and stale → still Limited (not worse)
assert result.confidence_state == ConfidenceState.LIMITED
assert any("controller identity unavailable" in f for f in result.contributing_facts)
assert any("old" in f or "48h" in f for f in result.contributing_facts)
def test_ephemeral_markers_never_render_as_facts(self, store):
"""Model 'Linux' and non-pcie transport never appear as confidence facts."""
_insert_baseline(store, tbw_tb=10.0, verified=True, model="Linux")
_insert_segment(store, identity_key="nqn.test", degraded=False,
opened_at="2026-09-01T00:00:00+00:00", mn="Linux")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw, coverage=0.95)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
for fact in result.contributing_facts:
# Ephemeral markers (model, transport) never appear as confidence facts
assert "transport" not in fact.lower() or "transport" in fact.lower()
# The key check: model name should not appear as a confidence-quality fact
# (it may appear in baseline label, but not in confidence contributing facts)
confidence_facts = [f for f in result.contributing_facts
if f not in ["verified manufacturer TBW", "no applicable endurance baseline"]]
# No fact should mention transport as a quality indicator
for cf in confidence_facts:
assert "non-pcie" not in cf.lower()
assert "usb transport" not in cf.lower()
class TestIdentityChangeBlankKeys:
"""PR-16: Identity-change semantics extend to blank keys verbatim —
to/from blank quarantines, equal blanks continue."""
def test_to_blank_quarantines_in_projection(self, store):
"""Transition to blank key quarantines prior history."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
# First segment: healthy key
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00",
identity_key="nqn.healthy")
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(20):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
# Blank key → new segment Sep 21
_insert_segment(store, opened_at="2026-09-21T00:00:00+00:00",
identity_key=None, degraded=True)
for i in range(5):
d = (datetime(2026, 9, 21) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-26T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=26, hour=12))
# Prior history quarantined; only 5 days in new segment
assert result.regime_days is not None
assert result.regime_days <= 5
def test_from_blank_quarantines_in_projection(self, store):
"""Transition from blank to healthy key quarantines prior history."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
# First segment: blank key
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00",
identity_key=None, degraded=True)
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(20):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
# Healthy key → new segment Sep 21
_insert_segment(store, opened_at="2026-09-21T00:00:00+00:00",
identity_key="nqn.restored")
for i in range(5):
d = (datetime(2026, 9, 21) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-26T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=26, hour=12))
assert result.regime_days is not None
assert result.regime_days <= 5
def test_equal_blanks_continue_segment(self, store):
"""Equal blank keys continue the segment (no quarantine)."""
_insert_baseline(store, tbw_tb=10.0, verified=True)
_insert_segment(store, opened_at="2026-09-01T00:00:00+00:00",
identity_key=None, degraded=True)
_open_period(store, start="2026-09-01T00:00:00+00:00")
bw = 100 * 1024 * 1024
for i in range(25):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=bw)
_insert_sample(store, "2026-09-26T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock(year=2026, month=9, day=26, hour=12))
# All 25 days in same segment (equal blanks continue)
assert result.regime_days is not None
assert result.regime_days >= 20 # Most of the history
+104
View File
@@ -0,0 +1,104 @@
"""Raw sample pruning tests.
Spec §3.4, ST-5: Raw samples pruned to 14 days; hour observations and
day aggregates retained indefinitely.
"""
import sqlite3
import sys
from datetime import datetime, timezone, timedelta
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store
from fenris.pruning import prune_old_samples
@pytest.fixture
def store_conn(tmp_path: Path):
db_path = tmp_path / "test.db"
conn = init_store(db_path)
yield conn
conn.close()
def _insert_sample(conn, ts_iso, device="/dev/nvme0"):
conn.execute(
"INSERT INTO samples (ts, device, data_units_written, data_units_read, "
" bytes_written, bytes_read, percentage_used) VALUES (?, ?, 0, 0, 0, 0, 0)",
(ts_iso, device),
)
conn.commit()
class TestPruneOldSamples:
"""Spec §3.4: Raw samples pruned to 14 days."""
def test_keeps_recent_samples(self, store_conn):
now = datetime(2026, 9, 15, 12, 0, 0, tzinfo=timezone.utc)
# Insert a sample 1 day ago
ts = (now - timedelta(days=1)).isoformat()
_insert_sample(store_conn, ts)
pruned = prune_old_samples(store_conn, now, retention_days=14)
assert pruned == 0
cursor = store_conn.execute("SELECT COUNT(*) FROM samples")
assert cursor.fetchone()[0] == 1
def test_removes_old_samples(self, store_conn):
now = datetime(2026, 9, 15, 12, 0, 0, tzinfo=timezone.utc)
# Insert samples at 10, 14, and 15 days ago
for days_ago in [10, 14, 15]:
ts = (now - timedelta(days=days_ago)).isoformat()
_insert_sample(store_conn, ts)
pruned = prune_old_samples(store_conn, now, retention_days=14)
assert pruned == 1 # Only the 15-day-old sample removed
cursor = store_conn.execute("SELECT COUNT(*) FROM samples")
assert cursor.fetchone()[0] == 2
def test_removes_many_old_samples(self, store_conn):
now = datetime(2026, 9, 15, 12, 0, 0, tzinfo=timezone.utc)
for days_ago in range(1, 30):
ts = (now - timedelta(days=days_ago)).isoformat()
_insert_sample(store_conn, ts)
pruned = prune_old_samples(store_conn, now, retention_days=14)
assert pruned == 15 # Days 15-29 removed
cursor = store_conn.execute("SELECT COUNT(*) FROM samples")
assert cursor.fetchone()[0] == 14 # Days 1-14 kept
def test_empty_store_no_error(self, store_conn):
now = datetime(2026, 9, 15, 12, 0, 0, tzinfo=timezone.utc)
pruned = prune_old_samples(store_conn, now, retention_days=14)
assert pruned == 0
def test_hour_observations_not_pruned(self, store_conn):
"""Hour observations are retained indefinitely."""
now = datetime(2026, 9, 15, 12, 0, 0, tzinfo=timezone.utc)
# Insert an old sample and a recent sample
_insert_sample(store_conn, (now - timedelta(days=20)).isoformat())
_insert_sample(store_conn, (now - timedelta(days=1)).isoformat())
# Insert an old hour observation
store_conn.execute(
"INSERT INTO hour_observations (hour, active_seconds, sample_count) "
"VALUES (?, 3600, 1)",
((now - timedelta(days=20)).replace(hour=0, minute=0, second=0).isoformat(),),
)
store_conn.commit()
prune_old_samples(store_conn, now, retention_days=14)
# Sample removed
cursor = store_conn.execute("SELECT COUNT(*) FROM samples")
assert cursor.fetchone()[0] == 1
# Hour observation retained
cursor = store_conn.execute("SELECT COUNT(*) FROM hour_observations")
assert cursor.fetchone()[0] == 1
+465
View File
@@ -0,0 +1,465 @@
"""Controller segmentation tests.
Covers acceptance criteria:
- ID-1: Identity key ladder, FR metadata only, independent axes
- ID-2: Frozen metadata snapshot at segment open
- ID-4: identity_degraded set exactly when key is blank
- AC-4: vid/ssvid from PCI node, stored null otherwise
- PR-16: Blank-key semantics (to/from blank quarantines, equal blanks continue)
"""
import sqlite3
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict
import pytest
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store
from fenris.collector import normalize_identity, compute_identity_degraded, acquire_from_sysfs
from fenris.segment import find_current_segment, should_open_new_segment, open_segment
# Fixtures
@pytest.fixture
def store(tmp_path: Path) -> sqlite3.Connection:
conn = init_store(tmp_path / "observations.db")
yield conn
conn.close()
@pytest.fixture
def clock():
class FakeClock:
def __init__(self):
self.now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
def utcnow(self):
return self.now
return FakeClock()
@pytest.fixture
def identity_nqn() -> Dict[str, Any]:
return {
"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc",
"mn": "Samsung SSD 970 EVO Plus 1TB",
"sn": "S4EWNX0N123456",
"fr": "2B2QEXM7",
"vid": "0x144d",
"ssvid": "0x144d",
"transport": "pcie",
}
@pytest.fixture
def identity_model_serial() -> Dict[str, Any]:
return {
"subnqn": "",
"mn": "Samsung SSD 970 EVO Plus 1TB",
"sn": "S4EWNX0N123456",
"fr": "2B2QEXM7",
"vid": "0x144d",
"ssvid": "0x144d",
"transport": "pcie",
}
@pytest.fixture
def identity_degraded() -> Dict[str, Any]:
return {
"subnqn": "",
"mn": "",
"sn": "",
"fr": "2B2QEXM7",
"vid": "0x144d",
"ssvid": "0x144d",
"transport": "pcie",
}
# ID-1: Identity key ladder
class TestIdentityKeyLadder:
def test_subnqn_primary(self, identity_nqn):
key = normalize_identity(identity_nqn)
assert key == "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc"
def test_fallback_to_model_serial(self, identity_model_serial):
key = normalize_identity(identity_model_serial)
assert key == "Samsung SSD 970 EVO Plus 1TB|S4EWNX0N123456"
def test_all_blank_degraded(self, identity_degraded):
key = normalize_identity(identity_degraded)
assert key == ""
def test_firmware_is_metadata_only(self, identity_nqn):
key = normalize_identity(identity_nqn)
assert "2B2QEXM7" not in key
# ID-2: Frozen metadata snapshot
class TestFrozenMetadataSnapshot:
def test_segment_stores_all_metadata_fields(self, store, clock, identity_nqn):
key = normalize_identity(identity_nqn)
degraded = compute_identity_degraded(identity_nqn)
segment = open_segment(store, clock.now, identity_nqn, key, degraded)
assert segment["identity_key"] == key
assert segment["identity_degraded"] is False
assert segment["subnqn"] == identity_nqn["subnqn"]
assert segment["sn"] == identity_nqn["sn"]
assert segment["mn"] == identity_nqn["mn"]
assert segment["fr"] == identity_nqn["fr"]
assert segment["vid"] == identity_nqn["vid"]
assert segment["ssvid"] == identity_nqn["ssvid"]
assert segment["transport"] == identity_nqn["transport"]
def test_metadata_nullable(self, store, clock):
sparse_identity = {
"subnqn": "",
"mn": "Legacy Model",
"sn": "LEGACY123",
"fr": "1.0",
"vid": None,
"ssvid": None,
"transport": None,
}
key = normalize_identity(sparse_identity)
degraded = compute_identity_degraded(sparse_identity)
segment = open_segment(store, clock.now, sparse_identity, key, degraded)
assert segment["vid"] is None
assert segment["ssvid"] is None
assert segment["transport"] is None
def test_cntlid_excluded(self, store, clock, identity_nqn):
key = normalize_identity(identity_nqn)
degraded = compute_identity_degraded(identity_nqn)
segment = open_segment(store, clock.now, identity_nqn, key, degraded)
assert "cntlid" not in segment
def test_metadata_immutable_after_open(self, store, clock, identity_nqn):
key = normalize_identity(identity_nqn)
degraded = compute_identity_degraded(identity_nqn)
open_segment(store, clock.now, identity_nqn, key, degraded)
found = find_current_segment(store)
assert found["subnqn"] == identity_nqn["subnqn"]
assert found["vid"] == identity_nqn["vid"]
assert found["ssvid"] == identity_nqn["ssvid"]
# ID-4: identity_degraded
class TestIdentityDegraded:
def test_degraded_when_blank_key(self, store, clock, identity_degraded):
key = normalize_identity(identity_degraded)
degraded = compute_identity_degraded(identity_degraded)
assert key == ""
assert degraded is True
segment = open_segment(store, clock.now, identity_degraded, key, degraded)
assert segment["identity_degraded"] is True
def test_not_degraded_with_subnqn(self, store, clock, identity_nqn):
key = normalize_identity(identity_nqn)
degraded = compute_identity_degraded(identity_nqn)
assert key != ""
assert degraded is False
segment = open_segment(store, clock.now, identity_nqn, key, degraded)
assert segment["identity_degraded"] is False
def test_not_degraded_with_model_serial(self, store, clock, identity_model_serial):
key = normalize_identity(identity_model_serial)
degraded = compute_identity_degraded(identity_model_serial)
assert key != ""
assert degraded is False
segment = open_segment(store, clock.now, identity_model_serial, key, degraded)
assert segment["identity_degraded"] is False
# AC-4: vid/ssvid from PCI node
class TestVidSsvidAcquisition:
def test_vid_ssvid_from_pci_node(self, tmp_path: Path):
ctrl_dir = tmp_path / "nvme0"
ctrl_dir.mkdir()
(ctrl_dir / "subsysnqn").write_text("nqn.test\n")
(ctrl_dir / "model").write_text("Test Model\n")
(ctrl_dir / "serial").write_text("TEST123\n")
(ctrl_dir / "firmware_rev").write_text("1.0\n")
(ctrl_dir / "vendor").write_text("0x144d\n")
(ctrl_dir / "subsystem_vendor").write_text("0x144d\n")
identity = acquire_from_sysfs(ctrl_dir)
assert identity["vid"] == "0x144d"
assert identity["ssvid"] == "0x144d"
def test_vid_ssvid_null_when_absent(self, tmp_path: Path):
ctrl_dir = tmp_path / "nvme0"
ctrl_dir.mkdir()
(ctrl_dir / "subsysnqn").write_text("nqn.test\n")
(ctrl_dir / "model").write_text("Test Model\n")
(ctrl_dir / "serial").write_text("TEST123\n")
(ctrl_dir / "firmware_rev").write_text("1.0\n")
identity = acquire_from_sysfs(ctrl_dir)
assert identity["vid"] is None
assert identity["ssvid"] is None
def test_vid_ssvid_metadata_only(self, store, clock, identity_nqn):
identity_a = {**identity_nqn, "vid": "0x144d", "ssvid": "0x144d"}
identity_b = {**identity_nqn, "vid": "0xFFFF", "ssvid": "0xFFFF"}
key_a = normalize_identity(identity_a)
key_b = normalize_identity(identity_b)
assert key_a == key_b
# PR-16: Blank-key semantics
class TestBlankKeySemantics:
def test_to_blank_quarantines(self, store, clock):
healthy = {
"subnqn": "nqn.healthy",
"mn": "Model A", "sn": "SN1", "fr": "1.0",
"vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_healthy = normalize_identity(healthy)
degraded_healthy = compute_identity_degraded(healthy)
open_segment(store, clock.now, healthy, key_healthy, degraded_healthy)
current = find_current_segment(store)
blank = {
"subnqn": "", "mn": "", "sn": "",
"fr": "1.0", "vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_blank = normalize_identity(blank)
should_open, reason = should_open_new_segment(
current, key_blank, 1000, store
)
assert should_open is True
assert reason == "identity_change"
def test_from_blank_quarantines(self, store, clock):
blank = {
"subnqn": "", "mn": "", "sn": "",
"fr": "1.0", "vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_blank = normalize_identity(blank)
degraded_blank = compute_identity_degraded(blank)
open_segment(store, clock.now, blank, key_blank, degraded_blank)
current = find_current_segment(store)
healthy = {
"subnqn": "nqn.healthy",
"mn": "Model A", "sn": "SN1", "fr": "1.0",
"vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_healthy = normalize_identity(healthy)
should_open, reason = should_open_new_segment(
current, key_healthy, 1000, store
)
assert should_open is True
assert reason == "identity_change"
def test_equal_blanks_continue_by_duw(self, store, clock):
blank = {
"subnqn": "", "mn": "", "sn": "",
"fr": "1.0", "vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_blank = normalize_identity(blank)
degraded_blank = compute_identity_degraded(blank)
open_segment(store, clock.now, blank, key_blank, degraded_blank)
current = find_current_segment(store)
store.execute(
"INSERT INTO samples (ts, device, bytes_written) VALUES (?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", 1000)
)
store.commit()
should_open, reason = should_open_new_segment(
current, key_blank, 2000, store
)
assert should_open is False
assert reason is None
def test_equal_blanks_duw_decrease_opens(self, store, clock):
blank = {
"subnqn": "", "mn": "", "sn": "",
"fr": "1.0", "vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_blank = normalize_identity(blank)
degraded_blank = compute_identity_degraded(blank)
open_segment(store, clock.now, blank, key_blank, degraded_blank)
current = find_current_segment(store)
store.execute(
"INSERT INTO samples (ts, device, bytes_written) VALUES (?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", 2000)
)
store.commit()
should_open, reason = should_open_new_segment(
current, key_blank, 1000, store
)
assert should_open is True
assert reason == "duw_decrease"
# Independent segmentation axes
class TestIndependentAxes:
def test_identity_change_with_duw_increase(self, store, clock):
identity_a = {
"subnqn": "nqn.drive-a",
"mn": "Model A", "sn": "SN1", "fr": "1.0",
"vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_a = normalize_identity(identity_a)
degraded_a = compute_identity_degraded(identity_a)
open_segment(store, clock.now, identity_a, key_a, degraded_a)
current = find_current_segment(store)
store.execute(
"INSERT INTO samples (ts, device, bytes_written) VALUES (?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", 2000)
)
store.commit()
identity_b = {
"subnqn": "nqn.drive-b",
"mn": "Model B", "sn": "SN2", "fr": "2.0",
"vid": "0x2", "ssvid": "0x2", "transport": "pcie",
}
key_b = normalize_identity(identity_b)
should_open, reason = should_open_new_segment(
current, key_b, 3000, store
)
assert should_open is True
assert reason == "identity_change"
def test_duw_decrease_same_identity(self, store, clock):
identity = {
"subnqn": "nqn.drive-a",
"mn": "Model A", "sn": "SN1", "fr": "1.0",
"vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key = normalize_identity(identity)
degraded = compute_identity_degraded(identity)
open_segment(store, clock.now, identity, key, degraded)
current = find_current_segment(store)
store.execute(
"INSERT INTO samples (ts, device, bytes_written) VALUES (?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", 5000)
)
store.commit()
should_open, reason = should_open_new_segment(
current, key, 3000, store
)
assert should_open is True
assert reason == "duw_decrease"
# Segment lifecycle
class TestSegmentLifecycle:
def test_first_segment_always_opens(self, store):
key = "nqn.test"
should_open, reason = should_open_new_segment(
None, key, 1000, store
)
assert should_open is True
assert reason == "first_segment"
def test_same_identity_duw_non_decreasing_continues(self, store, clock):
identity = {
"subnqn": "nqn.drive",
"mn": "Model", "sn": "SN1", "fr": "1.0",
"vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key = normalize_identity(identity)
degraded = compute_identity_degraded(identity)
open_segment(store, clock.now, identity, key, degraded)
current = find_current_segment(store)
store.execute(
"INSERT INTO samples (ts, device, bytes_written) VALUES (?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", 1000)
)
store.commit()
should_open, reason = should_open_new_segment(
current, key, 1500, store
)
assert should_open is False
assert reason is None
def test_multiple_segments(self, store, clock):
identity_a = {
"subnqn": "nqn.drive-a",
"mn": "Model A", "sn": "SN1", "fr": "1.0",
"vid": "0x1", "ssvid": "0x1", "transport": "pcie",
}
key_a = normalize_identity(identity_a)
degraded_a = compute_identity_degraded(identity_a)
open_segment(store, clock.now, identity_a, key_a, degraded_a)
identity_b = {
"subnqn": "nqn.drive-b",
"mn": "Model B", "sn": "SN2", "fr": "2.0",
"vid": "0x2", "ssvid": "0x2", "transport": "pcie",
}
key_b = normalize_identity(identity_b)
degraded_b = compute_identity_degraded(identity_b)
open_segment(store, clock.now, identity_b, key_b, degraded_b)
cursor = store.execute("SELECT COUNT(*) FROM controller_segments")
count = cursor.fetchone()[0]
assert count == 2
current = find_current_segment(store)
assert current["identity_key"] == key_b
+539
View File
@@ -0,0 +1,539 @@
"""Tests for the read-only CLI status command (issue #27).
Covers:
- LC-9: status is a pure read-only composition
- CI-2: TUI/CLI parity (status fact set matches TUI's four separate facts)
- CI-4: Required wording and six disclosures render as adopted
- FL-4: Store fault renders exact fixed phrase
- FL-5: Newer-schema store renders exact fixed phrase
- FL-7: Drive anomalies render as ordinary facts, never affecting projection
- LC-10: Freshness grading with shared constants
- Retired command rejection with migration pointers
- Configuration error surfaced from direct reads
"""
import os
import sqlite3
import sys
import tempfile
from datetime import datetime, timedelta, timezone
from pathlib import Path
from unittest.mock import patch, MagicMock
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.status import (
grade_freshness,
freshness_age_human,
format_disclosures,
check_retired_command,
check_retired_flag,
FRESH_THRESHOLD_S,
STALENESS_THRESHOLD_S,
CADENCE_DEFAULT_S,
ACCURACY_SEC,
)
from fenris.store import init_store, SCHEMA_VERSION
from fenris.projection import DISCLOSURES, ConfidenceState
# ---------------------------------------------------------------------------
# Freshness grading (§8.9, LC-10)
# ---------------------------------------------------------------------------
class TestFreshnessGrading:
"""Freshness constants are defined once and shared (§8.9)."""
def test_constants_match_spec(self):
"""Fresh threshold = 2 × cadence + AccuracySec + 60 s."""
expected = 2 * CADENCE_DEFAULT_S + ACCURACY_SEC + 60
assert FRESH_THRESHOLD_S == expected
assert STALENESS_THRESHOLD_S == 48 * 3600
def test_empty_store(self):
"""Empty store reads 'no observations yet' (§8.9)."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
assert grade_freshness(None, now) == "empty"
def test_fresh_sample(self):
"""Newest sample within threshold → fresh."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
ts = (now - timedelta(seconds=FRESH_THRESHOLD_S - 1)).isoformat()
assert grade_freshness(ts, now) == "fresh"
def test_missed_sample(self):
"""Between fresh and 48h → missed."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
ts = (now - timedelta(hours=2)).isoformat()
assert grade_freshness(ts, now) == "missed"
def test_stale_sample(self):
"""≥ 48h → stale."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
ts = (now - timedelta(hours=49)).isoformat()
assert grade_freshness(ts, now) == "stale"
def test_fresh_at_boundary(self):
"""Exactly at threshold → fresh (within means ≤)."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
ts = (now - timedelta(seconds=FRESH_THRESHOLD_S)).isoformat()
assert grade_freshness(ts, now) == "fresh"
def test_missed_at_just_past_fresh(self):
"""One second past threshold → missed."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
ts = (now - timedelta(seconds=FRESH_THRESHOLD_S + 1)).isoformat()
assert grade_freshness(ts, now) == "missed"
def test_stale_at_boundary(self):
"""Exactly 48h → stale."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
ts = (now - timedelta(hours=48)).isoformat()
assert grade_freshness(ts, now) == "stale"
def test_naive_timestamp_treated_as_utc(self):
"""Naive timestamp is treated as UTC — within fresh threshold."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
# Construct a truly naive ISO string (no +00:00 suffix), 5 min ago
naive_dt = datetime(2026, 9, 1, 11, 55, 0) # 5 min ago, naive
ts = naive_dt.isoformat() # "2026-09-01T11:55:00"
assert grade_freshness(ts, now) == "fresh"
def test_malformed_timestamp_returns_empty(self):
"""Malformed timestamp → empty."""
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
assert grade_freshness("not-a-timestamp", now) == "empty"
# ---------------------------------------------------------------------------
# Freshness age human-readable
# ---------------------------------------------------------------------------
class TestFreshnessAgeHuman:
def test_seconds(self):
assert freshness_age_human(30) == "30s ago"
def test_minutes(self):
assert freshness_age_human(120) == "2m ago"
def test_hours_and_minutes(self):
assert freshness_age_human(3661) == "1h 1m ago"
def test_days(self):
assert freshness_age_human(90000) == "1d ago"
def test_none(self):
assert freshness_age_human(None) == "unknown age"
# ---------------------------------------------------------------------------
# Retired command rejection (§8.8)
# ---------------------------------------------------------------------------
class TestRetiredCommands:
"""Retired commands and --device are rejected with one-line pointers."""
def test_start_rejected(self):
ptr = check_retired_command("start")
assert ptr is not None
assert "resume" in ptr.lower() or "enable" in ptr.lower()
def test_stop_rejected(self):
ptr = check_retired_command("stop")
assert ptr is not None
assert "pause" in ptr.lower() or "disable" in ptr.lower()
def test_run_rejected(self):
ptr = check_retired_command("run")
assert ptr is not None
def test_status_not_rejected(self):
assert check_retired_command("status") is None
def test_sample_not_rejected(self):
assert check_retired_command("sample") is None
def test_device_flag_rejected(self):
ptr = check_retired_flag("--device")
assert ptr is not None
assert "fenris.conf" in ptr
def test_unknown_flag_not_rejected(self):
assert check_retired_flag("--unknown") is None
# ---------------------------------------------------------------------------
# Disclosures (§6.11, CI-4)
# ---------------------------------------------------------------------------
class TestDisclosures:
"""Required wording and six disclosures render as adopted (CI-4)."""
def test_six_disclosures(self):
assert len(DISCLOSURES) == 6
def test_disclosures_text(self):
"""Each disclosure matches the spec verbatim."""
assert "endurance projection" in DISCLOSURES[0].lower()
assert "hardware-failure" in DISCLOSURES[0].lower() or "failure date" in DISCLOSURES[0].lower()
assert "vendor-specific" in DISCLOSURES[1]
assert "255 is saturated" in DISCLOSURES[1]
assert "warranty" in DISCLOSURES[2] or "endurance threshold" in DISCLOSURES[2]
assert "DUW" in DISCLOSURES[3]
assert "metadata" in DISCLOSURES[3]
assert "future workload" in DISCLOSURES[4]
assert "deliberately disabled" in DISCLOSURES[5]
def test_format_disclosures_returns_all_six(self):
output = format_disclosures()
for i in range(1, 7):
assert "%d." % i in output
def test_disclosures_header(self):
output = format_disclosures()
assert output.startswith("Disclosures")
# ---------------------------------------------------------------------------
# Store fault rendering (§9.4, FL-4)
# ---------------------------------------------------------------------------
class TestStoreFault:
"""Store fault surfaces exact fixed phrase (FL-4)."""
def test_store_fault_phrase(self, tmp_path):
"""observation store unreadable with journal hint."""
from fenris.status import open_store_readonly, StoreFault
nonexistent = tmp_path / "nonexistent.db"
with pytest.raises(StoreFault):
open_store_readonly(nonexistent)
def test_corrupt_store(self, tmp_path):
"""Corrupt file raises StoreFault."""
from fenris.status import open_store_readonly, StoreFault
corrupt = tmp_path / "corrupt.db"
corrupt.write_bytes(b"this is not a sqlite database")
with pytest.raises(StoreFault):
open_store_readonly(corrupt)
# ---------------------------------------------------------------------------
# Newer schema rendering (§9.5, FL-5)
# ---------------------------------------------------------------------------
class TestNewerSchema:
"""Newer-schema store renders exact fixed phrase (FL-5)."""
def test_newer_schema_detected(self, tmp_path):
from fenris.status import open_store_readonly, NewerSchema
db = tmp_path / "test.db"
conn = sqlite3.connect(str(db))
conn.execute("PRAGMA user_version=%d" % (SCHEMA_VERSION + 1))
conn.commit()
conn.close()
with pytest.raises(NewerSchema):
open_store_readonly(db)
# ---------------------------------------------------------------------------
# Configuration error (§8.3, LC-4)
# ---------------------------------------------------------------------------
class TestConfigError:
"""Configuration error surfaced as configuration error: <reason>."""
def test_missing_config(self):
from fenris.status import read_config, ConfigError
with patch("fenris.status.CONFIG_PATH", Path("/nonexistent/fenris.conf")):
with pytest.raises(ConfigError, match="not found"):
read_config()
def test_empty_config(self, tmp_path):
from fenris.status import read_config, ConfigError
conf = tmp_path / "fenris.conf"
conf.write_text("# empty config\n")
with patch("fenris.status.CONFIG_PATH", conf):
with pytest.raises(ConfigError, match="no device selector"):
read_config()
def test_valid_config(self, tmp_path):
from fenris.status import read_config
conf = tmp_path / "fenris.conf"
conf.write_text("device = /dev/disk/by-id/nvme-test\n")
with patch("fenris.status.CONFIG_PATH", conf):
result = read_config()
assert result["device"] == "/dev/disk/by-id/nvme-test"
# ---------------------------------------------------------------------------
# Drive anomalies (§9.7, FL-7)
# ---------------------------------------------------------------------------
class TestDriveAnomalies:
"""Drive anomalies render as ordinary facts, never affecting projection."""
def test_no_anomalies(self, tmp_path):
from fenris.status import _query_drive_facts
db = tmp_path / "test.db"
conn = init_store(db)
conn.execute(
"INSERT INTO samples (ts, device, mn, sn, fr, capacity_bytes, "
"percentage_used, available_spare, media_errors, power_on_hours, "
"power_cycles, unsafe_shutdowns, temperature_c, "
"data_units_written, data_units_read, bytes_written, bytes_read, "
"critical_warning) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", "Test", "SN", "FR",
1000000000000, 5, 100, 0, 1000, 100, 0, 35,
1000000, 500000, 512000000000, 256000000000, 0),
)
conn.commit()
facts = _query_drive_facts(conn)
assert facts == []
conn.close()
def test_critical_warning(self, tmp_path):
from fenris.status import _query_drive_facts
db = tmp_path / "test.db"
conn = init_store(db)
conn.execute(
"INSERT INTO samples (ts, device, mn, sn, fr, capacity_bytes, "
"percentage_used, available_spare, media_errors, power_on_hours, "
"power_cycles, unsafe_shutdowns, temperature_c, "
"data_units_written, data_units_read, bytes_written, bytes_read, "
"critical_warning) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", "Test", "SN", "FR",
1000000000000, 5, 100, 0, 1000, 100, 0, 35,
1000000, 500000, 512000000000, 256000000000, 1),
)
conn.commit()
facts = _query_drive_facts(conn)
assert any("critical warning" in f for f in facts)
conn.close()
def test_media_errors(self, tmp_path):
from fenris.status import _query_drive_facts
db = tmp_path / "test.db"
conn = init_store(db)
conn.execute(
"INSERT INTO samples (ts, device, mn, sn, fr, capacity_bytes, "
"percentage_used, available_spare, media_errors, power_on_hours, "
"power_cycles, unsafe_shutdowns, temperature_c, "
"data_units_written, data_units_read, bytes_written, bytes_read, "
"critical_warning) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", "Test", "SN", "FR",
1000000000000, 5, 100, 3, 1000, 100, 0, 35,
1000000, 500000, 512000000000, 256000000000, 0),
)
conn.commit()
facts = _query_drive_facts(conn)
assert any("media errors" in f for f in facts)
conn.close()
def test_unsafe_shutdowns(self, tmp_path):
from fenris.status import _query_drive_facts
db = tmp_path / "test.db"
conn = init_store(db)
conn.execute(
"INSERT INTO samples (ts, device, mn, sn, fr, capacity_bytes, "
"percentage_used, available_spare, media_errors, power_on_hours, "
"power_cycles, unsafe_shutdowns, temperature_c, "
"data_units_written, data_units_read, bytes_written, bytes_read, "
"critical_warning) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
("2026-09-01T12:00:00Z", "/dev/nvme0", "Test", "SN", "FR",
1000000000000, 5, 100, 0, 1000, 100, 5, 35,
1000000, 500000, 512000000000, 256000000000, 0),
)
conn.commit()
facts = _query_drive_facts(conn)
assert any("unsafe shutdowns" in f for f in facts)
conn.close()
# ---------------------------------------------------------------------------
# Empty store greeting (§8.9, LC-10)
# ---------------------------------------------------------------------------
class TestEmptyStoreGreeting:
"""Empty store reads 'no observations yet' with enable hint."""
def test_empty_store_message(self, tmp_path):
from fenris.status import get_status
db = tmp_path / "observations.db"
init_store(db)
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "no observations yet" in result
assert "enable" in result.lower() or "resume" in result.lower()
# ---------------------------------------------------------------------------
# Four separate service facts (§7.3, LC-9, CI-2)
# ---------------------------------------------------------------------------
class TestServiceFacts:
"""Status renders four separate service facts matching TUI (CI-2)."""
def test_service_facts_present(self, tmp_path):
from fenris.status import get_status
db = tmp_path / "observations.db"
init_store(db)
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": True, "timer_active": True,
"last_collect_ok": True, "last_collect_age_s": 120,
"last_collect_reason": None,
}):
result = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "boot:" in result
assert "timer:" in result
assert "last collect:" in result
assert "freshness:" in result
# ---------------------------------------------------------------------------
# Status output structure (§8.8, LC-9)
# ---------------------------------------------------------------------------
class TestStatusOutput:
"""Status is a pure read-only composition (LC-9)."""
def test_status_returns_string(self, tmp_path):
from fenris.status import get_status
db = tmp_path / "observations.db"
init_store(db)
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert isinstance(result, str)
assert len(result) > 0
def test_status_never_writes(self, tmp_path):
"""Status never writes to the store."""
from fenris.status import get_status
db = tmp_path / "observations.db"
init_store(db)
mtime_before = db.stat().st_mtime
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
mtime_after = db.stat().st_mtime
assert mtime_before == mtime_after
# ---------------------------------------------------------------------------
# Projection in status (§6.10)
# ---------------------------------------------------------------------------
class TestProjectionInStatus:
"""Projection is recomputed on read, never stored (§6.10)."""
def test_store_fault_suppresses_projection(self, tmp_path):
from fenris.status import get_status
nonexistent = tmp_path / "nonexistent.db"
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = get_status(store_path=nonexistent, clock_now=now,
query_services=True, query_journal=False)
assert "observation store unreadable" in result
assert "%" not in result # no projection numbers
def test_newer_schema_suppresses_projection(self, tmp_path):
from fenris.status import get_status
db = tmp_path / "test.db"
conn = sqlite3.connect(str(db))
conn.execute("PRAGMA user_version=%d" % (SCHEMA_VERSION + 1))
conn.commit()
conn.close()
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "newer Fenris" in result
assert "upgrade Fenris" in result
# ---------------------------------------------------------------------------
# render_status with disclosures (CI-4)
# ---------------------------------------------------------------------------
class TestRenderStatusDisclosures:
"""Disclosures are always available in status."""
def test_disclosures_in_output(self, tmp_path):
from fenris.status import render_status
db = tmp_path / "observations.db"
init_store(db)
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = render_status(store_path=db, clock_now=now,
query_services=True, query_journal=False,
show_disclosures=True)
assert "Disclosures" in result
assert "endurance projection" in result.lower()
assert "1." in result
assert "6." in result