Compare commits

...
11 changed files with 473 additions and 0 deletions
+1
View File
@@ -1,4 +1,5 @@
__pycache__/ __pycache__/
.pi/
*.pyc *.pyc
.commandcode/ .commandcode/
data/fenris.pid data/fenris.pid
+9
View File
@@ -0,0 +1,9 @@
## Agent skills
### Issue tracker
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
### Domain docs
This is a single-context repository. See `docs/agents/domain.md`.
+73
View File
@@ -0,0 +1,73 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Store fault**:
The condition where the observation store is present but cannot be read or trusted — unreadable, corrupt, or written by a newer Fenris — degrading every view that depends on it rather than crashing or guessing.
_Avoid_: Database error, corruption, broken data
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Endurance baseline**:
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
+42
View File
@@ -0,0 +1,42 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -0,0 +1,49 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -0,0 +1,31 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger and monitoring-period bookkeeping; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
@@ -0,0 +1,31 @@
# 4. Installation lifecycle: Makefile-delivered venv, dormant install, sanctioned teardown
## Status
Accepted — resolves [Define installation, upgrade, and removal behavior](https://git.bongbetic.com/xavierk/Fenris/issues/9) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris runs today from the source checkout (`fenris.py`, `fenris.sh`, `data/`): code, state, and control all live relative to wherever the checkout sits. The redesign fixes system artifacts — helpers in `/usr/libexec/fenris` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), the observation store at `/var/lib/fenris/observations.db` ([ADR 0001](0001-observation-store-sqlite.md)), configuration at `/etc/fenris/fenris.conf` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), Textual on Python 3.9+ ([framework decision](https://git.bongbetic.com/xavierk/Fenris/issues/6)) — but nothing says how those artifacts are delivered, upgraded, or removed, or what happens when the checkout moves or disappears.
## Decision
1. **Delivery.** `sudo make install` builds a wheel from the checkout and installs it, with pinned dependencies, into a dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only: after install, nothing references it.
2. **Layout and manifest.** Units in `/etc/systemd/system` (`fenris-collect.{timer,service}`, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)); helpers in `/usr/libexec/fenris`; polkit policy in `/usr/share/polkit-1/actions/`; configuration and observation store in their ADR-fixed locations. The installer records every file it places in an explicit manifest consumed by upgrade and uninstall.
3. **Privilege.** One root installer (`sudo make install`); at runtime, elevation is exclusively polkit (`auth_admin`, `fenris-monitor` only, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)). The installer never enables or starts units.
4. **Dormant install.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle (`fenris monitor resume [--now]`, or the first-run TUI prompt), which enables the timer and opens the first period in one step.
5. **Legacy import.** The installer detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs [ADR 0001](0001-observation-store-sqlite.md)'s idempotent single-transaction import, and reports imported counts — existing observations never depend on checkout survival. `fenris import <path>` remains available for later finds.
6. **Upgrade.** `sudo make upgrade` builds and installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed and it is active — safe with `Persistent=no`), leaves timer state untouched, and never kills an in-flight collection run: a running oneshot finishes on its mapped interpreter, so at worst one old-code run completes to the store and the next run uses the new code. It then applies forward-only observation-store schema migrations governed by a `schema_version` table. `/var/lib/fenris` is never rebuilt.
7. **Rollback.** Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade is explicitly unsupported.
8. **Removal.** `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. `make purge` additionally removes configuration and store. Journal entries age out naturally.
9. **Dependencies.** Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing.
10. **Scaffolding and floor.** The installer creates `/var/lib/fenris` with [ADR 0001](0001-observation-store-sqlite.md)'s root-written group-read permissions and verifies `python3 ≥ 3.9`, failing cleanly otherwise — the Textual contingency becomes an install-time gate rather than a runtime crash. The database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet.
## Consequences
- Installed Fenris survives checkout deletion; the checkout is only where builds happen.
- Teardown preserves monitoring-period semantics: deliberate removal excludes the uninstalled span from the usage habit instead of leaving it as unknown-inside.
- Installs are reproducible; dependency drift cannot ride in on an upgrade.
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
- Rollback support is exactly one generation deep, no further.
- The README documents install, upgrade, uninstall/purge, and legacy import alongside [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s menu-successor mapping.
@@ -0,0 +1,27 @@
# 5. Failure and recovery: visible degradation, never fabrication
## Status
Accepted — resolves [Define failure and recovery behavior](https://git.bongbetic.com/xavierk/Fenris/issues/10) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The observation store ([ADR 0001](0001-observation-store-sqlite.md)) and the service lifecycle ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) settled single-writer transactions, bounded single failed runs, freshness grading (fresh / missed / stale), and absent-hour classification through power-on-hours evidence. Left open by the [failure ticket](https://git.bongbetic.com/xavierk/Fenris/issues/10): behavior per failure class — malformed observations inside the store, missed observations, store faults (unreadable, corrupt, or newer-schema database), and repeated collector failures — and how the habit record re-anchors after the store itself is lost.
## Decision
1. **Malformed observations — refuse at the write boundary.** The collector validates every row it would write against the store's domain invariants (hour seconds sum to 3600, non-negative DUW delta within a controller segment, coverage consistent with sample count). A violating run writes nothing for that run, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact, but under a single trusted writer they should never see one. Store invariant: everything persisted is well-formed.
2. **Missed observations — never backfill.** Fenris never interpolates, estimates, or fabricates an hour. Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s power-on-hours classification is the only inference admitted.
3. **Store faults — degrade, never recreate over.** An unreadable or corrupt database is a store fault: readers surface a "observation store unreadable" fact with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside, the next run starts a fresh store, and if the legacy import never completed, the still-present `history.jsonl` is re-imported. No built-in destructive command exists.
4. **Newer schema — readers refuse symmetrically.** The TUI and `status` detect a `user_version` newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal in [ADR 0001](0001-observation-store-sqlite.md) and the forward-only upgrade rule of [ADR 0004](0004-install-upgrade-removal-lifecycle.md).
5. **Repeated collector failures — flat cadence, no escalation.** The timer's retry is the recovery path; the settled freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff and no notification machinery; a persistent failure reads as stale exactly like any other gap.
6. **Drive-reported anomalies — facts, not alerts.** `critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status`; no alerting or notification surface exists. Fenris observes and projects; it does not alarm. The projection is unaffected: endurance math consumes writes, not warnings.
7. **Orphaned samples — the collector re-anchors observed fact.** When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent: only the sanctioned path of [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) records a `user_disabled` close. Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
## Consequences
- Validation lives at one boundary — the collector — so the store's contract is "everything in it is well-formed" and readers only defend against the impossible.
- No synthetic data can ever enter the habit record; confidence categories can be trusted to reflect real evidence.
- Store-fault recovery can lose history; the mitigation is the one-file backup story of [ADR 0001](0001-observation-store-sqlite.md), kept human-sanctioned so loss is never silent.
- Period bookkeeping splits by epistemics: the helper records intent, the collector records observed fact.
- Fenris stays fully local and silent: no notification, escalation, or alerting machinery anywhere.
+26
View File
@@ -0,0 +1,26 @@
# Domain Docs
How engineering skills should consume this repository’s domain documentation.
## Layout
This is a single-context repository:
```text
/
├── CONTEXT.md
├── docs/adr/
└── ...
```
## Before exploring
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
## Use the glossary’s vocabulary
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
## Flag ADR conflicts
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
+85
View File
@@ -0,0 +1,85 @@
# Issue tracker: Gitea
Issues for this repository live in Gitea at:
https://git.bongbetic.com/xavierk/Fenris/issues
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
## General operations
- List: `tea issues list`
- Read: `tea issues <index> --comments`
- Create: `tea issues create --title "<title>" --description "<body>"`
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
- Assign: `tea issues edit <index> --add-assignees "<username>"`
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
- Comment: `tea comments add <index> --description "<comment>"`
- Close: `tea issues close <index>`
- Reopen: `tea issues reopen <index>`
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
## When a skill says “publish to the issue tracker”
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
## When a skill says “fetch the relevant ticket”
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
## Wayfinding operations
Wayfinder maps and decision tickets are Gitea issues.
### Map and ticket grouping
- A map has the label `wayfinder:map`.
- Create one milestone named `Wayfinder: <map title>` for the effort.
- Assign the map and all its tickets to that milestone.
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
### Blocking
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
```bash
tea api -X POST \
repos/{owner}/{repo}/issues/<blocked>/dependencies \
-F index=<blocker> \
-f owner=xavierk \
-f repo=Fenris
```
List blockers:
```bash
tea api repos/{owner}/{repo}/issues/<index>/dependencies
```
Remove the relationship with the same payload and `-X DELETE`.
### Frontier
List open issues in the map’s milestone. Exclude:
- the issue labelled `wayfinder:map`
- assigned tickets, because assignment is the claim
- tickets whose dependency query returns any open issue
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
### Claim
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
### Resolve
1. Add the answer as a resolution comment.
2. Close the ticket.
3. Re-fetch the map immediately before editing it.
4. Append a linked one-line context pointer to `Decisions so far`.
5. Create newly visible tickets, then wire dependencies in a second pass.
+99
View File
@@ -0,0 +1,99 @@
# Controller identity for observation-history segmentation
Research for [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11).
## Decision
Record the **subsystem NQN as exposed by the Linux kernel** — normalized, trailing-space stripped — as the controller-segment identity key:
```text
identity_key = strip(subnqn) # /sys/class/nvme-subsystem/…/subsysnqn,
# identical to /sys/class/nvme/nvmeX/subsysnqn
fallback: "nqn.2014.08.org.nvmexpress:" + hex4(vid) + hex4(ssvid)
+ raw20(sn) + raw40(mn) # byte-for-byte the kernel's synthesized NQN
last resort: strip(mn) + "|" + strip(sn) # when only smartctl-style fields exist
```
Store the raw `subnqn`, `sn`, `mn` strings (normalized) plus `fr` (firmware revision) as **segment metadata**, never `fr` inside the key: firmware revision is the one mandatory field that legitimately changes on the same drive ([Base Spec 2.0e §5.17.2.1: FR is the *currently active* firmware revision](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Namespace identifiers (NGUID, EUI-64, UUID) are excluded from the key: they are namespace-scoped while the SMART counters being segmented are controller-scoped, and each may be absent or reused. When the key is blank, record an empty key with a degraded marker and rely on the DUW-monotonic rule; a same-model same-serial replacement is unobservable by any identifier and is caught — as a boundary, not an identity change — by the DUW-decrease rule of ADR 0001.
This key satisfies the ticket's asymmetry: a drive replacement changes `subnqn` (real NQNs are unique per subsystem; kernel-generated ones embed serial+model), while firmware quirks and counter resets on the same drive leave it unchanged — counter discontinuities are already ADR 0001's second segmentation axis.
## What each candidate is
| Candidate | Spec definition | Scope | Mandatory | Verdict for the key |
|---|---|---|---|---|
| **SN + MN** | ASCII strings assigned by the vendor in Identify Controller, bytes 23:04 and 63:24; §4.3 shows them left-justified and space-padded ([2.0e §5.17.2.1, Identify Controller data structure](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf), [§4.3 Identifier Format and Layout](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | NVM subsystem | Mandatory for I/O and Admin controllers | Core of the fallback; uniqueness explicitly not guaranteed by the spec |
| **FR** | Currently active firmware revision, ASCII, bytes 71:64 ([2.0e §5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Domain (subsystem) | Mandatory | Never in the key — it is *meant* to change on the same drive |
| **SUBNQN** | NVM Subsystem NQN, UTF-8 null-terminated, bytes 1023:768; mandatory if the controller is ≥ 1.2.1, otherwise may be all zero ([2.0e §5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | NVM subsystem (shared by all its controllers) | Mandatory ≥ 1.2.1, optional below | **Chosen key**; the spec says hosts *should* use it as the subsystem's unique identifier |
| **NGUID / EUI-64** | IEEE-based identifiers in Identify Namespace, bytes 119:104 / 127:120 ([§4.3.4, §4.3.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Namespace | Optional; may both be zero | Rejected — wrong scope, may be missing, may be reused |
| **CNTLID** | Controller ID, unique only *within* a subsystem ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Controller | Mandatory | Rejected — not unique across subsystems (this host's drive reports 0) |
A note on the PDF: figure numbers in the table of contents of the 2.0e revision are offset from the body captions (e.g. the SN/MN figure is "Figure 128" in the TOC but "Figure 130" in the body), so this document cites section numbers, which are stable.
## Scope: the counters being segmented are controller-scoped
SMART / Health Information (LID 02h) is scope **Controller** (mandatory) with an optional namespace view ([2.0e §5.16.1, Get Log Page – Log Page Identifiers](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Section 5.16.1.3: "The information provided is over the life of the controller and is retained across power cycles"; hosts request the controller log page with NSID `FFFFFFFFh`/`0h`, the per-namespace view is optional (LPA bit 0), and "the controller log page and namespaces specific log page contain identical information" in 2.0e ([§5.16.1.3](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Data Units Written therefore accumulates per controller/subsystem, not per namespace — the identity key must be subsystem-scoped, and SN/MN/SUBNQN are all defined as NVM-subsystem fields ([§5.17.2.1: SN/MN are "the serial number/model number for the NVM subsystem"](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); the Persistent Event Log repeats that its SN/MN/SUBNQN copies are the same subsystem values).
NGUID and EUI-64 live in Identify **Namespace**, not the controller structure ([§4.3.4, §4.3.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); libnvme documents them on `struct nvme_id_ns`, while `sn`/`mn`/`fr` live on [`struct nvme_id_ctrl`](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L1467-L1477) and `subnqn` on the same structure ([types.h](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L1576))). Segmentation keyed on a namespace identifier would split or merge history whenever namespaces are attached, detached, formatted, or recreated, while the DUW counter — the thing being differenced — sails on unchanged. The spec's own namespace-identity guidance (§3.2.1.6: NSIDs "may change across power off conditions"; to detect the same namespace use UUID, NGUID, or EUI-64) addresses a different problem from ours.
## Stability verdicts (reboots, firmware, replacement)
1. **Reboots, same drive** — SN, MN, SUBNQN are stable: they are vendor-assigned subsystem fields, and an NQN "is permanent for the lifetime of the host or NVM subsystem" ([§4.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). SMART data "is retained across power cycles" ([§5.16.1.3](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
2. **Firmware update, same drive** — FR changes by definition (it reports the active revision). The spec guarantees persistence of SMART data across *power cycles*, and says nothing about firmware commits resetting counters — vendor behavior, which is precisely why ADR 0001 keeps the DUW-monotonic rule as an independent boundary. Identity (SN/MN/SUBNQN) is not specified to change with firmware; keeping FR out of the key means a firmware update never quarantines history as a "new drive", and a firmware-induced counter reset is caught by the DUW rule instead.
3. **Drive replacement, different model** — every candidate changes.
4. **Drive replacement, identical model** — MN unchanged; SN changes *if* vendor serials are unique; SUBNQN changes because both real NQNs (empirically this host's Micron embeds the serial: `nqn.2016-08.com.micron:nvme:nvm-subsystem-sn-233542F44436`) and kernel-generated NQNs (which concatenate SN and MN, [drivers/nvme/host/core.c nvme_init_subnqn](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3164-L3194)) derive from the serial.
5. **NGUID/EUI-64 under namespace churn** — not stable in the needed sense: if the UIDREUSE bit is 0 "a controller **may reuse** a non-zero NGUID/EUI64 value for a new namespace after the original namespace using the value has been deleted" ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)); libnvme's own field docs say the values hold only "throughout the life of the namespace", "preserved across namespace and controller operations" ([types.h nguid/eui64](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L2648-L2656)).
## Availability and known pathologies
- **SN/MN**: mandatory for I/O and Admin controllers ([§5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) and exposed by the kernel since 4.5 ([sysfs-nvme: /sys/class/nvme/nvmeX/{model,serial,firmware_rev}](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L1-L10)). But uniqueness is disclaimed: "The mechanism used by the vendor to assign Serial Number and Model Number values to ensure uniqueness is outside the scope of this specification" ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Duplicate serials across units are therefore spec-legal.
- **SUBNQN**: zero on pre-1.2.1 subsystems ([§5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); the Persistent Event Log likewise defines the not-supported case as all bytes cleared to 0h), and some real devices report garbage: the kernel carries `NVME_QUIRK_IGNORE_DEV_SUBNQN` for, among others, Intel P4500/P4600, Intel 760p/Pro 7600p, and a Silicon Motion device ([pci.c quirk table](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/pci.c#L4121-L4144), [nvme.h flag definition](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/nvme.h#L103-L106)). This is why Fenris should read the **kernel-exposed** value rather than the raw Identify bytes: when the device NQN is missing, invalid, or quirk-ignored, `nvme_init_subnqn` synthesizes `nqn.2014.08.org.nvmexpress:{vid}{ssvid}{sn}{mn}` — mirroring the spec's own construction for pre-1.2.1 subsystems ([§4.5.1, "NQN Construction for Older NVM Subsystems"](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf), which composes the NQN starting string, VID, SSVID, SN, MN) — so `/sys/.../subsysnqn` is populated on every kernel ≥ 4.8 for every controller ([sysfs-nvme subsysnqn entry, added 4.8](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L48-L60)). The kernel also uses the NQN as *the* subsystem key when building multipath heads ([core.c: subsystems are matched by `subsys->subnqn`](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3345-L3350)).
- **Virtual controllers / blank serials**: the kernel's own NVMe target (nvmet, the `loop` transport) sets model number to the literal `"Linux"` and generates a **random** serial per subsystem "as our controllers are ephemeral" ([target/core.c nvmet_subsys_alloc](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/core.c#L1840-L1856), [nvmet.h `NVMET_DEFAULT_CTRL_MODEL`](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/nvmet.h#L31)); Identify then reports those values verbatim ([target/admin-cmd.c](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/admin-cmd.c#L670-L677)). For such devices `model|serial` is unstable across target re-creation, while the NQN is the configured subsystem name. No identifier can make an ephemeral virtual drive look like stable hardware; the degraded marker covers it.
- **NGUID/EUI-64**: optional — the kernel sysfs attributes are documented as "Hidden if all zeros" ([sysfs-nvme nguid/eui entries](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L259-L293)), and the spec requires only that *at least one* of EUI64/NGUID/UUID be valid at namespace creation ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
## The key, normalization, and failure handling
1. **Primary key**: `strip(subnqn)` read from the kernel path (via libnvme; see next section). Values are ASCII/UTF-8 with code values 0x20–0x7E, left-justified and space-padded per the spec's string rules ([§1.4.2 ASCII/UTF-8 string conventions](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)); normalize by stripping trailing (and leading) spaces only. **No case folding**: NQNs are compared "as binary strings without any text processing (e.g., case folding)" ([§4.5, NQN processing rules](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)), and SN/MN have no canonical case either.
2. **Fallback ladder** (defensive; on Linux ≥ 4.8 the kernel fallback already fires before Fenris ever sees an empty value): (a) device-provided SUBNQN; (b) the kernel's composite `nqn.2014.08.org.nvmexpress:{vid}{ssvid}{sn}{mn}` built from Identify — the kernel spells the date with dots and concatenates the raw fixed-width SN and MN, "slightly different from the format specified" in §4.5.1's NQN construction "for historic reasons" ([core.c nvme_init_subnqn](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3164-L3194)); Fenris's fallback reproduces the kernel spelling so a fallback-built key equals the sysfs value byte-for-byte; (c) `strip(mn)|strip(sn)`. Every rung changes on drive replacement and survives firmware updates; the ladder exists so a missing rung never yields a blank key from a perfectly good physical drive.
3. **Blank key** (all components empty — e.g. a virtual controller reporting nothing): record `identity_key = ""` plus a `degraded` marker and segment by DUW monotonicity alone; surface as a fact, never fabricate uniqueness (matches ADR 0005's "facts, not alerts" posture).
4. **Duplicate keys across physical units** are undetectable by construction when both units report identical SN/MN/SUBNQN. The mitigation already exists in ADR 0001: a replacement drive almost certainly reports a **lower** DUW than the accumulated history, and any DUW decrease forces a segment boundary regardless of identity. Write deltas remain quarantined even though identity cannot distinguish the units.
5. **Recorded metadata per segment**: normalized `subnqn`, `sn`, `mn`, `fr`, plus `transport` — diagnostics for humans, not key components.
## How the collector obtains the key via libnvme
Three concrete paths, all first-party:
1. **libnvme Python bindings** (`from libnvme import nvme`, official SWIG bindings, ["python bindings for libnvme"](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/pyproject.toml#L1-L8)). `nvme.root()` scans sysfs ([nvme.i: nvme_root() calls nvme_scan](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L492-L501)); controller objects expose `model`, `serial`, `firmware`, `subsysnqn`, `name`, `sysfs_dir` as attributes ([nvme.i struct nvme_ctrl attributes](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L402-L424)); namespace objects expose `nsid`, `nguid`, `eui64`, `uuid` ([nvme.i struct nvme_ns](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L481-L490)). These getters read sysfs via `nvme_get_ctrl_attr` ([tree.c populates ctrl fields from sysfs attributes](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/tree.c#L2085-L2097)), and `__nvme_get_attr` **strips the trailing newline and trailing spaces and returns NULL when the result is empty** ([linux.c L526–552](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/linux.c#L526-L552)) — i.e., the bindings deliver exactly the normalization this decision requires, and a blank field arrives as `None`.
2. **Admin passthrough** for raw Identify: `nvme_ctrl_identify(c, &id)` fills `struct nvme_id_ctrl` ([man page: "Issues an 'identify controller' command"](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_ctrl_identify.2#L13-L17)), whose `sn[20]`/`mn[40]`/`fr[8]` and `subnqn[256]` members are documented in the [nvme_id_ctrl(2) man page](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_id_ctrl.2#L5-L15) ([subnqn member](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_id_ctrl.2#L616-L617)); namespace identifiers come from `nvme_ns_identify` filling `struct nvme_id_ns` ([tree.h nvme_ns_identify](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/tree.h#L823-L832)). Here the collector must strip trailing spaces itself.
3. **nvme-cli JSON** (built on libnvme): `nvme id-ctrl -o json` emits `sn`, `mn`, `fr`, `cntlid` with strings copied verbatim *including padding* ([nvme-print-json.c L409–L422](https://github.com/linux-nvme/nvme-cli/blob/c8ec7e41f3738b20828849228a549f4ed1d03fc0/src/nvme-print-json.c#L409-L422)) plus `subnqn`, which is included only when non-empty ([L522–L523](https://github.com/linux-nvme/nvme-cli/blob/c8ec7e41f3738b20828849228a549f4ed1d03fc0/src/nvme-print-json.c#L522-L523)); `nvme list -o json` emits per device `DevicePath`, `Firmware`, `ModelNumber`, `SerialNumber` ([v2.16 json_list_item_obj](https://github.com/linux-nvme/nvme-cli/blob/faf7326a2997dea91687fd3daa17fc405910a4c1/nvme-print-json.c#L4669-L4693)). So stripping trailing spaces is the collector's job in both nvme-cli paths.
What libnvme/nvme-cli offer beyond the current `smartctl -j` collector: `subnqn` (the chosen key), `cntlid`, `vid`/`ssvid`, and the namespace identifiers — smartmontools' NVMe JSON device section carries `model_name`, `serial_number`, `firmware_version` (from `id_ctrl.mn/sn/fr`, [nvmeprint.cpp print_drive_info](https://github.com/smartmontools/smartmontools/blob/9f83095a631ff71df44f8065c7a4a00134d3d404/smartmontools/nvmeprint.cpp#L108-L121)) and no NQN. smartmontools also trims the strings it copies ([utility.cpp format_char_array strips leading/trailing spaces](https://github.com/smartmontools/smartmontools/blob/9f83095a631ff71df44f8065c7a4a00134d3d404/smartmontools/utility.cpp#L692-L708)), so normalization is compatible across both collectors.
## Empirical check (one real drive, sysfs only)
```console
$ cat /sys/class/nvme/nvme0/{model,serial,firmware_rev,subsysnqn}
Micron_2400_MTFDKBA512QFM<spaces to 40> # sysfs preserves Identify padding
233542F44436<spaces to 20>
V3MA001<space> # FR padded to 8
nqn.2016-08.com.micron:nvme:nvm-subsystem-sn-233542F44436
$ cat /sys/block/nvme0n1/{nguid,eui,wwid}
00000000-0000-0001-00a0-752342f44436
00 a0 75 01 42 f4 44 36
eui.000000000000000100a0752342f44436
```
Confirms: raw sysfs keeps the spec's trailing-space padding (strip before storing); the vendor NQN embeds the serial; NGUID/EUI-64 are present here but are per-namespace; `/sys/class/nvme-subsystem/nvme-subsys0/{model,serial,firmware_rev,subsysnqn}` carries the same subsystem-level values ([sysfs-nvme nvme-subsystem entries, added 4.15](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L425-L434)). The kernel's own namespace `wwid` uses the priority ladder `uuid.{UUID}` → `eui.{NGUID}` → `eui.{EUI64}` → `nvme.{VID}-{SERIAL}-{MODEL}-{NSID}` ([sysfs-nvme wwid](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L277-L285)) — a first-party precedent for a serial+model fallback when better identifiers are absent.
## What ADR 0001's segment and migration rules must assume
1. **Legacy `history.jsonl` cannot carry a full identity.** The legacy collector records `device` (`/dev/nvme0`), `model` (smartctl `model_name`), `capacity_bytes`, and the counters — **no serial, no NQN, no firmware** ([fenris.py sample(): the identity-adjacent fields are `device` and `model` only](https://git.bongbetic.com/xavierk/Fenris/src/commit/d45d931a0deaaa4ff8c8ba0ff2454c0d2a211ef9/fenris.py#L73-L76)). Migration may therefore only import legacy samples under a **model-scoped legacy identity**, explicitly labeled incomplete. It cannot prove the legacy samples came from the current drive, cannot detect a same-model swap that happened before migration, and must not backfill serials retroactively.
2. **The first new-version collection run starts a new controller segment** for the monitored drive, recording full `identity_key` plus `sn`/`mn`/`fr`/`subnqn` metadata, because the legacy and new identities are not comparable — the identity change is an epistemic boundary, not a detected drive swap. Imported legacy rows keep their legacy segment; the DUW-monotonic rule already governs deltas within it.
3. **Two independent segmentation axes stay independent**: identity-key change ⇒ boundary (drive replacement); DUW decrease ⇒ boundary (counter reset, firmware quirk, or same-identity replacement). Neither implies the other; ADR 0001's `controller_segments` wording ("boundaries where controller identity changes or DUW decreases") already encodes this, and this research fixes "controller identity" to mean the normalized kernel-exposed subsystem NQN with the fallback ladder above.
4. **The key is a string with rules, not just a field choice**: normalization (space-stripping, no case folding) must be specified once and applied at write time, or the same drive could split its own history across a collector implementation change (smartctl, nvme-cli, and libnvme deliver differently-trimmed values, as shown above).
## Newly surfaced questions
- Should the store record `vid`/`ssvid`/`cntlid` alongside segment metadata now (cheap) for future composite needs, or keep segments minimal?
- Should a detected blank/degraded identity degrade the projection confidence category (it weakens the "identity/counter discontinuity" axis of the Unavailable state)?
- Does the collector pin to one acquisition path (libnvme bindings vs `nvme list -o json` vs continued `smartctl -j` plus a sysfs read for `subnqn`) — an operational choice this research does not settle?