Compare commits

...
8 changed files with 282 additions and 0 deletions
+1
View File
@@ -1,4 +1,5 @@
__pycache__/
.pi/
*.pyc
.commandcode/
data/fenris.pid
+9
View File
@@ -0,0 +1,9 @@
## Agent skills
### Issue tracker
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
### Domain docs
This is a single-context repository. See `docs/agents/domain.md`.
+4
View File
@@ -28,6 +28,10 @@ _Avoid_: Daemon uptime, calibration window
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Store fault**:
The condition where the observation store is present but cannot be read or trusted — unreadable, corrupt, or written by a newer Fenris — degrading every view that depends on it rather than crashing or guessing.
_Avoid_: Database error, corruption, broken data
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
@@ -0,0 +1,31 @@
# 4. Installation lifecycle: Makefile-delivered venv, dormant install, sanctioned teardown
## Status
Accepted — resolves [Define installation, upgrade, and removal behavior](https://git.bongbetic.com/xavierk/Fenris/issues/9) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris runs today from the source checkout (`fenris.py`, `fenris.sh`, `data/`): code, state, and control all live relative to wherever the checkout sits. The redesign fixes system artifacts — helpers in `/usr/libexec/fenris` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), the observation store at `/var/lib/fenris/observations.db` ([ADR 0001](0001-observation-store-sqlite.md)), configuration at `/etc/fenris/fenris.conf` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), Textual on Python 3.9+ ([framework decision](https://git.bongbetic.com/xavierk/Fenris/issues/6)) — but nothing says how those artifacts are delivered, upgraded, or removed, or what happens when the checkout moves or disappears.
## Decision
1. **Delivery.** `sudo make install` builds a wheel from the checkout and installs it, with pinned dependencies, into a dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only: after install, nothing references it.
2. **Layout and manifest.** Units in `/etc/systemd/system` (`fenris-collect.{timer,service}`, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)); helpers in `/usr/libexec/fenris`; polkit policy in `/usr/share/polkit-1/actions/`; configuration and observation store in their ADR-fixed locations. The installer records every file it places in an explicit manifest consumed by upgrade and uninstall.
3. **Privilege.** One root installer (`sudo make install`); at runtime, elevation is exclusively polkit (`auth_admin`, `fenris-monitor` only, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)). The installer never enables or starts units.
4. **Dormant install.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle (`fenris monitor resume [--now]`, or the first-run TUI prompt), which enables the timer and opens the first period in one step.
5. **Legacy import.** The installer detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs [ADR 0001](0001-observation-store-sqlite.md)'s idempotent single-transaction import, and reports imported counts — existing observations never depend on checkout survival. `fenris import <path>` remains available for later finds.
6. **Upgrade.** `sudo make upgrade` builds and installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed and it is active — safe with `Persistent=no`), leaves timer state untouched, and never kills an in-flight collection run: a running oneshot finishes on its mapped interpreter, so at worst one old-code run completes to the store and the next run uses the new code. It then applies forward-only observation-store schema migrations governed by a `schema_version` table. `/var/lib/fenris` is never rebuilt.
7. **Rollback.** Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade is explicitly unsupported.
8. **Removal.** `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. `make purge` additionally removes configuration and store. Journal entries age out naturally.
9. **Dependencies.** Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing.
10. **Scaffolding and floor.** The installer creates `/var/lib/fenris` with [ADR 0001](0001-observation-store-sqlite.md)'s root-written group-read permissions and verifies `python3 ≥ 3.9`, failing cleanly otherwise — the Textual contingency becomes an install-time gate rather than a runtime crash. The database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet.
## Consequences
- Installed Fenris survives checkout deletion; the checkout is only where builds happen.
- Teardown preserves monitoring-period semantics: deliberate removal excludes the uninstalled span from the usage habit instead of leaving it as unknown-inside.
- Installs are reproducible; dependency drift cannot ride in on an upgrade.
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
- Rollback support is exactly one generation deep, no further.
- The README documents install, upgrade, uninstall/purge, and legacy import alongside [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s menu-successor mapping.
@@ -0,0 +1,27 @@
# 5. Failure and recovery: visible degradation, never fabrication
## Status
Accepted — resolves [Define failure and recovery behavior](https://git.bongbetic.com/xavierk/Fenris/issues/10) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The observation store ([ADR 0001](0001-observation-store-sqlite.md)) and the service lifecycle ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) settled single-writer transactions, bounded single failed runs, freshness grading (fresh / missed / stale), and absent-hour classification through power-on-hours evidence. Left open by the [failure ticket](https://git.bongbetic.com/xavierk/Fenris/issues/10): behavior per failure class — malformed observations inside the store, missed observations, store faults (unreadable, corrupt, or newer-schema database), and repeated collector failures — and how the habit record re-anchors after the store itself is lost.
## Decision
1. **Malformed observations — refuse at the write boundary.** The collector validates every row it would write against the store's domain invariants (hour seconds sum to 3600, non-negative DUW delta within a controller segment, coverage consistent with sample count). A violating run writes nothing for that run, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact, but under a single trusted writer they should never see one. Store invariant: everything persisted is well-formed.
2. **Missed observations — never backfill.** Fenris never interpolates, estimates, or fabricates an hour. Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s power-on-hours classification is the only inference admitted.
3. **Store faults — degrade, never recreate over.** An unreadable or corrupt database is a store fault: readers surface a "observation store unreadable" fact with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside, the next run starts a fresh store, and if the legacy import never completed, the still-present `history.jsonl` is re-imported. No built-in destructive command exists.
4. **Newer schema — readers refuse symmetrically.** The TUI and `status` detect a `user_version` newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal in [ADR 0001](0001-observation-store-sqlite.md) and the forward-only upgrade rule of [ADR 0004](0004-install-upgrade-removal-lifecycle.md).
5. **Repeated collector failures — flat cadence, no escalation.** The timer's retry is the recovery path; the settled freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff and no notification machinery; a persistent failure reads as stale exactly like any other gap.
6. **Drive-reported anomalies — facts, not alerts.** `critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status`; no alerting or notification surface exists. Fenris observes and projects; it does not alarm. The projection is unaffected: endurance math consumes writes, not warnings.
7. **Orphaned samples — the collector re-anchors observed fact.** When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent: only the sanctioned path of [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) records a `user_disabled` close. Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
## Consequences
- Validation lives at one boundary — the collector — so the store's contract is "everything in it is well-formed" and readers only defend against the impossible.
- No synthetic data can ever enter the habit record; confidence categories can be trusted to reflect real evidence.
- Store-fault recovery can lose history; the mitigation is the one-file backup story of [ADR 0001](0001-observation-store-sqlite.md), kept human-sanctioned so loss is never silent.
- Period bookkeeping splits by epistemics: the helper records intent, the collector records observed fact.
- Fenris stays fully local and silent: no notification, escalation, or alerting machinery anywhere.
+26
View File
@@ -0,0 +1,26 @@
# Domain Docs
How engineering skills should consume this repository’s domain documentation.
## Layout
This is a single-context repository:
```text
/
├── CONTEXT.md
├── docs/adr/
└── ...
```
## Before exploring
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
## Use the glossary’s vocabulary
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
## Flag ADR conflicts
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
+85
View File
@@ -0,0 +1,85 @@
# Issue tracker: Gitea
Issues for this repository live in Gitea at:
https://git.bongbetic.com/xavierk/Fenris/issues
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
## General operations
- List: `tea issues list`
- Read: `tea issues <index> --comments`
- Create: `tea issues create --title "<title>" --description "<body>"`
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
- Assign: `tea issues edit <index> --add-assignees "<username>"`
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
- Comment: `tea comments add <index> --description "<comment>"`
- Close: `tea issues close <index>`
- Reopen: `tea issues reopen <index>`
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
## When a skill says “publish to the issue tracker”
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
## When a skill says “fetch the relevant ticket”
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
## Wayfinding operations
Wayfinder maps and decision tickets are Gitea issues.
### Map and ticket grouping
- A map has the label `wayfinder:map`.
- Create one milestone named `Wayfinder: <map title>` for the effort.
- Assign the map and all its tickets to that milestone.
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
### Blocking
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
```bash
tea api -X POST \
repos/{owner}/{repo}/issues/<blocked>/dependencies \
-F index=<blocker> \
-f owner=xavierk \
-f repo=Fenris
```
List blockers:
```bash
tea api repos/{owner}/{repo}/issues/<index>/dependencies
```
Remove the relationship with the same payload and `-X DELETE`.
### Frontier
List open issues in the map’s milestone. Exclude:
- the issue labelled `wayfinder:map`
- assigned tickets, because assignment is the claim
- tickets whose dependency query returns any open issue
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
### Claim
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
### Resolve
1. Add the answer as a resolution comment.
2. Close the ticket.
3. Re-fetch the map immediately before editing it.
4. Append a linked one-line context pointer to `Decisions so far`.
5. Create newly visible tickets, then wire dependencies in a second pass.
+99
View File
@@ -0,0 +1,99 @@
# Controller identity for observation-history segmentation
Research for [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11).
## Decision
Record the **subsystem NQN as exposed by the Linux kernel** — normalized, trailing-space stripped — as the controller-segment identity key:
```text
identity_key = strip(subnqn) # /sys/class/nvme-subsystem/…/subsysnqn,
# identical to /sys/class/nvme/nvmeX/subsysnqn
fallback: "nqn.2014.08.org.nvmexpress:" + hex4(vid) + hex4(ssvid)
+ raw20(sn) + raw40(mn) # byte-for-byte the kernel's synthesized NQN
last resort: strip(mn) + "|" + strip(sn) # when only smartctl-style fields exist
```
Store the raw `subnqn`, `sn`, `mn` strings (normalized) plus `fr` (firmware revision) as **segment metadata**, never `fr` inside the key: firmware revision is the one mandatory field that legitimately changes on the same drive ([Base Spec 2.0e §5.17.2.1: FR is the *currently active* firmware revision](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Namespace identifiers (NGUID, EUI-64, UUID) are excluded from the key: they are namespace-scoped while the SMART counters being segmented are controller-scoped, and each may be absent or reused. When the key is blank, record an empty key with a degraded marker and rely on the DUW-monotonic rule; a same-model same-serial replacement is unobservable by any identifier and is caught — as a boundary, not an identity change — by the DUW-decrease rule of ADR 0001.
This key satisfies the ticket's asymmetry: a drive replacement changes `subnqn` (real NQNs are unique per subsystem; kernel-generated ones embed serial+model), while firmware quirks and counter resets on the same drive leave it unchanged — counter discontinuities are already ADR 0001's second segmentation axis.
## What each candidate is
| Candidate | Spec definition | Scope | Mandatory | Verdict for the key |
|---|---|---|---|---|
| **SN + MN** | ASCII strings assigned by the vendor in Identify Controller, bytes 23:04 and 63:24; §4.3 shows them left-justified and space-padded ([2.0e §5.17.2.1, Identify Controller data structure](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf), [§4.3 Identifier Format and Layout](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | NVM subsystem | Mandatory for I/O and Admin controllers | Core of the fallback; uniqueness explicitly not guaranteed by the spec |
| **FR** | Currently active firmware revision, ASCII, bytes 71:64 ([2.0e §5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Domain (subsystem) | Mandatory | Never in the key — it is *meant* to change on the same drive |
| **SUBNQN** | NVM Subsystem NQN, UTF-8 null-terminated, bytes 1023:768; mandatory if the controller is ≥ 1.2.1, otherwise may be all zero ([2.0e §5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | NVM subsystem (shared by all its controllers) | Mandatory ≥ 1.2.1, optional below | **Chosen key**; the spec says hosts *should* use it as the subsystem's unique identifier |
| **NGUID / EUI-64** | IEEE-based identifiers in Identify Namespace, bytes 119:104 / 127:120 ([§4.3.4, §4.3.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Namespace | Optional; may both be zero | Rejected — wrong scope, may be missing, may be reused |
| **CNTLID** | Controller ID, unique only *within* a subsystem ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) | Controller | Mandatory | Rejected — not unique across subsystems (this host's drive reports 0) |
A note on the PDF: figure numbers in the table of contents of the 2.0e revision are offset from the body captions (e.g. the SN/MN figure is "Figure 128" in the TOC but "Figure 130" in the body), so this document cites section numbers, which are stable.
## Scope: the counters being segmented are controller-scoped
SMART / Health Information (LID 02h) is scope **Controller** (mandatory) with an optional namespace view ([2.0e §5.16.1, Get Log Page – Log Page Identifiers](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Section 5.16.1.3: "The information provided is over the life of the controller and is retained across power cycles"; hosts request the controller log page with NSID `FFFFFFFFh`/`0h`, the per-namespace view is optional (LPA bit 0), and "the controller log page and namespaces specific log page contain identical information" in 2.0e ([§5.16.1.3](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Data Units Written therefore accumulates per controller/subsystem, not per namespace — the identity key must be subsystem-scoped, and SN/MN/SUBNQN are all defined as NVM-subsystem fields ([§5.17.2.1: SN/MN are "the serial number/model number for the NVM subsystem"](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); the Persistent Event Log repeats that its SN/MN/SUBNQN copies are the same subsystem values).
NGUID and EUI-64 live in Identify **Namespace**, not the controller structure ([§4.3.4, §4.3.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); libnvme documents them on `struct nvme_id_ns`, while `sn`/`mn`/`fr` live on [`struct nvme_id_ctrl`](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L1467-L1477) and `subnqn` on the same structure ([types.h](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L1576))). Segmentation keyed on a namespace identifier would split or merge history whenever namespaces are attached, detached, formatted, or recreated, while the DUW counter — the thing being differenced — sails on unchanged. The spec's own namespace-identity guidance (§3.2.1.6: NSIDs "may change across power off conditions"; to detect the same namespace use UUID, NGUID, or EUI-64) addresses a different problem from ours.
## Stability verdicts (reboots, firmware, replacement)
1. **Reboots, same drive** — SN, MN, SUBNQN are stable: they are vendor-assigned subsystem fields, and an NQN "is permanent for the lifetime of the host or NVM subsystem" ([§4.5](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). SMART data "is retained across power cycles" ([§5.16.1.3](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
2. **Firmware update, same drive** — FR changes by definition (it reports the active revision). The spec guarantees persistence of SMART data across *power cycles*, and says nothing about firmware commits resetting counters — vendor behavior, which is precisely why ADR 0001 keeps the DUW-monotonic rule as an independent boundary. Identity (SN/MN/SUBNQN) is not specified to change with firmware; keeping FR out of the key means a firmware update never quarantines history as a "new drive", and a firmware-induced counter reset is caught by the DUW rule instead.
3. **Drive replacement, different model** — every candidate changes.
4. **Drive replacement, identical model** — MN unchanged; SN changes *if* vendor serials are unique; SUBNQN changes because both real NQNs (empirically this host's Micron embeds the serial: `nqn.2016-08.com.micron:nvme:nvm-subsystem-sn-233542F44436`) and kernel-generated NQNs (which concatenate SN and MN, [drivers/nvme/host/core.c nvme_init_subnqn](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3164-L3194)) derive from the serial.
5. **NGUID/EUI-64 under namespace churn** — not stable in the needed sense: if the UIDREUSE bit is 0 "a controller **may reuse** a non-zero NGUID/EUI64 value for a new namespace after the original namespace using the value has been deleted" ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)); libnvme's own field docs say the values hold only "throughout the life of the namespace", "preserved across namespace and controller operations" ([types.h nguid/eui64](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/types.h#L2648-L2656)).
## Availability and known pathologies
- **SN/MN**: mandatory for I/O and Admin controllers ([§5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)) and exposed by the kernel since 4.5 ([sysfs-nvme: /sys/class/nvme/nvmeX/{model,serial,firmware_rev}](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L1-L10)). But uniqueness is disclaimed: "The mechanism used by the vendor to assign Serial Number and Model Number values to ensure uniqueness is outside the scope of this specification" ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)). Duplicate serials across units are therefore spec-legal.
- **SUBNQN**: zero on pre-1.2.1 subsystems ([§5.17.2.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf); the Persistent Event Log likewise defines the not-supported case as all bytes cleared to 0h), and some real devices report garbage: the kernel carries `NVME_QUIRK_IGNORE_DEV_SUBNQN` for, among others, Intel P4500/P4600, Intel 760p/Pro 7600p, and a Silicon Motion device ([pci.c quirk table](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/pci.c#L4121-L4144), [nvme.h flag definition](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/nvme.h#L103-L106)). This is why Fenris should read the **kernel-exposed** value rather than the raw Identify bytes: when the device NQN is missing, invalid, or quirk-ignored, `nvme_init_subnqn` synthesizes `nqn.2014.08.org.nvmexpress:{vid}{ssvid}{sn}{mn}` — mirroring the spec's own construction for pre-1.2.1 subsystems ([§4.5.1, "NQN Construction for Older NVM Subsystems"](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf), which composes the NQN starting string, VID, SSVID, SN, MN) — so `/sys/.../subsysnqn` is populated on every kernel ≥ 4.8 for every controller ([sysfs-nvme subsysnqn entry, added 4.8](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L48-L60)). The kernel also uses the NQN as *the* subsystem key when building multipath heads ([core.c: subsystems are matched by `subsys->subnqn`](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3345-L3350)).
- **Virtual controllers / blank serials**: the kernel's own NVMe target (nvmet, the `loop` transport) sets model number to the literal `"Linux"` and generates a **random** serial per subsystem "as our controllers are ephemeral" ([target/core.c nvmet_subsys_alloc](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/core.c#L1840-L1856), [nvmet.h `NVMET_DEFAULT_CTRL_MODEL`](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/nvmet.h#L31)); Identify then reports those values verbatim ([target/admin-cmd.c](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/target/admin-cmd.c#L670-L677)). For such devices `model|serial` is unstable across target re-creation, while the NQN is the configured subsystem name. No identifier can make an ephemeral virtual drive look like stable hardware; the degraded marker covers it.
- **NGUID/EUI-64**: optional — the kernel sysfs attributes are documented as "Hidden if all zeros" ([sysfs-nvme nguid/eui entries](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L259-L293)), and the spec requires only that *at least one* of EUI64/NGUID/UUID be valid at namespace creation ([§4.5.1](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)).
## The key, normalization, and failure handling
1. **Primary key**: `strip(subnqn)` read from the kernel path (via libnvme; see next section). Values are ASCII/UTF-8 with code values 0x20–0x7E, left-justified and space-padded per the spec's string rules ([§1.4.2 ASCII/UTF-8 string conventions](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)); normalize by stripping trailing (and leading) spaces only. **No case folding**: NQNs are compared "as binary strings without any text processing (e.g., case folding)" ([§4.5, NQN processing rules](https://nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0e-2024.07.29-Ratified.pdf)), and SN/MN have no canonical case either.
2. **Fallback ladder** (defensive; on Linux ≥ 4.8 the kernel fallback already fires before Fenris ever sees an empty value): (a) device-provided SUBNQN; (b) the kernel's composite `nqn.2014.08.org.nvmexpress:{vid}{ssvid}{sn}{mn}` built from Identify — the kernel spells the date with dots and concatenates the raw fixed-width SN and MN, "slightly different from the format specified" in §4.5.1's NQN construction "for historic reasons" ([core.c nvme_init_subnqn](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/drivers/nvme/host/core.c#L3164-L3194)); Fenris's fallback reproduces the kernel spelling so a fallback-built key equals the sysfs value byte-for-byte; (c) `strip(mn)|strip(sn)`. Every rung changes on drive replacement and survives firmware updates; the ladder exists so a missing rung never yields a blank key from a perfectly good physical drive.
3. **Blank key** (all components empty — e.g. a virtual controller reporting nothing): record `identity_key = ""` plus a `degraded` marker and segment by DUW monotonicity alone; surface as a fact, never fabricate uniqueness (matches ADR 0005's "facts, not alerts" posture).
4. **Duplicate keys across physical units** are undetectable by construction when both units report identical SN/MN/SUBNQN. The mitigation already exists in ADR 0001: a replacement drive almost certainly reports a **lower** DUW than the accumulated history, and any DUW decrease forces a segment boundary regardless of identity. Write deltas remain quarantined even though identity cannot distinguish the units.
5. **Recorded metadata per segment**: normalized `subnqn`, `sn`, `mn`, `fr`, plus `transport` — diagnostics for humans, not key components.
## How the collector obtains the key via libnvme
Three concrete paths, all first-party:
1. **libnvme Python bindings** (`from libnvme import nvme`, official SWIG bindings, ["python bindings for libnvme"](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/pyproject.toml#L1-L8)). `nvme.root()` scans sysfs ([nvme.i: nvme_root() calls nvme_scan](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L492-L501)); controller objects expose `model`, `serial`, `firmware`, `subsysnqn`, `name`, `sysfs_dir` as attributes ([nvme.i struct nvme_ctrl attributes](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L402-L424)); namespace objects expose `nsid`, `nguid`, `eui64`, `uuid` ([nvme.i struct nvme_ns](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/libnvme/nvme.i#L481-L490)). These getters read sysfs via `nvme_get_ctrl_attr` ([tree.c populates ctrl fields from sysfs attributes](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/tree.c#L2085-L2097)), and `__nvme_get_attr` **strips the trailing newline and trailing spaces and returns NULL when the result is empty** ([linux.c L526–552](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/linux.c#L526-L552)) — i.e., the bindings deliver exactly the normalization this decision requires, and a blank field arrives as `None`.
2. **Admin passthrough** for raw Identify: `nvme_ctrl_identify(c, &id)` fills `struct nvme_id_ctrl` ([man page: "Issues an 'identify controller' command"](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_ctrl_identify.2#L13-L17)), whose `sn[20]`/`mn[40]`/`fr[8]` and `subnqn[256]` members are documented in the [nvme_id_ctrl(2) man page](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_id_ctrl.2#L5-L15) ([subnqn member](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/doc/man/nvme_id_ctrl.2#L616-L617)); namespace identifiers come from `nvme_ns_identify` filling `struct nvme_id_ns` ([tree.h nvme_ns_identify](https://github.com/linux-nvme/libnvme/blob/ad61ac8a319ad0823c1c9861eecbf66125f8b9a1/src/nvme/tree.h#L823-L832)). Here the collector must strip trailing spaces itself.
3. **nvme-cli JSON** (built on libnvme): `nvme id-ctrl -o json` emits `sn`, `mn`, `fr`, `cntlid` with strings copied verbatim *including padding* ([nvme-print-json.c L409–L422](https://github.com/linux-nvme/nvme-cli/blob/c8ec7e41f3738b20828849228a549f4ed1d03fc0/src/nvme-print-json.c#L409-L422)) plus `subnqn`, which is included only when non-empty ([L522–L523](https://github.com/linux-nvme/nvme-cli/blob/c8ec7e41f3738b20828849228a549f4ed1d03fc0/src/nvme-print-json.c#L522-L523)); `nvme list -o json` emits per device `DevicePath`, `Firmware`, `ModelNumber`, `SerialNumber` ([v2.16 json_list_item_obj](https://github.com/linux-nvme/nvme-cli/blob/faf7326a2997dea91687fd3daa17fc405910a4c1/nvme-print-json.c#L4669-L4693)). So stripping trailing spaces is the collector's job in both nvme-cli paths.
What libnvme/nvme-cli offer beyond the current `smartctl -j` collector: `subnqn` (the chosen key), `cntlid`, `vid`/`ssvid`, and the namespace identifiers — smartmontools' NVMe JSON device section carries `model_name`, `serial_number`, `firmware_version` (from `id_ctrl.mn/sn/fr`, [nvmeprint.cpp print_drive_info](https://github.com/smartmontools/smartmontools/blob/9f83095a631ff71df44f8065c7a4a00134d3d404/smartmontools/nvmeprint.cpp#L108-L121)) and no NQN. smartmontools also trims the strings it copies ([utility.cpp format_char_array strips leading/trailing spaces](https://github.com/smartmontools/smartmontools/blob/9f83095a631ff71df44f8065c7a4a00134d3d404/smartmontools/utility.cpp#L692-L708)), so normalization is compatible across both collectors.
## Empirical check (one real drive, sysfs only)
```console
$ cat /sys/class/nvme/nvme0/{model,serial,firmware_rev,subsysnqn}
Micron_2400_MTFDKBA512QFM<spaces to 40> # sysfs preserves Identify padding
233542F44436<spaces to 20>
V3MA001<space> # FR padded to 8
nqn.2016-08.com.micron:nvme:nvm-subsystem-sn-233542F44436
$ cat /sys/block/nvme0n1/{nguid,eui,wwid}
00000000-0000-0001-00a0-752342f44436
00 a0 75 01 42 f4 44 36
eui.000000000000000100a0752342f44436
```
Confirms: raw sysfs keeps the spec's trailing-space padding (strip before storing); the vendor NQN embeds the serial; NGUID/EUI-64 are present here but are per-namespace; `/sys/class/nvme-subsystem/nvme-subsys0/{model,serial,firmware_rev,subsysnqn}` carries the same subsystem-level values ([sysfs-nvme nvme-subsystem entries, added 4.15](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L425-L434)). The kernel's own namespace `wwid` uses the priority ladder `uuid.{UUID}` → `eui.{NGUID}` → `eui.{EUI64}` → `nvme.{VID}-{SERIAL}-{MODEL}-{NSID}` ([sysfs-nvme wwid](https://github.com/torvalds/linux/blob/cee9395acd8043be0644b25c34bfa86623f2b935/Documentation/ABI/stable/sysfs-nvme#L277-L285)) — a first-party precedent for a serial+model fallback when better identifiers are absent.
## What ADR 0001's segment and migration rules must assume
1. **Legacy `history.jsonl` cannot carry a full identity.** The legacy collector records `device` (`/dev/nvme0`), `model` (smartctl `model_name`), `capacity_bytes`, and the counters — **no serial, no NQN, no firmware** ([fenris.py sample(): the identity-adjacent fields are `device` and `model` only](https://git.bongbetic.com/xavierk/Fenris/src/commit/d45d931a0deaaa4ff8c8ba0ff2454c0d2a211ef9/fenris.py#L73-L76)). Migration may therefore only import legacy samples under a **model-scoped legacy identity**, explicitly labeled incomplete. It cannot prove the legacy samples came from the current drive, cannot detect a same-model swap that happened before migration, and must not backfill serials retroactively.
2. **The first new-version collection run starts a new controller segment** for the monitored drive, recording full `identity_key` plus `sn`/`mn`/`fr`/`subnqn` metadata, because the legacy and new identities are not comparable — the identity change is an epistemic boundary, not a detected drive swap. Imported legacy rows keep their legacy segment; the DUW-monotonic rule already governs deltas within it.
3. **Two independent segmentation axes stay independent**: identity-key change ⇒ boundary (drive replacement); DUW decrease ⇒ boundary (counter reset, firmware quirk, or same-identity replacement). Neither implies the other; ADR 0001's `controller_segments` wording ("boundaries where controller identity changes or DUW decreases") already encodes this, and this research fixes "controller identity" to mean the normalized kernel-exposed subsystem NQN with the fallback ladder above.
4. **The key is a string with rules, not just a field choice**: normalization (space-stripping, no case folding) must be specified once and applied at write time, or the same drive could split its own history across a collector implementation change (smartctl, nvme-cli, and libnvme deliver differently-trimmed values, as shown above).
## Newly surfaced questions
- Should the store record `vid`/`ssvid`/`cntlid` alongside segment metadata now (cheap) for future composite needs, or keep segments minimal?
- Should a detected blank/degraded identity degrade the projection confidence category (it weakens the "identity/counter discontinuity" axis of the Unavailable state)?
- Does the collector pin to one acquisition path (libnvme bindings vs `nvme list -o json` vs continued `smartctl -j` plus a sysfs read for `subnqn`) — an operational choice this research does not settle?