Compare commits

...
8 changed files with 312 additions and 0 deletions
+1
View File
@@ -1,4 +1,5 @@
__pycache__/ __pycache__/
.pi/
*.pyc *.pyc
.commandcode/ .commandcode/
data/fenris.pid data/fenris.pid
+9
View File
@@ -0,0 +1,9 @@
## Agent skills
### Issue tracker
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
### Domain docs
This is a single-context repository. See `docs/agents/domain.md`.
+69
View File
@@ -0,0 +1,69 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Endurance baseline**:
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
+42
View File
@@ -0,0 +1,42 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -0,0 +1,49 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -0,0 +1,31 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger and monitoring-period bookkeeping; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
+26
View File
@@ -0,0 +1,26 @@
# Domain Docs
How engineering skills should consume this repository’s domain documentation.
## Layout
This is a single-context repository:
```text
/
├── CONTEXT.md
├── docs/adr/
└── ...
```
## Before exploring
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
## Use the glossary’s vocabulary
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
## Flag ADR conflicts
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
+85
View File
@@ -0,0 +1,85 @@
# Issue tracker: Gitea
Issues for this repository live in Gitea at:
https://git.bongbetic.com/xavierk/Fenris/issues
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
## General operations
- List: `tea issues list`
- Read: `tea issues <index> --comments`
- Create: `tea issues create --title "<title>" --description "<body>"`
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
- Assign: `tea issues edit <index> --add-assignees "<username>"`
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
- Comment: `tea comments add <index> --description "<comment>"`
- Close: `tea issues close <index>`
- Reopen: `tea issues reopen <index>`
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
## When a skill says “publish to the issue tracker”
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
## When a skill says “fetch the relevant ticket”
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
## Wayfinding operations
Wayfinder maps and decision tickets are Gitea issues.
### Map and ticket grouping
- A map has the label `wayfinder:map`.
- Create one milestone named `Wayfinder: <map title>` for the effort.
- Assign the map and all its tickets to that milestone.
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
### Blocking
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
```bash
tea api -X POST \
repos/{owner}/{repo}/issues/<blocked>/dependencies \
-F index=<blocker> \
-f owner=xavierk \
-f repo=Fenris
```
List blockers:
```bash
tea api repos/{owner}/{repo}/issues/<index>/dependencies
```
Remove the relationship with the same payload and `-X DELETE`.
### Frontier
List open issues in the map’s milestone. Exclude:
- the issue labelled `wayfinder:map`
- assigned tickets, because assignment is the claim
- tickets whose dependency query returns any open issue
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
### Claim
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
### Resolve
1. Add the answer as a resolution comment.
2. Close the ticket.
3. Re-fetch the map immediately before editing it.
4. Append a linked one-line context pointer to `Decisions so far`.
5. Create newly visible tickets, then wire dependencies in a second pass.