Compare commits
4
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
de1e8c753b | ||
|
|
26a6703152 | ||
|
|
43467f8957 | ||
|
|
a63da74e44 |
@@ -1,4 +1,5 @@
|
|||||||
__pycache__/
|
__pycache__/
|
||||||
|
.pi/
|
||||||
*.pyc
|
*.pyc
|
||||||
.commandcode/
|
.commandcode/
|
||||||
data/fenris.pid
|
data/fenris.pid
|
||||||
|
|||||||
@@ -0,0 +1,9 @@
|
|||||||
|
## Agent skills
|
||||||
|
|
||||||
|
### Issue tracker
|
||||||
|
|
||||||
|
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
|
||||||
|
|
||||||
|
### Domain docs
|
||||||
|
|
||||||
|
This is a single-context repository. See `docs/agents/domain.md`.
|
||||||
+69
@@ -0,0 +1,69 @@
|
|||||||
|
# Fenris
|
||||||
|
|
||||||
|
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
|
||||||
|
|
||||||
|
## Language
|
||||||
|
|
||||||
|
**Observation history**:
|
||||||
|
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
|
||||||
|
_Avoid_: Calibration data, temporary history
|
||||||
|
|
||||||
|
**Observed usage habit**:
|
||||||
|
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
|
||||||
|
_Avoid_: Current usage, benchmark workload
|
||||||
|
|
||||||
|
**Usage-adjusted theoretical lifespan**:
|
||||||
|
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
|
||||||
|
_Avoid_: Future life, actual lifespan, failure date
|
||||||
|
|
||||||
|
**Projection confidence**:
|
||||||
|
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
|
||||||
|
_Avoid_: Accuracy percentage, certainty
|
||||||
|
|
||||||
|
**Monitoring period**:
|
||||||
|
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
|
||||||
|
_Avoid_: Daemon uptime, calibration window
|
||||||
|
|
||||||
|
**Observation store**:
|
||||||
|
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
|
||||||
|
_Avoid_: Data directory, history.jsonl, the database (generic)
|
||||||
|
|
||||||
|
**Hour observation**:
|
||||||
|
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
|
||||||
|
_Avoid_: Hourly record, hourly.jsonl entry
|
||||||
|
|
||||||
|
**Day aggregate**:
|
||||||
|
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
|
||||||
|
_Avoid_: Daily summary, daily stats
|
||||||
|
|
||||||
|
**Controller segment**:
|
||||||
|
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
|
||||||
|
_Avoid_: Counter reset handling, drive swap detection
|
||||||
|
|
||||||
|
**Endurance baseline**:
|
||||||
|
The write-endurance value a projection consumes: a verified rated-TBW override stored with provenance when one exists, otherwise a coarse implied baseline derived from vendor wear and labeled as such.
|
||||||
|
_Avoid_: TBW value, failure threshold, max writes
|
||||||
|
|
||||||
|
**Sustained regime**:
|
||||||
|
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
|
||||||
|
_Avoid_: Current window, detection period
|
||||||
|
|
||||||
|
**Habit change**:
|
||||||
|
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
|
||||||
|
_Avoid_: Spike, anomaly
|
||||||
|
|
||||||
|
**Scenario range**:
|
||||||
|
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
|
||||||
|
_Avoid_: Confidence interval, error bar
|
||||||
|
|
||||||
|
**Coverage**:
|
||||||
|
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
|
||||||
|
_Avoid_: Uptime, sample count
|
||||||
|
|
||||||
|
**Collection run**:
|
||||||
|
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
|
||||||
|
_Avoid_: Poll, daemon tick
|
||||||
|
|
||||||
|
**Deliberate disable**:
|
||||||
|
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
|
||||||
|
_Avoid_: Manual stop, service stop
|
||||||
@@ -0,0 +1,42 @@
|
|||||||
|
# 1. Observation store: a single SQLite database
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
|
||||||
|
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
|
||||||
|
3. **Entities**:
|
||||||
|
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
|
||||||
|
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
|
||||||
|
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
|
||||||
|
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
|
||||||
|
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment.
|
||||||
|
- `endurance_baseline` — verified rated-TBW override in bytes plus provenance (source URL, document revision, entry date).
|
||||||
|
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
|
||||||
|
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
|
||||||
|
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
|
||||||
|
6. **Migration** (first new-version collection run):
|
||||||
|
1. If the database already carries the legacy-import marker, do nothing.
|
||||||
|
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
|
||||||
|
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
|
||||||
|
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
|
||||||
|
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
|
||||||
|
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
|
||||||
|
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
|
||||||
|
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
|
||||||
|
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- Backups and state migration are copying one file (plus its WAL sidecars).
|
||||||
|
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
|
||||||
|
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
|
||||||
|
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
|
||||||
|
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# 2. Projection model: sustained-regime rate with categorical confidence
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
|
||||||
|
2. **Headline rate from the sustained regime.**
|
||||||
|
```text
|
||||||
|
rate = regime DUW delta bytes / in-period wall-clock seconds
|
||||||
|
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
|
||||||
|
E_rated = entered_TBW × 10¹² bytes
|
||||||
|
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
|
||||||
|
```
|
||||||
|
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
|
||||||
|
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
|
||||||
|
4. **Hour classification** (named constants, no configuration surface):
|
||||||
|
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
|
||||||
|
- **Active**: DUW delta ≥ 256 MiB in the hour.
|
||||||
|
- **Idle**: powered on, sampled, below the active threshold.
|
||||||
|
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
|
||||||
|
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
|
||||||
|
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
|
||||||
|
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
|
||||||
|
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
|
||||||
|
8. **Confidence rule table.**
|
||||||
|
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
|
||||||
|
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old.
|
||||||
|
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
|
||||||
|
- Confidence always renders as state plus contributing facts, never a percentage.
|
||||||
|
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive.
|
||||||
|
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
|
||||||
|
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
|
||||||
|
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
|
||||||
|
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts, the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
|
||||||
|
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
|
||||||
|
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
|
||||||
|
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
|
||||||
|
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
|
||||||
|
2. **Cadence.** Default five minutes: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
|
||||||
|
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
|
||||||
|
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger and monitoring-period bookkeeping; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
|
||||||
|
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
|
||||||
|
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
|
||||||
|
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
|
||||||
|
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
|
||||||
|
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
|
||||||
|
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
|
||||||
|
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
|
||||||
|
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
|
||||||
|
- Headless administration has full parity: every TUI action has a CLI twin.
|
||||||
|
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
|
||||||
|
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
# Domain Docs
|
||||||
|
|
||||||
|
How engineering skills should consume this repository’s domain documentation.
|
||||||
|
|
||||||
|
## Layout
|
||||||
|
|
||||||
|
This is a single-context repository:
|
||||||
|
|
||||||
|
```text
|
||||||
|
/
|
||||||
|
├── CONTEXT.md
|
||||||
|
├── docs/adr/
|
||||||
|
└── ...
|
||||||
|
```
|
||||||
|
|
||||||
|
## Before exploring
|
||||||
|
|
||||||
|
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
|
||||||
|
|
||||||
|
## Use the glossary’s vocabulary
|
||||||
|
|
||||||
|
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
|
||||||
|
|
||||||
|
## Flag ADR conflicts
|
||||||
|
|
||||||
|
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
|
||||||
@@ -0,0 +1,85 @@
|
|||||||
|
# Issue tracker: Gitea
|
||||||
|
|
||||||
|
Issues for this repository live in Gitea at:
|
||||||
|
|
||||||
|
https://git.bongbetic.com/xavierk/Fenris/issues
|
||||||
|
|
||||||
|
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
|
||||||
|
|
||||||
|
## General operations
|
||||||
|
|
||||||
|
- List: `tea issues list`
|
||||||
|
- Read: `tea issues <index> --comments`
|
||||||
|
- Create: `tea issues create --title "<title>" --description "<body>"`
|
||||||
|
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
|
||||||
|
- Assign: `tea issues edit <index> --add-assignees "<username>"`
|
||||||
|
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
|
||||||
|
- Comment: `tea comments add <index> --description "<comment>"`
|
||||||
|
- Close: `tea issues close <index>`
|
||||||
|
- Reopen: `tea issues reopen <index>`
|
||||||
|
|
||||||
|
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
|
||||||
|
|
||||||
|
## When a skill says “publish to the issue tracker”
|
||||||
|
|
||||||
|
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
|
||||||
|
|
||||||
|
## When a skill says “fetch the relevant ticket”
|
||||||
|
|
||||||
|
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
|
||||||
|
|
||||||
|
## Wayfinding operations
|
||||||
|
|
||||||
|
Wayfinder maps and decision tickets are Gitea issues.
|
||||||
|
|
||||||
|
### Map and ticket grouping
|
||||||
|
|
||||||
|
- A map has the label `wayfinder:map`.
|
||||||
|
- Create one milestone named `Wayfinder: <map title>` for the effort.
|
||||||
|
- Assign the map and all its tickets to that milestone.
|
||||||
|
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
|
||||||
|
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
|
||||||
|
|
||||||
|
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
|
||||||
|
|
||||||
|
### Blocking
|
||||||
|
|
||||||
|
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
tea api -X POST \
|
||||||
|
repos/{owner}/{repo}/issues/<blocked>/dependencies \
|
||||||
|
-F index=<blocker> \
|
||||||
|
-f owner=xavierk \
|
||||||
|
-f repo=Fenris
|
||||||
|
```
|
||||||
|
|
||||||
|
List blockers:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
tea api repos/{owner}/{repo}/issues/<index>/dependencies
|
||||||
|
```
|
||||||
|
|
||||||
|
Remove the relationship with the same payload and `-X DELETE`.
|
||||||
|
|
||||||
|
### Frontier
|
||||||
|
|
||||||
|
List open issues in the map’s milestone. Exclude:
|
||||||
|
|
||||||
|
- the issue labelled `wayfinder:map`
|
||||||
|
- assigned tickets, because assignment is the claim
|
||||||
|
- tickets whose dependency query returns any open issue
|
||||||
|
|
||||||
|
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
|
||||||
|
|
||||||
|
### Claim
|
||||||
|
|
||||||
|
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
|
||||||
|
|
||||||
|
### Resolve
|
||||||
|
|
||||||
|
1. Add the answer as a resolution comment.
|
||||||
|
2. Close the ticket.
|
||||||
|
3. Re-fetch the map immediately before editing it.
|
||||||
|
4. Append a linked one-line context pointer to `Decisions so far`.
|
||||||
|
5. Create newly visible tickets, then wire dependencies in a second pass.
|
||||||
@@ -1,69 +0,0 @@
|
|||||||
# Python TUI frameworks for Fenris
|
|
||||||
|
|
||||||
_Research snapshot: 2026-08-31. Planning only; no product implementation._
|
|
||||||
|
|
||||||
## Decision
|
|
||||||
|
|
||||||
Use **Textual** for Fenris's later TUI prototype and, if the prototype checks below pass, for the persistent TUI.
|
|
||||||
|
|
||||||
Textual is the best fit for Fenris's keyboard-first Overview, Usage History, Drive Health, and Settings views because one framework supplies flexible grid/horizontal/vertical layouts, dashboard and form widgets, background workers, and a headless interaction driver ([layout](https://textual.textualize.io/guide/layout/), [widget gallery](https://textual.textualize.io/widget_gallery/), [workers](https://textual.textualize.io/guide/workers/), [testing](https://textual.textualize.io/guide/testing/)). Its principal costs are the largest direct dependency set in this shortlist and only a compact `Sparkline` as built-in charting; detailed historical plots would need a custom widget or the first-party `textual-plotext` integration ([PyPI metadata](https://pypi.org/pypi/textual/json), [Sparkline](https://textual.textualize.io/widgets/sparkline/), [`textual-plotext`](https://github.com/Textualize/textual-plotext)).
|
|
||||||
|
|
||||||
This recommendation is **conditional on raising Fenris's Python floor from its currently documented Python 3.7+ to Python 3.9+**: Textual 8.2.8 and Urwid 4.0.13 require Python 3.9+, while prompt_toolkit 3.0.53 requires Python 3.10+ ([Fenris README](../../README.md#you-need), [Textual metadata](https://pypi.org/pypi/textual/json), [Urwid metadata](https://pypi.org/pypi/urwid/json), [prompt_toolkit metadata](https://pypi.org/pypi/prompt-toolkit/json)). If retaining Python 3.7 is mandatory, none of the current versions evaluated here qualifies.
|
|
||||||
|
|
||||||
**Fallback:** choose **Urwid** if prototype evidence shows that explicit low-level terminal/display control or a smaller direct dependency set matters more than Textual's higher-level layouts, forms, workers, and test driver ([display modules](https://urwid.org/manual/displaymodules.html), [package metadata](https://pypi.org/pypi/urwid/json)).
|
|
||||||
|
|
||||||
## Scope and method
|
|
||||||
|
|
||||||
The shortlist covers maintained, full-screen-capable Python projects with first-party application/widget documentation: Textual, Urwid, and prompt_toolkit. The evaluation uses official documentation, repository release APIs, and PyPI package metadata. Versions, Python floors, dependencies, and release activity are a point-in-time snapshot and should be rechecked when dependencies are locked.
|
|
||||||
|
|
||||||
## Comparison
|
|
||||||
|
|
||||||
| Criterion | Textual 8.2.8 | Urwid 4.0.13 | prompt_toolkit 3.0.53 |
|
|
||||||
|---|---|---|---|
|
|
||||||
| **Terminal compatibility** | PyPI classifies Linux, macOS, and Windows 10/11 support; this exceeds Fenris's Linux-only scope ([metadata](https://pypi.org/pypi/textual/json)). | Offers pure-Python raw and OS curses displays. The manual compares UTF-8, color, mouse, and external-event-loop capabilities; curses is described as broadly terminal-compatible, while raw supports 88/256/24-bit color and external loops ([display modules](https://urwid.org/manual/displaymodules.html)). | Provides full-screen applications and platform-specific POSIX/Windows test input; official input docs describe cross-platform one-key-at-a-time input and Windows behavior ([full-screen apps](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/full_screen_apps.html), [input](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/asking_for_input.html), [testing](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/advanced_topics/unit_testing.html)). |
|
|
||||||
| **Responsive layout** | Vertical, horizontal, and grid layouts support fractional, percentage, and automatic sizing, nesting, spans, overflow, docking, and runtime layout changes ([layout](https://textual.textualize.io/guide/layout/)). | `Pile`, `Columns`, `GridFlow`, `Overlay`, `ListBox`, padding, and flow/box/fixed sizing provide procedural composition ([widgets](https://urwid.org/manual/widgets.html)). | `VSplit`, `HSplit`, `FloatContainer`, `ConditionalContainer`, and scrollable panes compose full-screen regions, but the reviewed guide does not document breakpoint-style responsiveness ([full-screen apps](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/full_screen_apps.html)). |
|
|
||||||
| **Dashboard, charts, history** | Built-ins include `DataTable`, `ProgressBar`, `Digits`, and `Sparkline`; Sparkline summarizes reactive numerical data into bars according to available widget width ([gallery](https://textual.textualize.io/widget_gallery/), [Sparkline](https://textual.textualize.io/widgets/sparkline/)). `textual-plotext` adds a `PlotextPlot` wrapper for richer plotting ([repository](https://github.com/Textualize/textual-plotext)). | Includes `BarGraph`, `GraphVScale`, and `ProgressBar`; richer history views still require composition or custom rendering ([graph widgets](https://urwid.org/reference/widget.html#graph-widgets)). | The documented reusable widget set supports text areas, buttons, frames, dialogs, and menus, but does not list a chart widget; Fenris charts would therefore be custom controls/rendering ([widget reference](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/reference.html#module-prompt_toolkit.widgets)). |
|
|
||||||
| **Forms and keyboard use** | Built-ins include `Input`, `MaskedInput`, `Select`, checkbox/radio/selection controls, switches, buttons, and text areas ([gallery](https://textual.textualize.io/widget_gallery/)). `Input` supports validation on change, blur, or submit and emits results with its messages ([Input](https://textual.textualize.io/widgets/input/)). | Supplies `Edit`, `Button`, `CheckBox`, `RadioButton`, focus handling, and keyboard/mouse event support; application code owns more form orchestration ([widgets](https://urwid.org/manual/widgets.html)). | Strong key-binding, editing, completion, history, suggestion, and validation facilities; full-screen reusable components include `TextArea` and `Button` ([input](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/asking_for_input.html), [full-screen apps](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/full_screen_apps.html)). |
|
|
||||||
| **Async/background updates** | Async workers run coroutines concurrently; thread workers cover blocking APIs. Workers support cancellation and exclusivity, while thread-originated UI updates use `call_from_thread()` or thread-safe messages ([workers](https://textual.textualize.io/guide/workers/)). | Supports Select, asyncio, Twisted, GLib, Tornado, and ZMQ event loops plus alarms and watched files; executor use varies by loop ([main loop](https://urwid.org/manual/mainloop.html)). | Uses asyncio natively and supports `Application.run_async()` inside an existing event loop ([asyncio](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/advanced_topics/asyncio.html)). |
|
|
||||||
| **Testability** | `App.run_test()` runs headlessly and returns a `Pilot` that can press keys, click, wait for queued messages, and run at specified terminal sizes ([testing](https://textual.textualize.io/guide/testing/)). | The public manuals reviewed do not expose a comparable application pilot; tests would be built around widgets, rendering, callbacks, and event-loop seams ([widgets](https://urwid.org/manual/widgets.html), [main loop](https://urwid.org/manual/mainloop.html)). | Official testing guidance supplies platform-specific pipe input, `DummyOutput`, and app sessions for asserting results or data changes, while warning against brittle stdout-byte assertions ([unit testing](https://python-prompt-toolkit.readthedocs.io/en/stable/pages/advanced_topics/unit_testing.html)). |
|
|
||||||
| **Direct dependency cost** | Five required distributions: `markdown-it-py[linkify]`, `mdit-py-plugins`, `platformdirs`, `pygments`, and `rich`; syntax-tree packages are optional-extra dependencies ([metadata](https://pypi.org/pypi/textual/json)). | Two required distributions: `wcwidth` and `typing-extensions`; curses, GLib, Tornado, Trio, Twisted, ZMQ, serial, and LCD integrations are optional extras ([metadata](https://pypi.org/pypi/urwid/json)). | One required distribution: `wcwidth` ([metadata](https://pypi.org/pypi/prompt-toolkit/json)). |
|
|
||||||
| **Python floor** | `>=3.9,<4.0` ([metadata](https://pypi.org/pypi/textual/json)). | `>=3.9.0` ([metadata](https://pypi.org/pypi/urwid/json)). | `>=3.10` ([metadata](https://pypi.org/pypi/prompt-toolkit/json)). |
|
|
||||||
| **Project health signal** | Five non-draft releases from 8.2.4 through 8.2.8 were published between 2026-04-19 and 2026-06-30 ([release API](https://api.github.com/repos/Textualize/textual/releases?per_page=5)). | Five releases from 4.0.9 through 4.0.13 were published between 2026-08-14 and 2026-08-25 ([release API](https://api.github.com/repos/urwid/urwid/releases?per_page=5)). | Release 3.0.53 was published 2026-07-26, following 3.0.52 on 2025-08-27 ([release API](https://api.github.com/repos/prompt-toolkit/python-prompt-toolkit/releases?per_page=5)). |
|
|
||||||
|
|
||||||
All three show recent release activity, so maintenance does not eliminate a candidate. Textual wins on integrated application-level capabilities rather than on release recency alone.
|
|
||||||
|
|
||||||
## Fit to the four views
|
|
||||||
|
|
||||||
- **Overview:** use Textual layout containers with `Digits`, labels, progress indicators, and compact sparklines for monitoring status, usage-adjusted theoretical lifespan, projection confidence, and recent activity ([layout](https://textual.textualize.io/guide/layout/), [gallery](https://textual.textualize.io/widget_gallery/)).
|
|
||||||
- **Usage History:** prototype a `DataTable` plus `Sparkline` first. Add `textual-plotext` only if user testing establishes a need for axes, multiple series, or richer plots, keeping plotting out of the core dependency decision until then ([DataTable](https://textual.textualize.io/widgets/data_table/), [Sparkline](https://textual.textualize.io/widgets/sparkline/), [`textual-plotext`](https://github.com/Textualize/textual-plotext)).
|
|
||||||
- **Drive Health:** use a table and status/progress widgets for NVMe attributes; wording must continue to distinguish endurance projection from a hardware-failure prediction ([gallery](https://textual.textualize.io/widget_gallery/), [parent-map destination and terminology](https://git.bongbetic.com/xavierk/Fenris/issues/1)).
|
|
||||||
- **Settings:** Textual's selection, boolean, and validated text-entry controls cover a keyboard-only form without inventing controls ([gallery](https://textual.textualize.io/widget_gallery/), [Input validation](https://textual.textualize.io/widgets/input/)).
|
|
||||||
|
|
||||||
The TUI should consume snapshots from the independent collector rather than run NVMe collection in its UI loop; the ticket's parent map makes that independent-collector architecture part of the destination ([parent map](https://git.bongbetic.com/xavierk/Fenris/issues/1)). Textual workers remain useful for non-blocking database/IPC reads and cancellable refreshes ([workers](https://textual.textualize.io/guide/workers/)).
|
|
||||||
|
|
||||||
## Prototype acceptance checks
|
|
||||||
|
|
||||||
The later prototype should verify:
|
|
||||||
|
|
||||||
1. Complete keyboard-only operation and predictable focus order across all four views.
|
|
||||||
2. Reflow or deliberate scrolling at 80×24 and a representative wide terminal; Textual's test runner accepts explicit terminal sizes ([testing](https://textual.textualize.io/guide/testing/)).
|
|
||||||
3. Smooth refresh while database/IPC reads are delayed, plus clean cancellation and shutdown ([workers](https://textual.textualize.io/guide/workers/)).
|
|
||||||
4. Headless tests for view switching, settings validation, stale/error states, and collector disappearance ([testing](https://textual.textualize.io/guide/testing/), [Input](https://textual.textualize.io/widgets/input/)).
|
|
||||||
5. Readable behavior in the actual Linux terminal matrix, including SSH and tmux/screen if those are required; framework platform labels alone do not establish every terminal's behavior.
|
|
||||||
6. Usage History performance with realistic observation volume and gaps.
|
|
||||||
7. A table/text equivalent for every chart so observation gaps and projection confidence are not encoded by color or graphics alone.
|
|
||||||
8. Installation with Fenris's chosen packaging mechanism and the agreed Python floor.
|
|
||||||
|
|
||||||
## Newly surfaced questions
|
|
||||||
|
|
||||||
Do not open follow-up tickets during this research resolution; carry these into the next Wayfinder session:
|
|
||||||
|
|
||||||
- May the redesign raise Fenris's minimum Python version from 3.7 to 3.9?
|
|
||||||
- What is the minimum supported terminal contract: modern color terminals only, or also Linux virtual console, monochrome/16-color, SSH, tmux, and screen?
|
|
||||||
- Is a table plus sparkline sufficient for Usage History, or are axes, multiple series, zoom, and interval selection required?
|
|
||||||
- Should richer plotting be an optional install extra?
|
|
||||||
- What is the small-terminal policy: reflow, hide secondary content, scroll, or reject below a documented size?
|
|
||||||
- Which accessibility targets are required, including screen readers, reduced color, high contrast, and reduced animation?
|
|
||||||
- Does the TUI read observation storage directly or use an IPC snapshot API from the collector?
|
|
||||||
- Which settings are user-owned versus system/collector-owned, and which changes require privilege escalation?
|
|
||||||
- Are semantic interaction assertions sufficient, or are deterministic visual snapshots also required?
|
|
||||||
Reference in New Issue
Block a user