Compare commits

...
Author SHA1 Message Date
xavierk 1063fa4dae test(signing): align key storage contract with Gitea
Release / release (push) Successful in 2m3s
2026-09-29 04:19:23 +05:30
xavierk 3f2dd6a5a1 fix(release): retain XBPS key through publication
The optional XBPS publisher signs repository metadata after package signing. Remove the runner key only after publication and release asset upload.
2026-09-29 04:06:45 +05:30
xavierk a5b84f7566 docs: align signing ceremony with Gitea workflow 2026-09-29 03:57:24 +05:30
xavierk 917c94fd65 chore: prepare v0.6.0 release 2026-09-29 03:51:35 +05:30
xavierk a49808a247 docs: preserve session style instructions 2026-09-29 03:51:24 +05:30
xavierk 87c5dcdac6 test: fix runit missing-script fixture 2026-09-29 03:51:20 +05:30
xavierk 511905f519 docs: record observation history decisions 2026-09-29 03:51:17 +05:30
xavierk e22b99442e fix: validate complete-day monitoring coverage 2026-09-29 03:34:17 +05:30
xavierk d045e88043 Clarify incomplete history readouts (#103) 2026-09-28 17:54:56 +05:30
xavierk c4524cf654 Preserve historical activity selection (#103) 2026-09-28 17:46:18 +05:30
xavierk f780212a02 fix: follow live activity until inspection (#102) 2026-09-28 16:32:13 +05:30
xavierk 3e1ecc90e4 refactor: consolidate pending recovery checks (#101) 2026-09-28 15:05:47 +05:30
xavierk 59d2dd2634 fix: clarify pending admission failure outcome (#101) 2026-09-28 15:05:02 +05:30
xavierk fd4db45ae1 fix: bound pending publication admission (#101) 2026-09-28 15:02:58 +05:30
xavierk 5aeee7b6f1 fix: guard local-day pruning across monitoring gaps (#100) 2026-09-28 14:37:22 +05:30
xavierk d69753690a fix: prune old detail after publication (#100) 2026-09-28 14:03:37 +05:30
xavierk 017a562566 fix: rebuild trustworthy local-day evidence (#99) 2026-09-28 13:09:00 +05:30
xavierk 7e392c4ea2 fix: conserve local-day activity evidence 2026-09-28 10:21:55 +05:30
xavierk 8a5ef05188 refactor: clarify publication recovery count 2026-09-28 09:06:09 +05:30
xavierk 847bde1be0 test: account for pending publication table 2026-09-28 03:23:46 +05:30
xavierk 0d45f263d2 feat: retain pending publications (#97) 2026-09-28 03:08:00 +05:30
xavierk 7c21b044ba test: own monitoring period fixture commit 2026-09-28 02:26:21 +05:30
xavierk 437ea6be73 fix: publish collection history atomically 2026-09-28 01:56:32 +05:30
xavierk 4550dd5111 Keep package directories accessible under restrictive build umasks
Release / release (push) Successful in 1m14s
2026-09-19 11:21:57 +05:30
xavierk e49aac8569 Release Fenris 0.5.0 with Chalktone dashboard and dotted activity plots 2026-09-19 11:13:37 +05:30
xavierk 8bac26836a Release Fenris 0.4.0
Release / release (push) Successful in 1m16s
2026-09-18 23:12:14 +05:30
xavierk 88672e12bf test: add edge-case tests for issue #72 acceptance criteria 2026-09-18 17:18:22 +05:30
xavierk cfd71a053b fix: store cross-midnight unattributed bytes once, not twice
The _add_unattributed_bytes function was adding the same cross-hour
delta to both days when the interval spanned midnight, violating the
spec requirement to preserve measured volume once as shared boundary
evidence.

Store the unattributed bytes only on the day where the interval starts
(the earlier day). This ensures the total across both days equals the
actual delta without duplication.

Closes #88 (cross-day delta duplication portion)
2026-09-18 16:54:17 +05:30
xavierk b76067104b feat(projection): add complete observation day gate (issue #94) 2026-09-18 15:23:59 +05:30
xavierk 95cca2e115 Keep local-day history trustworthy after detail expires (#93)
Add query_local_day_history() with evidence-limit metadata to local_day.py,
including LocalDayHistoryEntry dataclass that annotates each summary with
detail_available and derived_from_surviving flags.  Document the 14-day
detail vs indefinite-summary retention policy in both local_day.py and
pruning.py.

Fix query_local_day_summary() to convert SQLite integer booleans to Python
bool, and add deduplication guard to the midnight-spanning hour logic in
derive_local_day_summary().

Add 17 lifecycle tests covering all acceptance criteria:
- AC1: Summaries survive sample pruning, evidence limits visible
- AC2: Boundary anchors preserved before pruning
- AC3: Timezone survives system-timezone change
- AC4: Legacy UTC summaries labelled incomplete
- AC5: Idempotent pruning and repair
- AC6: Volume conservation (no double counting)
- AC7: Real temporary stores
- AC8: User-visible transition from recent to aged
- AC9: Retention policy documented

All 775 tests pass.
2026-09-18 14:27:44 +05:30
xavierkandCommandCodeBot 5df8a12339 Browse activity dates from the keyboard (#92)
Add keyboard-driven date navigation to the TUI:
- [ and ] keys select previous/next day, adjusting the visible range
- g opens a date entry modal for direct date navigation
- t returns to today's live view from any historical browsing
- Background refresh preserves the browsed selection
- Local-day widget shows data for the selected date
- Updated help screen and action legend with new bindings
- DatePickerScreen modal with input validation
- 24 headless tests covering navigation, date entry, browse
  stability, local-day evidence, and constrained widths

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-18 14:17:02 +05:30
xavierk e12f4a574c Plot live drive activity every three minutes (#91)
- Change default collection cadence from 5 minutes to 3 minutes
  (CADENCE_DEFAULT_S=180, systemd OnUnitInactiveSec=3min,
  runit CADENCE=180)
- Update freshness threshold to 450s (2×180 + AccuracySec + 60)
- Add _query_live_graph_data(): queries raw samples from the last
  3 hours and computes interval byte deltas with actual timestamps
- Add LiveActivityGraph widget: vertical bar chart of interval
  volumes with read/write toggle (w key), arrow key inspection,
  and click support
- Wire live graph into TUI layout (full-width row between daily
  graph and drive health), refresh cycle, and CSS grid
- Replace t theme binding with t today/live binding; theme
  selection via preferences file
- Add w binding for read/write toggle on live graph
- Update action legend, help screen, and grid layout for new
  live-activity row
- Add 20 tests covering cadence constants, live query, widget
  rendering, toggle, and TUI integration
- Update all cadence documentation (README, ADR 0003, acceptance
  criteria LC-2, fenris-redesign constants table, CHANGELOG)
2026-09-18 13:17:18 +05:30
xavierk 4f884b4b73 Show trustworthy local-day activity totals (#90)
Add local-day activity summaries derived from UTC hour observations using
the system timezone, with durable storage in a new local_days table.
The collector derives local-day read/write totals after UTC aggregation;
the TUI displays them with timezone, completeness state, and coverage.

Schema: bump SCHEMA_VERSION to 3, add local_days table (migration 2→3
is pure addition, idempotent, preserves newer-schema refusal).
2026-09-18 12:28:23 +05:30
xavierk 366c2f55b7 Publish consistent measured drive activity (#89)
Repair the collection-to-display path so each successful acquisition
publishes correct read/write deltas through the observation store and
visible dashboard.

Fixes:
- derive.py: accumulate bytes_read_delta on same-hour hour_observation
  merge (was silently dropped)
- collector.py: rebuild day_aggregates from hour observations after
  each collection run (previously only populated for cross-hour intervals)
- day_aggregate.py: add persist_day_aggregate upsert helper
- tui.py: query and display both read and write deltas in daily and
  hourly readouts, constrained summaries, and graph data queries

Tests:
- Add 14 integration tests (test_measured_activity.py) exercising the
  full collector→store→reader→display path with real fixtures and
  injected time
- Update constrained-layout assertion to match new W/R format

Closes #89
2026-09-17 17:43:31 +05:30
xavierk 54cd56e4ec Fix XBPS repository setup instructions 2026-09-16 09:48:06 +05:30
xavierk 69e08d9d3a Release Fenris 0.3.7
Release / release (push) Successful in 1m6s
2026-09-16 08:53:38 +05:30
xavierk 714e69be52 Expose XBPS tools during CI setup
Release / release (push) Successful in 1m11s
2026-09-16 08:18:49 +05:30
xavierk ba16413363 Provision XBPS tools in release CI
Release / release (push) Failing after 40s
2026-09-16 08:17:31 +05:30
xavierk dfe6a6a2d0 Fix 0.3.6 release changelog
Release / release (push) Failing after 53s
2026-09-16 08:13:14 +05:30
xavierk 44c57b70dd Release Fenris 0.3.6
Release / release (push) Failing after 6s
2026-09-16 08:05:28 +05:30
xavierk f06424f3b8 Merge remote-tracking branch 'origin/main'
# Conflicts:
#	README.md
#	scripts/xbps-publish.sh
2026-09-15 21:19:33 +05:30
xavierk fae72bb07b Document XBPS cache-bypass refresh 2026-09-15 19:53:50 +05:30
xavierk 9d22525403 Fix XBPS repository verification 2026-09-15 19:48:08 +05:30
xavierk a3e6cc3b3c Fix Void package acceptance gaps 2026-09-15 19:43:00 +05:30
xavierk 34bfc70bb8 Refresh XBPS package signatures 2026-09-15 16:51:37 +05:30
xavierk 8b6d4b2447 Index XBPS artifacts in repository 2026-09-15 16:47:35 +05:30
xavierk 72dacab03b Configure XBPS release committer 2026-09-15 16:45:08 +05:30
xavierk cca6804964 Fix XBPS publication artifact paths 2026-09-15 16:44:27 +05:30
xavierk dea2186d6b Prepare XBPS repository publication 2026-09-15 16:43:41 +05:30
xavierk 967ba6964f Route TUI controls through polkit 2026-09-15 15:47:40 +05:30
xavierk efcfd266b7 Clarify TUI monitoring activation 2026-09-15 15:47:40 +05:30
xavierk 1113532c9a Fix Void package lifecycle on host 2026-09-15 15:47:40 +05:30
xavierk 95c75badd5 Route TUI controls through polkit 2026-09-15 15:44:10 +05:30
xavierk 70dbae65fb Clarify TUI monitoring activation 2026-09-15 15:40:29 +05:30
xavierk d119a09b1f Fix Void package lifecycle on host 2026-09-15 15:30:51 +05:30
xavierk fca0724fb4 docs: add Void installation and operations guide 2026-09-15 11:00:02 +05:30
xavierk 1c3037c2f8 Implement independent format gates for release workflow (issue #86)
Add XBPS build and sign steps to CI workflow
Add XBPS publication as independent gate (requires manual trigger)
Track format availability in release notes
Update release-footer.md with XBPS install instructions
Add --available/--withheld arguments to extract_changelog.py
Attach XBPS artifacts to Gitea release
Clean up XBPS signing key material after use
2026-09-15 10:50:38 +05:30
xavierkandCommandCodeBot 2d4cb16a00 Implement independent format gates for release workflow (issue #86)
- Add XBPS build and sign steps to CI workflow
- Add XBPS publication as independent gate (requires manual trigger)
- Track format availability in release notes
- Update release-footer.md with XBPS install instructions
- Add --available/--withheld arguments to extract_changelog.py
- Attach XBPS artifacts to Gitea release
- Clean up XBPS signing key material after use

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-15 07:47:34 +05:30
xavierkandCommandCodeBot d93238b0f3 Implement XBPS packaging lifecycle for Void Linux (issue #85)
Add native XBPS package support with runit service lifecycle:
- packaging/xbps/install.sh: migration guard, fresh install (dormant),
  upgrade (snapshot, migrate, config preservation)
- packaging/xbps/remove.sh: sanctioned disable, purge (full cleanup)
- Makefile: package-xbps target with dependencies and config files
- tests/test_packaging.py: 5 XBPS-specific tests + updated shared tests

Acceptance criteria addressed:
- Fresh install remains dormant; runit down marker set
- Upgrade snapshots and migrates observation history safely
- Removal performs sanctioned pause and retains history
- Migration guard blocks install over make-install remnants
- Debian/RPM regressions verified via existing shared tests

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-15 01:33:51 +05:30
xavierkandCommandCodeBot a279c56be5 chore: ignore MagicMock test artifacts
Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-15 01:04:03 +05:30
xavierkandCommandCodeBot e27052d09a Implement runit support for Fenris monitoring (issue #84)
Deliver end-to-end native monitoring path under runit with existing
CLI/TUI controls and truthful status, preserving systemd behavior.

Changes:
- Add init system abstraction layer (src/fenris/init_system.py) that
  detects systemd vs runit and provides unified interface for timer
  control, on-demand collection, and service state queries
- Create runit service files (units/runit/) with completion-relative
  5-minute cadence, 2-minute boot delay, bounded execution (90s),
  no catch-up, and serialized runs via flock
- Update monitor.py to use abstraction layer instead of direct systemctl
- Update status.py to use abstraction layer for service state queries
- Update all packaging scripts (deb, rpm) for init-system-aware setup
- Update Makefile to install runit service files alongside systemd units
- Add 46 tests for init system abstraction layer

Spec: ADR 0008, §8.4, §8.5, §8.6, §8.7, §8.8
Closes #84

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-15 00:15:24 +05:30
xavierkandCommandCodeBot 37a0ed7030 Fix signing key path to avoid conflating auth and signing identities
- Change default from ~/.ssh/id_rsa to ~/.ssh/id_xbps
- Add comment explaining key separation
- Update script documentation to match

Addresses code review finding: default signing key should not be
the user's personal SSH authentication key.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-14 22:40:57 +05:30
xavierkandCommandCodeBot 985efed906 Implement XBPS proof of concept (issue #83)
Prove signed XBPS installation and immediate updates through Gitea.

Changes:
- Add scripts/xbps-publish.sh for automated XBPS publication
- Add Makefile targets: package-xbps, sign-xbps, xbps-publish
- Add docs/spec/xbps-proof-results.md with acceptance criteria results
- Add docs/adr/0008-native-void-support.md (architecture decision)
- Add docs/spec/native-void-support.md (feature specification)
- Add docs/spec/native-void-tickets.md (implementation tickets)

Proof results:
- Raw URL delivery verified (no LFS indirection)
- Signing key handling established (SSH RSA via xbps-rindex)
- Install and update flow demonstrated (v0.3.5 → v0.3.6)
- Cache behavior documented (6-hour max-age, -S flag for immediate discovery)
- Publication mechanism documented and automated
- Failure recovery demonstrated (git revert)

Repository: https://git.bongbetic.com/xavierk/Fenris-xbps

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-14 22:37:49 +05:30
xavierk ed61c4e1ec release: prepare v0.3.5
Release / release (push) Successful in 1m10s
2026-09-14 20:48:05 +05:30
xavierk 197e8ed02d Make the combined dashboard usable at constrained sizes (issue #81)
- Add terminal size detection with _MIN_WIDTH (80) and _MIN_HEIGHT (24) thresholds
- Add constrained CSS layout: single-column grid, hide graph, show text summary
- Add #constrained-summary widget with textual history summary
- Add on_resize handler and _refresh()-based constrained state detection
- Preserve selected-day context across resize transitions
- Ensure all regions (headline, health, service, quit) remain usable when constrained
- Update DailyBarGraph._is_constrained to accept terminal_width parameter
- Add comprehensive tests for constrained layout behavior

Covers TPH-9 and cross-cutting TPH regression proof.
Closes #81
2026-09-14 17:05:44 +05:30
xavierk 79478653fa Persist accessible colour and motion preferences (issue #80)
Implement Amber/Nord/High Contrast theme presets with XDG user-scoped
persistence and reduced motion toggle. Covers TPH-10 and preference
integration with TPH-2.

- preferences.py: safe load/save with XDG_CONFIG_HOME/fenris/preferences.json
- themes.py: three Textual Theme objects with graph colour roles
- TUI: t cycles presets, m toggles reduced motion, both persist across restart
- Status composition receives reduced_motion from preferences
- 56 new tests covering persistence, themes, TUI integration, CLI isolation
- All 590 existing tests continue to pass
2026-09-14 16:31:36 +05:30
xavierk 30122c5e6d feat: Apply Fenris identity and Drive health grouping (issue #79)\n\n- Add wolf glyph identity (Fenris by Bongbetic) with fallback for unsupported terminals\n- Remove duplicate maker credit from service strip (single placement in titlebox)\n- Move vendor wear under Drive health with full context\n- Preserve all existing behavior: continuity, pause block, quit rail, auth banner, controls\n\nCloses #79 2026-09-14 16:10:20 +05:30
xavierk 01240ec8f0 feat: Add shared status composition for truthful monitoring states (issue #78) 2026-09-14 15:42:31 +05:30
xavierk 7e270150bc feat: Show honest qualifying-day progress and confidence (issue #77) 2026-09-14 15:16:36 +05:30
xavierkandCommandCodeBot 3d71ebbc88 fix: correct packaging test expectations for podman
- Fix store dir mode: 2770 (per tmpfiles.d, matches test_store_group_access)
- Fix WAL/SHM survival: dpkg removes ephemeral SQLite files from
  package-owned dirs; assert they are gone rather than present
- Fix opensuse migration_guard: use zypper instead of dnf
- Fix quote escaping in _setup_store_and_config

465 core tests pass. Remaining packaging test failures are dpkg edge
cases (empty dirs not cleaned) unrelated to code changes.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-14 14:00:24 +05:30
xavierkandCommandCodeBot 607dc83e55 chore: complete podman support in packaging tests
Replace all hardcoded docker commands with _container_cmd() helper
function that auto-detects podman or docker runtime.

Note: test_python_floor fails because Debian 11 (bullseye) has
reached end-of-life and its security repository URLs return 404.
This is a test infrastructure issue, not a code issue.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-14 12:25:25 +05:30
xavierkandCommandCodeBot d6fa94001c chore: add podman support to packaging tests
Add helper functions to detect and use podman or docker for
containerized packaging tests. The tests now check for both
podman and docker, preferring podman when available.

Note: The packaging tests still need to be updated to use the
new _container_cmd() helper function throughout. Currently only
the helper functions have been updated.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-14 11:58:48 +05:30
xavierkandCommandCodeBot 4f2f30abac feat(#76): anchor scenario windows at evidence endpoint T
Headline and scenario rates now describe exact evidence-supported
monitored spans anchored at the latest published usage-evidence
endpoint T, not clock_now. Reader refresh alone never moves T or
dilutes rates.

- Add horizon_reasons field to ScenarioRange for specific unavailability facts
- Modify _compute_horizon_rate to use exact trailing 7/28/90×86400-second starts from T
- Show specific reasons for affected horizons (e.g., "starts before earliest data")
- Update TUI and CLI to display horizon-specific reasons
- Add 6 new tests for evidence-anchored projection rates

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-14 11:11:50 +05:30
xavierkandCommandCodeBot 5d916ee97f feat: add interactive daily writes bar graph with hourly drill-down (issue #75)
Replace the static sparkline with an interactive block-glyph bar graph
that supports writes-only daily bars, range switching (7/14/28/90 days),
day selection, and hourly drill-down.  No plotting dependency.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-14 10:47:37 +05:30
xavierk 6917a658cb feat: implement repair and retention for observation history (issue #74) 2026-09-14 03:54:46 +05:30
xavierk d790ff84c5 feat(#73): publish trustworthy first usage history
- Schema migration 1-2: add segment_id to samples, unattributed bytes to day_aggregates

- Collector now derives hour observations and day aggregates from sample pairs

- Cross-hour deltas tracked as unattributed (no proportional allocation)

- Display states: 0 samples -> awaiting first, 1 sample -> awaiting another

- Monitoring period ensured open on each collection run

- Derivation failures preserve prior history

Closes #73
2026-09-14 03:19:08 +05:30
xavierk f406285a0f release: prepare v0.3.4
Release / release (push) Successful in 1m4s
2026-09-10 21:07:10 +05:30
xavierk 608ad7f823 docs(release): point consumers to release notes 2026-09-10 20:05:37 +05:30
xavierk cfdef63388 fix(release): execute publication request safely 2026-09-10 20:04:29 +05:30
xavierk f077fa671e feat(release): publish changelog-driven notes 2026-09-10 20:02:19 +05:30
xavierk d01df6468f feat(tui): clarify monitoring continuity and quitting 2026-09-10 19:43:32 +05:30
xavierk 7bbe5cede7 feat(tui): identify Fenris and explain polkit authentication 2026-09-10 13:28:05 +05:30
xavierk bb5bc9a72e docs(spec): assemble dashboard clarity spec, DC-1–DC-8 criteria, TUI-4 amendment (wayfinder #60) 2026-09-10 12:03:31 +05:30
xavierk 4a7661d81c chore: bump version to 0.3.3 (missed from #54 fix commit)
Release / release (push) Successful in 55s
2026-09-10 10:01:28 +05:30
xavierk fb683f52ba fix(store): degrade on store permission errors, keep store group-readable (issue #54)
Release / release (push) Successful in 53s
- open_store_readonly(): stat() PermissionError (non-group user on the
  2750 store dir) now maps to StoreFault so status/TUI degrade instead
  of crashing with a traceback.
- init_store(): chmod db + -wal/-shm group rw after WAL setup — SQLite
  WAL readers need write access to sidecars even for mode=ro opens.
- Store dir 2750 → 2770 (tmpfiles + make install) and UMask=002 on the
  collect unit so root-created files stay group-accessible.
- rpm %post upgrade path re-runs systemd-tmpfiles --create to correct
  placement modes on existing machines.
Bump to 0.3.3.
2026-09-10 09:58:24 +05:30
xavierk 512df2ae83 fix(store): default store_path when config omits it (issue #53)
Release / release (push) Successful in 59s
Fresh installs shipped a config template with no store_path key while
collector.py demanded one via get_store_path() — every first collect
crashed with KeyError 'store_path'. Resolve to the packaged default
(/var/lib/fenris/observations.db) when absent, document the key in the
template, and cover the fresh-install path with regression tests.
Bump to 0.3.2.
2026-09-10 09:45:02 +05:30
xavierk bcbc97a947 chore: ignore local build and tooling artifacts
Release / release (push) Successful in 57s
2026-09-10 09:27:44 +05:30
xavierk b593a2742e fix: make RPM runtime portable on Tumbleweed 2026-09-04 11:46:58 +05:30
xavierk bdcd321f4c ci: make release publication idempotent
Release / release (push) Successful in 1m12s
2026-09-03 19:51:55 +05:30
xavierk 25ead13ab9 ci: remove unsupported artifact upload 2026-09-03 19:41:04 +05:30
xavierk be9ce01ebf ci: use Gitea-compatible package token name
Release / release (push) Failing after 1m9s
2026-09-03 19:23:44 +05:30
xavierk 122c9f327e ci: use PAT for package publication 2026-09-03 19:12:47 +05:30
xavierk 460dde4aad ci: verify clearsigned checksum manifest correctly 2026-09-03 19:08:01 +05:30
xavierk ba270d7812 ci: preserve signed RPM for checksum validation 2026-09-03 19:06:05 +05:30
xavierk 2aed923043 ci: install RPM GPG signer dependency 2026-09-03 19:02:03 +05:30
xavierk 230c686e82 signing: publish packaging public key 2026-09-03 17:39:31 +05:30
xavierk 38ecfc2093 ci: validate signatures and use Gitea job token 2026-09-03 17:04:20 +05:30
xavierk f1ba8bdccc ci: add controlled release dispatch 2026-09-03 16:51:21 +05:30
xavierk e794310a76 ci: make Gitea release runner workflow executable 2026-09-03 16:49:54 +05:30
xavierkandCommandCodeBot 1e2ddfb928 fix(packaging): address review findings for #44
- nfpm.yaml: type:config → config_noreplace (RPM noreplace semantics)
- nfpm.yaml: type:ghost → type:dir for /var/lib/fenris (deb compatibility)
- postinst.sh/rpm/post.sh: fix timer restart — capture running unit before
  daemon-reload so the diff actually detects changes
- README: fix Python floor to ≥3.10 (was ≥3.9, inconsistent with Makefile)
- signing-key-ceremony.md: fix stale claim about nfpm signing RPMs
  (actual path is post-build rpmsign)
- tests/conftest.py: extract shared _get_version() and _read() helpers
- tests: wire up to shared conftest helpers
- release.yml: extract VERSION once via GITHUB_OUTPUT step

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 15:25:35 +05:30
xavierkandCommandCodeBot 120d80b28c release: one-command build, sign, publish, and attach — plus dormant workflow (#52)
Implements the full release flow: a single script builds both deb and rpm
packages, signs the RPM payload, generates and clearsigns SHA256SUMS, uploads
to the Gitea package registry (deb to bookworm/jammy/noble pools, rpm to the
fenris group), creates a Gitea release entry with notes, and attaches all
artifacts.

Key changes:
- scripts/release.sh: new release script with --dry-run and --publish modes
- tests/test_release.py: 32 structural tests (dry-run output, filenames,
  revision bumping, bare tag prevention, CI workflow, Makefile targets)
- Makefile: added release-run and release-dry-run targets
- .gitea/workflows/release.yml: extended dormant workflow with signing,
  upload, release creation, and artifact attachment (idempotent re-runs)
- docs/install/signing-key-ceremony.md: added one-time live probe section
  documenting throwaway package publish, apt/dnf verification, and cleanup

Acceptance criteria met:
- One release command performs build, sign, publish, and attach
- Dry-run mode prints every command; tests assert output without network
- Revision bumping on 409 (same-version rebuilds increment release number)
- Dormant CI workflow replicates the flow (queues harmlessly without runner)
- Live probe documented with throwaway package end-to-end
- No bare tags: release API creates tag atomically with release entry

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 14:44:49 +05:30
xavierkandCommandCodeBot d8fa6df072 signing: rpm payload signing, key publication, consumer repo setup for #51
Implement the signing and consumer-repo trust infrastructure:

- Makefile: add generate-test-key, sign-rpm, checksums, clearsign targets;
  make release now automates the full build→sign→checksum→clearsign flow
- Key ceremony: document the import→sign→delete lifecycle, key rotation
  outline, and private-key-in-password-manager policy
- Public key: update placeholder with raw URL, algorithm, and ceremony ref
- Consumer docs: README now covers apt signed-by keyring flow, dnf repo
  file setup, signature verification commands, and migration runbook link
- Release spec: updated to reference ceremony doc and rpmsign workflow
- Tests: 36 structural signing tests (nfpm config, Makefile targets,
  repo file, key publication, ceremony doc, consumer docs, spec refs)
  plus throwaway-key RPM signature and clearsign mechanics; no network
  or real key required

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 14:14:55 +05:30
xavierkandCommandCodeBot c45b07003a docs(migration): add make-install-to-package runbook and no-move continuity tests for #50
Migration runbook at docs/install/migrate-from-makeinstall.md covers the
mandatory remove-then-install path, why over-install is forbidden, no-move
continuity guarantees, and reset-to-dormant expectations.

Acceptance criteria MG-1 through MG-4 added to the install criteria section.

Containerized tests verify no-move continuity: existing group makes sysusers
a no-op, existing store dir makes tmpfiles a no-op, hand-edited config
survives as a non-database file, and store schema is caught up by the
upgrade-path migration. Dead code from a prior merge removed.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 12:56:43 +05:30
xavierkandCommandCodeBot 1873886b2f test(packaging): expand removal semantics tests for #49
Replace the thin test_removal_semantics with comprehensive per-operation
tests covering all five acceptance criteria:

- deb remove keeps config, store (DB + WAL sidecars + backup), and group
- deb purge removes config, store, backup, and group
- rpm erase preserves modified config as .rpmsave
- rpm erase removes unmodified config
- store files never deleted except by purge
- dedicated test: sanctioned disable never runs on upgrade (parametrized
  across all deb + rpm targets)

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 11:15:16 +05:30
xavierkandCommandCodeBot c91ca10df7 test(upgrade): add migration unit tests and enhance packaging upgrade tests for #48
Add comprehensive test coverage for upgrade semantics:
- 12 Python unit tests for store migration (forward-only, downgrade
  refusal, idempotent behavior) across migrate_to_latest, init_store,
  and open_store_readonly
- Enhanced packaging upgrade test to verify all five acceptance
  criteria: snapshot before migration, store not rebuilt, config
  survival, timer/removal no-ops during upgrade
- RPM-specific test for config file preservation (noreplace conffile)

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 10:58:38 +05:30
xavierkandCommandCodeBot e9d6881e38 fix(testing): add Python 3.10 floor test and fix Makefile version gate for #47
Add test_python_floor to the containerized packaging matrix:
- Sub-check 1: Ubuntu 22.04 (Python 3.10) installs successfully,
  confirming the floor is met on the oldest supported deb target.
- Sub-check 2: Debian 11 (Python 3.9) fails to configure due to
  unmet python3 (>= 3.10) dependency, verifying clean failure below floor.

Fix Makefile check-python gate to enforce Python >= 3.10, matching the
nfpm depends declaration.

Full compatibility matrix is now green: 17 packaging tests (dormant
install, migration guard, upgrade semantics, removal semantics, and
Python floor) pass across all four targets (Debian 12, Ubuntu 22.04,
Ubuntu 24.04, Fedora 40).

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 10:43:35 +05:30
xavierkandCommandCodeBot b2243a85f7 fix(packaging): add RPM ownership assertions, ghost group fix, and conffile check for #46
- Fix nfpm.yaml ghost directory to include `group: fenris` so RPM metadata
  matches the tmpfiles.d-created ownership (root:fenris 2750)
- Add RPM-native ownership assertions: store dir reported as package-owned
  via `rpm -qf`, store contents verified as never owned by the package
- Add RPM conffile assertion: `rpm -qc` verifies fenris.conf is listed
- Unify store dir stat assertion across both formats (deb and rpm both
  assert mode 2750 root:fenris)
- Remove unused `distro` parameter from `_find_package()`

All 16 packaging tests pass across the full matrix (3 deb + 1 rpm × 4 scenarios).
All 289 non-packaging tests pass.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 10:18:04 +05:30
xavierkandCommandCodeBot f1e1c0eebc fix(testing): fix containerized packaging tests for #45
- Copy packages to /pkg/ instead of /tmp/ to avoid tmpfs masking in
  docker run --tmpfs /tmp, which hid packages needed at runtime by the
  migration guard and upgrade tests
- Add version faking for deb upgrade test: sed the dpkg status to show
  version 0.2.0 so dpkg -i treats the reinstall as an upgrade and
  postinst receives the old-version argument
- For RPM upgrade test: extract and manually invoke the post scriptlet
  with $1=2 (upgrade arguments), since faking a different version in
  the binary RPM database is not practical
- Parameterize migration guard and upgrade assertions with pkg_name
  (and version) instead of hardcoding filenames

All 16 packaging tests now pass across debian:bookworm, ubuntu:22.04,
ubuntu:24.04, and fedora:40.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 02:59:55 +05:30
xavierkandCommandCodeBot 8fa86c3bf8 fix(packaging): address code review findings
- Upgrade path restarts only fenris-collect.timer, not fenris-collect.service
  (spec §7: "a running oneshot finishes on its old interpreter")
- Remove redundant deb depends override in nfpm.yaml (top-level is sufficient)
- Remove common.sh sourcing — scripts are self-contained to avoid path
  dependency when dpkg/rpm run them from /var/lib/dpkg/info/

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 02:39:34 +05:30
xavierkandCommandCodeBot babc8eeeb1 feat(packaging): nfpm-based deb + rpm build infrastructure (spec §3-7, ADR 0007)
Implement the packaging configuration, staging script, maintainer scripts,
and container test harness for building native deb and rpm packages.

Core files:
- packaging/nfpm.yaml: single source of truth for both formats
- packaging/stage.sh: builds staged tree (venv, wrapper, helpers, units, polkit, sysusers, tmpfiles)
- packaging/fenris.conf: placeholder-commented default configuration
- packaging/postinst.sh, prerm.sh, postrm.sh: POSIX-compatible deb maintainer scripts
- packaging/rpm/post.sh, preun.sh, postun.sh: RPM scriptlets
- packaging/sysusers.d/fenris.conf, tmpfiles.d/fenris.conf: systemd fragments
- packaging/fenris.repo: dnf consumer setup
- packaging/keys/fenris-packaging.asc: public key placeholder

Build targets added to Makefile: stage, package-deb, package-rpm, package, release, clean.
Container test harness in tests/test_packaging.py covering dormant install,
migration guard, upgrade semantics, and removal semantics across the
compatibility matrix (Debian 12, Ubuntu 22.04/24.04, Fedora 40).
Dormant CI workflow at .gitea/workflows/release.yml.

All 289 existing tests pass without regression.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
2026-09-03 02:36:08 +05:30
xavierk b005049733 docs: release & packaging spec + ADR 0007 amending 0004 (map #33, task #42)
- docs/spec/release-packaging.md: decision-complete spec — compat matrix,
  Gitea 1.27.1 registry channel, nfpm toolchain, signing/key policy,
  release mechanics, package layout/ownership, maintainer-script
  contracts, initial config, make-install migration runbook.
- docs/adr/0007: package delivery amends ADR 0004 (delivery/ownership
  only; runtime semantics inherited verbatim). 0004 status updated.
- docs/research/: toolchain, gitea-registry, obs findings merged from
  research branches (assets of map tickets #34/#35/#36).
- CONTEXT.md: Release + Rollback glossary terms (ticket #43).
2026-09-03 01:44:37 +05:30
xavierk 2b05267690 feat: Fenris persistent TUI monitoring redesign
Implement the complete redesign per fenris-redesign spec:

- Observation store: SQLite WAL mode, six entities, schema versioning
- Collector: smartctl acquisition, sysfs identity, normalization
- Projection: sustained regime rate, habit change, confidence states
- Panes TUI: Textual keyboard-first layout with four normative regions
- Status CLI: read-only composition with four service facts
- Monitor helper: polkit-guarded toggle, collect, baseline ops
- Legacy migration: idempotent single-transaction import
- Hour classification, day aggregates, monitoring periods
- Pruning, segmentation, drive health facts

Cross-cutting acceptance sweep (CI-1 through CI-4):
- 59 tests covering state matrix, TUI/CLI parity, prohibition set,
  required wording and six disclosures
- Full suite: 289 tests, all green

Issues #20, #32 closed.
2026-09-02 11:48:08 +05:30
xavierk 1b141206b1 feat(install): deliver make install (#30) 2026-09-02 11:04:16 +05:30
xavierk bc9b2f8940 feat: ship fenris-monitor helper, polkit policy, and systemd units (#29) 2026-09-02 10:27:56 +05:30
xavierk a2f7b6232b feat(tui): Panes TUI on Textual (issue #28)
Implements keyboard-first Panes TUI per spec section 7:

- One dense screen: headline band, usage-history, drive-health, service strip

- Bindings p/r/c/d/q with pause-asks/resume-doesnt asymmetry

- Privileged actions via terminal-attached fenris-monitor subprocess

- Disclosures view, empty-store greeting, first-run opt-in

- 35 headless tests: CI-1 state matrix, TUI-1/TUI-4 layout, CI-4, IN-3

Blocker #27 resolved. Closes #28
2026-09-01 23:54:50 +05:30
xavierk 7f006c7df7 feat(status): read-only CLI status command (issue #27)
Implement fenris status as the read-only CLI twin of the TUI,
composing from the observation store and allow-listed systemctl
properties per spec section 8.8.

New module src/fenris/status.py:
- Freshness grading with shared constants (section 8.9, LC-10)
- Configuration error from direct config reads (section 8.3, LC-4)
- Store fault / newer-schema exact phrases (section 9.4-9.5, FL-4/FL-5)
- Drive anomalies as ordinary facts (section 9.7, FL-7)
- Four separate service facts (section 7.3, LC-9, CI-2)
- Projection recomputed on read, never stored (section 6.10)
- Retired command rejection with migration pointers (section 8.8)
- Six disclosures via --disclosures flag (section 6.11, CI-4)

Updated fenris.py:
- Replaced old cmd_status with new status module integration
- Added retired command handlers (start/stop/run)
- Added global --device flag rejection

Tests: 43 new, 182 total passing, zero regressions.
Closes #27.
2026-09-01 23:49:23 +05:30
xavierk b99ebfe9dd fix(projection): regime dynamics, warming gate, habit change detection, horizon coverage
Fixes #26.

Changes:
- Fix warming gate: check total_days < 14 OR days_below_coverage > 2
- Rewrite _detect_habit_change: correct consecutive-day scanning
- Fix _compute_horizon_rate: require history spans full horizon

Tests: 29 new tests covering PR-2,3,6,7,9,15,16. 139 total passing.
2026-09-01 23:32:00 +05:30
xavierk 7802a72606 feat: implement projection core pure function (closes #25) 2026-09-01 23:16:33 +05:30
xavierk d005412c0d feat: implement legacy history migration (closes #24) 2026-09-01 23:09:05 +05:30
xavierk 6a1841c447 feat: segment observation history by controller identity (closes #23) 2026-09-01 23:03:42 +05:30
xavierk 7d219c4697 feat(hour/day derivation): hour classification, monitoring periods, day aggregates, pruning
Hour classification (PR-4):
- Powered-off: POH delta < 90% of wall-clock span
- Active: DUW delta >= 256 MiB
- Idle: powered on + sampled + below active threshold
- Unknown: unsampled without POH evidence
- Four splits sum to exactly wall_clock_seconds
- Disabled time is never an hour state

Monitoring periods (FL-8):
- ensure_period_open: opens period at run moment if none exists
- close_period: closes with end cause
- is_inside_period: checks timestamp against period bounds
- Never backdated; wall-clock outside periods excluded from denominator

Day aggregates (ST-4, PR-5):
- Derived monotonically from hour rows
- UTC-bounded; no 23/25-hour days
- Coverage: known seconds / period wall-clock
- Gap hours inside periods contribute unknown seconds
- Hours outside periods excluded entirely
- No absent hour interpolated/estimated/fabricated (FL-3)

Raw sample pruning (ST-5):
- Prunes samples older than 14 days
- Hour observations and day aggregates retained indefinitely

Closes #22
2026-09-01 22:58:03 +05:30
xavierk 2217b00ff6 feat: collector tracer bullet (#21)\n\nSmartctl acquisition with validation\nSysfs controller identity acquisition\nIdentity normalization (strip, no case fold, blank handling)\nSQLite store in WAL mode with six entities\nSchema versioning via PRAGMA user_version\nInvariant validation (negative bytes, etc.)\n17 passing tests across two test files\n\nCloses #21 2026-09-01 22:33:26 +05:30
xavierk 566d1c81b1 docs(spec): fenris-redesign.md — implementation-ready specification from ADRs 0001-0006 with two-way traceability matrix 2026-08-31 23:43:22 +05:30
xavierk ff51be2d8a docs(spec): fill assembly traceability gaps — PR-17 projection arithmetic, TUI-4 Panes layout, IN-10 artifact placement, CI-3 /run prohibition 2026-08-31 23:43:22 +05:30
xavierk b2ef1308b0 docs(adr): 0006 collector acquisition path — smartctl counters + sysfs identity, hard pin, no partial samples; SLOT-B filled (AC-1–5) 2026-08-31 23:02:25 +05:30
xavierk 580fab26be docs(adr): amend 0002 — blank identity key caps confidence at Limited, blank-key change semantics; SLOT-A filled (PR-15/16, ID-4); glossary term degraded identity 2026-08-31 22:30:40 +05:30
xavierk bf83ac5481 docs(spec): acceptance criteria for the redesign — subsystem gates, A/P/M evidence classes, ADR traceability; state-matrix and parity gates, prohibition set 2026-08-31 22:13:20 +05:30
xavierk 305bdd4779 docs(adr): amend 0001 — controller-segment metadata snapshot: normalized identity diagnostics, vid/ssvid/transport, frozen at open, nullable 2026-08-31 21:57:54 +05:30
xavierk 4d332bef53 docs(adr): amend 0001 + 0003 — baseline provenance and validation; glossary terms for verified and unverified override 2026-08-31 21:02:09 +05:30
xavierk d45d931a0d docs(adr): 0005 failure and recovery — refuse bad writes, never backfill, degrade store faults; glossary term for store fault 2026-08-31 20:14:23 +05:30
xavierk ffa6f22777 docs(adr): 0004 installation lifecycle — Makefile venv install, dormant install, sanctioned teardown 2026-08-31 18:58:05 +05:30
xavierk de1e8c753b docs(agents): issue-tracker and domain guidance for agents; ignore .pi session state 2026-08-31 17:25:10 +05:30
xavierk 26a6703152 docs(adr): 0003 service lifecycle — timer-driven collector, sanctioned control helper; glossary terms for collection run, deliberate disable 2026-08-31 16:14:52 +05:30
xavierk 43467f8957 docs(adr): 0002 projection model — sustained regime, categorical confidence; glossary terms for regime, habit change, scenario range, coverage 2026-08-31 15:38:41 +05:30
xavierk a63da74e44 docs(adr): 0001 observation store in SQLite; glossary terms for store entities 2026-08-31 14:48:14 +05:30
142 changed files with 36667 additions and 1295 deletions
+350
View File
@@ -0,0 +1,350 @@
# Fenris release workflow — release path on Coolify-hosted Gitea runner.
# The runner is repository-scoped and executes package build, signing, validation,
# registry publication, and release attachment. Spec: §5, issue #52
name: Release
on:
push:
tags:
- 'v*'
workflow_dispatch:
inputs:
publish_xbps:
description: 'Publish XBPS package to distribution repository (requires host acceptance)'
required: false
default: false
type: boolean
# Built-in Gitea token needs write access for release assets and package registry.
permissions:
contents: read
releases: write
packages: write
jobs:
release:
runs-on: [self-hosted]
steps:
- uses: actions/checkout@v4
- name: Validate release tag and notes
run: |
set -euo pipefail
VERSION="$(sed -n 's/^version = "\(.*\)"/\1/p' pyproject.toml)"
if [ -z "${VERSION}" ]; then
echo "::error::could not determine the project version"
exit 1
fi
if [ "${GITHUB_EVENT_NAME}" != "workflow_dispatch" ]; then
EXPECTED_TAG="v${VERSION}"
ACTUAL_TAG="${GITHUB_REF#refs/tags/}"
if [ "${ACTUAL_TAG}" != "${EXPECTED_TAG}" ]; then
echo "::error::tag ${ACTUAL_TAG} does not match ${EXPECTED_TAG}"
exit 1
fi
fi
python3 scripts/extract_changelog.py CHANGELOG.md "${VERSION}" \
--footer packaging/release-footer.md > "${RUNNER_TEMP}/release-body.md"
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.12'
- name: Install build dependencies
run: |
sudo apt-get update
sudo apt-get install -y gnupg2 rpm python3-venv
python3 -m venv /tmp/fenris-ci
/tmp/fenris-ci/bin/pip install --quiet build
echo "/tmp/fenris-ci/bin" >> "$GITHUB_PATH"
NFPM_VERSION=2.47.0
curl --fail --silent --show-error --location \
"https://github.com/goreleaser/nfpm/releases/download/v${NFPM_VERSION}/nfpm_${NFPM_VERSION}_Linux_x86_64.tar.gz" \
-o /tmp/nfpm.tar.gz
sudo tar -xzf /tmp/nfpm.tar.gz -C /usr/local/bin nfpm
nfpm --version
# Ubuntu does not package the XBPS build tools. Use Void's static
# toolchain, pinned and checksum-verified before it reaches PATH.
XBPS_STATIC_VERSION=0.60.4_1
XBPS_STATIC_ARCHIVE="xbps-static-static-${XBPS_STATIC_VERSION}.x86_64-musl.tar.xz"
XBPS_STATIC_SHA256=603b3c55e9cabd5af79b461b929b14e1556a443c97b5714d188681c2172d9e28
curl --fail --silent --show-error --location \
"https://repo-default.voidlinux.org/static/${XBPS_STATIC_ARCHIVE}" \
-o "/tmp/${XBPS_STATIC_ARCHIVE}"
echo "${XBPS_STATIC_SHA256} /tmp/${XBPS_STATIC_ARCHIVE}" | sha256sum --check --strict
mkdir -p /tmp/xbps-static
tar -xJf "/tmp/${XBPS_STATIC_ARCHIVE}" -C /tmp/xbps-static
export PATH="/tmp/xbps-static/usr/bin:${PATH}"
echo "/tmp/xbps-static/usr/bin" >> "$GITHUB_PATH"
xbps-create --version
- name: Build packages
run: make package
- name: Build XBPS package
run: make package-xbps
- name: Import packaging key
env:
GPG_PRIVATE_KEY: ${{ secrets.GPG_PRIVATE_KEY }}
run: |
set -euo pipefail
if [ -z "${GPG_PRIVATE_KEY}" ]; then
echo "::error::GPG_PRIVATE_KEY repository secret is not configured"
exit 1
fi
printf '%s\n' "${GPG_PRIVATE_KEY}" | gpg --batch --import
SECRET_FINGERPRINT="$(gpg --batch --list-secret-keys --with-colons 'packaging@bongbetic.com' | awk -F: '$1 == "fpr" { print $10; exit }')"
PUBLIC_FINGERPRINT="$(gpg --batch --show-keys --with-colons packaging/keys/fenris-packaging.asc | awk -F: '$1 == "fpr" { print $10; exit }')"
if [ -z "${PUBLIC_FINGERPRINT}" ]; then
echo "::error::packaging/keys/fenris-packaging.asc has no OpenPGP key"
exit 1
fi
if [ "${SECRET_FINGERPRINT}" != "${PUBLIC_FINGERPRINT}" ]; then
echo "::error::packaging public key does not match imported private key"
exit 1
fi
echo "Packaging key fingerprint verified: ${PUBLIC_FINGERPRINT}"
- name: Sign RPM payload
run: make sign-rpm
- name: Import XBPS signing key
id: import-xbps-key
env:
XBPS_SIGNING_KEY: ${{ secrets.XBPS_SIGNING_KEY }}
run: |
set -euo pipefail
if [ -z "${XBPS_SIGNING_KEY}" ]; then
echo "::warning::XBPS_SIGNING_KEY secret not configured; XBPS signing skipped"
echo "xbps_signed=false" >> "$GITHUB_OUTPUT"
exit 0
fi
mkdir -p ~/.ssh
printf '%s\n' "${XBPS_SIGNING_KEY}" > ~/.ssh/id_xbps
chmod 600 ~/.ssh/id_xbps
echo "xbps_signed=true" >> "$GITHUB_OUTPUT"
- name: Sign XBPS package
if: steps.import-xbps-key.outputs.xbps_signed == 'true'
run: make sign-xbps
- name: Generate and clearsign SHA256SUMS
run: |
set -euo pipefail
VERSION="$(sed -n 's/^version = "\(.*\)"/\1/p' pyproject.toml)"
cd dist
sha256sum "fenris_${VERSION}_amd64.deb" \
"fenris-${VERSION}-1.x86_64.rpm" > SHA256SUMS
gpg --batch --yes --clearsign --local-user packaging@bongbetic.com SHA256SUMS
- name: Validate signatures and checksums
run: |
set -euo pipefail
VERSION="$(sed -n 's/^version = "\(.*\)"/\1/p' pyproject.toml)"
RPM="fenris-${VERSION}-1.x86_64.rpm"
RPM_VERIFY="$(rpm -Kv "dist/${RPM}" 2>&1)"
printf '%s\n' "${RPM_VERIFY}"
printf '%s\n' "${RPM_VERIFY}" | grep -Eiq 'signature.*: *ok'
gpg --batch --verify dist/SHA256SUMS.asc
(cd dist && sha256sum -c SHA256SUMS)
- name: Remove packaging key material
if: always()
run: |
set +e
FINGERPRINT="$(gpg --batch --list-secret-keys --with-colons 'packaging@bongbetic.com' 2>/dev/null | awk -F: '$1 == "fpr" { print $10; exit }')"
if [ -n "${FINGERPRINT}" ]; then
gpg --batch --yes --delete-secret-keys "${FINGERPRINT}"
gpg --batch --yes --delete-keys "${FINGERPRINT}"
fi
- name: Determine version
id: version
run: echo "version=$(sed -n 's/^version = "\(.*\)"/\1/p' pyproject.toml)" >> "$GITHUB_OUTPUT"
- name: Upload deb packages to registry
env:
GITEA_PUBLISH_TOKEN: ${{ secrets.GITEAPACKAGETOKEN }}
run: |
set -euo pipefail
if [ -z "${GITEA_PUBLISH_TOKEN}" ]; then
echo "::error::GITEAPACKAGETOKEN repository secret is not configured"
exit 1
fi
VERSION=${{ steps.version.outputs.version }}
DEB="fenris_${VERSION}_amd64.deb"
for CODENAME in bookworm jammy noble; do
STATUS=$(curl --silent --show-error --user "xavierk:${GITEA_PUBLISH_TOKEN}" -X PUT \
-T "dist/${DEB}" -o /dev/null -w '%{http_code}' \
"https://git.bongbetic.com/api/packages/xavierk/debian/pool/${CODENAME}/main/upload" || true)
case "${STATUS}" in
200|201|204) echo "Debian ${CODENAME}: uploaded" ;;
409) echo "Debian ${CODENAME}: already exists, kept existing package" ;;
*) echo "::error::Debian ${CODENAME} upload failed with HTTP ${STATUS}"; exit 1 ;;
esac
done
- name: Upload RPM to registry
env:
GITEA_PUBLISH_TOKEN: ${{ secrets.GITEAPACKAGETOKEN }}
run: |
set -euo pipefail
VERSION=${{ steps.version.outputs.version }}
RPM="fenris-${VERSION}-1.x86_64.rpm"
STATUS=$(curl --silent --show-error --user "xavierk:${GITEA_PUBLISH_TOKEN}" -X PUT \
-T "dist/${RPM}" -o /dev/null -w '%{http_code}' \
"https://git.bongbetic.com/api/packages/xavierk/rpm/fenris/upload" || true)
case "${STATUS}" in
200|201|204) echo "RPM: uploaded" ;;
409) echo "RPM: already exists, kept existing package" ;;
*) echo "::error::RPM upload failed with HTTP ${STATUS}"; exit 1 ;;
esac
- name: Publish XBPS to distribution repository
if: github.event.inputs.publish_xbps == 'true'
run: |
set -euo pipefail
VERSION=${{ steps.version.outputs.version }}
XBPS_FILE="fenris-${VERSION}_1.x86_64.xbps"
if [ ! -f "${XBPS_FILE}" ]; then
echo "::error::XBPS package not found: ${XBPS_FILE}"
exit 1
fi
if [ ! -f "${XBPS_FILE}.sig2" ]; then
echo "::error::XBPS signature not found: ${XBPS_FILE}.sig2"
exit 1
fi
bash scripts/xbps-publish.sh --publish
- name: Track format availability
id: formats
run: |
set -euo pipefail
VERSION=${{ steps.version.outputs.version }}
DEB_EXISTS=$([ -f "dist/fenris_${VERSION}_amd64.deb" ] && echo "true" || echo "false")
RPM_EXISTS=$([ -f "dist/fenris-${VERSION}-1.x86_64.rpm" ] && echo "true" || echo "false")
XBPS_EXISTS=$([ -f "fenris-${VERSION}_1.x86_64.xbps" ] && echo "true" || echo "false")
XBPS_PUBLISHED=$([ "${{ github.event.inputs.publish_xbps }}" = "true" ] && echo "true" || echo "false")
echo "deb_available=${DEB_EXISTS}" >> "$GITHUB_OUTPUT"
echo "rpm_available=${RPM_EXISTS}" >> "$GITHUB_OUTPUT"
echo "xbps_available=${XBPS_EXISTS}" >> "$GITHUB_OUTPUT"
echo "xbps_published=${XBPS_PUBLISHED}" >> "$GITHUB_OUTPUT"
# Build format availability summary for release notes
AVAILABLE_FORMATS=""
WITHHELD_FORMATS=""
if [ "${DEB_EXISTS}" = "true" ]; then
AVAILABLE_FORMATS="${AVAILABLE_FORMATS}Debian/Ubuntu (deb), "
fi
if [ "${RPM_EXISTS}" = "true" ]; then
AVAILABLE_FORMATS="${AVAILABLE_FORMATS}Fedora/openSUSE (rpm), "
fi
if [ "${XBPS_EXISTS}" = "true" ] && [ "${XBPS_PUBLISHED}" = "true" ]; then
AVAILABLE_FORMATS="${AVAILABLE_FORMATS}Void Linux (xbps)"
elif [ "${XBPS_EXISTS}" = "true" ]; then
WITHHELD_FORMATS="Void Linux (xbps) — pending host acceptance"
fi
# Remove trailing comma and space
AVAILABLE_FORMATS=$(echo "${AVAILABLE_FORMATS}" | sed 's/, $//')
echo "available_formats=${AVAILABLE_FORMATS}" >> "$GITHUB_OUTPUT"
echo "withheld_formats=${WITHHELD_FORMATS}" >> "$GITHUB_OUTPUT"
- name: Create Gitea release
env:
GITEA_PUBLISH_TOKEN: ${{ secrets.GITEAPACKAGETOKEN }}
run: |
set -euo pipefail
if [ -z "${GITEA_PUBLISH_TOKEN}" ]; then
echo "::error::GITEAPACKAGETOKEN repository secret is not configured"
exit 1
fi
VERSION=${{ steps.version.outputs.version }}
RELEASE_BODY="${RUNNER_TEMP}/release-body.md"
if [ ! -s "${RELEASE_BODY}" ]; then
echo "::error::validated release body is missing or empty"
exit 1
fi
# Append format availability to release notes
AVAILABLE_FORMATS="${{ steps.formats.outputs.available_formats }}"
WITHHELD_FORMATS="${{ steps.formats.outputs.withheld_formats }}"
RELEASE_BODY_WITH_FORMATS="${RUNNER_TEMP}/release-body-formats.md"
cp "${RELEASE_BODY}" "${RELEASE_BODY_WITH_FORMATS}"
echo "" >> "${RELEASE_BODY_WITH_FORMATS}"
echo "## Package formats" >> "${RELEASE_BODY_WITH_FORMATS}"
echo "" >> "${RELEASE_BODY_WITH_FORMATS}"
echo "Available: ${AVAILABLE_FORMATS}" >> "${RELEASE_BODY_WITH_FORMATS}"
if [ -n "${WITHHELD_FORMATS}" ]; then
echo "Withheld: ${WITHHELD_FORMATS}" >> "${RELEASE_BODY_WITH_FORMATS}"
fi
EXISTING_RELEASE="${RUNNER_TEMP}/existing-release.json"
EXISTING=$(curl --silent --show-error -o "${EXISTING_RELEASE}" -w '%{http_code}' \
-H "Authorization: token ${GITEA_PUBLISH_TOKEN}" \
"https://git.bongbetic.com/api/v1/repos/xavierk/Fenris/releases/tags/v${VERSION}" || true)
case "${EXISTING}" in
200)
echo "Release v${VERSION} exists; resynchronizing its notes"
REQUEST="$(python3 scripts/release_request.py --version "${VERSION}" \
--body-file "${RELEASE_BODY_WITH_FORMATS}" --existing-release "${EXISTING_RELEASE}")"
;;
404)
REQUEST="$(python3 scripts/release_request.py --version "${VERSION}" \
--body-file "${RELEASE_BODY_WITH_FORMATS}")"
;;
*)
echo "::error::release lookup failed with HTTP ${EXISTING}"
exit 1
;;
esac
METHOD="$(printf '%s' "${REQUEST}" | python3 -c "import json,sys; print(json.load(sys.stdin)['method'])")"
RELEASE_PATH="$(printf '%s' "${REQUEST}" | python3 -c "import json,sys; print(json.load(sys.stdin)['path'])")"
PAYLOAD="$(printf '%s' "${REQUEST}" | python3 -c "import json,sys; print(json.dumps(json.load(sys.stdin)['payload']))")"
curl --fail --silent --show-error -X "${METHOD}" \
-H "Authorization: token ${GITEA_PUBLISH_TOKEN}" \
-H "Content-Type: application/json" \
-d "${PAYLOAD}" \
"https://git.bongbetic.com/api/v1/repos/xavierk/Fenris${RELEASE_PATH}"
- name: Attach artifacts to release
env:
GITEA_PUBLISH_TOKEN: ${{ secrets.GITEAPACKAGETOKEN }}
run: |
set -euo pipefail
VERSION=${{ steps.version.outputs.version }}
# Get release ID for this tag
RELEASE_JSON=$(curl --fail --silent --show-error \
-H "Authorization: token ${GITEA_PUBLISH_TOKEN}" \
"https://git.bongbetic.com/api/v1/repos/xavierk/Fenris/releases/tags/v${VERSION}")
RELEASE_ID=$(printf '%s' "${RELEASE_JSON}" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['id'])")
# Attach deb, rpm, clearsigned checksums, and XBPS artifacts once.
ARTIFACTS=(
"dist/fenris_${VERSION}_amd64.deb"
"dist/fenris-${VERSION}-1.x86_64.rpm"
"dist/SHA256SUMS.asc"
)
# Add XBPS artifacts if they exist
XBPS_FILE="fenris-${VERSION}_1.x86_64.xbps"
if [ -f "${XBPS_FILE}" ]; then
if [ -f "${XBPS_FILE}.sig2" ]; then
ARTIFACTS+=("${XBPS_FILE}")
ARTIFACTS+=("${XBPS_FILE}.sig2")
else
echo "::warning::Unsigned XBPS artifact omitted from release assets"
fi
fi
for FILE in "${ARTIFACTS[@]}"; do
ASSET_NAME="${FILE##*/}"
if python3 -c 'import json,sys; name=sys.argv[1]; sys.exit(0 if any(a.get("name") == name for a in json.load(sys.stdin).get("assets", [])) else 1)' "${ASSET_NAME}" <<<"${RELEASE_JSON}"; then
echo "${ASSET_NAME}: already attached"
else
curl --fail --silent --show-error -X POST \
-H "Authorization: token ${GITEA_PUBLISH_TOKEN}" \
-F "attachment=@${FILE}" \
"https://git.bongbetic.com/api/v1/repos/xavierk/Fenris/releases/${RELEASE_ID}/assets"
fi
done
- name: Remove XBPS signing key
if: always()
run: rm -f ~/.ssh/id_xbps
+14
View File
@@ -1,4 +1,5 @@
__pycache__/
.pi/
*.pyc
.commandcode/
data/fenris.pid
@@ -6,3 +7,16 @@ data/fenris.log
data/history.jsonl
data/hourly.jsonl
plan-dash-changes.md
# Packaging build artifacts
build/
dist/
# Local tooling
graphify-out/
json
src/fenris.egg-info/
.pytest_cache/
.venv/
MagicMock*
<MagicMock*
+18
View File
@@ -0,0 +1,18 @@
## Session style
After the first user message in each session, load the global `caveman` skill and activate `ultra` mode. Keep `ultra` mode active until the user changes or stops it.
## Agent skills
## Commit messages
Do not add `Co-authored-by: CommandCodeBot <noreply@commandcode.ai>` or other
CommandCodeBot attribution trailers to commits.
### Issue tracker
Issues are tracked in Gitea using the authenticated `tea` CLI. See `docs/agents/issue-tracker.md`.
### Domain docs
This is a single-context repository. See `docs/agents/domain.md`.
+93
View File
@@ -0,0 +1,93 @@
# Changelog
<!--
Maintainers add one user-facing entry to Unreleased with each change. A release
commit bumps pyproject.toml, renames Unreleased to that bare-semver version and
an ISO date, then restores an empty Unreleased section; tag that commit. Do not
backfill releases from before this changelog.
-->
## [Unreleased]
## [0.6.0] - 2026-09-29
### Added
- Preserve local-day read and write evidence separately, including ambiguous midnight-spanning volume as shared evidence instead of assigning it to either day.
- Keep valid observations pending when publication fails, and recover them without exposing partial history or counting activity twice.
- Preserve live-point inspection and historical activity selection through graph refresh.
### Changed
- Prune old detail only after durable summaries and required boundary evidence are published.
- Require a complete local day to fit within one monitoring period before it can open the endurance outlook.
## [0.5.0] - 2026-09-19
### Changed
- Make drive activity the focus of a Chalktone dashboard with Live, Day, and History tabs, panel zoom, and fixed monitoring controls.
- Replace block bars with labelled dotted volume plots that fit the terminal and preserve gaps, partial evidence, and point inspection.
### Fixed
- Preserve historical selections and store-fault messages through graph refresh and resize.
- Show hourly drill-down results without overwriting them with a loading placeholder.
- Keep known unallocated daily volume visible across coverage gaps, and distinguish missing hourly evidence from zero on small terminals.
- Stage package directories with consistent public permissions, avoiding openSUSE RPM conflicts and inaccessible runtime paths when built with a restrictive umask.
## [0.4.0] - 2026-09-18
### Added
- Show local-day read and write totals with their recorded timezone, labelled as totals so far for the current day.
- Open the dashboard on a live three-hour written-volume graph with a read/write toggle and selected-point inspection.
- Browse history by day from the keyboard with `[`, `]`, `g` date entry, and `t` for today and the live view.
### Changed
- Collect every three minutes on systemd and runit, plot live points at their actual timestamps, and retain three-minute detail for 14 days before durable summaries.
- Withhold the endurance outlook until one full local observation day has usable evidence, then show it with categorical confidence.
### Fixed
- Keep midnight-spanning activity once as shared boundary evidence instead of adding it to both days.
- Publish derived hour and day history consistently with each collection run.
## [0.3.7] - 2026-09-16
### Changed
- Make usage-history axes, units, UTC boundaries, active 7/14/30/90-day window, gaps, partial periods, stacked write attribution, and hourly drill-down explicit; keep live graph data refreshed on the current three-minute cadence.
## [0.3.6] - 2026-09-16
### Changed
- Use sentence case throughout the dashboard and add persistent keyboard and sudo guidance.
- Share read-only status acquisition between the CLI and TUI, preserving unknown monitoring state and store-fault recovery guidance.
- Use one authenticated action path for the CLI and TUI; let valid collection runs finish without the former 30-second dashboard cutoff.
- Remove the unused sparkline and habit-bar rendering path while retaining the interactive history graph.
- Align all package formats on the MIT license and Python 3.10 minimum; include the license text in native packages.
### Fixed
- Correct observation-store directory permissions in native packages, including repair of older runit installations during upgrade.
## [0.3.5] - 2026-09-14
### Added
- Show trustworthy first-use history, qualifying-day progress, and confidence in the dashboard.
- Add an interactive daily-writes graph with hourly drill-down and evidence-anchored scenario windows.
- Repair retained observation history safely and keep its retention state visible.
- Present a unified monitoring status, drive-health context, and Fenris identity in the dashboard.
- Remember accessible colour and motion preferences and keep the dashboard usable at constrained terminal sizes.
## [0.3.4] - 2026-09-10
### Added
- Add Fenris identity and a polkit authentication notice to the dashboard.
- Clarify monitoring continuity, deliberate pauses, and quitting in the dashboard and status output.
- Add per-release notes with installation, verification, and rollback guidance.
+117
View File
@@ -0,0 +1,117 @@
# Fenris
Fenris observes an NVMe drive’s real-world use and translates that history into an understandable endurance outlook.
## Language
**Observation history**:
The persisted record of drive activity gathered while Fenris monitoring is enabled, retained across restarts and reboots.
_Avoid_: Calibration data, temporary history
**Observed usage habit**:
The pattern of active, idle, and powered-off hours represented by the observation history, with recent sustained behavior carrying more relevance than distant behavior.
_Avoid_: Current usage, benchmark workload
**Usage-adjusted theoretical lifespan**:
The theoretical time until the drive’s write endurance is exhausted if its observed usage habit continues; it is an endurance projection, not a predicted hardware-failure date.
_Avoid_: Future life, actual lifespan, failure date
**Projection confidence**:
The degree to which the observation history is sufficiently long, complete, and stable to support the usage-adjusted theoretical lifespan.
_Avoid_: Accuracy percentage, certainty
**Monitoring period**:
A span during which Fenris monitoring is enabled; powered-off time remains part of the usage habit, while deliberately disabled time does not.
_Avoid_: Daemon uptime, calibration window
**Observation store**:
The single SQLite database at `/var/lib/fenris/observations.db` that persists the observation history, monitoring periods, hour observations, day aggregates, and endurance baseline.
_Avoid_: Data directory, history.jsonl, the database (generic)
**Pending publication**:
The condition where valid acquired observations are retained for recovery but their dependent evidence has not yet been published consistently. Those observations are not part of the reader-visible observation history until publication succeeds; readers retain the last consistent evidence with the pending condition made explicit.
_Avoid_: Successful collection, fresh published evidence
**Store fault**:
The condition where the observation store is present but cannot be read or trusted — unreadable, corrupt, or written by a newer Fenris — degrading every view that depends on it rather than crashing or guessing.
_Avoid_: Database error, corruption, broken data
**Hour observation**:
One row per UTC hour in the observation store, recording that hour's usage-habit split into active, idle, powered-off, and unknown seconds, plus write/read deltas, thermal evidence, and coverage.
_Avoid_: Hourly record, hourly.jsonl entry
**Day aggregate**:
One row per UTC day derived from hour observations; the grain at which usage-habit evidence is judged.
_Avoid_: Daily summary, daily stats
**Local-day evidence**:
Measured read and write activity attributable to a local calendar day with its recorded timezone and midnight boundaries. Known volumes remain incomplete when gaps or shared local-day evidence prevent an exact total; they are not estimates of the missing activity.
_Avoid_: Estimated daily total, localised UTC day aggregate
**Shared local-day evidence**:
Measured read or write volume spanning a local midnight that cannot honestly be allocated to either adjacent day. It is retained once, separately from either day's known volume, rather than prorated or counted in both days.
_Avoid_: Missing bytes, estimated midnight split
**Activity selection**:
The read/write measurement, date and point being inspected in live or historical drive activity. Live activity follows the newest point until deliberate inspection pins a point; historical selection survives refresh, while an expired live pin is explained before following resumes.
_Avoid_: Projection window, current usage habit
**Usage-history window**:
An exact consecutive span of UTC calendar days ending today, shown from day aggregates; a day without trustworthy evidence remains an explicit gap rather than disappearing or being estimated.
_Avoid_: Available records, dataset range
**Unallocated write evidence**:
Writes known to belong to a UTC day but which cannot be assigned honestly to a particular hour; they contribute to that day's total but are never distributed across hourly bars.
_Avoid_: Missing writes, estimated hourly writes
**Controller segment**:
A span of observation history within which the drive's controller identity is unchanged and counters are monotonic; write deltas are never computed across a segment boundary.
_Avoid_: Counter reset handling, drive swap detection
**Degraded identity**:
The condition where a controller segment's identity key is blank because no identifier rung produced a value; replacement detection then relies on write-counter continuity alone, and projection confidence is capped.
_Avoid_: Identity error, unknown device, virtual drive
**Endurance baseline**:
The write-endurance value a projection consumes, chosen by precedence: a verified override when one exists, otherwise an unverified override, otherwise a coarse implied baseline derived from vendor wear — each labeled as such.
_Avoid_: TBW value, failure threshold, max writes
**Verified override**:
A rated-TBW override with complete provenance whose applicability to the detected drive was confirmed by machine match or explicit user attestation; the strongest endurance baseline.
_Avoid_: Confirmed TBW, trusted value
**Unverified override**:
A rated-TBW override knowingly stored with incomplete provenance; always presented as user-supplied, never as verified.
_Avoid_: Forced entry, fallback baseline
**Sustained regime**:
The most recent stretch of the observation history over which the observed usage habit has been stable; the interval whose write rate the usage-adjusted theoretical lifespan consumes.
_Avoid_: Current window, detection period
**Habit change**:
A sustained divergence between recent and earlier daily write rates that starts a new sustained regime.
_Avoid_: Spike, anomaly
**Scenario range**:
The spread of lifespan projections computed from the 7-, 28-, and 90-day horizons of the observation history, shown in place of a statistical interval.
_Avoid_: Confidence interval, error bar
**Coverage**:
The share of wall-clock seconds inside monitoring periods whose usage-habit classification is known rather than unknown.
_Avoid_: Uptime, sample count
**Collection run**:
One scheduled or on-demand execution of the collector that interrogates the drive and extends the observation history.
_Avoid_: Poll, daemon tick
**Release**:
A published version of Fenris: a version tag, its packages in the channel, and its human-readable change notes, all together; a bare tag is not one.
_Avoid_: Tag, upload, build
**Rollback**:
Returning to an earlier release by restoring an observation-store snapshot and then installing that release; installing an older package over a newer store is unsupported.
_Avoid_: Downgrade, version pinning (as a promise)
**Deliberate disable**:
A monitoring pause made through Fenris's own control path, closing the monitoring period so the paused time is excluded from the usage habit.
_Avoid_: Manual stop, service stop
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Fenris contributors
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+399
View File
@@ -0,0 +1,399 @@
# Fenris Makefile
# Spec: §10.1-10.6
SHELL := /bin/bash
PYTHON := python3
VENV_DIR := /opt/fenris
VENDOR_DIR := $(VENV_DIR)/vendor
BIN_DIR := /usr/local/bin
LIBEXEC_DIR := /usr/libexec/fenris
UNIT_DIR := /etc/systemd/system
POLKIT_DIR := /usr/share/polkit-1/actions
CONF_DIR := /etc/fenris
DATA_DIR := /var/lib/fenris
# Placement manifest
MANIFEST := $(DATA_DIR)/manifest.txt
# Legacy history path (IN-4)
LEGACY_HISTORY := ./data/history.jsonl
.PHONY: help install upgrade uninstall purge update-deps test lint check-python check-smartctl import-legacy stage package-deb package-rpm package-xbps package generate-test-key sign-rpm sign-xbps checksums clearsign release release-run release-dry-run xbps-publish xbps-publish-dry-run clean
help:
@echo "Fenris NVMe endurance monitor"
@echo ""
@echo "Targets:"
@echo " install - Install Fenris (builds wheel, installs to /opt/fenris)"
@echo " upgrade - Upgrade Fenris (reinstall wheel, sync units)"
@echo " uninstall - Uninstall Fenris (preserves config and store)"
@echo " purge - Remove everything including config and store"
@echo " test - Run tests"
@echo " lint - Run linter"
@echo " update-deps - Update dependency pins"
@echo " stage - Stage packaging tree for nfpm"
@echo " package - Build deb + rpm packages"
@echo " package-deb - Build deb package only"
@echo " package-rpm - Build rpm package only"
@echo " package-xbps - Build XBPS package only"
@echo " sign-xbps - Sign XBPS package with signing key"
@echo " generate-test-key - Create throwaway GPG key for CI/testing"
@echo " sign-rpm - Sign RPM payload with packaging key"
@echo " checksums - Generate SHA256SUMS manifest"
@echo " clearsign - Clearsign SHA256SUMS with packaging key"
@echo " release - Full release (build, sign, checksum, print upload steps)"
@echo " release-run - Execute the full release flow via scripts/release.sh"
@echo " release-dry-run - Dry-run of the release flow (prints commands only)"
@echo " xbps-publish - Publish XBPS package to distribution repository"
@echo " xbps-publish-dry-run - Dry-run of XBPS publication (prints commands only)"
@echo " clean - Remove build artifacts"
# ─── Pre-install gates ──────────────────────────────────────────────────────
check-python:
@echo "=== Verifying Python ≥ 3.10 ==="
@$(PYTHON) -c "import sys; v=sys.version_info; exit(0 if (v>=(3,10)) else 1)" || { echo "Error: Python 3.10+ required (found $$($(PYTHON) --version 2>&1))"; exit 1; }
check-smartctl:
@echo "=== Verifying smartctl ==="
@smartctl --version 2>/dev/null | head -1 || { echo "Error: smartctl not found (install smartmontools)"; exit 1; }
# ─── Build ──────────────────────────────────────────────────────────────────
dist/fenris-*.whl: pyproject.toml src/fenris/*.py
@mkdir -p dist
$(PYTHON) -m pip wheel --no-deps --wheel-dir dist .
# ─── Install ────────────────────────────────────────────────────────────────
install: check-python check-smartctl dist/fenris-*.whl
@echo "=== Creating directories ==="
@sudo mkdir -p $(LIBEXEC_DIR)
@sudo mkdir -p $(CONF_DIR)
@sudo mkdir -p $(POLKIT_DIR)
@echo "=== Creating data directory (root-written, group-read) ==="
@sudo groupadd -f fenris
@sudo install -d -o root -g fenris -m 2770 $(DATA_DIR)
@echo "=== Installing version-neutral runtime packages ==="
@sudo rm -rf $(VENV_DIR)
@sudo install -d -m 0755 $(VENDOR_DIR)
@sudo $(PYTHON) -m pip install --disable-pip-version-check --no-compile --target $(VENDOR_DIR) -r requirements.txt dist/fenris-*.whl --quiet
@echo "=== Installing wrapper ==="
@sudo install -m 0755 scripts/fenris $(BIN_DIR)/fenris
@echo "=== Installing helpers ==="
@sudo install -m 0755 src/fenris/monitor.py $(LIBEXEC_DIR)/fenris-monitor
@sudo install -m 0755 src/fenris/collect.py $(LIBEXEC_DIR)/fenris-collect
@echo "=== Installing systemd units (dormant — not enabled/started) ==="
@sudo install -m 0644 units/fenris-collect.timer $(UNIT_DIR)/
@sudo install -m 0644 units/fenris-collect.service $(UNIT_DIR)/
@sudo systemctl daemon-reload
@echo "=== Installing runit service files (dormant — not enabled) ==="
@sudo install -d -m 0755 /etc/sv/fenris-collect/log
@sudo install -d -o root -g fenris -m 2770 /var/log/fenris-collect
@sudo install -m 0755 units/runit/fenris-collect/run /etc/sv/fenris-collect/run
@sudo install -m 0755 units/runit/fenris-collect/log/run /etc/sv/fenris-collect/log/run
@sudo touch /etc/sv/fenris-collect/down
@echo "=== Installing polkit policy ==="
@sudo install -m 0644 polkit/com.bongbetic.fenris.monitor.policy $(POLKIT_DIR)/
@echo "=== Recording manifest (IN-2, IN-10) ==="
@echo "# Fenris placement manifest — do not edit" | sudo tee $(MANIFEST) > /dev/null
@echo "# Generated by: sudo make install" | sudo tee -a $(MANIFEST) > /dev/null
@echo "# Timestamp: $$(date -u +%%Y-%%m-%%dT%%H:%%M:%%SZ)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(BIN_DIR)/fenris" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(LIBEXEC_DIR)/fenris-monitor" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(LIBEXEC_DIR)/fenris-collect" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(UNIT_DIR)/fenris-collect.timer" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(UNIT_DIR)/fenris-collect.service" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/etc/sv/fenris-collect/run" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/etc/sv/fenris-collect/log/run" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/etc/sv/fenris-collect/down" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(POLKIT_DIR)/com.bongbetic.fenris.monitor.policy" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(VENV_DIR)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(DATA_DIR)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(CONF_DIR)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/var/log/fenris-collect" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(MANIFEST)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "=== Install complete ==="
@echo "Units installed but NOT enabled or started (dormant — IN-3)."
@echo "To start monitoring: fenris monitor resume"
@$(MAKE) --no-print-directory import-legacy
# ─── Legacy import (IN-4, ST-7) ────────────────────────────────────────────
import-legacy:
@if [ -f "$(LEGACY_HISTORY)" ]; then \
echo "=== Detected legacy history: $(LEGACY_HISTORY) ==="; \
echo "Running idempotent import..."; \
PYTHONPATH=$(VENDOR_DIR) $(PYTHON) -c "from fenris.legacy import import_legacy_history; from fenris.store import init_store; from pathlib import Path; conn = init_store(Path('$(DATA_DIR)/observations.db')); r = import_legacy_history(conn, Path('$(LEGACY_HISTORY)')); conn.close(); print(f' Samples imported: {r.get(\"samples_imported\", 0)}'); print(f' Hours imported: {r.get(\"hours_imported\", 0)}'); print(f' Malformed lines: {r.get(\"malformed_lines\", 0)}') if not r.get('skipped') else print(' Skipped: already imported')" || echo " Warning: import failed (non-fatal)"; \
else \
echo "=== No legacy history found at $(LEGACY_HISTORY) ==="; \
fi
# ─── Upgrade ────────────────────────────────────────────────────────────────
upgrade: dist/fenris-*.whl
@echo "=== Upgrading Fenris ==="
@echo "=== Snapshotting database (IN-6) ==="
@sudo cp $(DATA_DIR)/observations.db $(DATA_DIR)/observations.db.bak 2>/dev/null || true
@echo "=== Installing new version-neutral runtime packages ==="
@sudo rm -rf $(VENDOR_DIR)
@sudo install -d -m 0755 $(VENDOR_DIR)
@sudo $(PYTHON) -m pip install --disable-pip-version-check --no-compile --target $(VENDOR_DIR) -r requirements.txt dist/fenris-*.whl --quiet
@echo "=== Syncing units against manifest ==="
@sudo install -m 0644 units/fenris-collect.timer $(UNIT_DIR)/
@sudo install -m 0644 units/fenris-collect.service $(UNIT_DIR)/
@sudo install -d -m 0755 /etc/sv/fenris-collect/log
@sudo groupadd -f fenris
@sudo install -d -o root -g fenris -m 2770 /var/log/fenris-collect
@sudo install -m 0755 units/runit/fenris-collect/run /etc/sv/fenris-collect/run
@sudo install -m 0755 units/runit/fenris-collect/log/run /etc/sv/fenris-collect/log/run
@sudo install -m 0644 polkit/com.bongbetic.fenris.monitor.policy $(POLKIT_DIR)/
@sudo install -m 0755 scripts/fenris $(BIN_DIR)/fenris
@sudo install -m 0755 src/fenris/monitor.py $(LIBEXEC_DIR)/fenris-monitor
@sudo install -m 0755 src/fenris/collect.py $(LIBEXEC_DIR)/fenris-collect
@sudo systemctl daemon-reload
@echo "=== Updating manifest ==="
@echo "# Fenris placement manifest — do not edit" | sudo tee $(MANIFEST) > /dev/null
@echo "# Generated by: sudo make upgrade" | sudo tee -a $(MANIFEST) > /dev/null
@echo "# Timestamp: $$(date -u +%%Y-%%m-%%dT%%H:%%M:%%SZ)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(BIN_DIR)/fenris" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(LIBEXEC_DIR)/fenris-monitor" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(LIBEXEC_DIR)/fenris-collect" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(UNIT_DIR)/fenris-collect.timer" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(UNIT_DIR)/fenris-collect.service" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/etc/sv/fenris-collect/run" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/etc/sv/fenris-collect/log/run" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/etc/sv/fenris-collect/down" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(POLKIT_DIR)/com.bongbetic.fenris.monitor.policy" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(VENV_DIR)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(DATA_DIR)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(CONF_DIR)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "/var/log/fenris-collect" | sudo tee -a $(MANIFEST) > /dev/null
@echo "$(MANIFEST)" | sudo tee -a $(MANIFEST) > /dev/null
@echo "=== Restarting timer only if contents changed and active (IN-5) ==="
@for unit in fenris-collect.timer fenris-collect.service; do \
TMPFILE=$$(mktemp); \
sudo systemctl cat $$unit > $$TMPFILE 2>/dev/null || true; \
if ! diff -q $$TMPFILE $(UNIT_DIR)/$$unit > /dev/null 2>&1; then \
if systemctl is-active --quiet $$unit; then \
echo " $$unit changed and active — restarting"; \
sudo systemctl restart $$unit; \
fi; \
fi; \
rm -f $$TMPFILE; \
done
@echo "=== Applying forward-only schema migrations (IN-5, IN-6) ==="
@sudo env PYTHONPATH=$(VENDOR_DIR) $(PYTHON) -c "from fenris.store import migrate_to_latest; from pathlib import Path; n = migrate_to_latest(Path('$(DATA_DIR)/observations.db')); print(f' Migration steps applied: {n}') if n else print(' Schema already current')"
@echo "=== Upgrade complete ==="
# ─── Uninstall (IN-7) ──────────────────────────────────────────────────────
uninstall:
@echo "=== Uninstalling Fenris ==="
@if [ -x $(LIBEXEC_DIR)/fenris-monitor ]; then \
echo "=== Performing sanctioned disable (§10.4) ==="; \
sudo $(LIBEXEC_DIR)/fenris-monitor disable --now || true; \
fi
@echo "=== Stopping units ==="
@sudo systemctl stop fenris-collect.timer 2>/dev/null || true
@sudo systemctl disable fenris-collect.timer 2>/dev/null || true
@sudo systemctl daemon-reload
@-sudo rm -f /var/service/fenris-collect 2>/dev/null || true
@-sudo rm -rf /etc/sv/fenris-collect 2>/dev/null || true
@-sudo rm -rf /var/log/fenris-collect 2>/dev/null || true
@echo "=== Removing installed files (preserving config and store) ==="
@-rm -f $(BIN_DIR)/fenris
@-rm -f $(LIBEXEC_DIR)/fenris-monitor
@-rm -f $(LIBEXEC_DIR)/fenris-collect
@-rmdir $(LIBEXEC_DIR) 2>/dev/null || true
@-rm -f $(UNIT_DIR)/fenris-collect.timer
@-rm -f $(UNIT_DIR)/fenris-collect.service
@-rm -f $(POLKIT_DIR)/com.bongbetic.fenris.monitor.policy
@sudo rm -rf $(VENV_DIR)
@-sudo rm -f $(MANIFEST)
@echo "=== Uninstall complete ==="
@echo "Config preserved at $(CONF_DIR)"
@echo "Store preserved at $(DATA_DIR)"
# ─── Purge ──────────────────────────────────────────────────────────────────
purge: uninstall
@echo "=== Purging Fenris ==="
@-sudo rm -rf $(CONF_DIR)
@-sudo rm -rf $(DATA_DIR)
@echo "=== Purge complete ==="
# ─── Test ───────────────────────────────────────────────────────────────────
test:
$(PYTHON) -m pytest tests/ -v
# ─── Lint ───────────────────────────────────────────────────────────────────
lint:
$(PYTHON) -m ruff check src/ tests/
# ─── Dependencies (IN-8) ───────────────────────────────────────────────────
update-deps:
$(PYTHON) -m pip compile pyproject.toml -o requirements.txt
# ─── Packaging (spec §3, §4, §5) ────────────────────────────────────────────
# Version is sourced from pyproject.toml for both formats
FENRIS_VERSION := $(shell sed -n 's/^version = "\(.*\)"/\1/p' pyproject.toml)
# GPG signing — packaging key UID (spec §4)
PACKAGING_KEY ?= packaging@bongbetic.com
stage:
@echo "=== Staging packaging tree (v$(FENRIS_VERSION)) ==="
$(PYTHON) -m pip wheel --no-deps --wheel-dir dist .
bash packaging/stage.sh "$(FENRIS_VERSION)"
package-deb: stage
@echo "=== Building deb package ==="
VERSION="$(FENRIS_VERSION)" nfpm pkg -f packaging/nfpm.yaml -p deb -t dist/
@echo "=== deb package built: dist/fenris_$(FENRIS_VERSION)_amd64.deb ==="
package-rpm: stage
@echo "=== Building rpm package ==="
VERSION="$(FENRIS_VERSION)" nfpm pkg -f packaging/nfpm.yaml -p rpm -t dist/
@echo "=== rpm package built: dist/fenris-$(FENRIS_VERSION)-1.x86_64.rpm ==="
package: package-deb package-rpm
@echo "=== Both packages built in dist/ ==="
# ─── XBPS packaging (ADR 0008) ─────────────────────────────────────────────
# XBPS signing key (separate from SSH authentication key)
XBPS_SIGNING_KEY ?= $(HOME)/.ssh/id_xbps
XBPS_REVISION ?= 1
package-xbps: stage
@echo "=== Building XBPS package ==="
cp packaging/xbps/install.sh build/stage/INSTALL
cp packaging/xbps/remove.sh build/stage/REMOVE
install -D -m 0644 packaging/fenris.conf build/stage/etc/fenris/fenris.conf
chmod 755 build/stage/INSTALL build/stage/REMOVE
xbps-create -A x86_64 \
-n fenris-$(FENRIS_VERSION)_$(XBPS_REVISION) \
-s "Fenris NVMe wear monitor" \
-S "NVMe wear monitor with persistent TUI" \
-m "Fenris Packaging <packaging@bongbetic.com>" \
-H "https://git.bongbetic.com/xavierk/Fenris" \
-l "MIT" \
-D "python3>=3.10 smartmontools>=0" \
-F "/etc/fenris/fenris.conf" \
build/stage
@echo "=== XBPS package built: fenris-$(FENRIS_VERSION)_$(XBPS_REVISION).x86_64.xbps ==="
sign-xbps: package-xbps
@echo "=== Signing XBPS package ==="
xbps-rindex --sign-pkg --privkey $(XBPS_SIGNING_KEY) \
fenris-$(FENRIS_VERSION)_$(XBPS_REVISION).x86_64.xbps
@echo "=== XBPS package signed ==="
xbps-publish:
bash scripts/xbps-publish.sh --publish
xbps-publish-dry-run:
bash scripts/xbps-publish.sh --dry-run
# ─── GPG key management ─────────────────────────────────────────────────────
generate-test-key:
@echo "=== Generating throwaway test GPG key ==="
@echo "This key is for CI/testing only — never use for real releases."
printf '%%no-protection\nKey-Type: RSA\nKey-Length: 3072\nName-Real: Fenris Packaging (TESTING ONLY)\nName-Email: packaging-test@bongbetic.com\nExpire-Date: 0\n%%commit\n' | \
gpg --batch --gen-key
@echo "=== Test key created. Fingerprint: ==="
@gpg --fingerprint packaging-test@bongbetic.com
# ─── Signing ────────────────────────────────────────────────────────────────
sign-rpm: package-rpm
@echo "=== Signing RPM payload ==="
@rpm --import packaging/keys/fenris-packaging.asc 2>/dev/null || true
rpmsign --addsign --define "_gpg_name $(PACKAGING_KEY)" \
dist/fenris-$(FENRIS_VERSION)-1.x86_64.rpm
@echo "=== RPM signed ==="
@rpm -Kv dist/fenris-$(FENRIS_VERSION)-1.x86_64.rpm
checksums: package
@echo "=== Generating SHA256SUMS ==="
cd dist && sha256sum fenris_$(FENRIS_VERSION)_amd64.deb \
fenris-$(FENRIS_VERSION)-1.x86_64.rpm > SHA256SUMS
@echo "=== SHA256SUMS written ==="
@cat dist/SHA256SUMS
clearsign: checksums
@echo "=== Clearsigning SHA256SUMS ==="
gpg --batch --yes --clearsign --local-user $(PACKAGING_KEY) \
dist/SHA256SUMS
@echo "=== SHA256SUMS.asc written ==="
# ─── Release (spec §5) ──────────────────────────────────────────────────────
release: package sign-rpm clearsign
@echo ""
@echo "=== Release v$(FENRIS_VERSION) ==="
@echo ""
@echo "Artifacts:"
@ls -la dist/fenris_$(FENRIS_VERSION)_amd64.deb \
dist/fenris-$(FENRIS_VERSION)-1.x86_64.rpm \
dist/SHA256SUMS.asc 2>/dev/null
@echo ""
@echo "Verify signing (manual):"
@echo " rpm -Kv dist/fenris-$(FENRIS_VERSION)-1.x86_64.rpm"
@echo " gpg --verify dist/SHA256SUMS.asc dist/SHA256SUMS"
@echo ""
@echo "Upload to registry:"
@echo " curl -X PUT -u user:token -T dist/fenris_$(FENRIS_VERSION)_amd64.deb \\"
@echo " 'https://git.bongbetic.com/api/packages/xavierk/debian/pool/bookworm/main/upload'"
@echo " curl -X PUT -u user:token -T dist/fenris_$(FENRIS_VERSION)_amd64.deb \\"
@echo " 'https://git.bongbetic.com/api/packages/xavierk/debian/pool/jammy/main/upload'"
@echo " curl -X PUT -u user:token -T dist/fenris_$(FENRIS_VERSION)_amd64.deb \\"
@echo " 'https://git.bongbetic.com/api/packages/xavierk/debian/pool/noble/main/upload'"
@echo " curl -X PUT -u user:token -T dist/fenris-$(FENRIS_VERSION)-1.x86_64.rpm \\"
@echo " 'https://git.bongbetic.com/api/packages/xavierk/rpm/fenris/upload'"
@echo ""
@echo "Create Gitea release with notes and attach:"
@echo " dist/fenris_$(FENRIS_VERSION)_amd64.deb"
@echo " dist/fenris-$(FENRIS_VERSION)-1.x86_64.rpm"
@echo " dist/SHA256SUMS.asc"
@echo ""
@echo "Key ceremony: delete the private key after upload."
@echo " See docs/install/signing-key-ceremony.md"
# ─── Automated release flow (issue #52) ──────────────────────────────────────
release-run:
bash scripts/release.sh --publish
release-dry-run:
bash scripts/release.sh --dry-run
clean:
@echo "=== Cleaning build artifacts ==="
rm -rf build/stage dist/fenris-*.deb dist/fenris-*.rpm dist/SHA256SUMS*
+309 -115
View File
@@ -1,157 +1,351 @@
<p align="center">
<picture>
<source srcset="assets/bongbetic-brand/wordmark-light.png" media="(prefers-color-scheme: dark)">
<img src="assets/bongbetic-brand/wordmark-dark.png" alt="Bongbetic" width="260">
</picture>
<br>
<sub>crafted with stubborn curiosity by <a href="https://bongbetic.com">Bongbetic</a></sub>
</p>
# Fenris 🐺
<p align="center">
<img src="assets/bongbetic-brand/b_glyph.svg" width="48" alt="Fenris glyph">
</p>
*Observes an NVMe drive's real-world use and translates that history into an understandable endurance outlook.*
<h1 align="center">Fenris 🐺 — Your SSD's Tell-All Diary</h1>
<p align="center">
<em>Your NVMe drive has been keeping secrets. Fenris makes it confess — in real time.</em>
<br>
<em>How much did you write today? How long until it taps out? No fairy dust — just your actual bytes.</em>
</p>
Fenris is a persistent TUI monitor backed by a short-lived privileged collector on the host's native scheduler. It reads SMART data every few minutes, stores compact observation history in SQLite, and recomputes a usage-adjusted theoretical lifespan on every screen render — no fairy dust, just your actual bytes.
---
Fenris is a tiny, stubborn daemon that eavesdrops on your NVMe drive's SMART gossip, writes it down every few minutes, and serves you a live dashboard that actually means something. Not "vibes" — **real GB written in the last 24 hours, real GB/hour, and a real countdown in hours, days, and years until your drive's endurance runs out**.
## Requirements
> Think of it as a Fitbit for your SSD. Except it doesn't nag you to drink water.
- **Python ≥ 3.10** (verified at install time)
- **smartmontools** (`smartctl` — verified at install time)
- **systemd** or **runit**, with a polkit agent (the collector runs as root; elevation is exclusively polkit)
## What it actually does (no hand-waving)
No other OS packages or Python dependencies beyond [Textual](https://textual.textualize.io/) (pinned in the lockfile).
- **Listens** — polls `smartctl -j` on your NVMe device (default every 5 minutes, you pick).
- **Remembers** — appends every sample to `data/history.jsonl` and rolls up per-hour totals into `data/hourly.jsonl` (survives restarts, rebuilds itself if you yank the power).
- **Calculates** — rolling 24-hour window: *exact* bytes written in the last 24h, GB/h, GB/day, implied total TBW from `percentage_used`, remaining TB, and a projected life-remaining breakdown. Warming-up badge until it has 24h of coverage — no fake confidence.
- **Shows off** — dense, live dashboard with wear-over-time + trailing-24h per-hour bars, sticky header, live countdown, and stale warnings if the daemon dozes off.
## Install from package (recommended)
## You need
### Debian / Ubuntu (apt)
- **Python 3.7+**
- **smartmontools** (`smartctl`)
- Root-ish access to read NVMe SMART (passwordless `smartctl` or just run with `sudo` — your call)
The Gitea instance Debian registry signs metadata with its own key. Verify the
instance key fingerprint (TOFU hardening):
### The sudo dance (one time)
```text
Fingerprint: F937E81D2FB0736B15BC611884BEBD586DFAC010
```
Fenris runs `sudo -n smartctl ...` so it doesn't get stuck asking for a password mid-nap:
Add the instance key and repository:
```bash
sudo visudo
# add this line (swap in your username):
youruser ALL=(root) NOPASSWD: /usr/sbin/smartctl
sudo mkdir -p /etc/apt/keyrings
sudo curl -fsSL -o /etc/apt/keyrings/gitea-xavierk.asc \
https://git.bongbetic.com/api/packages/xavierk/debian/repository.key
echo "deb [signed-by=/etc/apt/keyrings/gitea-xavierk.asc] \
https://git.bongbetic.com/api/packages/xavierk/debian bookworm main" \
| sudo tee /etc/apt/sources.list.d/fenris.list
sudo apt update && sudo apt install fenris
```
No sudo? Run the whole thing with `sudo` and it'll still behave.
Replace `bookworm` with your distribution codename (`bookworm`, `jammy`, or
`noble`).
## Get it running — 30 seconds
### Fedora / openSUSE Tumbleweed (RPM)
### The cozy way
Use the Fenris-owned repo file (not Gitea's auto-generated one):
```bash
./fenris.sh
# pick 1) Start monitoring → choose device / interval / port → done
sudo dnf config-manager --add-repo \
https://git.bongbetic.com/xavierk/Fenris/raw/branch/main/packaging/fenris.repo
sudo dnf install fenris
```
### The no-nonsense way
On openSUSE Tumbleweed, add the same standard RPM repository file and install
with zypper:
```bash
python3 fenris.py start # defaults: /dev/nvme0, every 300s, port 8420
python3 fenris.py start --interval 60 --port 9000 # if you're impatient
python3 fenris.py status # "are we live? how's the drive?"
python3 fenris.py sample # one sneaky sample right now
python3 fenris.py stop # tuck it back in
sudo zypper addrepo --refresh \
https://git.bongbetic.com/xavierk/Fenris/raw/branch/main/packaging/fenris.repo fenris
sudo zypper install fenris
```
Dashboard lives at **http://localhost:8420** (or whatever port you chose).
The repo file sets `gpgcheck=1` against the Fenris packaging key (downloaded
from the raw URL in `gpgkey`) and `repo_gpgcheck=0` (metadata check left to
TLS).
## The menu, demystified
### Void Linux (XBPS)
Run `./fenris.sh` and you'll get:
Void x86_64 with glibc and runit is the native target. Its signed XBPS channel
is available from the permanent repository below. Add it, refresh its
metadata, and install the released package:
```
1) Start monitoring (background daemon + dashboard)
2) Stop monitoring
3) Status / current wear stats
4) Take one sample right now
5) Open dashboard URL
---
h) Help / how this works
q) Exit (go touch grass)
```bash
sudo install -d -m 0755 /etc/xbps.d
echo 'repository=https://git.bongbetic.com/xavierk/Fenris-xbps/raw/branch/stable/x86_64' \
| sudo tee /etc/xbps.d/fenris.conf
sudo xbps-install -M -S fenris
```
## What Fenris jots down
XBPS requires remote repositories to be signed. On the first refresh it
displays the repository signing key embedded in the signed metadata; accept it
only when its RSA SHA256 fingerprint is
`SHA256:AvPMRlKMikPg75u0iKr8AUkxlfU/Ad4k/S4o2M9W4/w`. The public key is also
available at
`https://git.bongbetic.com/xavierk/Fenris-xbps/raw/branch/stable/keys/fenris-xbps-signing.pub`.
For later updates, always refresh first so XBPS fetches the current index:
| Field | What's the gossip? |
|-------|---------------------|
| `percentage_used` | The drive's own wear-o-meter (0–100%) |
| `bytes_written` / `bytes_read` | Lifetime totals — the receipts |
| `available_spare` | Spare blocks left (%) |
| `media_errors` | Uncorrectable boo-boos |
| `power_on_hours` | How long it's been awake |
| `temperature_c` | Is it sweating? |
| `critical_warning` | NVMe's panic flags |
```bash
sudo xbps-install -M -Syu
```
Hourly rollups also stash `bytes_written` per hour, `pct_start`/`pct_end`, and temp peaks — so the 24h math stays honest.
The `-M` flag bypasses XBPS's on-disk repodata cache. It is required when
checking for a newly published package through Gitea's cached raw-file URL.
## The dashboard — what's on screen
The runit service remains dormant after installation. `fenris monitor resume`
creates `/var/service/fenris-collect`; pause removes that link and records a
deliberate disable in the observation history.
- **Hero card: Projected life remaining** — big, friendly `361 d 2 h` (plus `≈ 361 days · ≈ 8666 hours · ≈ 0.99 years`), backed by `~280 GB/day` and `~101 TB left of ~202 TB total` on the test box.
- **Data written (24h)** — exact GB in the rolling window + coverage (`10.4h of 24h` until warmed up).
- **Write rate** — GB/h and GB/day, live.
- **Wear, spare, temp, errors, power-on** — the usual suspects, with progress bars and polite color-coding.
- **Two charts, side by side:** wear over time + trailing-24h hourly write bars (with a cheeky "now" bar for the current partial hour).
- **Live plumbing:** polling synced to your interval, ETag-cached, countdown to next sample, warming-up + stale banners, pauses when you hide the tab (saves your battery, you're welcome).
Fenris keeps the observation store root-written and readable by the `fenris`
group. Add each TUI user to that group, then start a new login session before
running Fenris:
**API for the tinkerers:** `GET /api/data` · `/api/hourly` · `/api/summary` · `/api/config` · `/api/status` — all JSON, all friendly.
```bash
sudo usermod -aG fenris "$USER"
```
### Package signature verification
The RPM payload is signed with the Fenris packaging key (RSA 3072).
Verification happens automatically via dnf's `gpgcheck=1`. For manual
verification of downloaded assets:
```bash
rpm -Kv fenris-*.x86_64.rpm # RPM payload signature
gpg --verify SHA256SUMS.asc SHA256SUMS # Clearsigned checksum manifest
sha256sum -c SHA256SUMS # Checksum match
```
The packaging public key is published in-repo — no keyservers. See
`packaging/keys/fenris-packaging.asc` and
`docs/install/signing-key-ceremony.md` for key lifecycle details.
### Dormant install
A fresh package install is fully dormant. Its native scheduler is present but
disabled; nothing runs. The only opt-in is the sanctioned toggle:
```bash
fenris monitor resume # enable scheduling + open first monitoring period
fenris monitor pause # close the period, disable scheduling
```
## Development install (make install)
For contributors building from source:
```bash
sudo make install
```
This builds a wheel, installs its locked pure-Python runtime packages into
`/opt/fenris/vendor`,
and places helpers, units, and the polkit policy. Units are dormant by default.
```bash
sudo make upgrade # re-sync wheel, units, schema
make uninstall # removes artifacts, preserves config and store
make purge # also removes /etc/fenris and /var/lib/fenris
```
## Upgrade
### Package upgrade
```bash
sudo apt update && sudo apt upgrade fenris # Debian/Ubuntu
sudo dnf upgrade fenris # Fedora
sudo xbps-install -Syu # Void Linux
```
### Development upgrade
```bash
sudo make upgrade
```
What it does:
1. Snapshots `observations.db` to a one-generation backup (`.bak`).
2. Replaces the locked runtime packages under `/opt/fenris/vendor`.
3. Syncs units and polkit against the manifest; runs `daemon-reload`.
4. Restarts the timer **only** if unit contents changed **and** it is active — a running collection run finishes on its mapped interpreter; the next run uses the new code.
5. Applies forward-only schema migrations (the store directory is never rebuilt; automatic downgrade does not exist).
Rollback: reinstall the previous version and restore `observations.db.bak`.
On Void, pause monitoring first, copy the compatible snapshot back to
`/var/lib/fenris/observations.db`, then force-install the matching older
package version. If that version is no longer indexed, add its retained XBPS
archive to a local repository with `xbps-rindex -a` and use
`xbps-install -R <local-repository> -f fenris-<version>`. Installing an older
package over a newer observation store is unsupported.
## Migration from make install
If Fenris was previously installed with `sudo make uninstall` first, then
installed from the package, existing config, store, and group survive by path
continuity. Over-installing the package over a `make install` is
**forbidden** — stale units shadow vendor placement. See
[docs/install/migrate-from-makeinstall.md](docs/install/migrate-from-makeinstall.md).
## Uninstall and purge
### Package removal
```bash
sudo apt remove fenris # preserves config and store
sudo apt purge fenris # also removes config and store
sudo dnf remove fenris # preserves config and store
sudo xbps-remove fenris # preserves config and store
```
### Development removal
```bash
make uninstall # removes artifacts, preserves config and observation history
make purge # also removes /etc/fenris and /var/lib/fenris
```
Uninstall performs the sanctioned disable first (`fenris-monitor disable --now`) — an open period closes `user_disabled` — then removes the runtime packages, helpers, units, polkit policy, and wrapper while keeping `/etc/fenris` and the observation store. Reinstalling resumes from the preserved store.
## Cadence drop-ins
The default collection cadence is **3 minutes** (`OnUnitInactiveSec=3min` in the timer unit). To change it, place a systemd drop-in:
```bash
sudo systemctl edit fenris-collect.timer
# Add:
# [Timer]
# OnUnitInactiveSec=10min
```
No interval key exists in `/etc/fenris/fenris.conf`. Cadence is a systemd concern, not a Fenris configuration key.
On Void, Fenris uses its native runit service instead: its initial collection
is delayed by two minutes and later collections run three minutes after the
previous run finishes. Inspect its state and diagnostics with:
```bash
sv status fenris-collect
sudo tail -n 50 /var/log/fenris-collect/current
```
`fenris status` also reports the separate boot-enabled, runtime-active,
collection outcome, and observation-store freshness facts. A failed collection
is retried at the next interval; it never fabricates missing observations.
## CLI reference
| Command | Behavior |
|---|---|
| `fenris` | Opens the TUI (no arguments). |
| `fenris status` | Projection facts, enabled/active state, last collect outcome, journal hint on failure or staleness. Never auto-samples. |
| `fenris sample` | On-demand collection via the privileged helper. Blocks until the run completes. |
| `fenris monitor pause` | Sanctioned disable — asks for confirmation, then disables native scheduling and closes the monitoring period. |
| `fenris monitor resume` | Sanctioned enable — enables native scheduling and opens a monitoring period. No confirmation. |
| `fenris baseline set <json>` | CLI-side validation, then polkit-guarded persistence. |
| `fenris baseline clear` | Remove the endurance baseline. |
| `fenris import <path>` | Idempotent single-transaction legacy import. |
| `fenris start` / `stop` / `run` | Rejected with a one-line migration pointer — never aliased. |
| `fenris --device` | Rejected with a pointer to the configuration file. |
## Reading the dashboard
![Chalktone dashboard with a dotted activity plot](assets/dashboard-chalktone.png)
Preview uses synthetic observations, not measurements from a real drive.
The Chalktone dashboard opens with a large **Live** activity plot. **Day** shows
hourly evidence for a selected date, labelled UTC, and **History** shows daily
evidence. Local-day totals keep their recorded timezone. Dotted traces show
measured read/write volumes, not transfer speed; missing evidence breaks the
trace. `?` marks a gap, `~` a partial total, and `u` unallocated daily volume.
Use the selected-point readout for exact values and evidence state.
Local-day totals use measured intervals inside each recorded local date. An
interval that crosses local midnight appears once as shared evidence and stays
outside both known totals. The dashboard labels current totals as so far and
shows incomplete or unavailable dates without treating them as zero. Older
UTC-only summaries cannot establish exact local-day totals.
- `v` cycles Live / Day / History; the tabs are also clickable.
- `←` / `→` inspect points; `w` switches read/write volume in every view.
- `[` / `]` browse dates, `g` enters a date, and `t` returns to today/live.
- `Tab` / `Shift+Tab` move focus; `z` expands the focused panel, and `z` or
`Esc` restores it. Monitoring status and controls remain visible.
- `s` cycles Chalktone, Amber, Nord, and High Contrast; saved theme preferences
survive upgrades. `m` toggles reduced motion.
On smaller terminals, textual summaries and scrollable panels keep evidence
accessible. Pause, resume, collect, disclosures, help, and quit remain available
in the fixed control row.
Run `fenris` as your normal user to open the TUI dashboard. The dashboard does
not need `sudo`. Use `sudo` for package installation and system configuration;
pause, resume, and collect-now actions normally authenticate through polkit.
If the observation store is inaccessible, add your login user to the `fenris`
group with `sudo usermod -aG fenris "$USER"`, then log out and back in.
If polkit authentication is unavailable, quit the dashboard and run only the
required administrative action in your terminal:
```bash
sudo fenris monitor resume # Enable monitoring now and across reboots
sudo fenris monitor pause # Confirm a deliberate monitoring pause
sudo fenris sample # Request one collection run
```
Reopen the dashboard with `fenris` afterward. Press `?` for these instructions
and keyboard controls at any time; use the arrow keys to scroll and `Esc` to close.
The CLI and TUI share the same authenticated action path. Authentication and
waiting for collection have no separate dashboard deadline; the native collector
enforces its 90-second runtime limit. An interrupted or failed action is not
automatically retried—check `fenris status` before retrying.
- **Continuity** — the service strip's continuity line (and `fenris status`) reports whether monitoring survives reboots: `monitoring: active in background · persists across reboots`, or `monitoring: does not start on next boot`.
- **Paused vs. quit** — a full-width `monitoring: paused — deliberate disable` block means collection is stopped (`fenris monitor pause`); resume with `fenris monitor resume`. Pressing `q` only leaves the screen — monitoring keeps running in the background.
- **Auth banner** — the launch notice explains normal-user startup and polkit authentication, then clears on the first refresh. The `?` help screen remains available.
Per-release notes live on the [releases page](https://git.bongbetic.com/xavierk/Fenris/releases): each entry is the version's `CHANGELOG.md` section — what was added, changed, and fixed — plus standing install and verification instructions.
## Retired menu options
The legacy `fenris.sh` menu script and the `fenris.py` monolith have been removed. Here's where the old options went:
| Legacy option | Successor |
|---|---|
| 1) Start monitoring | `fenris monitor resume` |
| 2) Stop monitoring | `fenris monitor pause` |
| 3) Status / current wear stats | `fenris status` |
| 4) Take one sample right now | `fenris sample` |
| 5) Open dashboard URL | Removed — the HTML dashboard and HTTP server are gone; the TUI is the primary interface. |
## Configuration
`/etc/fenris/fenris.conf` holds exactly one key — the device selector:
```
device = /dev/disk/by-id/nvme-Samsung_SSD_980_PRO_2TB_S6BENS0Txxxxx
```
Use a stable `/dev/disk/by-id/` path. Raw `/dev/nvmeX` paths are warned against. The file is re-read every collection run.
## Where's my stuff?
```
fenris/
├── fenris.py # the whole show — daemon + server + math
├── fenris.sh # the cozy menu
├── README.md # hi — you're here
├── assets/bongbetic-brand/ # Bongbetic wordmarks & glyphs (for Gitea + dashboard)
└── data/
├── history.jsonl # raw samples (JSONL, append-only)
├── hourly.jsonl # per-hour rollups (auto-rebuilt on restart)
├── fenris.pid # daemon PID
└── fenris.log # daemon chatter
```
## CLI cheat sheet
```bash
python3 fenris.py start [--device /dev/nvme0] [--interval 300] [--port 8420]
python3 fenris.py stop
python3 fenris.py status
python3 fenris.py sample [--device /dev/nvme0]
python3 fenris.py run # foreground mode — what `start` spawns internally
```
## Oops — troubleshooting without the tears
**"smartctl not found"**
```bash
sudo apt install smartmontools # Debian/Ubuntu
sudo pacman -S smartmontools # Arch — you already knew
```
**"needs root" / permission denied**
Set up the passwordless line above, or just `sudo ./fenris.sh`.
**Dashboard says "stale"**
Daemon napped or crashed. `python3 fenris.py status` will tell you. Kick it again with `start`.
**Only 10 hours of data and it says "preliminary"?**
That's honesty, not a bug. It needs 24h of real writes to give a tight estimate. Let it simmer — the number gets sharper every hour.
| Artifact | Package install | make install |
|---|---|---|
| Wrapper | `/usr/bin/fenris` | `/usr/local/bin/fenris` |
| Helpers | `/usr/libexec/fenris/` | `/usr/libexec/fenris/` |
| Units | `/usr/lib/systemd/system/` (vendor) | `/etc/systemd/system/` |
| Polkit policy | `/usr/share/polkit-1/actions/` | `/usr/share/polkit-1/actions/` |
| sysusers/tmpfiles | `/usr/lib/{sysusers,tmpfiles}.d/fenris.conf` | managed by Makefile |
| Configuration | `/etc/fenris/fenris.conf` | `/etc/fenris/fenris.conf` |
| Observation store | `/var/lib/fenris/observations.db` | `/var/lib/fenris/observations.db` |
| Runtime packages | `/opt/fenris/vendor` | `/opt/fenris/vendor` |
| Legacy history | — | `./data/history.jsonl` (auto-imported) |
---
Binary file not shown.

After

Width:  |  Height:  |  Size: 296 KiB

+46
View File
@@ -0,0 +1,46 @@
# 1. Observation store: a single SQLite database
## Status
Accepted — resolves [Define the persistent observation store and legacy migration](https://git.bongbetic.com/xavierk/Fenris/issues/2) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
Amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12): the `endurance_baseline` field set and validation contract — derived verification, entry-time unprivileged sysfs validation, and read-time controller-segment applicability.
Amended by [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14): the `controller_segments` metadata snapshot — normalized identity diagnostics plus `vid`/`ssvid`/`transport` and a degraded flag, frozen at segment open, all nullable.
## Context
Fenris today persists full SMART samples to an append-only `data/history.jsonl` beside a derived `data/hourly.jsonl`, both in the checkout, with no schema versioning and silent skipping of malformed lines. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer and an unprivileged TUI ([lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md)), and projects a usage-adjusted theoretical lifespan from Data Units Written over wall-clock time with categorical confidence ([endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md)). The store must support a root writer appearing every few minutes while an unprivileged reader queries concurrently, must migrate the legacy observation history idempotently and interruption-safely, and must version its schema.
## Decision
1. **Substrate**: one SQLite database in WAL mode at `/var/lib/fenris/observations.db`. WAL gives the unprivileged reader a consistent snapshot while the collector writes; migration and schema changes are single transactions.
2. **Access**: the database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI opens it read-only. No `/run` snapshot or export layer.
3. **Entities**:
- `samples` — recent raw SMART samples: timestamp, controller identity, raw `data_units_written`/`data_units_read` integers, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning`.
- `hour_observations` — one row per UTC hour: the usage-habit split (`seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown`), DUW/DUR deltas, temperature min/avg/max, sample count, coverage flag. Classification thresholds belong to the projection model, not the store.
- `day_aggregates` — one row per UTC day; the habit-evidence grain.
- `monitoring_periods` — `started_at`, `ended_at` (NULL = open), `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not.
- `controller_segments` — boundaries where controller identity changes or DUW decreases; write deltas are never computed across a segment. Each row carries a metadata snapshot frozen when the segment opens and immutable thereafter ([Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14)): normalized `subnqn`, `sn`, `mn`, `fr`, plus `vid`, `ssvid`, `transport`, and an `identity_degraded` flag — human diagnostics, never key components (`cntlid` excluded: it distinguishes controllers within one subsystem, out of scope for a single-drive monitor). `fr` may go stale after a mid-segment firmware update; counter discontinuities belong to the DUW-monotonic axis. Every metadata column is nullable — legacy-imported segments carry `mn` with NULLs, degraded segments whatever was observed — so incompleteness stays explicit.
- `endurance_baseline` — one active row, replaced on edit ([Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12)): the rated-TBW value in bytes (`E_rated = entered_TBW × 10¹²`) plus mandatory provenance — source URL, document revision, entry date, model string, nominal capacity — and frozen validation facts (detected model, detected capacity bytes, `validated_by` `machine`/`user`, `validated_at`). Verification is derived at read — complete provenance and a drive match (machine or attested), never a stored boolean; incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier. Entry validation is an unprivileged live sysfs read of the configured device (normalized model containment with an interactive confirm recorded as `validated_by = user`; capacity within ±1%); at projection time applicability is a model match against the current controller segment, and a mismatch is retained — never auto-deleted — leaving the projection Unavailable.
- Projections are not stored; they are recomputed on read. There is no separate latest-status table.
4. **Day boundary**: UTC, matching hours, so day derivation from hour rows is monotonic and DST-ambiguous or 23/25-hour days never exist in the store.
5. **Retention**: raw samples are kept 14 days and pruned opportunistically by the collector; hour observations and day aggregates are retained indefinitely.
6. **Migration** (first new-version collection run):
1. If the database already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples, derive hour observations and day aggregates from them, and ignore `hourly.jsonl` as derived data (diff and log mismatches; do not trust).
3. One implicit `monitoring_periods` row opens at the first legacy sample and closes with `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts (samples present ⇒ powered on; DUW deltas ⇒ writes occurred).
4. The import is a single transaction: interruption leaves the database fully pre- or post-migration.
5. Only after commit are legacy files renamed to `*.migrated` (never deleted).
6. Malformed legacy lines are quarantined with a logged count, never silently dropped.
7. **Projection inputs**: the `endurance_baseline` table lives in the database and is edited via the CLI; `/etc/fenris/` holds only operational configuration.
8. **Versioning**: `PRAGMA user_version` plus ordered migration steps in code, each in its own transaction; the collector refuses to run against an unknown newer version.
9. **Collector health**: not stored. Failures go to the journal (per the lifecycle decision); the freshest sample timestamp is the store's own staleness signal.
## Consequences
- Backups and state migration are copying one file (plus its WAL sidecars).
- SQLite becomes a runtime dependency of both the collector and the TUI (Python `sqlite3` stdlib suffices; no server).
- The collector's prune, import, and version steps are all transactional, so a killed timer run cannot leave partial state.
- Legacy checkout-relative `data/` files stop being authoritative at migration; the migration ticket's rename-after-commit rule keeps them as a recovery trail.
- The active/idle/powered-off classification contract with the projection model is the `hour_observations` column set, keeping storage and model decisions separable.
@@ -0,0 +1,52 @@
# 2. Projection model: sustained-regime rate with categorical confidence
## Status
Accepted — resolves [Define the lifespan projection and confidence model](https://git.bongbetic.com/xavierk/Fenris/issues/4) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
Amended by [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15): a blank (degraded) identity key caps confidence at Limited evidence, and identity-change semantics extend verbatim to blank keys.
## Context
Fenris's current `compute_summary` projects from a single trailing-24-hour write rate against endurance inferred as `DUW / Percentage Used` or synthesized as `capacity × 600`, alongside a second linear regression of Percentage Used toward 100. The [endurance research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/nvme-endurance-signals/docs/research/nvme-endurance-signals.md) established which signals can defensibly support a projection, and [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store while leaving classification thresholds and every projection rule to this model. This decision defines the algorithm and the user-facing contract the TUI consumes.
## Decision
1. **One projection.** The usage-adjusted theoretical lifespan is computed once, against the endurance baseline chosen by precedence (verified rated TBW → unverified manual override → Percentage-Used-implied → projection unavailable). Percentage Used is context, never a second projection: it renders as a vendor wear line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The current PU-slope regression (`wear_days`) and the `capacity × 600` synthesis are dropped.
2. **Headline rate from the sustained regime.**
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
The default regime is the full observation history capped at 90 days. The 7-, 28-, and 90-day rates are computed independently of the regime and shown as a **scenario range**; only horizons the history actually covers appear (no placeholders).
3. **Habit change.** A change is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for 3 consecutive days. The new regime starts at the first day of divergence and is adopted automatically, labeled "usage habit changed N days ago"; the scenario range keeps the longer horizons visible. A regime younger than 7 days caps projection confidence at Limited evidence.
4. **Hour classification** (named constants, no configuration surface):
- **Powered-off**: the hour's power-on-hours delta is below 90% of its wall-clock span.
- **Active**: DUW delta ≥ 256 MiB in the hour.
- **Idle**: powered on, sampled, below the active threshold.
- **Unknown**: everything else — unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable), or inconsistent counters.
- Disabled time is not an hour state: it is wall-clock outside monitoring periods.
5. **Denominator.** Wall-clock seconds inside monitoring periods, including powered-off and unknown time. Disabled periods are excluded from numerator and denominator. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage.
6. **Minimum evidence.** Warming up until there are 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage. The projection still renders while warming up, labeled with its facts. Unavailable conditions (no baseline, unsupported DUW, zero rate over the regime, identity change) render no lifespan number.
7. **Staleness.** A newest day aggregate older than 48 hours drops confidence one level (Supported → Limited) and is shown as a contributing fact.
8. **Confidence rule table.**
- **Unavailable**: no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported**: verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80% **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50% of trailing 28-day bytes **and** regime ≥ 7 days old **and** the current controller segment's identity key is not degraded.
- **Limited**: every other case with a baseline and a positive rate; the failing facts are shown.
- **Degraded identity** ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15)): a controller segment whose identity key is blank — every rung of the key ladder empty — is identity-degraded. Supported is unreachable while the current segment is degraded, because a blank key cannot detect a replacement; the fact "controller identity unavailable — replacement detection relies on write-counter continuity only" renders with every state, and the cap combines idempotently with the staleness drop (both land at Limited). Ephemeral markers (model "Linux", non-pcie transport) are segment metadata, never confidence facts.
- Confidence always renders as state plus contributing facts, never a percentage.
9. **Segment breaks.** A DUW decrease with unchanged controller identity quarantines nothing: prior day aggregates remain habit evidence and the projection is Unavailable only until the new segment re-warms. A controller-identity change quarantines prior history from projection entirely — it describes a different drive. Degraded keys get no special casing ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15)): a blank-key segment is marked identity-degraded and segmented by DUW monotonicity alone, any visible change of the recorded key — including to or from blank — is a controller-identity change and quarantines, and equal blank keys continue the segment. Since even a degraded→healthy transition quarantines, the projection window only ever spans segments sharing one key, so a degraded current segment needs no cross-segment propagation rule.
10. **Implied-baseline eligibility.** The Percentage-Used-implied baseline is computed only after ≥ 2 Percentage Used increments within the current controller segment; until then the projection is Unavailable with "vendor wear estimate too coarse to imply endurance".
11. **Uncertainty.** The scenario range is the only spread shown; no statistical confidence interval appears anywhere. Zero rate → "no finite projection from this history", never infinity or zero.
12. **Language.** The endurance research's required wording and six disclosures are adopted verbatim as the specification's language section.
13. **Contract.** The projection function hands the TUI: the confidence state, the contributing facts — including the degraded-identity fact when the current segment's key is blank — the headline remaining time when one exists, the scenario range, the Percentage-Used context line, and the disclosure text. Projections are recomputed on read, never stored.
## Consequences
- The TUI information-architecture prototype (its ticket) consumes a fixed contract rather than inventing presentation states.
- `compute_summary`'s wear-slope regression and capacity-synthesized endurance disappear; migration must not synthesize baselines for legacy history.
- Coverage becomes a first-class displayed fact rather than an internal heuristic.
- All guardrail thresholds live as documented constants in one projection module; tuning demand, if it ever appears, is a future decision rather than a config surface.
- Two follow-on decisions surfaced and are ticketed separately: the controller-identity key that segments history, and endurance-baseline provenance validation.
@@ -0,0 +1,33 @@
# 3. Service lifecycle: timer-driven collection with a sanctioned control path
## Status
Accepted — resolves [Define the collector, service, and CLI lifecycle](https://git.bongbetic.com/xavierk/Fenris/issues/8) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amends the toggle mechanism of [Verify systemd lifecycle and privilege constraints](https://git.bongbetic.com/xavierk/Fenris/issues/7); its spirit — scoped, explicit, authenticated, no generic `manage-unit-files` grant — is intact.
Amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12): the helper gains a `baseline` verb that persists the CLI-validated endurance-baseline row — same fixed-operation, polkit-mediated pattern.
## Context
Fenris's current single process combines daemonization, a PID file, an HTTP dashboard, and control (`fenris.py start/stop/status/sample`) over checkout-relative state. [ADR 0001](0001-observation-store-sqlite.md) fixed the observation store, including `monitoring_periods` whose `user_disabled` end cause records deliberate pauses, and the [systemd lifecycle research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/systemd-privilege-lifecycle/docs/research/systemd-privilege-lifecycle.md) fixed the timer + oneshot architecture, standard paths, journal diagnostics, allow-listed status reads, and polkit-mediated startup toggles — while leaving cadence mechanics, the configuration surface, CLI compatibility, staleness thresholds, and the mechanism that records a deliberate disable open. In particular, `systemctl enable`/`disable` cannot write a monitoring-period row, so a direct-systemctl toggle cannot satisfy the store's semantics.
## Decision
1. **Units.** Two system units only: `fenris-collect.timer` (`WantedBy=timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code). The TUI and CLI are ordinary unprivileged processes and never units. There is no `/run/fenris` coordination surface: systemd serializes runs, the observation store holds state, and failures go to the journal per [ADR 0001](0001-observation-store-sqlite.md).
2. **Cadence.** Default three minutes: `OnBootSec=2min`, `OnUnitInactiveSec=3min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no`, no suspend catch-up (absent hours classify through power-on-hours evidence), `TimeoutStartSec=90s` so a hung interrogation fails visibly. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); no interval key exists in configuration.
3. **Configuration.** `/etc/fenris/fenris.conf` holds exactly one key: the device selector, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run, so there is no reload path to design. An invalid selector is a bounded failed run — journal plus failed unit result, retried next interval; `status` and the TUI also read the world-readable file directly and surface a `configuration error: <reason>` fact.
4. **Entry points.** Two privileged binaries: `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable` and `disable` with optional `--now`, plus the collect trigger, monitoring-period bookkeeping, and `baseline set`/`baseline clear` persistence for the CLI-validated endurance baseline; the only binary the polkit policy authorizes). One unprivileged `fenris` for humans: no arguments opens the TUI; subcommands (`status`, `sample`, `monitor pause`, `monitor resume`) are the CLI.
5. **Sanctioned toggle.** Pause = `disable --now`; Resume = `enable --now`; both executed by `fenris-monitor`, which performs the systemctl operation and the monitoring-period bookkeeping in one step, under polkit action `com.bongbetic.fenris.monitor` (`auth_admin`, covering the collect trigger too). Root invokes the helpers directly; where no polkit agent exists the operation fails cleanly and prints the root equivalent. This amends the research's direct-systemctl toggle: a period boundary cannot be recorded by systemctl, so the toggle must be Fenris's own fixed operation.
6. **Period rows.** Idempotent matrix: a first-ever enable opens a period at the enable moment (hours before the first successful sample are unknown-but-inside, correctly so when the device errors); a resume with an open period — a raw `systemctl stop` intervened — changes no row, the gap remaining inside as unknown seconds; a resume with no open period opens a new row at the resume moment; a pause with an open period closes it `user_disabled` at the pause moment; a pause otherwise is a no-op. A raw stop or disable outside the helper is an unexplained gap, never `user_disabled`: only the sanctioned path can record intent.
7. **On-demand collection.** `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits, and the outcome (freshness line or journal hint) is reported synchronously. No code path outside `fenris-collect` touches the device; the TUI never samples in-process; no confirmation is required.
8. **TUI controls.** Pause asks for confirmation; Resume does not (benign — friction invites raw-systemctl escapes). Boot enablement and current runtime activity are always displayed as separate facts, next to last collect outcome and freshness. No bare start/stop exists anywhere.
9. **CLI compatibility.** `status` is a pure read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness; it never auto-samples and never prompts. `sample` is retained via the helper path; `--device` is rejected with a pointer to the configuration file. `start`, `stop`, and `run` are rejected with one-line migration pointers, not aliased — an alias would silently change meaning. `fenris.sh` is retired: not shipped, removed from the repository, and the README maps its five menu options to their successors.
10. **Freshness constants.** Documented once, consumed by TUI and CLI alike: fresh means the newest sample is within 2× cadence + `AccuracySec` + 60 s; between that and 48 h the store is missed (a contributing fact); at ≥ 48 h it is stale, matching [ADR 0002](0002-projection-model-sustained-regime.md)'s evidence gate; an empty store reads "no observations yet" with an enable hint.
## Consequences
- Polkit ships one Fenris-specific policy authorizing exactly one fixed-operation binary; the collector itself is never polkit-reachable.
- Monitoring-period boundaries are exact at toggle moments; approximation never enters the habit record.
- Interval tuning is a systemd drop-in documented in the README; `/etc/fenris` stays a one-key file.
- Headless administration has full parity: every TUI action has a CLI twin.
- The TUI must run privileged operations through a terminal-attached subprocess so the platform polkit agent can prompt; the TUI prototype ticket validates this in practice.
- Nothing survives of the prototype's daemonization, PID files, or HTTP server; their commands fail with pointers instead of quiet behavior changes.
@@ -0,0 +1,31 @@
# 4. Installation lifecycle: Makefile-delivered venv, dormant install, sanctioned teardown
## Status
Accepted — resolves [Define installation, upgrade, and removal behavior](https://git.bongbetic.com/xavierk/Fenris/issues/9) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). Amended by [ADR 0007](0007-package-delivery-amends-0004.md): package delivery replaces `make install` as primary; layout/ownership and maintainer-script mechanics per 0007. Runtime semantics (dormant install, polkit-only elevation, snapshot + forward-only migration, one-generation rollback) unchanged.
## Context
Fenris runs today from the source checkout (`fenris.py`, `fenris.sh`, `data/`): code, state, and control all live relative to wherever the checkout sits. The redesign fixes system artifacts — helpers in `/usr/libexec/fenris` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), the observation store at `/var/lib/fenris/observations.db` ([ADR 0001](0001-observation-store-sqlite.md)), configuration at `/etc/fenris/fenris.conf` ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)), Textual on Python 3.9+ ([framework decision](https://git.bongbetic.com/xavierk/Fenris/issues/6)) — but nothing says how those artifacts are delivered, upgraded, or removed, or what happens when the checkout moves or disappears.
## Decision
1. **Delivery.** `sudo make install` builds a wheel from the checkout and installs it, with pinned dependencies, into a dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only: after install, nothing references it.
2. **Layout and manifest.** Units in `/etc/systemd/system` (`fenris-collect.{timer,service}`, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)); helpers in `/usr/libexec/fenris`; polkit policy in `/usr/share/polkit-1/actions/`; configuration and observation store in their ADR-fixed locations. The installer records every file it places in an explicit manifest consumed by upgrade and uninstall.
3. **Privilege.** One root installer (`sudo make install`); at runtime, elevation is exclusively polkit (`auth_admin`, `fenris-monitor` only, [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)). The installer never enables or starts units.
4. **Dormant install.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle (`fenris monitor resume [--now]`, or the first-run TUI prompt), which enables the timer and opens the first period in one step.
5. **Legacy import.** The installer detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs [ADR 0001](0001-observation-store-sqlite.md)'s idempotent single-transaction import, and reports imported counts — existing observations never depend on checkout survival. `fenris import <path>` remains available for later finds.
6. **Upgrade.** `sudo make upgrade` builds and installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed and it is active — safe with `Persistent=no`), leaves timer state untouched, and never kills an in-flight collection run: a running oneshot finishes on its mapped interpreter, so at worst one old-code run completes to the store and the next run uses the new code. It then applies forward-only observation-store schema migrations governed by a `schema_version` table. `/var/lib/fenris` is never rebuilt.
7. **Rollback.** Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade is explicitly unsupported.
8. **Removal.** `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. `make purge` additionally removes configuration and store. Journal entries age out naturally.
9. **Dependencies.** Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing.
10. **Scaffolding and floor.** The installer creates `/var/lib/fenris` with [ADR 0001](0001-observation-store-sqlite.md)'s root-written group-read permissions and verifies `python3 ≥ 3.9`, failing cleanly otherwise — the Textual contingency becomes an install-time gate rather than a runtime crash. The database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet.
## Consequences
- Installed Fenris survives checkout deletion; the checkout is only where builds happen.
- Teardown preserves monitoring-period semantics: deliberate removal excludes the uninstalled span from the usage habit instead of leaving it as unknown-inside.
- Installs are reproducible; dependency drift cannot ride in on an upgrade.
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
- Rollback support is exactly one generation deep, no further.
- The README documents install, upgrade, uninstall/purge, and legacy import alongside [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s menu-successor mapping.
@@ -0,0 +1,27 @@
# 5. Failure and recovery: visible degradation, never fabrication
## Status
Accepted — resolves [Define failure and recovery behavior](https://git.bongbetic.com/xavierk/Fenris/issues/10) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The observation store ([ADR 0001](0001-observation-store-sqlite.md)) and the service lifecycle ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) settled single-writer transactions, bounded single failed runs, freshness grading (fresh / missed / stale), and absent-hour classification through power-on-hours evidence. Left open by the [failure ticket](https://git.bongbetic.com/xavierk/Fenris/issues/10): behavior per failure class — malformed observations inside the store, missed observations, store faults (unreadable, corrupt, or newer-schema database), and repeated collector failures — and how the habit record re-anchors after the store itself is lost.
## Decision
1. **Malformed observations — refuse at the write boundary.** The collector validates every row it would write against the store's domain invariants (hour seconds sum to 3600, non-negative DUW delta within a controller segment, coverage consistent with sample count). A violating run writes nothing for that run, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact, but under a single trusted writer they should never see one. Store invariant: everything persisted is well-formed.
2. **Missed observations — never backfill.** Fenris never interpolates, estimates, or fabricates an hour. Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)'s power-on-hours classification is the only inference admitted.
3. **Store faults — degrade, never recreate over.** An unreadable or corrupt database is a store fault: readers surface a "observation store unreadable" fact with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and never recreates or overwrites an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside, the next run starts a fresh store, and if the legacy import never completed, the still-present `history.jsonl` is re-imported. No built-in destructive command exists.
4. **Newer schema — readers refuse symmetrically.** The TUI and `status` detect a `user_version` newer than they understand and display "observation store written by a newer Fenris — upgrade Fenris" without partial interpretation, matching the collector's refusal in [ADR 0001](0001-observation-store-sqlite.md) and the forward-only upgrade rule of [ADR 0004](0004-install-upgrade-removal-lifecycle.md).
5. **Repeated collector failures — flat cadence, no escalation.** The timer's retry is the recovery path; the settled freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff and no notification machinery; a persistent failure reads as stale exactly like any other gap.
6. **Drive-reported anomalies — facts, not alerts.** `critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status`; no alerting or notification surface exists. Fenris observes and projects; it does not alarm. The projection is unaffected: endurance math consumes writes, not warnings.
7. **Orphaned samples — the collector re-anchors observed fact.** When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the run moment, never backdated. This records observed fact, not intent: only the sanctioned path of [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) records a `user_disabled` close. Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
## Consequences
- Validation lives at one boundary — the collector — so the store's contract is "everything in it is well-formed" and readers only defend against the impossible.
- No synthetic data can ever enter the habit record; confidence categories can be trusted to reflect real evidence.
- Store-fault recovery can lose history; the mitigation is the one-file backup story of [ADR 0001](0001-observation-store-sqlite.md), kept human-sanctioned so loss is never silent.
- Period bookkeeping splits by epistemics: the helper records intent, the collector records observed fact.
- Fenris stays fully local and silent: no notification, escalation, or alerting machinery anywhere.
@@ -0,0 +1,28 @@
# 6. Collector acquisition path: smartctl counters, sysfs identity
## Status
Accepted — resolves [Choose the collector's NVMe acquisition path](https://git.bongbetic.com/xavierk/Fenris/issues/16) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1).
## Context
The collector ([ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md)) must acquire SMART/Health counters, thermal evidence, and controller identity each run. The [controller-identity research](https://git.bongbetic.com/xavierk/Fenris/src/branch/research/controller-identity/docs/research/controller-identity.md) fixed the identity key to the normalized, kernel-exposed subsystem NQN and warned that normalization must be specified once and applied at write time — or a collector implementation change can split a drive's own history. [ADR 0004](0004-install-upgrade-removal-lifecycle.md) pins exact Python dependencies in a dedicated venv, and the [segment-metadata decision](https://git.bongbetic.com/xavierk/Fenris/issues/14) froze nullable `vid`/`ssvid`/`transport` alongside the identity fields. Three first-party paths were candidates: the official libnvme Python bindings (SWIG; sysfs-backed attribute getters delivering normalized values), `nvme` CLI JSON output, and the incumbent `smartctl -j` plus sysfs reads.
## Decision
1. **Pin.** Every collection run acquires counters and thermal evidence solely from `smartctl -a -j <device>` and controller identity (`subnqn`, `sn`, `mn`, `fr`, `transport`) solely from sysfs (`/sys/class/nvme/<ctrl>/`). No other acquisition path exists anywhere in the codebase.
2. **Hard pin, no fallback.** Any acquisition failure — missing binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the whole collection run; [ADR 0005](0005-failure-detection-and-recovery.md)'s flat retry and freshness grading absorb the miss. A partial sample (identity without counters, or counters without identity) is never written: a transient read failure must not push a healthy drive down the degraded-identity path.
3. **Normalization once, at write time.** One collector-side function normalizes every identity field: trailing spaces and newlines stripped, no case folding, empty-after-strip stored blank. `smartctl` counter and thermal fields are consumed as-is (smartmontools already trims the strings it copies). Padded and unpadded renderings of the same field therefore yield byte-identical stored values.
4. **Segment metadata sourcing.** `transport` comes from the NVMe class sysfs directory; `vid`/`ssvid` from the PCI node (`/sys/class/nvme/<ctrl>/device/{vendor,subsystem_vendor}`) when present, null otherwise — metadata only, never key components.
5. **Prerequisites.** `make install` verifies `smartctl` is present and fails cleanly otherwise. The acquisition path adds no Python dependency and no OS package beyond smartmontools; the [ADR 0004](0004-install-upgrade-removal-lifecycle.md) lockfile is untouched.
## Considered options
- **libnvme Python bindings** — the purest API and natively-normalized getters, but the SWIG module is not on PyPI: entering the venv requires the distro's `python3-libnvme` through `--system-site-packages` or a from-source build, coupling the exact-lockfile venv to the system Python and the distro's shipping choices. Rejected on dependency weight for one privileged five-minute oneshot.
- **`nvme` CLI JSON** — one binary covers counters and identity, but it adds an OS package for what smartmontools already provides, emits untrimmed strings, and reports `subnqn` from Identify data rather than the kernel: when a controller reports an empty NQN the kernel synthesizes one for sysfs while `id-ctrl` JSON omits the field, so the identity ladder would drop a rung depending on the drive. Rejected on packaging and identity-key consistency.
## Consequences
- The venv stays pure-Python; the two acquisition channels per run (subprocess JSON plus sysfs reads) hide behind one acquisition function, gated by acceptance criteria AC-1–AC-5.
- Identity is read from exactly the source the identity key names; libnvme's getters wrap the same sysfs attributes, so the values agree byte-for-byte where both exist.
- Switching acquisition path later is history-sensitive: a future path must deliver byte-identical normalized identity values, or the change itself forces a controller-segment boundary.
@@ -0,0 +1,34 @@
# 7. Package delivery: native deb + rpm packages, amending the installation lifecycle
## Status
Accepted — resolves [Task: Compose release spec + ADR amending 0004](https://git.bongbetic.com/xavierk/Fenris/issues/42) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/33). This ADR **amends [ADR 0004](0004-install-upgrade-removal-lifecycle.md)** on delivery and file ownership only; every runtime semantic of 0004 — dormant install, polkit-only elevation, observation-store snapshot + forward-only migration, one-generation rollback — is inherited verbatim, restated below where the package delivery changes *who* performs it.
## Context
ADR 0004 fixed delivery as `sudo make install` from a source checkout: wheel into a Fenris-owned runtime directory at `/opt/fenris`, a hand-rolled placement manifest, units in `/etc/systemd/system`. The release plan ([map](https://git.bongbetic.com/xavierk/Fenris/issues/33), decisions [Lock channel + toolchain](https://git.bongbetic.com/xavierk/Fenris/issues/38), [Signing + key policy](https://git.bongbetic.com/xavierk/Fenris/issues/39), [Package ownership](https://git.bongbetic.com/xavierk/Fenris/issues/40), [Migration path](https://git.bongbetic.com/xavierk/Fenris/issues/41), [Release cadence](https://git.bongbetic.com/xavierk/Fenris/issues/43)) now ships Fenris as native deb + rpm packages built by nfpm and published to the self-hosted Gitea 1.27.1 package registry, for Debian 12, Ubuntu 22.04/24.04, Fedora 40+, and openSUSE Tumbleweed (x86_64), with locked pure-Python dependencies vendored at `/opt/fenris/vendor`. Packages become the primary delivery; ADR 0004's delivery model demotes to a dev fallback.
The implementation-ready operative contracts live in the [release and packaging specification](../spec/release-packaging.md); this ADR records the decisions and their rationale.
## Decision
Amendments to ADR 0004, section by section:
1. **Delivery (amended).** Packages are primary: one deb per codename pool (`bookworm`, `jammy`, `noble`) and one rpm (group `fenris`, Fedora 40+ and openSUSE Tumbleweed), built by nfpm from a single `packaging/nfpm.yaml` over locked pure-Python runtime packages staged at `/opt/fenris/vendor`, published to the Gitea Debian/RPM registry and installed with `apt`, `dnf`, or `zypper`. `sudo make install` remains as the dev fallback for machines without packages; the two deliveries are mutually exclusive per machine. Version scheme `<pyproject-version>-1`, revision bump on rebuild.
2. **Layout and manifest (amended).** The hand-rolled manifest model is retired: the dpkg/rpm database **is** the manifest, and nothing like `manifest.txt` ships. Package-owned layout: units in `/usr/lib/systemd/system` (vendor placement; `/etc/systemd/system` is admin-only for drop-ins and enable state); helpers stay in `/usr/libexec/fenris` (exactly `fenris-monitor` and `fenris-collect` — no new polkit-reachable binaries); polkit policy in `/usr/share/polkit-1/actions/`; wrapper at `/usr/bin/fenris` (FHS; `/usr/local/bin` remains `make install`'s). The `fenris` group is declared in `/usr/lib/sysusers.d/fenris.conf` (`g fenris -`) and `/var/lib/fenris` in `/usr/lib/tmpfiles.d/fenris.conf` (`d /var/lib/fenris 2750 root fenris -`), both invoked from the maintainer scripts. The package owns the `/var/lib/fenris` directory only; `observations.db`, WAL sidecars, and `.bak` are never owned and never ghosted — ghost-erase would delete the store, violating 0004 §8.
3. **Privilege (unchanged).** Root acts through maintainer scripts at install/upgrade/removal time; at runtime, elevation is exclusively polkit, exactly as 0004 §3 and [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) §5 fix it.
4. **Dormant install (restated for packages).** A fresh package install is fully dormant: postinst/%post performs `systemctl daemon-reload` (plus `systemd-sysusers` and `systemd-tmpfiles --create`) and nothing else — never enable, never preset, never start; no preset file ships. The sanctioned toggle (`fenris monitor resume`) remains the only opt-in.
5. **Legacy import (narrowed).** Auto-detection of `./data/history.jsonl` is scoped to `make install` only — a package install has no checkout to inspect. `fenris import <path>` remains available as the only import path from packages.
6. **Upgrade (inherited, maintainer-script mechanics).** Upgrades arrive as packages from the single registry channel. postinst/%post on upgrade: snapshot `observations.db` → one-generation `.bak`, run forward-only schema migrations through the target `python3` with `/opt/fenris/vendor` on its import path (no new binaries), `daemon-reload`, and restart `fenris-collect.timer` only if unit contents changed **and** it is active. `/var/lib/fenris` is never rebuilt; a running oneshot finishes on its old interpreter.
7. **Rollback (unchanged, plus one hard edge).** One-generation `.bak` semantics are unchanged. Package downgrade is additionally unsupported: forward-only store-version refusal means installing an older package over a newer store fails by design; documented rollback = restore the snapshot, then install the old release.
8. **Removal (mapped).** deb `remove` ≈ `make uninstall` (conffile and store survive); deb `purge` ≈ `make purge` (plus `.bak` and group cleanup); rpm erase ≈ `make uninstall` (unmodified config removed, modified survives as `.rpmsave`; purge is a documented manual command). prerm/%preun performs the sanctioned disable — `fenris-monitor disable --now`, closing the period `user_disabled` — on remove/erase **only, never on upgrade** (deb prerm upgrade case is a no-op; rpm `%preun` gated on `$1 -eq 0`).
9. **Conffile semantics (new).** `/etc/fenris/fenris.conf` ships as a placeholder-commented default with no active device selector — deb conffile, rpm `%config(noreplace)`. The device selector is entered by hand (root edits the file), as in both prior deliveries; no configuration verb is added to `fenris-monitor`, and [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) §3's read-and-validate-at-collection-time semantics are untouched. On upgrade, local edits survive as-is; a changed package default lands beside them as `.dpkg-new`/`.rpmnew`.
10. **Migration from make-install systems (new).** Remove-then-install via runbook only — no migration script, no auto-clean. preinst/%pre aborts with a pointer to the runbook if make-install remnants are detected (`/var/lib/fenris/manifest.txt` or `/etc/systemd/system/fenris-collect.timer`). Store and config survive by path continuity; the migration resets the system to dormant and the user opts back in with `fenris monitor resume`.
## Consequences
- Package installs, upgrades, and removals carry dpkg/rpm-native semantics; nothing in Fenris's own tooling duplicates them.
- The manifest was 0004's answer to "what did the installer place"; the package database answers it better, and uninstall-keeps-store now holds by package ownership rather than by manifest discipline.
- `make install` and packages are mutually exclusive per machine; over-install is blocked, not repaired (stale `/etc` units would silently shadow vendor units).
- Hand-edited configuration remains the model: the device selector is a root-edited file in every delivery, keeping the polkit surface at exactly one binary.
- Release mechanics — channel, signing, cadence, rollback documentation — are fixed in the [release and packaging specification](../spec/release-packaging.md) and the tickets it cites; this ADR deliberately stops at lifecycle semantics.
+21
View File
@@ -0,0 +1,21 @@
# 8. Native Void Linux support and XBPS delivery
Status: Accepted — implementation and host acceptance completed; evidence is recorded in [issue #87](https://git.bongbetic.com/xavierk/Fenris/issues/87).
Fenris will support Void Linux natively with runit and full application feature parity, while retaining its existing Debian/RPM and systemd support. This extends the platform boundary in [ADR 0003](0003-service-lifecycle-and-sanctioned-toggle.md) and the delivery scope in [ADR 0007](0007-package-delivery-amends-0004.md): requiring Void users to replace their init system would not meet the native-support goal.
Delivery will include a Fenris-maintained, signed XBPS repository that users configure once for subsequent installation and updates through XBPS, plus versioned release artifacts and notes on Gitea. All downloads must be served directly by Gitea itself; a separate static HTTP repository, even alongside Gitea, does not satisfy this requirement.
Use the dedicated public Gitea repository `xavierk/Fenris-xbps` with its permanent `stable` branch. Its raw-file URL serves the XBPS index, versioned packages, and package signatures as ordinary Git blobs without LFS. Keeping binaries in a separate repository avoids increasing application source-clone size. This accepts growth in distribution-repository Git history in exchange for publishing index and artifacts together through one branch update, without the generic registry's delete-and-upload index replacement gap.
Serialize publication, commit the signed index and its new artifacts together, and retain older versioned artifacts in the current tree so clients with cached older indexes can still download them. Native XBPS installation and update tests against the actual endpoint are required before release validation. The first release acceptance recorded in issue #87 verified direct artifact delivery, signed metadata, retained packages, and immediate discovery after an explicit memory-synchronized refresh. The Gitea raw endpoint advertises six-hour HTTP caching; users should use `xbps-install -M -S` when looking for updates so XBPS bypasses its on-disk repodata cache.
Immediate availability is required: after successful XBPS publication, an explicit repository refresh against the permanent URL must discover the newly published version without a cache-expiry wait or a URL change. This does not promise automatic installation on client machines. Acceptance must exercise a client that fetched the previous index before publication and verify that refresh retrieves the new index and its signed package afterward. Resolve and document actual client and intermediary cache behavior; if the selected Gitea route cannot meet this requirement, hold XBPS publication and revisit its delivery mechanics rather than silently accepting delayed availability.
Hosting evidence: [Gitea generic registry](https://docs.gitea.com/usage/packages/generic), [raw download routing](https://github.com/go-gitea/gitea/blob/main/routers/web/web.go), [download handler](https://github.com/go-gitea/gitea/blob/main/routers/web/repo/download.go), and [Void repository signing](https://docs.voidlinux.org/xbps/repositories/signing.html). Source inspection and an existing raw-file GET establish feasibility, not end-to-end XBPS validation.
The first supported Void target is x86_64 with glibc, matching the inspected development machine. Release validation must cover installation, real collection, pause/resume, reboot persistence, upgrade, and removal on that machine, preserving existing observation history. Debian/RPM compatibility remains part of the acceptance scope. Additional architectures and musl support are outside this first release.
Package formats have independent publication gates: publish each validated format, and hold only formats that have not passed release validation. A failure or pending validation in XBPS must not prevent a validated Debian or RPM package from shipping, and vice versa. Release notes must identify available formats and those still withheld; publication must not imply validation of a missing format.
After successful native acceptance testing, leave the released XBPS package installed on this machine and monitoring the selected NVMe drive. Preserve the observation history collected during testing. Coordinate the reboot test with the user so it can occur at a suitable interruption point. Issue #87 records that this final state, including reboot persistence, was achieved for the first supported release.
+8
View File
@@ -0,0 +1,8 @@
# One MIT license across source and packages
The native package metadata previously disagreed: deb/rpm declared Proprietary,
while XBPS declared MIT and the repository carried no license text. On
2026-09-16 the maintainer chose MIT for Fenris. The repository now includes the
standard MIT license, and Python, deb, rpm, and XBPS distributions must preserve
that same licensing decision; bundled third-party dependencies retain their own
license notices.
@@ -0,0 +1,11 @@
# 10. Preserve local-day activity history with its recorded timezone
Status: Accepted; legacy-history migration and repair implemented in issue #99.
Fenris will retain local-day read/write summaries and the boundary evidence needed to interpret them in the existing observation store before three-minute detail expires after 14 days. Each historical summary retains its recorded timezone and day boundaries; this extends [ADR 0001](0001-observation-store-sqlite.md) while preserving the UTC hour/day evidence used for endurance projections. Collection derives new measured intervals. Ordered schema migration and explicit repair rebuild legacy summaries from surviving evidence; TUI and CLI readers remain read-only.
UTC hourly summaries alone cannot recover a local day whose midnight falls inside a UTC hour. Preserving labelled local summaries trades arbitrary future timezone reinterpretation for bounded detailed-history retention; retaining all fine-grained intervals indefinitely or estimating a split from UTC summaries would violate the agreed retention or evidence semantics.
Only surviving evidence may be used to derive older local summaries. Dates without enough evidence remain incomplete or unavailable; a measured interval crossing midnight is retained once as shared boundary evidence, never prorated or counted in full on both days. A later timezone change does not silently rewrite historical day boundaries.
The agreed presentation and validation requirements are in the [live drive activity specification](../spec/live-drive-activity.md).
@@ -0,0 +1,20 @@
# 11. Preserve valid unpublished observations without exposing partial history
Status: Accepted; implemented on `main` for the planned v0.6.0 release (commit `e22b994`).
A non-invariant derivation failure must not discard valid acquired observations or report a successful collection: retain the observations as pending publication in the existing observation store and retry through the normal scheduled collection path. Readers continue to see the last consistent published evidence, with an explicit pending-publication explanation, rather than combining newly acquired counters with older derived totals. This trades recovery bookkeeping for preservation of measured evidence and consistent read views; neither discarding every valid acquisition on derivation failure nor exposing partially derived history satisfies both requirements.
## Constraints
- The publication distinction applies to every dependent reader, including activity, freshness, controller-segment interpretation and projection inputs, not just the graph. ADR 0001's freshest-sample signal means the freshest published sample; an unpublished observation must not make published evidence appear fresh.
- Pending-work bookkeeping qualifies ADR 0005's recovery-without-new-state wording: it records unfinished evidence publication, not a new health grade or escalation mechanism. Last-collection outcomes still come from the existing native monitoring and logging paths.
- Invariant violations retain [ADR 0005](0005-failure-detection-and-recovery.md)'s write-nothing rule; retaining recoverable work is not permission to persist invalid observations. Actual store faults retain their existing refusal and degradation behavior.
- Recovery and retention remain collector-owned. Repeated recovery must not duplicate measured volume, and required source evidence cannot be pruned before trustworthy derived evidence is durable.
- This extends [ADR 0001](0001-observation-store-sqlite.md)'s consistent read model and [ADR 0010](0010-local-day-activity-history.md)'s preservation rule without another observation store, sampler, background process or retry cadence. Pending-work metadata is not a persisted projection or a replacement for journalled collection outcomes.
- Pending observations use private staging separate from published samples and derived evidence inside the same observation store. This makes exclusion from ordinary reader and projection queries structural, rather than requiring each query to remember a publication filter; the exact table layout remains an implementation detail.
## Bounded pending work
Pending publication has a fixed initial admission capacity of 6,720 observations, equivalent to 14 days at the default three-minute cadence. This is a count limit, not an expiry rule: older pending observations are never discarded merely to admit newer ones.
The collection module attempts recovery and checks capacity before invoking acquisition. If capacity remains exhausted, the collection is unsuccessful and visibly explains why no new observation was acquired; normal scheduled recovery attempts continue. This deliberately accepts a gap in new observations rather than unlimited pending growth or loss of already retained evidence. It is not a deliberate disable, changes no monitoring intent, and never permits fabricated activity across the gap.
+26
View File
@@ -0,0 +1,26 @@
# Domain Docs
How engineering skills should consume this repository’s domain documentation.
## Layout
This is a single-context repository:
```text
/
├── CONTEXT.md
├── docs/adr/
└── ...
```
## Before exploring
Read `CONTEXT.md` and relevant ADRs under `docs/adr/` when they exist. If they do not exist, proceed silently. Domain-modeling skills create them lazily when terminology or durable architectural decisions are resolved.
## Use the glossary’s vocabulary
Use terminology defined in `CONTEXT.md` consistently. If required terminology is missing or contradictory, raise it through domain modeling rather than silently inventing synonyms.
## Flag ADR conflicts
If proposed work contradicts an existing ADR, identify the conflict explicitly instead of silently overriding it.
+85
View File
@@ -0,0 +1,85 @@
# Issue tracker: Gitea
Issues for this repository live in Gitea at:
https://git.bongbetic.com/xavierk/Fenris/issues
Use the authenticated `tea` CLI from the repository root. The configured login is `xavierk`.
## General operations
- List: `tea issues list`
- Read: `tea issues <index> --comments`
- Create: `tea issues create --title "<title>" --description "<body>"`
- Edit: `tea issues edit <index> --title "<title>" --description "<body>"`
- Assign: `tea issues edit <index> --add-assignees "<username>"`
- Add labels: `tea issues edit <index> --add-labels "<labels>"`
- Comment: `tea comments add <index> --description "<comment>"`
- Close: `tea issues close <index>`
- Reopen: `tea issues reopen <index>`
Use `--output json` for machine-readable list and read operations. Use `tea api` when the high-level issue commands do not expose a native Gitea operation.
## When a skill says “publish to the issue tracker”
Create a Gitea issue in this repository. Preserve Markdown formatting in its body and apply any labels required by the invoking skill.
## When a skill says “fetch the relevant ticket”
Read the named issue with comments. The user may provide its URL, title, or index. In user-facing output, refer to issues by their linked titles rather than bare indices.
## Wayfinding operations
Wayfinder maps and decision tickets are Gitea issues.
### Map and ticket grouping
- A map has the label `wayfinder:map`.
- Create one milestone named `Wayfinder: <map title>` for the effort.
- Assign the map and all its tickets to that milestone.
- Every ticket links its parent by name near the top: `Parent map: [<map title>](<map URL>)`.
- Every ticket has exactly one type label: `wayfinder:research`, `wayfinder:prototype`, `wayfinder:grilling`, or `wayfinder:task`.
The shared milestone and explicit parent link express the child relationship, because this Gitea version has no native parent/child issue API.
### Blocking
Use Gitea’s native issue-dependency relationship. To make `<blocked>` depend on `<blocker>`:
```bash
tea api -X POST \
repos/{owner}/{repo}/issues/<blocked>/dependencies \
-F index=<blocker> \
-f owner=xavierk \
-f repo=Fenris
```
List blockers:
```bash
tea api repos/{owner}/{repo}/issues/<index>/dependencies
```
Remove the relationship with the same payload and `-X DELETE`.
### Frontier
List open issues in the map’s milestone. Exclude:
- the issue labelled `wayfinder:map`
- assigned tickets, because assignment is the claim
- tickets whose dependency query returns any open issue
The remaining open, unassigned, unblocked tickets are the frontier. Choose the oldest first unless the user names one.
### Claim
Before doing any ticket work, assign it to the current `tea whoami` user. An open ticket without an assignee is unclaimed.
### Resolve
1. Add the answer as a resolution comment.
2. Close the ticket.
3. Re-fetch the map immediately before editing it.
4. Append a linked one-line context pointer to `Decisions so far`.
5. Create newly visible tickets, then wire dependencies in a second pass.
+87
View File
@@ -0,0 +1,87 @@
# Migrating from make-install to packages
This runbook covers the transition from a `sudo make install` system to the native deb or rpm package. Packages are the primary delivery; `make install` remains as the dev fallback. The two deliveries are **mutually exclusive** per machine.
## Why over-install is forbidden
Installing a package over a make-install system silently breaks things:
- **Stale admin units shadow vendor units.** `make install` places `fenris-collect.timer` and `fenris-collect.service` in `/etc/systemd/system/`. The package installs them in `/usr/lib/systemd/system/` (vendor placement). Systemd loads admin units first — the stale copy takes precedence, and the package update never reaches the running system.
- **The local wrapper shadows the package wrapper.** `make install` places the `fenris` wrapper at `/usr/local/bin/fenris`. The package places it at `/usr/bin/fenris`. The shell finds `/usr/local/bin` first on PATH — the old checkout-relative wrapper runs instead of the package wrapper.
Neither condition is reversible by reinstalling the package. The only safe path is remove-then-install.
## Pre-migration checklist
1. Confirm no monitoring period is actively running that you want to preserve across the gap:
```
fenris status
```
The migration resets the system to dormant (see [No-move continuity](#no-move-continuity) below). You opt back in with `fenris monitor resume`.
2. If you have hand-edited configuration at `/etc/fenris/fenris.conf`, note it. The config survives the migration in place (see below).
## Remove step
```
sudo make uninstall
```
This performs the **sanctioned disable** (`fenris-monitor disable --now`), closing the current monitoring period as `user_disabled`. It then removes all make-install artifacts: the venv at `/opt/fenris`, the wrapper at `/usr/local/bin/fenris`, the helpers at `/usr/libexec/fenris/`, the units in `/etc/systemd/system/`, and the polkit policy. The placement manifest at `/var/lib/fenris/manifest.txt` is removed.
**What survives the remove:**
- `/var/lib/fenris/observations.db` (and WAL sidecars, `.bak`) — the observation store
- `/var/lib/fenris/` directory itself — root-written, group-read
- `/etc/fenris/fenris.conf` — your hand-written configuration
- The `fenris` system group — created by `groupadd -f` during make-install
- Journal entries — age out naturally
## Install step
```
sudo apt install fenris # Debian/Ubuntu
sudo dnf install fenris # Fedora
```
The package installs into its own layout without touching the surviving store, config, or group.
## No-move continuity
These invariants are verified by the containerized acceptance tests (issue #50):
| Asset | Make-install state | Package post-install | Mechanism |
|---|---|---|---|
| `fenris` group | Exists (`groupadd -f`) | Unchanged | `systemd-sysusers` is a no-op when the group already exists |
| `/var/lib/fenris` directory | Exists (mode 2750, root:fenris) | Unchanged | `systemd-tmpfiles --create` is a no-op when the directory already exists |
| `observations.db` + sidecars | Present from prior monitoring | Unchanged, never owned by the package | Package owns the directory only; store contents are never ghosted |
| `/etc/fenris/fenris.conf` | Hand-edited device selector | Survives in place; package default lands as `.dpkg-new` / `.rpmnew` | dpkg conffile / rpm `%config(noreplace)` semantics |
| Store schema | Version from prior Fenris release | Caught up by the upgrade-path migration | `postinst` / `%post` runs `migrate_to_latest()` on upgrade |
The package detects the make-install system has been removed by the absence of the two markers:
- `/var/lib/fenris/manifest.txt` (the placement manifest)
- `/etc/systemd/system/fenris-collect.timer` (pre-manifest make installs)
If either marker exists, the package installation aborts with a pointer to this runbook.
## Reset-to-dormant
`make uninstall`'s sanctioned disable closes the open monitoring period as `user_disabled`. After the package install, the system is dormant — the timer is installed but disabled, nothing is running, no monitoring period is open.
To resume monitoring:
```
fenris monitor resume
```
This is the sanctioned opt-in. It enables the timer and opens the first monitoring period in one step. The migration costs at most one short sample gap (the interval between `make uninstall` and `fenris monitor resume`), honestly recorded in the endurance timeline.
## Verification
After migration, confirm the package is correctly installed:
```
fenris status
```
The status command should show the dormant state: timer disabled, no active monitoring period, and the observation store intact from the prior make-install system.
+231
View File
@@ -0,0 +1,231 @@
# Signing key ceremony
The Fenris packaging key signs RPM payloads and clearsigns SHA256SUMS manifests.
This document describes the key's lifecycle: creation, per-release use, rotation,
and destruction.
## Key specification
| Property | Value |
|---|---|
| Algorithm | RSA 3072 |
| UID | `Fenris Packaging <packaging@bongbetic.com>` |
| Expiry | 2 years from creation |
| Hierarchy | Single key — no master/subkey split (single maintainer, manual builds) |
| Private key storage | Gitea repository Actions secret `GPG_PRIVATE_KEY` |
| Public key storage | `packaging/keys/fenris-packaging.asc` in-repo, release notes, docs |
| Keyservers | Never — TOFU-over-TLS via raw URL |
## First release: key creation
```bash
# Generate the dedicated RSA-3072 packaging key
gpg --batch --gen-key <<EOF
%no-protection
Key-Type: RSA
Key-Length: 3072
Name-Real: Fenris Packaging
Name-Email: packaging@bongbetic.com
Expire-Date: 2y
%commit
EOF
# Export the public half — this file is committed to the repo
gpg --armor --export packaging@bongbetic.com > packaging/keys/fenris-packaging.asc
# Print the fingerprint for docs and release notes
gpg --fingerprint packaging@bongbetic.com
```
Provision the **private key** as the repository Actions secret `GPG_PRIVATE_KEY`.
Run the export on the trusted key-generation machine, then enter its output in
the Gitea repository's Actions secret settings. Do not save it in the checkout,
logs, or a runner directory. The release workflow checks its fingerprint
against the committed public key before signing.
```bash
gpg --armor --export-secret-keys packaging@bongbetic.com
```
After provisioning the secret, delete the private key from the key-generation
keyring:
```bash
gpg --batch --yes --delete-secret-keys packaging@bongbetic.com
gpg --batch --yes --delete-keys packaging@bongbetic.com
```
The committed `fenris-packaging.asc` must contain the real public key (replace
the placeholder comments).
## XBPS signing key
XBPS uses a separate RSA 3072 key. Its private half is stored as the Gitea
repository Actions secret `XBPS_SIGNING_KEY`. The corresponding public key is
published at
`https://git.bongbetic.com/xavierk/Fenris-xbps/raw/branch/stable/keys/fenris-xbps-signing.pub`,
with fingerprint `SHA256:AvPMRlKMikPg75u0iKr8AUkxlfU/Ad4k/S4o2M9W4/w`.
The secret must match that public key.
The release workflow writes the key to `~/.ssh/id_xbps` to sign the XBPS
package. A requested XBPS publication also uses the key to sign repository
metadata. A final `always()` cleanup removes the runner copy after publication
and release asset upload, including when an earlier step fails.
## Per-release signing flow
Each tagged release performs: **import → verify → sign → delete** on the
repository-scoped Gitea Actions runner. The Gitea secret remains configured;
the runner's keyring copy is removed after the job.
### Step 1: Push the release tag
After updating the version and dated changelog section, push the matching tag:
```bash
git push origin v<version>
```
### Step 2: Build, verify, and sign packages
The release workflow imports `GPG_PRIVATE_KEY`, checks it against
`packaging/keys/fenris-packaging.asc`, builds packages, signs the RPM and
clearsigned checksum manifest, validates both, and publishes the release. The
workflow imports `XBPS_SIGNING_KEY` separately and signs the XBPS package.
XBPS publication is optional and also signs repository metadata; it requires
host acceptance and explicit selection during workflow dispatch.
### Step 3: Verify runner cleanup
The workflow's `always()` cleanup removes the GPG key from the runner's keyring
and deletes `~/.ssh/id_xbps`, including after a failed job. Confirm no signing
key remains on the runner after the release job.
The Gitea Actions secret remains the approved signing source. Do not copy it to
the runner or repository outside the workflow.
## Key rotation (outline)
When the key approaches expiry, or if it is compromised:
1. **Generate a new key** using the same procedure as first release.
2. **Publish the new public key** alongside the old one in-repo:
```text
packaging/keys/fenris-packaging.asc # new key (primary)
packaging/keys/fenris-packaging-previous.asc # old key (one cycle)
```
3. **Sign the next RPM** with the new key.
4. **Update `fenris.repo`** to list both `gpgkey` URLs (dnf accepts multiple):
```ini
gpgkey=https://git.bongbetic.com/xavierk/Fenris/raw/branch/main/packaging/keys/fenris-packaging.asc
https://git.bongbetic.com/xavierk/Fenris/raw/branch/main/packaging/keys/fenris-packaging-previous.asc
```
5. **Drop the old key** from the repo after one release cycle. Delete
`fenris-packaging-previous.asc` and revert `gpgkey` to the single URL.
## Verification
Consumers verify the RPM payload signature via dnf (gpgcheck=1 in
`fenris.repo` points at the published public key). The SHA256SUMS manifest
verification is manual for downloaded assets:
```bash
gpg --verify SHA256SUMS.asc SHA256SUMS
sha256sum -c SHA256SUMS
```
## One-time live probe
Before the first real release, verify the full registry path end-to-end with a
throwaway package. This confirms apt/dnf metadata generation, signature
verification, and consumer setup work as a real consumer would experience them.
### Setup
```bash
# Create a throwaway package name to avoid polluting fenris metadata
PROBE_NAME="fenris-regtest"
PROBE_VERSION="0.0.1"
```
### Publish
```bash
# Build a throwaway deb and rpm (use the existing nfpm config with a dummy name)
# Or use a pre-built package — the probe tests the registry path, not the build
# Upload deb to all codename pools
for CODENAME in bookworm jammy noble; do
curl --fail -X PUT \
-u "xavierk:${GITEA_TOKEN}" \
-T "dist/${PROBE_NAME}_${PROBE_VERSION}_amd64.deb" \
"https://git.bongbetic.com/api/packages/xavierk/debian/pool/${CODENAME}/main/upload"
done
# Upload rpm
curl --fail -X PUT \
-u "xavierk:${GITEA_TOKEN}" \
-T "dist/${PROBE_NAME}-${PROBE_VERSION}-1.x86_64.rpm" \
"https://git.bongbetic.com/api/packages/xavierk/rpm/fenris/upload"
```
### Verify apt metadata (Debian/Ubuntu consumer perspective)
```bash
# On a Debian/Ubuntu machine:
sudo mkdir -p /etc/apt/keyrings
sudo curl -fsSL https://git.bongbetic.com/api/packages/xavierk/debian/repository.key \
| sudo gpg --dearmor -o /etc/apt/keyrings/gitea-xavierk.asc
echo "deb [signed-by=/etc/apt/keyrings/gitea-xavierk.asc] https://git.bongbetic.com/api/packages/xavierk/debian bookworm main" \
| sudo tee /etc/apt/sources.list.d/fenris.list
sudo apt update
apt show ${PROBE_NAME} # metadata present, correct version
apt install --dry-run ${PROBE_NAME} # dependency resolution works
# Verify InRelease signature
apt-key list 2>/dev/null || gpg --no-default-keyring --keyring /etc/apt/keyrings/gitea-xavierk.asc --list-keys
```
### Verify dnf metadata (Fedora consumer perspective)
```bash
# On a Fedora machine:
sudo dnf config-manager --add-repo https://git.bongbetic.com/xavierk/Fenris/raw/branch/main/packaging/fenris.repo
# Or use Gitea's auto-generated repo for the probe:
sudo dnf config-manager --add-repo https://git.bongbetic.com/api/packages/xavierk/rpm/fenris.repo
dnf info ${PROBE_NAME} # metadata present, correct version
dnf install --assumeno ${PROBE_NAME} # dependency resolution works
# Verify rpm signature
rpm -q --scripts ${PROBE_NAME} # no scripts (throwaway)
```
### Verify checksums and clearsign
```bash
# Download from release assets or local build
gpg --verify SHA256SUMS.asc SHA256SUMS
sha256sum -c SHA256SUMS
```
### Cleanup
```bash
# Delete the throwaway packages from the registry
for CODENAME in bookworm jammy noble; do
curl --fail -X DELETE \
-u "xavierk:${GITEA_TOKEN}" \
"https://git.bongbetic.com/api/packages/xavierk/debian/pool/${CODENAME}/main/${PROBE_NAME}/${PROBE_VERSION}/amd64"
done
curl --fail -X DELETE \
-u "xavierk:${GITEA_TOKEN}" \
"https://git.bongbetic.com/api/packages/xavierk/rpm/fenris/${PROBE_NAME}/${PROBE_VERSION}/x86_64"
# Remove test source list on consumer machines
sudo rm /etc/apt/sources.list.d/fenris.list
sudo apt update
```
+239
View File
@@ -0,0 +1,239 @@
# Research: deb + rpm packaging toolchain for bundled-venv builds
Issue: #34 (parent plan: #33) — branch `research/toolchain`
Date: 2026-09-03 · target: Fenris 0.3.0, x86_64, Debian 12 / Ubuntu 22.04+24.04 / Fedora 40+
## TL;DR
**Recommended: nfpm** with a build script that stages a `--copies` venv at
`/opt/fenris`. One `nfpm.yaml` is the single source of truth for both formats;
`nfpm pkg -p deb && nfpm pkg -p rpm` (one invocation per format — `-p` takes a
single string, verified in `internal/cmd/package.go`). Actively maintained
(releases v2.47.0, 2026-06-20; repo pushed 2026-08-31). Runner-up: fpm (active,
v1.18.0 gem 2026-08-26), but its "config" is a long CLI invocation per format —
the single source of truth degrades into a shell script. dh-virtualenv is
deb-only and its last upstream release is 2020-10 (effectively dormant);
rpmbuild spec is rpm-only and cannot share file lists with a deb build without
external generation.
## Constraint evidence: distro textual is unusable (mostly)
| Distro | python3-textual | Source |
|---|---|---|
| Debian 12 (bookworm) | **0.1.13-1** | https://packages.debian.org/bookworm/python3-textual |
| Ubuntu 22.04 (jammy) | **0.1.13-1** | https://packages.ubuntu.com/jammy/python3-textual |
| Ubuntu 24.04 (noble) | **0.1.13-1** | https://packages.ubuntu.com/noble/python3-textual |
| Fedora 40 | 0.48.1 | https://src.fedoraproject.org/rpms/python-textual (f40 spec) |
| Fedora 41 / 42 / 43 | 0.69.0 / 1.0.0 / 4.0.0 | same spec, f41–f43 branches |
Fenris declares `textual>=0.40.0` (pyproject) but pins `textual==8.2.8`
(requirements.txt). Debian 12 + Ubuntu 22.04/24.04 are ~1 major era behind even
the *floor*; Fedora 40 technically meets `>=0.40` but not the pin. Verdict
unchanged: **vendor deps inside the package for all targets**; per-format
`depends:` only on `python3 (>= 3.9)`, `smartmontools`, `systemd`.
## Tool-by-tool
### 1. nfpm (goreleaser) — RECOMMENDED
- **Route:** Makefile target builds staging tree → one `nfpm.yaml` →
`nfpm package -p deb` + `nfpm package -p rpm`. (Goreleaser release pipeline
can wrap both later.)
- **Shared assets:** version, description, maintainer, `depends`,
`contents:` file list, `scripts:` all live once in `nfpm.yaml`; per-format
deltas via `overrides: { deb: ..., rpm: ... }` and `packager:`-scoped
content entries (https://nfpm.goreleaser.com/configuration/, source
`www/content/docs/configuration.md`).
- **Prerequisites:** single static Go binary (`go install
github.com/goreleaser/nfpm/v2/cmd/nfpm@latest`, Homebrew, or release
tarball — https://nfpm.goreleaser.com/install/). No toolchain per distro,
no root, no containers required (build same tree for both formats).
- **Venv → file list:** stage with `python3 -m venv --copies staging/opt/fenris
&& staging/opt/fenris/bin/pip install dist/fenris-*.whl`; map in one entry:
`contents: [{ src: staging/opt/fenris/, dst: /opt/fenris, type: tree }]`.
Shebangs point at fixed absolute `/opt/fenris/bin/python` → no relocation
issues. `--copies` avoids symlink-to-/usr breakage. Config file →
`type: config|noreplace` (=%config(noreplace) on rpm, conffile semantics on
deb). `/var/lib/fenris` store → `type: ghost` (rpm: owned-but-not-packed;
deb: ignored → create in `postinstall` script instead).
- **Systemd/polkit/libexec:** plain `contents:` entries —
`/usr/lib/systemd/system/fenris-collect.{service,timer}` (or
`/etc/systemd/system` to match current Makefile), polkit action at
`/usr/share/polkit-1/actions/`, helpers under `/usr/libexec/fenris/`.
`scripts:` supports `postinstall` (deb maintainer script / rpm scriptlet) —
run `systemctl daemon-reload`, create `/var/lib/fenris` root:fenris 2750,
group creation.
- **Upgrade/removal:** deb — dpkg replaces all non-conffile files, conffile
prompts/preserves (`.dpkg-new`) per Debian Policy ch-files
(https://www.debian.org/doc/debian-policy/ch-files.html); removal keeps
conffiles + unowned store; purge cleans. rpm — `rpm -U` replaces,
`%config(noreplace)` keeps local edits as `.rpmnew`; only owned dirs are
removed on erase (nfpm `type: dir` exists precisely to claim ownership —
docs warn not to claim distro-owned dirs).
- **Maintenance:** very active. goreleaser/nfpm, 2.6k stars, last push
2026-08-31, v2.47.0 released 2026-06-20 (GitHub API).
### 2. fpm — viable, weaker single-source-of-truth
- **Route:** staging tree (same as above) then
`fpm -s dir -t deb ... staging/=/ ; fpm -s dir -t rpm ...`.
- **Shared assets:** none declarative — everything is CLI flags
(`-n`, `-v`, `--config-files`, `--deb-systemd`, `--directories`,
`--after-install`, `--rpm-posttrans`, …). Flag list:
https://fpm.readthedocs.io/en/latest/cli-reference.html. The two
invocations *will* drift unless wrapped in a Makefile that shares variables;
the "single source" is then a shell script, not a checked declarative file.
(`--deb-systemd` exists; no rpm-native unit macro — you hand it the unit
file plus `--rpm-posttrans` for daemon-reload.)
- **Prerequisites:** Ruby + gem (`gem install fpm`) or distro package;
building rpm side needs `rpmbuild` present for some features.
- **Venv → file list:** `-s dir` maps a directory into the package verbatim —
same staging-tree trick as nfpm. `--config-files /etc/fenris` marks
conffiles (deb) / %config (rpm).
- **Upgrade/removal:** identical downstream semantics to nfpm (native dpkg/rpm
behavior); differences are only in how metadata/scripts land in the
package.
- **Maintenance:** active — releases v1.16.0 (2024-12), v1.17.0 (2025-10),
v1.18.0 (2026-08-26); gem 1.18.0 on rubygems; ~11.5k stars. But docs are
openly "work in progress" (https://fpm.readthedocs.io/en/latest/).
### 3. dh-virtualenv (Spotify) — deb-only, dorms
- **Route:** debhelper add-on: `debian/rules` with
`dh $@ --with python-virtualenv --buildsystem=python_distutils`;
produces a .deb containing venv at `/opt/venvs/<package>`
(`DH_VIRTUALENV_INSTALL_ROOT` overridable, `--builtin-venv` for `python -m
venv`). Docs: repo `doc/usage.rst`, `doc/tutorial.rst`
(https://github.com/spotify/dh-virtualenv).
- **Shared assets:** none with rpm — it cannot emit .rpm at all. Would still
need a second toolchain for Fedora → fails the criterion outright.
- **Prerequisites:** `build-essential debhelper devscripts equivs` +
`dh-virtualenv` (tutorial.rst); Debian 12 still ships it as
`dh-virtualenv 1.2.2-1.3` (https://packages.debian.org/bookworm/dh-virtualenv).
- **Venv → file list:** automatic — it builds the venv during the debhelper
sequence and rewrites shebangs; the .deb owns the whole venv tree. Least
manual work of all four, for deb alone.
- **Upgrade/removal:** standard dpkg; whole venv tree is package-owned, so
`apt remove` deletes it cleanly; `--pypi-url`/requirements handled by tool.
- **Maintenance:** last upstream release **1.2.2, 2020-10-22** (GitHub tag);
repo last pushed 2024-04-27, RTD docs 404. Effectively dormant upstream —
fine via Debian's own packaging, but risky as strategic dependency.
- **Bonus fact:** PyPI `dh-virtualenv` project now returns 404 — install only
from Debian repo / git.
### 4. rpmbuild spec + vendored venv — rpm-native, no deb
- **Route:** hand-written `fenris.spec`: `%install` stage builds venv into
`%{buildroot}/opt/fenris`, `%files` lists it plus units/polkit/libexec,
`%ghost %attr(2750,root,fenris) /var/lib/fenris`, `%config(noreplace)` for
`/etc/fenris`, `systemd_post/preun` macros for the timer. Reference style:
https://docs.fedoraproject.org/en-US/packaging-guidelines/.
- **Shared assets:** the spec is a second, parallel description of the same
file list — nothing is shared with any deb build without generating one
side from the other (e.g. generate spec + debian/control from a manifest).
Worst single-source-of-truth score.
- **Prerequisites:** `rpm-build`, mock/koji for cleanroots; Fedora toolchain
knowledge; per-distro `Release:`/dist tag handling.
- **Venv → file list:** `%files` line `%{buildroot}/opt/fenris/...` — venv
becomes ordinary payload; shebangs already absolute.
- **Upgrade/removal:** canonical rpm semantics (same as above) plus real
systemd scriptlet macros — the *best-behaved* rpm integration of the four,
at the cost of hand-maintained spec.
- **Maintenance:** rpmbuild itself is maintained forever (part of RPM), but
*your* spec is 100% hand-maintained duplication.
## Comparison matrix
| Criterion | nfpm | fpm | dh-virtualenv | rpmbuild spec |
|---|---|---|---|---|
| deb + rpm from one config | ✅ one YAML (2 invocations) | ⚠️ flags per invocation | ❌ deb only | ❌ rpm only |
| File list shared across formats | ✅ `contents:` | ⚠️ per-invocation args | n/a | ❌ |
| Vendored venv supported | ✅ staging `type: tree` | ✅ `-s dir` | ✅✅ automatic (deb) | ✅ `%files` |
| conffile / %config(noreplace) | ✅ `type: config\|noreplace` | ✅ `--config-files` | ✅ (debhelper) | ✅ `%config(noreplace)` |
| ghost store dir | ✅ `type: ghost` | ⚠️ `--rpm-ghost`? (no deb equiv) | ❌ | ✅ `%ghost` |
| systemd scriptlets | ✅ `scripts:` + macros? (plain scripts) | ✅ `--deb-systemd`, `--rpm-posttrans` | ✅ (deb) | ✅✅ native macros |
| Prereqs on build host | Go binary (or brew/apt tarball) | Ruby gem | debhelper stack | rpm-build + mock |
| Maintenance (2026) | 🟢 active (v2.47.0) | 🟢 active (v1.18.0) | 🔴 dormant since 2020 (Debian carries it) | 🟢 tool yes / 🔴 your spec |
| Risk | young-ish config schema churn | docs thin | dead upstream | duplication forever |
## Proposed pipeline (sketch)
```make
# Makefile additions (build only — install target stays for source installs)
stage: dist/fenris-*.whl
rm -rf build/stage
python3 -m venv --copies build/stage/opt/fenris
build/stage/opt/fenris/bin/pip install --no-compile dist/fenris-*.whl
install -D -m 0755 scripts/fenris build/stage/usr/bin/fenris
install -D -m 0755 src/fenris/monitor.py build/stage/usr/libexec/fenris/fenris-monitor
install -D -m 0755 src/fenris/collect.py build/stage/usr/libexec/fenris/fenris-collect
install -D -m 0644 units/fenris-collect.timer build/stage/usr/lib/systemd/system/fenris-collect.timer
install -D -m 0644 units/fenris-collect.service build/stage/usr/lib/systemd/system/fenris-collect.service
install -D -m 0644 polkit/com.bongbetic.fenris.monitor.policy \
build/stage/usr/share/polkit-1/actions/com.bongbetic.fenris.monitor.policy
package-deb package-rpm: stage
nfpm pkg -f packaging/nfpm.yaml -p deb -t dist/
nfpm pkg -f packaging/nfpm.yaml -p rpm -t dist/
```
```yaml
# packaging/nfpm.yaml (excerpt)
name: fenris
arch: amd64
platform: linux
version: ${VERSION} # env expansion, documented feature
maintainer: Fenris Maintainers <ops@bongbetic.com>
description: SMART drive observation daemon with persistent TUI
homepage: https://git.bongbetic.com/xavierk/Fenris
depends: [smartmontools]
contents:
- src: build/stage/ # everything above
dst: /
type: tree
- dst: /etc/fenris # config dir; ship fenris.conf as config|noreplace
type: dir
- src: packaging/fenris.conf
dst: /etc/fenris/fenris.conf
type: config|noreplace
- dst: /var/lib/fenris # rpm: %ghost ownership; deb: create in postinst
type: ghost
scripts:
postinstall: packaging/postinst.sh # groupadd fenris; install -d -o root -g fenris -m 2750 /var/lib/fenris; systemctl daemon-reload (units shipped dormant)
preremove: packaging/prerm.sh # stop timer if running
overrides:
deb:
depends: [python3 (>= 3.9), smartmontools]
rpm:
depends: [python3 >= 3.9, smartmontools]
```
## Recommendation
Adopt **nfpm + staged `--copies` venv**: closest to single source of truth
(one YAML for both formats), smallest prerequisite surface (one static binary),
actively maintained, and every Fenris constraint (units, polkit, libexec,
`/etc/fenris` conffile, `/var/lib/fenris` ghost/store) has a first-class
mapping. Keep fpm as documented fallback (identical staging tree, works
anywhere Ruby exists). Do not build the release pipeline on dh-virtualenv
(dormant, deb-only) or on a hand-maintained spec file (duplication, deb side
unaddressed).
## Sources
- nfpm config reference: https://nfpm.goreleaser.com/configuration/ (source:
goreleaser/nfpm `www/content/docs/configuration.md`, accessed 2026-09-03)
- nfpm CLI (single `-p`): goreleaser/nfpm `internal/cmd/package.go`
- nfpm releases/status: GitHub API, repo pushed 2026-08-31, v2.47.0 2026-06-20
- fpm README + CLI reference: https://github.com/jordansissel/fpm,
https://fpm.readthedocs.io/en/latest/cli-reference.html; releases v1.18.0
(2026-08-26), gem 1.18.0
- dh-virtualenv docs: `doc/usage.rst`, `doc/tutorial.rst` @ master; tag 1.2.2
dated 2020-10-22 (GitHub commits API); PyPI project 404;
Debian 12 package 1.2.2-1.3 (packages.debian.org)
- Distro textual versions: packages.debian.org, packages.ubuntu.com,
src.fedoraproject.org `python-textual.spec` f40–f43
- Upgrade semantics: Debian Policy ch-files
(https://www.debian.org/doc/debian-policy/ch-files.html); Fedora packaging
guidelines (https://docs.fedoraproject.org/en-US/packaging-guidelines/);
RPM directive behavior quoted in nfpm config docs (%ghost, %config(noreplace))
+116
View File
@@ -0,0 +1,116 @@
# Research: Gitea 1.27 Debian + RPM package registry feasibility
Issue: [Fenris deb + rpm release plan](https://git.bongbetic.com/xavierk/Fenris/issues/33) →
[Research: Gitea 1.27 Debian + RPM package registry feasibility](https://git.bongbetic.com/xavierk/Fenris/issues/35)
Verified 2026-09-03 against live instance `https://git.bongbetic.com` (reports `1.27.1` via `/api/v1/version`)
and primary sources: docs.gitea.com 1.27 Debian/RPM registry pages and Gitea `v1.27.1` source (go-gitea/gitea tag).
**Verdict: feasible.** Every publish/consume path tested live with throwaway packages `fenris-regtest` (all deleted afterward; package list verified empty).
## 1. Publish paths (verified live, HTTP 201)
### Debian (`.deb`)
```bash
curl --user xavierk:$TOKEN --upload-file fenris_0.3.0_amd64.deb \
"https://git.bongbetic.com/api/packages/xavierk/debian/pool/{distribution}/{component}/upload"
```
- `distribution` and `component` are free-form path segments chosen at upload time (e.g. `bookworm/main`, `noble/main`). Gitea derives apt suites from what was uploaded — verified: same .deb published to `pool/bookworm/main` and `pool/noble/main` (both 201), both then served in `dists/bookworm/` and `dists/noble/` with correct `Suite:`/`Codename:` headers.
- Republish of identical name+version+distribution+component+architecture → **409 Conflict** (verified). Must delete first.
### RPM (`.rpm`)
```bash
# no group (flat repo)
curl --user xavierk:$TOKEN --upload-file fenris-0.3.0-1.el9.x86_64.rpm \
"https://git.bongbetic.com/api/packages/xavierk/rpm/upload"
# with group (distro tag, nestable)
curl --user xavierk:$TOKEN --upload-file fenris-0.3.0-1.fc40.x86_64.rpm \
"https://git.bongbetic.com/api/packages/xavierk/rpm/el9/upload" # e.g. el9, rocky/el9, fc40
```
- Group = free-form nesting used to partition repos per distro/track. Verified: publish to root group and `el9` group (both 201), duplicate → 409.
- Owner can be the user (`xavierk`) or an org; packages under a public owner are readable anonymously (verified: metadata fetches without auth succeeded).
## 2. Consumer setup (exact commands)
### apt clients
```bash
sudo mkdir -p /etc/apt/keyrings
sudo curl -o /etc/apt/keyrings/gitea-xavierk.asc \
https://git.bongbetic.com/api/packages/xavierk/debian/repository.key
echo "deb [signed-by=/etc/apt/keyrings/gitea-xavierk.asc] https://git.bongbetic.com/api/packages/xavierk/debian bookworm main" \
| sudo tee /etc/apt/sources.list.d/gitea.list # one line per distribution
sudo apt update
apt install fenris # or fenris=0.3.0
# private owner variant: https://{user}:{token}@git.bongbetic.com/api/packages/... in the URL
```
### dnf clients
```bash
sudo dnf config-manager --add-repo https://git.bongbetic.com/api/packages/xavierk/rpm/el9.repo
# private owner: add user:token into the baseurl inside /etc/yum.repos.d/gitea-xavierk-el9.repo afterwards
sudo dnf install fenris # or fenris-0.3.0
```
The served `.repo` (verified live) sets `gpgcheck=1` and points `gpgkey` at `…/rpm/repository.key`, so `dnf` auto-imports on first use.
## 3. Metadata signing: native, not passthrough
Gitea **signs generated metadata itself** with per-instance auto-generated PGP keys. Client-side signing config is limited to trusting the served keys.
- Debian: `dists/{suite}/InRelease` is clearsigned; `Release.gpg` detached sig also served. Key (RSA) fetched from `…/debian/repository.key`, uid literally `(Automatically generated Debian Registry Key; created …)`.
- RPM: `repodata/repomd.xml.asc` detached ASCII-armored signature, uid `(RPM Registry)`. Key from `…/rpm/repository.key`.
- Both verified with `gpg --verify` → **Good signature** (keys are self-generated; the "not certified" warning is expected and handled by the signed-by/keyring flow above).
- The apt `Release` also advertises `Acquire-By-Hash: yes` with MD5/SHA1/SHA256/SHA512 indexes of `Packages`/`.gz`/`.xz` (verified live). RPM repomd carries sha256 checksums for `primary/filelists/other.xml.gz`.
There is **no bring-your-own-signing-key config** for these registries in 1.27 — trust anchor is the instance's auto keys. For Fenris this is acceptable; TOFU over TLS via the key URLs above.
## 4. Multi-distro metadata
- Debian: distributions/suites are implicit — whatever `{distribution}` path segments appear on upload become `dists/{distribution}/` trees with `Suite:`/`Codename:` set to the segment. No server-side list to maintain; adding a new distro = upload with new segment + one more `deb …` sources line. Components likewise (`main`, etc.). Architectures come from each `.deb`'s control stanza (index served as `dists/{dist}/{component}/binary-{arch}/Packages`).
- RPM: same via `{group}` path segments (`el9`, `rocky/el9`, …); each group gets its own `repodata/`. No `basearch` filtering — clients pick the group; Gitea publishes whatever RPM arch was uploaded.
## 5. Version retention
- Default: **all versions retained indefinitely**; nothing auto-deletes. Old versions stay installable (`apt install fenris=0.2.9`, `dnf install fenris-0.2.9`).
- Republishing an existing name+version (deb: same dist/component/arch; rpm: same file name in group) → 409; overwrite requires delete-then-upload.
- Optional cleanup rules exist (per owner + package type): `KeepCount`, `KeepPattern`, `RemoveDays`, `RemovePattern`, `MatchFullName` (source: `models/packages/package_cleanup_rule.go`, executed by scheduled `CleanupTask` in `services/packages/cleanup/cleanup.go`). In 1.27.1 they are configurable **only in the web UI** (owner → Packages → Cleanup Rules); no v1 REST route (verified by route table grep of `routers/api/v1/api.go` — probes of `/api/v1/packages/{owner}/cleanuprules…` return 404/409-style errors).
- Deletes: format-specific `DELETE …/debian/pool/{dist}/{component}/{name}/{version}/{arch}` and `DELETE …/rpm/{group}/package/{name}/{version}/{arch}` (both verified, 204). Deleting last file removes the version. Generic fallback: `DELETE /api/v1/packages/{owner}/{type}/{name}/{version}`.
## 6. Release attachment: not supported
Gitea 1.27.1 has **no package↔release linkage**. Release assets (`…/releases/{id}/assets`) are standalone file uploads; the package model has no release field and no route links them (verified against `v1.27.1` source: `routers/api/v1/repo/release_attachment.go`, `models/packages/`). Options for Fenris releases:
1. Publish `.deb`/`.rpm` to the registry (real apt/dnf install UX) and reference the registry URLs in release notes.
2. Additionally upload tarballs/SHA256SUMS as plain release attachments.
3. Generic registry (`PUT /api/packages/{owner}/generic/{name}/{version}/{filename}`) if an untyped artifact store is needed.
## 7. Caveats for the release plan
- Owner choice matters: publish under an **org** (e.g. `fenris`) if multiple maintainers need write; `xavierk` user owner works today (token owner is admin).
- Metadata access follows owner visibility — public owner → anonymous consumers, no token in URLs (current state, verified). Keep owner public for frictionless installs, or embed `user:token` in sources/baseurl.
- apt distro naming should match OS release names (`bookworm`, `trixie`, `noble`) purely for client convention; server accepts anything.
- RPM groups should mirror `$distver` (e.g. `el9`, `fc40`) so `.repo` selection is obvious per target.
## 8. Test log (live, 2026-09-03)
| Step | Result |
|---|---|
| `PUT debian/pool/bookworm/main/upload` | 201 |
| `PUT debian/pool/noble/main/upload` (multi-dist) | 201 |
| `PUT debian` duplicate | 409 (expected) |
| `PUT rpm/upload` (no group) | 201 |
| `PUT rpm/el9/upload` (group) | 201 |
| `PUT rpm` duplicate | 409 (expected) |
| `GET debian/repository.key` / `rpm/repository.key` | PGP public keys (200) |
| `GET dists/bookworm/{Release,InRelease,Packages}` | correct; `gpg --verify` Good signature |
| `GET rpm{,/el9}/repodata/repomd.xml{,.asc}` | 200; Good signature |
| `GET rpm{,/el9}.repo` | generated repo files with `gpgcheck=1` |
| Cleanup-rules REST probes | 404 (not in v1 API — UI only) |
| `DELETE` all four test entries | 204 ×4; package list then empty |
Sources: [docs.gitea.com 1.27 Debian registry](https://docs.gitea.com/1.27/usage/packages/debian), [docs.gitea.com 1.27 RPM registry](https://docs.gitea.com/1.27/usage/packages/rpm), Gitea source tag `v1.27.1` (`routers/api/v1/api.go`, `models/packages/package_cleanup_rule.go`, `services/packages/cleanup/cleanup.go`), live instance `git.bongbetic.com`.
+67
View File
@@ -0,0 +1,67 @@
# Glint dashboard design adoption
Research date: 2026-09-19. Status: feasibility findings with interview decisions in progress. This note does not authorize implementation or change Fenris's accepted behavior.
## Finding
Fenris can adopt a Glint-inspired terminal dashboard through an independently authored redesign of its existing Textual presentation layer. Glint is a Rust terminal application, not an HTML/CSS dashboard or a drop-in Textual component library. Its most relevant ideas are a compact pane grid, consistent title rows and metadata, obvious focus, restrained colors, and enlarging a focused pane. A wholesale Glint integration would introduce a different UI runtime, unrelated application infrastructure, and a licensing decision without being necessary to achieve this visual direction.
The user has selected **Chalktone / screenshot 3**, **a large activity chart as the primary panel**, and **panel styling plus keyboard focus and zoom**. The activity panel has **Live / Day / History tabs**; enlarging a panel retains a **fixed monitoring status, freshness, and control strip**. A general dashboard builder is outside the selected scope. The developing product specification is [Glint dashboard design](../spec/glint-dashboard-design.md).
## Primary-source snapshot
Inspected Glint's default branch at commit [`c1d73d3e8ead2f4069630b2a237306af8f6e69c8`](https://github.com/ntrospect0/glint/tree/c1d73d3e8ead2f4069630b2a237306af8f6e69c8), committed 2026-07-19. A shallow reference checkout was created outside Fenris at `/tmp/fenris-glint-reference`.
- [README](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/README.md) contains three dashboard screenshots, a setup screenshot, and a live-capture link. The first two dashboard screenshots use `tokyonight`; the third uses `chalktone`. They show example compositions, not one mandatory layout.
- [Screenshot 1](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/docs/screenshots/glint-demo1.png), [screenshot 2](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/docs/screenshots/glint-demo2.png), [screenshot 3](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/docs/screenshots/glint-demo3.png), and [setup screenshot](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/docs/screenshots/glint-setup.png) are versioned in the repository. This source investigation does not substitute for visual inspection of those images.
- [Cargo.toml](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/Cargo.toml) declares Rust edition 2021, package version 0.5.0, Ratatui 0.28, Crossterm 0.28, and a `glint` binary. It also brings Tokio, HTTP clients, configuration/watch infrastructure, and optional widget dependencies.
- [GitHub repository API](https://api.github.com/repos/ntrospect0/glint), checked on the research date, reports an unarchived repository created 2026-05-27. [GitHub releases API](https://api.github.com/repos/ntrospect0/glint/releases) returned no releases. The [changelog](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/CHANGELOG.md) labels 0.5.0 and 0.4.0 unreleased, and README installation is from source. This is evidence of a young project and its distribution state, not proof that it is abandoned.
- [CI configuration](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/.github/workflows/ci.yml) runs Cargo tests and Clippy on Ubuntu; Clippy warnings do not fail CI. Source review alone does not establish that current CI passes. No Glint build or test run was performed for this assessment.
## What can transfer
| Glint pattern and source | Fit for Fenris | Scope implication |
| --- | --- | --- |
| Pane grid with row/column spans: [layout model](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/config/layout.rs), [default configuration](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/config/defaults/config.toml). | Fenris already uses a Textual grid and rounded bordered panes. A considered re-layout can use that existing ownership. | Decide which Fenris facts deserve persistent space. A fixed layout is much smaller than Glint's user-composable dashboard system. |
| Title integrated into the border, right-aligned metadata, focus treatment, shortcut indication: [UI title helpers](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/ui/mod.rs). | Useful for clearly named activity, history, drive, and monitoring panes. Date/range/freshness can become concise pane metadata where legible. | Independently implement the visible behavior in Fenris; do not translate or copy these GPL helpers. Ensure essential evidence is not merely truncated away. |
| Semantic theme roles separating border, title, metadata, and text: [theme model](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/theme/mod.rs), [bundled schemes](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/config/defaults/colorschemes.toml). | Fenris already owns themes in `src/fenris/themes.py`. Extend that role system for a chosen visual direction while keeping semantic status colors and text. | Select the preferred screenshot/palette before making the default. High Contrast and reduced-motion preferences already exist and must remain usable. |
| Tab/click focus, keyboard shortcuts, and focused-pane enlargement: [README controls](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/README.md), [app focus/zoom ownership](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/app.rs). | Particularly useful for graphs when a small terminal limits detail. | Zoom is new interaction behavior, not just styling. Selection/date/focus must survive entry, exit, resize, and refresh; essential fault/pause state needs a visibility decision. |
| Multiple widgets stacked in one cell, cycled with `.`/`,`: [stack implementation](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/widgets/stack.rs). | A possible way to expose secondary details without permanent screen cost. | Hiding monitoring state, confidence, or evidence behind inactive tabs would weaken Fenris's current contract. Choose deliberately; not assumed needed. |
| Responsive views exposing more detail at larger sizes: [view tiers](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/widgets/view_tier.rs), [widget author guide](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/docs/widget-sdk.md). | Useful principle, but Glint's breakpoints describe its content and are not Fenris requirements. | Preserve Fenris's constrained-terminal facts and controls, then design larger views. Do not simply reuse Glint's threshold values. |
| Compact bottom status bar: [status bar](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/ui/status_bar.rs). | Supports a quieter action/focus rail. | Fenris has real monitoring continuity, freshness, collection outcome, authentication, and quit semantics. Its footer cannot be reduced to Glint's version/clock/theme bar. |
Glint's calendar, stocks, news, email, weather, notes, galleries, external credentials, and profile/setup machinery are outside the stated Fenris redesign. Its widget SDK adds widgets inside the Rust application via a trait and registry; it is not an existing Python integration seam ([widget SDK](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/docs/widget-sdk.md), [widget interface](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/src/widgets/mod.rs)).
## Fenris ownership and constraints
The present application is Python with Textual, pinned to Textual 8.2.8 in [`requirements.txt`](../../requirements.txt). [`FenrisTuiApp`](../../src/fenris/tui.py) already owns the grid, header, bordered panes, compact layout, help, action bindings, and theme registration; `DailyBarGraph` and `LiveActivityGraph` already own graph selection/rendering. [`themes.py`](../../src/fenris/themes.py) owns Amber, Nord, High Contrast, and status color precedence. These are the natural extension points; a new backend or service is unnecessary for a visual redesign.
The design must preserve these existing responsibilities and evidence semantics:
- The collector owns privileged drive acquisition and observation-store writes; the TUI is an unprivileged reader and dispatches administrative actions through the sanctioned helper. A visual redesign does not imply a second sampler or store writer. Sources: [`CONTEXT.md`](../../CONTEXT.md), [redesign spec](../spec/fenris-redesign.md), [`tui.py`](../../src/fenris/tui.py).
- Usage-adjusted theoretical lifespan remains a write-endurance projection, accompanied by projection confidence and supporting facts. It must not become a physical failure countdown. Unknown, unavailable, warming, stale, and store-fault states require meaningful presentation. Sources: [`CONTEXT.md`](../../CONTEXT.md), [projection ADR](../adr/0002-projection-model-sustained-regime.md), [live activity specification](../spec/live-drive-activity.md).
- Gaps must not become zeros, and unavailable measurements must not become decorative smooth curves. Local-day labels, read/write units, incomplete evidence, selected dates, and timezone context matter. Sources: [local-day history ADR](../adr/0010-local-day-activity-history.md), [live activity specification](../spec/live-drive-activity.md).
- Pause is a deliberate disable; quitting only leaves the TUI. Monitoring continuity, last collection outcome, and freshness must remain understandable. Sources: [service lifecycle ADR](../adr/0003-service-lifecycle-and-sanctioned-toggle.md), [dashboard clarity specification](../spec/dashboard-clarity.md), [`status_composition.py`](../../src/fenris/status_composition.py).
- The current source switches to constrained presentation below 80 columns or 24 rows; the accepted live-activity specification requires daily totals, dates, evidence labels, confidence, and controls to remain accessible on narrow terminals. A new design must be reviewed against real target dimensions. Sources: [`tui.py`](../../src/fenris/tui.py), [live activity specification](../spec/live-drive-activity.md).
The fixed Panes information architecture is normative in the older [redesign specification, section 7](../spec/fenris-redesign.md), and [dashboard clarity](../spec/dashboard-clarity.md) says it preserves that architecture and prescribes a separate quit rail. A significantly different pane hierarchy or footer should explicitly amend those presentation decisions when the user selects the new design, rather than silently treating old contracts as irrelevant. The underlying evidence and action semantics can be retained.
## License boundary
Glint declares **GPL-3.0-or-later** in [Cargo.toml](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/Cargo.toml) and source SPDX headers. Its [README license section](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/README.md) says distributed modified versions must remain GPL-licensed with copyright notices, and its [LICENSE](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/LICENSE) specifies the distribution obligations. Fenris deliberately uses MIT across source and packages under [ADR 0009](../adr/0009-mit-license.md).
Copying or translating Glint implementation into an integrated distributed Fenris would therefore raise GPL compliance and licensing choices that conflict with keeping the combined implementation solely MIT. It is not an adoption route to take implicitly. The recommended route is independently authored Fenris UI code using common layout/interaction ideas observed in Glint, without copying Glint code or assets. If direct implementation reuse becomes a requirement, resolve licensing or separate permission first; this research does not establish that permission.
## Interview state
The selected direction is an independently implemented Chalktone-inspired dashboard with a fixed Fenris layout, an activity-first hierarchy, keyboard focus and enlargement, one activity panel with Live / Day / History tabs, and a persistent monitoring status/freshness/control strip during enlargement. Existing 80×24 support and constrained reflow are retained rather than reopened as a new minimum-size decision. The [developing design specification](../spec/glint-dashboard-design.md) owns the precise requirements and any remaining decisions; this research note is supporting evidence.
No product code, glossary, ADR, or existing acceptance criteria were changed by this investigation. Keyboard focus/zoom feasibility does not establish that exact interaction behavior has been implemented or runtime-validated in Fenris.
## Validation scope
Evidence comes from the pinned Glint source, its checked-in documentation/screenshots, GitHub's first-party API, and Fenris's current source and accepted documents. The main design investigation visually inspected screenshots 1 and 3 and separately verified Textual grid documentation.
The main investigation also reviewed synthetic renders of current Fenris at 140×44 and 80×24, after refresh cleared the launch authentication banner, using installed Textual 8.2.7 (the repository lock is 8.2.8). The captures are `/tmp/fenris-current-capture-3xdv46jm/fenris-current-140x44.png` and `/tmp/fenris-current-capture-3xdv46jm/fenris-current-80x24.png`. They showed a large blank vertical area; at 80×24 the drive, monitoring status, and actions fell below the initial viewport. This supports rearranging the panels and keeping the status/control strip visible, but is not a diagnosis of the deployed drive or proof across every state.
No Glint runtime was executed, and no Glint asset was copied into the product. No accessibility audit was conducted. This research establishes architectural feasibility and records decisions; it does not claim pixel fidelity, a working integration, an upstream test result, or a measured performance result.
+73
View File
@@ -0,0 +1,73 @@
# Research: OBS as an alternative build + distribution route
Resolves [Research: OBS as alternative build + distribution route](https://git.bongbetic.com/xavierk/Fenris/issues/36) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/33).
- Date: 2026-09-03
- Verdict: **Reject OBS now; ship via the self-hosted Gitea 1.27.1 registry** (deb + rpm), and revisit OBS only if publishing reach becomes a goal.
## Question
Evaluate openSUSE Open Build Service (OBS) as the build + distribution route — deb build quality, vendoring `textual>=0.40` via source services (offline sandbox), signing, publishing reach, account/maintenance cost, build latency — against the Gitea registry on our matrix: Debian 12, Ubuntu 22.04/24.04, Fedora 40+, x86_64.
## Findings
### 1. deb build support quality — real, with quirks
- OBS builds deb via the classic recipe trio: `debian.control`, `debian.rules`, `PACKAGE.dsc` (OBS User Guide §2.3 "Debian: Dsc"). The build phase runs `dpkg-buildpackage` on Debian-based distributions (§25.1.3 "Package Build"); Debian build environments can alternatively use the `debootstrap` build engine (§"Configuration File Syntax", `BuildEngine`).
- Quirk: release numbers are **not** auto-incremented across rebuilds unless the dsc carries `DEBTRANSFORM-RELEASE` (§2.3) — a packaging decision we'd own either way.
- Upstream build deps are available: `dh-virtualenv` and `dh-python` exist in Debian 12 (packages.debian.org, checked 2026-09-03), so the ADR-0004 venv/lockfile design maps onto an OBS dsc without patching the build root.
- All five matrix targets exist as public OBS build roots: `Debian:12`, `Ubuntu:22.04`, `Ubuntu:24.04`, `Fedora:40`, `Fedora:41` — each project `_meta` answered HTTP 200 on build.opensuse.org (checked 2026-09-03).
- Live proof of deb publishing quality: `isv:ownCloud:desktop/Debian_10` on download.opensuse.org serves a proper Debian archive (`Release`, `Release.gpg`, `InRelease` all HTTP 200, checked 2026-09-03).
### 2. Vendoring textual≥0.40 — the offline sandbox forces the same work we already planned
- The build environment has **no network**: "services requiring external network access are likely to fail in [buildtime] mode, because such access is not available if the build workers are running in secure mode (as is always the case at https://build.opensuse.org)" (User Guide §7.2, "Modes of Source Services"); Dockerfile builds likewise run "in a safe build environment without network access" (§29.3).
- Vendoring must therefore happen **before** the build, via source services that run server-side on commit (`default`/`trylocal` modes, §7.2) or via files committed to the package. The standard services are per-file fetchers — `download_url` (§22.1.3), `download_files`, `obs_scm`/`tar`/`set_version` (§8 SCM integration) — there is no "pip resolve" service, so a pinned dependency tree like Fenris's means either N `download_url` entries mirroring the committed lockfile, or simply committing the vendored wheel/sdist tree.
- Conclusion: OBS does not remove the vendoring step; it reproduces ADR-0004's committed-lockfile design with extra XML. Since distro `python3-textual` is 0.1.13 on Debian 12, Ubuntu 22.04 and 24.04 (packages.debian.org / packages.ubuntu.com, checked 2026-09-03) — far below the `>=0.40` floor — vendoring is unavoidable on any route.
### 3. Signing — OBS key, not ours; Gitea deb repo is our key
- OBS signs published repositories with the **instance's** key: one signer per partition "calls an external tool to execute the signing" (User Guide §23 "OBS Architecture", Signer); consumers accept the OBS repo key ("When prompted, accept the GPG key of the download repository", §1.10). A build.opensuse.org user cannot upload a personal signing key. Trust therefore flows to openSUSE infra, and the signature says nothing about Fenris's maintainers.
- Gitea 1.27.1's Debian registry serves apt metadata signed with the Gitea instance's PGP key (`repository.key` endpoint, `signed-by` in sources.list — docs.gitea.com, "Debian Package Registry"), i.e. **our** host and **our** key. The RPM registry serves a `.repo` endpoint but documents no GPG signing of repodata; rpm-file signing stays our choice at build time.
### 4. Publishing reach — OBS wins reach; reach is not our bottleneck
- OBS publishes home-project results to `https://download.opensuse.org/repositories/home:USER/<dist>` (§1.10) and offers generated download pages on software.opensuse.org (§17.4). That is genuine CDN-class reach.
- Caveats from the same docs: branched projects are **not** published by default (§1.10), and the repo is a live view of the project state — no release artefact pinning; deleting the project or flag disables distribution.
- The Gitea route's reach is exactly `git.bongbetic.com` plus whatever the README says — adequate for a named four-distro matrix whose users follow our instructions, and it keeps the release artefact under versioned control on the same host as the source.
### 5. Account and maintenance cost — strictly additive
- Using build.opensuse.org requires an openSUSE account (single sign-on; the web UI's "Sign up!") and work happens in `home:USERNAME` plus permitted subprojects (§"Setting Up Your Home Project for the First Time"; §23 "OBS Concepts" on home projects).
- Day-to-day: `osc` + `_service` XML + dsc/spec recipes maintained in OBS's own package VCS, kept in sync with Fenris's git. The SCM bridge (`scmsync`) does support self-hosted Gitea ("We also support Self-Hosted instances from GitHub, GitLab and Gitea", §8.1.3; setup in §28.1.2 — build descriptions must live in the repo's top level), but it also disables OBS-side workflows (no `_link` merging, limited workflows, §28.1.1).
- No published quota/SLA for the public instance; capacity and availability are a shared commons. The Gitea route needs zero new accounts, zero new artefact formats beyond the two package recipes we must write anyway, and reuses the existing release host.
### 6. Build latency — shared queue vs. deterministic local
- OBS routes every commit through scheduler → dispatcher → shared workers; the dispatcher "tries to assign jobs fairly between the project repositories" using a per-repository load model (§23, Scheduler/Dispatcher). For our five tiny x86_64 jobs this is typically minutes, but there is no documented SLA and the queue is global — worst case is unbounded (estimate; the docs guarantee only fairness, not latency).
- The Gitea route builds wherever `make` runs and publishes with one authenticated `PUT` per artefact (docs.gitea.com: Debian `PUT .../pool/{distribution}/{component}/upload`; RPM `PUT .../rpm/{group}/upload`). Latency = build time, fully under our control.
## Comparison on the 4-distro matrix
| Axis | OBS (build.opensuse.org) | Gitea 1.27.1 registry |
|---|---|---|
| Debian 12 / Ubuntu 22.04/24.04 deb | dsc + dpkg-buildpackage; DEBTRANSFORM-RELEASE quirk | we build the same deb locally, upload via PUT |
| Fedora 40+ rpm | spec + rpmbuild in Fedora roots | same spec built locally, `.repo` grouping (`fedora/40`) |
| Vendoring textual≥0.40 | offline sandbox forces committed vendored tree (no pip service) | same committed vendored tree (ADR-0004 lockfile) |
| Signing | OBS instance key (not ours) | deb repo signed with our key; rpm repodata unsigned |
| Reach | download.opensuse.org CDN + software.o.o pages | our domain only |
| Accounts/infra | new openSUSE account, osc workflow, commons SLA-free | zero new infra |
| Latency | global shared queue, minutes typical, no SLA | deterministic (local build) |
## Recommendation
**Reject OBS as the build + distribution route for Fenris.** The offline sandbox forces the exact vendoring work the Gitea route already requires, so OBS adds cost (account, osc/source-service maintenance, external commons in the release path, queue latency) without removing any; its one real advantage — CDN and software.o.o reach — does not matter for a hobby project whose four target distros are served by one signed apt repo and one rpm repo on the existing Gitea host, under our own key.
Revisit trigger: if Fenris later wants one-click installs via software.opensuse.org, architectures beyond x86_64, or many more distro targets — the deb publishing quality (verified live) and self-hosted-Gitea SCM bridge make OBS a viable amplifier then.
## Sources
- OBS User Guide (openbuildservice.org/help/manuals/obs-user-guide/, PDF): §2.3 Debian: Dsc; §7 Using Source Services (offline buildtime services, modes); §8.1.3 Supported SCMs; §17.4 download pages; §22.1.3 download_url; §23 OBS Architecture (Scheduler/Dispatcher/Signer); §25.1.3 Package Build; §28.1 SCM bridge; §29.3 Dockerfile builds (no network); §1.10 Installing Packages from OBS; "Configuration File Syntax" (BuildEngine, Repotype: debian).
- Live checks (2026-09-03): `Debian:12`/`Ubuntu:22.04`/`Ubuntu:24.04`/`Fedora:40`/`Fedora:41` project `_meta` on build.opensuse.org (all 200); `isv:ownCloud:desktop/Debian_10` `Release`/`Release.gpg`/`InRelease` on download.opensuse.org (all 200).
- packages.debian.org / packages.ubuntu.com (2026-09-03): `python3-textual` 0.1.13 on bookworm, jammy, noble; `dh-virtualenv`, `dh-python` present in bookworm.
- docs.gitea.com, "Debian Package Registry" and "RPM Package Registry" (1.27 line): apt sources with `signed-by` + `repository.key`, `PUT` upload endpoints, `.repo` groups.
+146
View File
@@ -0,0 +1,146 @@
# Acceptance criteria: the Fenris redesign
Status: Accepted — resolves [Define cross-cutting acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/issues/13) on the [Wayfinder map](https://git.bongbetic.com/xavierk/Fenris/issues/1). These criteria are the accepted definition of done for the finished redesign; the implementation-ready specification assembles them with ADRs 0001–0006 at handoff.
## Framework
- **Canonical term**: *acceptance criterion* — one testable behavioral statement. "Behavioral gate" is avoided as a synonym. Wording follows the repository glossary (`CONTEXT.md`).
- **Evidence classes** — every criterion carries exactly one:
- **A** — automated test (unit/integration, fixture-driven).
- **P** — scripted system probe on a host with systemd, polkit, and the configured NVMe device.
- **M** — manual checklist, reserved for interactions a fixture cannot capture (live polkit agent prompts, TUI keyboard feel).
- Whatever can be automated must be; **M** only where automation cannot reach.
- **Traceability-only**: every number and behavior cites the ADR or ticket that fixed it. Nothing undecided enters here; new demands become new tickets, never criteria.
- **Organization**: criteria are grouped by subsystem, with cross-cutting invariants spanning them. Coverage spans all decided areas — observation store and migration, collector lifecycle and privileges, projection and confidence, controller identity, the Panes TUI, failure paths, and installation lifecycle.
- **Placeholders**: none remain. SLOT-B was filled by [Choose the collector's NVMe acquisition path](https://git.bongbetic.com/xavierk/Fenris/issues/16) as AC-1–AC-5; SLOT-A was filled by [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15) as PR-15, PR-16, and ID-4.
- **Test-plan boundary**: Given/When/Then test specs are derived by the implementer at implementation time. This effort produces criteria only.
## Cross-cutting invariants
- **CI-1** (A; ADR 0002 §§6–8, ADR 0003 §10) *Exhaustive state matrix*: from a synthetic observation store, the TUI and `fenris status` render every realizable combination of confidence state (Unavailable, Limited, Supported) × freshness grade (fresh, missed, stale, empty store) × endurance-baseline tier (verified override, unverified override, implied, none) exactly as the ADR 0002 rule table and ADR 0003 freshness constants dictate — headline number present only when the rules allow it, contributing facts always, never a percentage.
- **CI-2** (P/A; ADR 0003 §§4–9) *TUI/CLI parity*: every TUI action has a CLI twin — pause (`fenris monitor pause`), resume (`fenris monitor resume`), collect-now (`fenris sample`), baseline set/clear, and the status fact set — with identical outcomes and wording.
- **CI-3** *Prohibition set* (A = test, P = probe; each cites its clause):
- No code path outside `fenris-collect` interrogates the device (ADR 0003 §7).
- No `/run/fenris` coordination surface or export layer exists anywhere; state lives in the observation store and coordination in systemd (ADR 0003 §1, ADR 0001 §2).
- Polkit authorizes exactly one binary, `fenris-monitor`, under `com.bongbetic.fenris.monitor` `auth_admin` (ADR 0003 §5, ADR 0004 §3).
- No absent hour is ever interpolated, estimated, or fabricated (ADR 0005 §2).
- No alerting, notification, or escalation machinery exists anywhere (ADR 0005 §§5–6).
- `/etc/fenris/fenris.conf` holds exactly one key — the device selector (ADR 0003 §3).
- No synthetic or capacity-derived baseline is ever created, including for legacy history (ADR 0002, Consequences).
- Readers never partially interpret a newer-schema store (ADR 0001 §8, ADR 0005 §4).
- Projections are never stored; always recomputed on read (ADR 0001 §3, ADR 0002 §13).
- **CI-4** (A; ADR 0002 §§11–12) *User-facing language*: TUI and status render the endurance research's required wording and six disclosures as adopted; zero-rate and unavailable cases use their exact phrasing; the scenario range is the only spread shown anywhere.
## Observation store and legacy migration (ADR 0001)
- **ST-1** (P) One SQLite database in WAL mode at `/var/lib/fenris/observations.db`, root-owned and group-readable through the `fenris` read group; the TUI opens it read-only.
- **ST-2** (A) An unprivileged reader querying during a collector write sees a consistent snapshot.
- **ST-3** (A) The schema carries `samples`, `hour_observations`, `day_aggregates`, `monitoring_periods`, `controller_segments`, and `endurance_baseline` with the ADR 0001 column sets as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) and [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14).
- **ST-4** (A) Hours and days are UTC-bounded; day derivation from hours is monotonic; no 23- or 25-hour days exist.
- **ST-5** (A) Raw samples are pruned opportunistically to 14 days; hour observations and day aggregates are retained indefinitely.
- **ST-6** (P) Legacy import is one transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
- **ST-7** (A) Import is idempotent: a second run no-ops on the legacy-import marker.
- **ST-8** (A) Legacy files are renamed `*.migrated` only after commit and never deleted.
- **ST-9** (A) Malformed legacy lines are quarantined with a logged count, never silently dropped.
- **ST-10** (A) `hourly.jsonl` is never trusted: mismatches against derived data are diffed and logged.
- **ST-11** (A) Migration opens one implicit monitoring period at the first legacy sample, closed `end_cause = migrated` at the migration moment; pre-migration hours carry an unknown activity split except directly evidenced facts.
- **ST-12** (A) Schema versioning: `PRAGMA user_version` with ordered, per-step transactional migrations; the collector refuses an unknown newer version.
## Collector lifecycle and privilege boundaries (ADR 0003)
- **LC-1** (P) Exactly two system units exist — `fenris-collect.timer` (`timers.target`) and `fenris-collect.service` (`Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`); the TUI and CLI are ordinary unprivileged processes and never units.
- **LC-2** (P) Timer defaults ship as `OnBootSec=2min`, `OnUnitInactiveSec=3min`, `AccuracySec=30s`, `Persistent=no`, `TimeoutStartSec=90s`; cadence changes are documented drop-ins and no interval key exists in configuration.
- **LC-3** (P) A hung device interrogation fails visibly within `TimeoutStartSec=90s` as a bounded failed run retried next interval.
- **LC-4** (A/P) `/etc/fenris/fenris.conf` holds exactly the device selector (stable `/dev/disk/by-id/…` path; raw nodes warned), re-read every run; an invalid selector is a bounded failed run surfaced as `configuration error: <reason>` in `status` and the TUI.
- **LC-5** (P) Two privileged binaries ship at `/usr/libexec/fenris/fenris-collect` and `/usr/libexec/fenris/fenris-monitor`; the unprivileged `fenris` wrapper opens the TUI with no arguments.
- **LC-6** (P+M) Pause = `fenris-monitor disable --now` asks for confirmation; Resume = `enable --now` does not; both perform the systemctl operation and period bookkeeping in one step under polkit `com.bongbetic.fenris.monitor` (`auth_admin`), failing cleanly with the printed root equivalent where no polkit agent exists. (M covers the live agent prompt.)
- **LC-7** (A) The period-row idempotent matrix of ADR 0003 §6 holds exactly: first-ever enable opens; resume with an open period changes nothing; resume without one opens anew; pause with an open period closes `user_disabled`; pause otherwise no-ops; a raw systemctl stop/disable never records `user_disabled`.
- **LC-8** (P) `fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, block until exit, and report the outcome (freshness line or journal hint) synchronously; the TUI never samples in-process.
- **LC-9** (A/P) CLI compatibility: `status` is a read-only composition (projection facts, enabled/active, last collect outcome, `journalctl` hint on failure or staleness) that never auto-samples and never prompts; `sample` is retained via the helper; `--device` is rejected with a pointer to the configuration file; `start`, `stop`, and `run` are rejected with one-line migration pointers; `fenris.sh` is not shipped and is removed from the repository; the README maps its five menu options to successors.
- **LC-10** (A) Freshness constants are defined once and shared by TUI and CLI: fresh = newest sample within 2× cadence + `AccuracySec` + 60 s; missed between that and 48 h; stale ≥ 48 h; an empty store reads "no observations yet" with an enable hint; freshness derives from the newest sample timestamp, never a stored health flag.
## Projection contract and confidence (ADR 0002)
- **PR-1** (A) Exactly one projection, from the precedence-chosen baseline (verified override → unverified override → implied → unavailable); Percentage Used renders as a vendor-wear context line, with a note when it disagrees with the observed write rate by more than a factor of 2; the PU-slope regression and `capacity × 600` synthesis are gone.
- **PR-2** (A) The headline rate is the sustained-regime rate (regime DUW bytes ÷ in-period wall-clock seconds), default regime = full history capped at 90 days; the 7/28/90-day scenario range is computed independently and shows only covered horizons, with no placeholders.
- **PR-3** (A) Habit change: trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean for 3 consecutive days starts a new regime at the first divergence day, adopted automatically and labeled "usage habit changed N days ago"; a regime younger than 7 days caps confidence at Limited evidence.
- **PR-4** (A) Hour classification uses the named constants: powered-off below 90% of power-on-hours span; active at ≥ 256 MiB DUW; idle below it while powered on and sampled; unknown otherwise; disabled time is wall-clock outside monitoring periods, never an hour state.
- **PR-5** (A) The denominator is wall-clock seconds inside monitoring periods including powered-off and unknown time; disabled periods are excluded from numerator and denominator; unexplained gaps keep the aggregate counter delta, remain as unknown seconds, and reduce coverage.
- **PR-6** (A) Warming up until 14 distinct UTC day aggregates of which at most 2 fall below 50% coverage; the projection still renders with its facts while warming; every Unavailable condition renders no lifespan number.
- **PR-7** (A) A newest day aggregate older than 48 h drops confidence one level and is shown as a contributing fact.
- **PR-8** (A) The confidence rule table of ADR 0002 §8 holds verbatim, rendering state plus contributing facts and never a percentage.
- **PR-9** (A) Segment breaks: a DUW decrease with unchanged identity keeps prior day aggregates as habit evidence with the projection Unavailable until re-warm; a controller-identity change quarantines prior history from projection entirely.
- **PR-10** (A) The implied baseline is eligible only after ≥ 2 Percentage-Used increments within the current controller segment; until then, Unavailable with "vendor wear estimate too coarse to imply endurance".
- **PR-11** (A) Zero rate renders "no finite projection from this history" — never infinity or zero; no statistical confidence interval appears anywhere.
- **PR-12** (A) The projection contract hands the TUI exactly: confidence state, contributing facts, headline remaining time when one exists, scenario range, Percentage-Used context line, disclosure text — recomputed on read, never stored.
- **PR-13** (A) Baseline provenance and validation per [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12): mandatory provenance (URL, revision, entry date, model, nominal capacity); one active row replaced on edit; verification derived at read (machine match or recorded attestation), never a stored boolean; incomplete provenance stores only behind explicit acknowledgment as the unverified tier; entry-time unprivileged sysfs validation (normalized model containment; capacity within ±1%; interactive confirm recorded as `validated_by = user`); read-time applicability is a model match against the current controller segment, with a mismatch retained — never auto-deleted — leaving the projection Unavailable.
- **PR-14** (P) `baseline set` / `baseline clear` persist through the polkit-guarded `fenris-monitor` verb after CLI-side validation.
- **PR-15** (A) An identity-degraded controller segment (blank identity key — every rung of the key ladder empty) caps projection confidence at Limited evidence, with the contributing fact "controller identity unavailable — replacement detection relies on write-counter continuity only" rendered in every state; the cap combines idempotently with the 48-hour staleness drop, and ephemeral markers (model "Linux", non-pcie transport) never render as confidence facts ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15); ADR 0002 §§8–9 as amended).
- **PR-16** (A) Identity-change semantics extend to blank keys verbatim: any visible change of the recorded identity key — including to or from a blank key — quarantines prior history from projection as a controller-identity change, while equal blank keys continue the segment segmented by DUW monotonicity alone (ADR 0002 §9 as amended).
- **PR-17** (A) Projection arithmetic is exactly `E_rated = entered_TBW × 10¹²` bytes, `E_implied = 100 · W_t / p` computed only for `1 ≤ p ≤ 254` (Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable), and `projected = max(E_baseline − W_t, 0) / rate` for `rate > 0` (ADR 0002 §2).
## Controller identity ([Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11), [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14); ADR 0001 §3 as amended)
- **ID-1** (A) The controller-segment identity key is the normalized kernel-exposed subsystem NQN, with the kernel composite then model|serial as fallbacks; FR is metadata only; identity change and DUW decrease act as independent axes.
- **ID-2** (A) Segments freeze a fully nullable metadata snapshot at open — normalized `subnqn`/`sn`/`mn`/`fr` plus `vid`/`ssvid`/`transport` and `identity_degraded` — immutable thereafter, with `cntlid` excluded.
- **ID-3** (A) Legacy history imports under a labeled model-scoped legacy identity (mn-only segments).
- **ID-4** (A) `identity_degraded` is set at segment open exactly when the identity key is blank; keys from the kernel-composite or `model|serial` rungs are not degraded ([Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15)).
## Panes TUI ([Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3), [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6); ADR 0003 §§8, 10; ADR 0004 §10)
- **TUI-1** (A) Variant A "Panes": one dense keyboard-first screen; confidence rendered as evidence (state + contributing facts); boot enablement, runtime activity, last collect outcome, and freshness displayed as four separate facts.
- **TUI-2** (M) Pause/resume asymmetry and polkit tty passthrough work in a live terminal: pause confirms, resume does not, and the platform agent prompts without breaking the TUI.
- **TUI-3** (P) Textual runs on Python 3.9+, gated at install time, never a runtime crash.
- **TUI-4** (A) The Panes screen layout is normative: a full-width headline band (lifespan headline or its no-projection wording, confidence state with contributing facts, scenario range); a usage-history pane on the left (write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus legend, habit-split bar with active/idle/powered-off/unknown shares); a drive-health and settings pane on the right (health facts, vendor-wear context line, read-only settings with the endurance baseline and its provenance label); a full-width service strip at the bottom (the four separate service facts, the monitoring-period line, the action legend). Production bindings are the footer `p pause · r resume · c collect · d disclosures` — pause asks, resume does not — plus a bordered quit rail `q QUIT TUI` visually separate from monitoring state; the rail owns quit and the footer carries no quit entry (bindings amended by [Lock the dashboard wording strings](https://git.bongbetic.com/xavierk/Fenris/issues/57); original [Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3)); the prototype branch is visual reference only.
## Failure and recovery (ADR 0005)
- **FL-1** (A) The collector validates every row against the store invariants (hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count); a violating run writes nothing, logs the refused row, and fails visibly.
- **FL-2** (A) Readers defensively exclude and count malformed rows as a contributing fact.
- **FL-3** (A) No backfill ever: gaps remain unknown seconds; degradation flows only through coverage, freshness facts, and confidence categories.
- **FL-4** (P/A) A store fault surfaces "observation store unreadable" with a journal hint and suppresses everything else store-dependent; the collector treats it as a bounded failed run and never recreates or overwrites the file; recovery is the documented human-sanctioned move-aside (with `history.jsonl` re-import if the legacy import never completed); no built-in destructive command exists.
- **FL-5** (A) A newer-schema store renders "observation store written by a newer Fenris — upgrade Fenris" in TUI and status, with no partial interpretation.
- **FL-6** (A/P) Repeated collector failures retry at flat cadence with no backoff or notification; persistence reads as stale exactly like any other gap.
- **FL-7** (A) `critical_warning`, media errors, and unsafe shutdowns render as ordinary facts in TUI and status and never affect the projection.
- **FL-8** (A) A collection run finding no open monitoring period opens one at the run moment, never backdated.
## Installation, upgrade, and removal (ADR 0004)
- **IN-1** (P) `sudo make install` builds a wheel from the checkout and installs pinned dependencies into the dedicated venv at `/opt/fenris`, with a `/usr/local/bin/fenris` wrapper; after install nothing references the checkout.
- **IN-2** (P) The installer records every placed file in an explicit manifest consumed by upgrade and uninstall.
- **IN-3** (P) The installer never enables or starts units: a fresh install is dormant (units disabled, nothing running, no monitoring period); the only opt-in is the sanctioned toggle — `fenris monitor resume [--now]` or the first-run TUI prompt — enabling the timer and opening the first period in one step.
- **IN-4** (P) Install-time legacy import detects `./data/history.jsonl` (or an explicit path), runs the idempotent single-transaction import, and reports imported counts; `fenris import <path>` remains available.
- **IN-5** (P) `sudo make upgrade` installs into the same venv, syncs units and polkit against the manifest (`daemon-reload`; timer restarted only if unit contents changed and it is active), never kills an in-flight collection run, then applies forward-only schema migrations; `/var/lib/fenris` is never rebuilt.
- **IN-6** (P) Before migrations, `observations.db` is snapshotted to a one-generation `.bak`; rollback is reinstall-previous plus restore; automatic schema downgrade does not exist.
- **IN-7** (P) `make uninstall` performs the sanctioned disable first (open period closes `user_disabled`), then removes venv, helpers, units, polkit policy, and wrapper while keeping `/etc/fenris` and the observation store; `make purge` additionally removes configuration and store.
- **IN-8** (P) Dependencies are exact pins in a committed lockfile installed by both install and upgrade; refreshing pins is an explicit `make update-deps` step, never an install side effect.
- **IN-9** (P) The installer verifies `python3 ≥ 3.9` and fails cleanly otherwise; `/var/lib/fenris` is created with root-written group-read permissions; the database file is created lazily by the first write.
- **IN-10** (P) Installed artifacts sit only at their fixed locations — units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/`, configuration at `/etc/fenris`, observation store under `/var/lib/fenris` — and every placed file is recorded in the manifest (ADR 0004 §2; ADR 0003 §4).
## Migration from make-install systems (ADR 0007 §10, spec §9)
- **MG-1** (M) The migration runbook is published in the install docs (`docs/install/migrate-from-makeinstall.md`): mandatory remove-then-install steps, why over-install is forbidden (stale admin-directory units silently shadow vendor units; the local wrapper shadows the package wrapper), no-move continuity, and the reset-to-dormant expectation (the user opts back in with the sanctioned resume).
- **MG-2** (A) The install guard is verified across the matrix: either make-install marker (the legacy placement manifest, or a unit file under the admin unit directory) causes an abort with a runbook pointer — never auto-clean. Tested by `test_migration_guard` on all four targets (Debian 12, Ubuntu 22.04, Ubuntu 24.04, Fedora 40).
- **MG-3** (A) No-move continuity is verified in a container seeded with a make-install-shaped system: existing group makes sysusers a no-op, existing store directory makes tmpfiles a no-op, the hand-written configuration survives as a non-database file (package default lands beside it), and the store schema is caught up by the upgrade-path migration. Tested by `test_no_move_continuity_deb` and `test_no_move_continuity_rpm`.
- **MG-4** (M) The migration costs at most one short sample gap, honestly recorded in the endurance timeline: `make uninstall`'s sanctioned disable closes the open period `user_disabled`; after migration the user opts back in with `fenris monitor resume`.
## Collector acquisition path (ADR 0006)
- **AC-1** (P) Each collection run acquires counters and thermal evidence solely from `smartctl -a -j <device>` and controller identity (`subnqn`, `sn`, `mn`, `fr`, `transport`) solely from sysfs; no other acquisition path exists anywhere in the codebase.
- **AC-2** (A) Identity normalization is applied exactly once, at write time — trailing spaces and newlines stripped, no case folding, empty-after-strip stored blank — so padded and unpadded renderings of the same field yield byte-identical stored values.
- **AC-3** (A) Any acquisition failure — missing binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the whole collection run; no partial sample (identity without counters, or counters without identity) is ever written; the miss surfaces through ADR 0005 freshness, never as degraded identity.
- **AC-4** (P) `vid`/`ssvid` are read from the PCI sysfs node when present and stored null otherwise; they are segment metadata only, never key components.
- **AC-5** (P) `make install` verifies `smartctl` and fails cleanly otherwise; the acquisition path adds no Python dependency and no OS package beyond smartmontools (ADR 0004 §9).
## Dashboard clarity and release notes ([Chart Fenris dashboard clarity](https://git.bongbetic.com/xavierk/Fenris/issues/55))
Decided in [Write the dashboard clarity acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/issues/59), from [Prototype the dashboard clarity additions](https://git.bongbetic.com/xavierk/Fenris/issues/56), [Lock the dashboard wording strings](https://git.bongbetic.com/xavierk/Fenris/issues/57), and [Specify the changelog and release-notes mechanism](https://git.bongbetic.com/xavierk/Fenris/issues/58).
- **DC-1** (A) TUI branding: the header bar renders `Fenris — NVMe endurance monitor`; a dimmed `by Bongbetic` sits inline with service facts in the bottom service strip; neither string appears in `fenris status` (TUI-only identity surfaces).
- **DC-2** (A) Continuity parity, keyed to the boot fact as-is: active + boot-enabled renders `monitoring: active in background · persists across reboots`; boot-disabled renders `monitoring: does not start on next boot` — identical lowercase source strings in the TUI service strip and `fenris status`, including while paused (paused implies boot-disabled; the row still reports the fact). Test impact: feeds the CI-2 sweep (lowercase source-string comparison).
- **DC-3** (A) Paused presentation (Deliberate disable): the TUI shows a strong state block titled `monitoring: paused — deliberate disable` with subline `paused time is excluded from your usage habit · resume: fenris monitor resume`; `fenris status` prints the same two lines with identical wording. Test impact: feeds the CI-2 sweep (lowercase source-string comparison).
- **DC-4** (A) Quit affordance distinct from monitoring state: a bordered labelled rail `q QUIT TUI` visually separate from the paused state block; the footer reads `p pause · r resume · c collect · d disclosures` with no quit entry (the rail owns quit); quitting the TUI never alters monitoring state. Amends TUI-4's binding parenthetical.
- **DC-5** (A) Launch auth banner: `privileged actions will prompt for authentication (polkit)` renders full-width under the header at TUI launch, clears on the first refresh tick, and never reappears in the session; no user-facing string uses "sudo" (polkit-accurate elevation wording only).
- **DC-6** (A) CHANGELOG.md shape (Keep a Changelog 1.1): `## [Unreleased]` always present at top, even empty; version headings `## [X.Y.Z] - YYYY-MM-DD` with strict ISO date; categories Added/Changed/Fixed only, security folding into Fixed; entries are single `- ` bullets, imperative mood, user-facing, no commit hashes or issue numbers.
- **DC-7** (A) Extraction fails closed: `scripts/extract_changelog.py` slices the requested version's section verbatim and never reads `[Unreleased]`; a missing or empty section or a malformed date produces `::error::` and a nonzero exit; the release workflow fails when the pushed tag ≠ `v{version from pyproject.toml}` (guard skipped on `workflow_dispatch`).
- **DC-8** (A/P) Release body: the body is the extracted section verbatim plus the standing footer from `packaging/release-footer.md`; a re-run against an existing release PATCHes the body (re-sync is a feature) while uploaded assets skip idempotently. A covers assembly/PATCH-logic unit tests; P is one scripted `workflow_dispatch` verification of body assembly.
+128
View File
@@ -0,0 +1,128 @@
# Fenris dashboard clarity specification
**Presentation amendment (2026-09-19):** The accepted
[Glint-inspired dashboard design](glint-dashboard-design.md) supersedes the
preserved Panes arrangement and standalone heavy quit rail. Identity, continuity,
paused-state explanations, and the distinction between quitting and pausing
remain required; quit now has an explicit entry in the fixed controls.
**Status: decision-complete.** Assembled by [Assemble the dashboard clarity specification and close the map](https://git.bongbetic.com/xavierk/Fenris/issues/60) from the closed tickets of the Wayfinder map [Chart Fenris dashboard clarity](https://git.bongbetic.com/xavierk/Fenris/issues/55). This document is normative for the follow-up **execution effort**; nothing here is implemented by the map.
**Canonical roles.** The [redesign specification](fenris-redesign.md) (frozen) and [ADRs 0001–0007](../adr/) remain authoritative and untouched — this is a companion spec covering five dashboard clarity additions plus the changelog-driven release-notes mechanism. The [criteria register](acceptance-criteria.md) carries the testable statements: **DC-1–DC-8**, appended by this assembly, with **TUI-4's binding list amended** (§4). Terminology follows the glossary in [`CONTEXT.md`](../../CONTEXT.md), including *Deliberate disable* and *Release*.
**Binding language.** *Must*, *exactly*, and *never* are normative.
## How to read this document
Five screen additions (§1–§5), one release-notes mechanism (§6), the verbatim string register (§7), and the README section to add at execution (§8). Each section cites its criteria. Source strings are lowercase; the TUI may render uppercase via styling only. Typography, governing every string: em-dash `—` separates a title from its qualifier; middle dot `·` joins facts within a line; UTF-8 is assumed. CI parity sweeps compare lowercase source strings — rendering case is styling, not wording.
## 1. Header bar and Bongbetic credit — DC-1
Visual base is treatment A, quiet integration: the existing Panes information architecture is preserved.
- The header bar reads `Fenris — NVMe endurance monitor`.
- The credit `by Bongbetic` renders dimmed, inline with service facts in the bottom service strip — never in the action row.
- Both are TUI-only identity surfaces: `fenris status` never renders them.
## 2. Continuity line — DC-2
A labelled `CONTINUITY` row in the service strip (treatment B), mirrored by `fenris status` — the TUI/CLI parity anchor. The row is keyed to the boot fact as-is, independently of run state (Deliberate disable runs `systemctl disable --now`, so paused implies boot-disabled; the row still reports the fact):
- Active + boot enabled: `monitoring: active in background · persists across reboots`
- Boot disabled: `monitoring: does not start on next boot`
Identical lowercase source strings in the TUI service strip and `fenris status`, including while paused.
## 3. Paused state block — DC-3
Treatment C, strong state blocks: when monitoring is paused, a full-width, high-contrast banner clearly identifying Deliberate disable:
- Title: `monitoring: paused — deliberate disable`
- Subline: `paused time is excluded from your usage habit · resume: fenris monitor resume`
`fenris status` prints the same two lines with identical wording (state line + consequence line). The resume hint uses the CLI form only; the footer owns key hints — no duplication.
## 4. Quit rail — DC-4 (amends TUI-4)
Treatment B, labelled rails: a prominent bordered `q QUIT TUI` rail, visually separate from the monitoring-state block and the paused banner. The footer becomes `p pause · r resume · c collect · d disclosures` — the rail owns quit; the footer carries no quit entry. Quitting the TUI never alters monitoring state. The register's TUI-4 binding parenthetical is amended accordingly by this assembly.
## 5. Launch auth banner — DC-5
A quiet informational line (treatment A) that never competes with drive state:
- Text: `privileged actions will prompt for authentication (polkit)`
- Full-width under the header at TUI launch; clears on the first refresh tick; never reappears in the session.
- TUI-only; `fenris status` never shows it.
- Elevation wording is polkit-accurate everywhere: no user-facing string uses "sudo" (sudo belongs to install/upgrade docs).
- Evidence class A: a Textual pilot drives refresh ticks headlessly.
## 6. Changelog and release notes — DC-6, DC-7, DC-8
Implements the existing glossary term *Release* (tag + packages + change notes together). No new glossary terms; no ADR (reversible mechanism).
### 6.1 CHANGELOG.md (source of truth, repo root)
- Keep a Changelog 1.1 shape. `## [Unreleased]` is always present at top, even empty. Version headings are `## [X.Y.Z] - YYYY-MM-DD` — bracketed bare semver, strict ISO date.
- Categories are `### Added`, `### Changed`, `### Fixed` only; security fixes fold into Fixed.
- Entries are single `- ` bullets, imperative mood, user-facing phrasing; no commit hashes or issue numbers.
### 6.2 Extraction (release.yml, tag time)
- `scripts/extract_changelog.py` (checked in, unit-tested): takes the changelog path and a version; slices that version's section verbatim; never reads `[Unreleased]`. Fails closed — `::error::` plus nonzero exit — when the section is missing or empty or the date is malformed.
- Guard: the workflow fails when the pushed tag ≠ `v{version from pyproject.toml}` (guard skipped on `workflow_dispatch`).
### 6.3 Release body
- Body = extracted version section verbatim + standing footer from `packaging/release-footer.md` (channel install one-liners, `sha256sum -c SHA256SUMS.asc` verify, rollback pointer). The footer is standing text; only the changelog section varies.
- Re-run against an existing release: PATCH the body (changelog re-sync is a feature); uploaded assets/packages keep their current idempotent-skip.
### 6.4 Discipline
- All entries land in `[Unreleased]` as part of the fixing change — no notes-later step.
- One release commit bumps the pyproject version, renames `[Unreleased]` → the version heading, and restores an empty `[Unreleased]`; the tag points at that commit (tag ↔ pyproject ↔ changelog triple-match, enforced fail-closed by DC-7).
- No backfill: per-release notes begin with the release shipping this mechanism; `CHANGELOG.md` starts with empty `[Unreleased]`.
## 7. String register (verbatim)
### TUI-only strings (launch/identity surfaces)
| Surface | String |
|---|---|
| Header bar | `Fenris — NVMe endurance monitor` |
| Credit (dimmed, inline with service facts) | `by Bongbetic` |
| Auth banner (full-width under header at launch, clears on first refresh tick, never reappears) | `privileged actions will prompt for authentication (polkit)` |
| Quit rail (bordered, labelled) | `q QUIT TUI` |
| Footer (owns key hints; no quit entry) | `p pause · r resume · c collect · d disclosures` |
### Parity strings (TUI and `fenris status` identical — CI-2)
| Surface | String |
|---|---|
| Continuity, active + boot enabled | `monitoring: active in background · persists across reboots` |
| Continuity, boot disabled | `monitoring: does not start on next boot` |
| Paused state line | `monitoring: paused — deliberate disable` |
| Paused consequence line | `paused time is excluded from your usage habit · resume: fenris monitor resume` |
`fenris status` prints the paused state line + consequence line when paused, identical wording to the banner title + subline.
## 8. README section (add at execution)
The README gains a "Reading the dashboard" section after the CLI reference. Verbatim text:
```markdown
## Reading the dashboard
`fenris` opens the TUI dashboard. Three things it tells you:
- **Continuity** — the service strip's continuity line (and `fenris status`) reports whether monitoring survives reboots: `monitoring: active in background · persists across reboots`, or `monitoring: does not start on next boot`.
- **Paused vs. quit** — a full-width `monitoring: paused — deliberate disable` block means collection is stopped (`fenris monitor pause`); resume with `fenris monitor resume`. Pressing `q` only leaves the screen — monitoring keeps running in the background.
- **Auth banner** — at launch, `privileged actions will prompt for authentication (polkit)` shows once and clears on the first refresh. Privileged actions elevate via polkit; Fenris never asks for sudo.
Per-release notes live on the [releases page](https://git.bongbetic.com/xavierk/Fenris/releases): each entry is the version's `CHANGELOG.md` section — what was added, changed, and fixed — plus standing install and verification instructions.
```
This resolves the map's README-wording fog: the wording is decided here; the actual README edit is execution.
## 9. Out of scope
Executing any of this — code, tests, releases — and any TUI layout or information-architecture redesign beyond the five additions named above. Execution is a fresh effort after handoff.
+660
View File
@@ -0,0 +1,660 @@
# Fenris redesign specification
**Status: implementation-ready.** Assembled by [Write the Fenris redesign specification and close the map](https://git.bongbetic.com/xavierk/Fenris/issues/19), executing the assembly decision [Assemble the implementation-ready specification](https://git.bongbetic.com/xavierk/Fenris/issues/17) (all seven recommendations accepted) on the Wayfinder map [Chart Fenris's persistent TUI monitoring redesign](https://git.bongbetic.com/xavierk/Fenris/issues/1).
**Canonical roles.** [ADRs 0001–0006](../adr/) are the immutable rationale records — the *why*. [Acceptance criteria](acceptance-criteria.md) are the single register of testable statements — the *definition of done*. This document normatively restates every **operative contract** — the *what* — so an implementer never needs Wayfinder-ticket access: schema column sets, constants, rule tables, unit and CLI definitions, and the Panes TUI layout. Nothing here overrides an ADR or restates a criterion as a criterion.
## How to read this document
- **Binding language.** *Must*, *exactly*, and *never* are normative. Terminology follows the glossary in [`CONTEXT.md`](../../CONTEXT.md): *observation history*, *usage-adjusted theoretical lifespan*, *projection confidence*, *monitoring period*, *observation store*, *hour observation*, *day aggregate*, *controller segment*, *degraded identity*, *endurance baseline*, *verified override*, *unverified override*, *sustained regime*, *habit change*, *scenario range*, *coverage*, *collection run*, *deliberate disable*, *store fault*.
- **Ordering.** Sections follow data flow: system context → collector acquisition → observation store → controller identity & segmentation → hour/day derivation → projection & confidence → Panes TUI → service lifecycle & sanctioned toggle → failure & recovery → installation. Each section opens with its ADR links and criterion-ID block.
- **Implementation boundary.** This specification plans the redesign; it does not implement it. The complete handoff is this document + the [criteria register](acceptance-criteria.md) + [ADRs 0001–0006](../adr/) + the glossary. Given/When/Then test specs are derived by the implementer at implementation time.
### Normative constants index
Every constant is defined once, in the section named below; other sections cite, never redefine. All are named constants in code, not configuration.
| Constant | Value | Defined in |
|---|---|---|
| Collection cadence (default) | 3 min (`OnUnitInactiveSec`) | §8.2 |
| First-boot delay | 2 min (`OnBootSec`) | §8.2 |
| Timer accuracy window | 30 s (`AccuracySec`) | §8.2 |
| Collection-run timeout | 90 s (`TimeoutStartSec`) | §8.2 |
| Fresh threshold | newest sample within 2 × cadence + `AccuracySec` + 60 s | §8.9 |
| Missed → stale boundary | 48 h | §8.9, §6.7 |
| Powered-off hour threshold | power-on-hours delta < 90 % of the hour's wall-clock span | §5.1 |
| Active hour threshold | DUW delta ≥ 256 MiB in the hour | §5.1 |
| Raw-sample retention | 14 days | §3.4 |
| Warming gate | 14 distinct UTC day aggregates, ≤ 2 below 50 % coverage | §6.6 |
| Supported coverage floor | 80 % | §6.7 |
| Horizon agreement | 7/28/90-day rates within a factor of 2 | §6.7 |
| Burst guard | no single day ≥ 50 % of trailing 28-day bytes | §6.7 |
| Young-regime cap | regime < 7 days old → Limited | §6.4 |
| Habit-change trigger | trailing 7-day mean ≥ 2× or ≤ 0.5× the preceding 28-day mean, 3 consecutive days | §6.4 |
| Regime span cap (default) | full observation history capped at 90 days | §6.4 |
| Scenario horizons | 7 / 28 / 90 days | §6.5 |
| Implied-baseline eligibility | ≥ 2 Percentage-Used increments within the current controller segment | §6.3 |
| Rated-TBW conversion | `E_rated = entered_TBW × 10¹²` bytes | §6.3 |
| Implied-baseline validity window | 1 ≤ p ≤ 254 | §6.3 |
| Wear-disagreement note | vendor wear vs. observed write rate by more than a factor of 2 | §6.1 |
| Capacity validation tolerance | ± 1 % | §6.2 |
---
## 1. System context
**ADRs:** [0001](../adr/0001-observation-store-sqlite.md), [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md), [0006](../adr/0006-collector-acquisition-path.md). **Criteria:** CI-3, LC-1, LC-5, ST-1.
### 1.1 Scope
Fenris observes one configured NVMe drive's real-world use and translates the observation history into a usage-adjusted theoretical lifespan. The redesign replaces the HTML dashboard with a keyboard-first TUI backed by a short-lived privileged collector on a systemd timer, persistent compact observation storage, and categorical projection confidence reflecting the length, completeness, and stability of real usage history.
Standing constraints, binding on every section:
- Linux with systemd and polkit only; no other init system is supported.
- Exactly one configured NVMe drive — the device named by `/etc/fenris/fenris.conf` (§8.3).
- Fully local: no telemetry, no network fetching, no automatic vendor-data retrieval.
- The HTML dashboard and HTTP server are gone; nothing of the daemonization, PID files, or `/run` state survives.
- CLI `status` and `sample` are retained (§8.8).
### 1.2 Components and privilege boundaries
| Component | Privilege | Path | Role |
|---|---|---|---|
| `fenris-collect.service` | root oneshot unit | `/usr/libexec/fenris/fenris-collect` | The only code path that interrogates the device and writes the observation store. |
| `fenris-collect.timer` | system timer | — | Schedules collection runs; `WantedBy=timers.target`. |
| `fenris-monitor` | root helper | `/usr/libexec/fenris/fenris-monitor` | Fixed privileged operations: `enable`/`disable` (optional `--now`), the collect trigger, monitoring-period bookkeeping, and baseline persistence. The only binary polkit authorizes. |
| `fenris` | unprivileged | `/usr/local/bin/fenris` | Human entry point: no arguments opens the TUI; subcommands are the CLI (§8.8). Never a unit. |
| Observation store | root-written, group-read | `/var/lib/fenris/observations.db` | Single SQLite database in WAL mode (§3). The TUI and `status` open it read-only. |
| Configuration | world-readable | `/etc/fenris/fenris.conf` | Exactly one key: the device selector (§8.3). |
The TUI and CLI are ordinary unprivileged processes. Elevation is exclusively polkit, exclusively for `fenris-monitor` (§8.5). There is no `/run/fenris` coordination surface and no export layer: systemd serializes collection runs, the observation store holds state, failures go to the journal.
### 1.3 Data flow
1. The timer fires; `fenris-collect.service` runs `fenris-collect`.
2. The collector acquires counters and thermal evidence from `smartctl -a -j` and controller identity from sysfs (§2), normalizes identity exactly once (§2.3), and either fails the whole run or writes one complete sample.
3. The collector derives and validates hour observations and day aggregates, advances controller segmentation and period bookkeeping, prunes raw samples, and commits (§3–§5, §9.1).
4. Readers — the TUI and `fenris status` — open the store read-only and **recompute the projection on every read** (§6); nothing derived is ever stored (§3.7).
Control flow is separate: the human drives the TUI/CLI; privileged operations route through `fenris-monitor` under polkit to `systemctl`; period rows record *intent* (only the sanctioned path), while the collector records *observed fact* (§8.5–§8.6, §9.8).
### 1.4 Cross-cutting prohibitions
These are operative contracts; each is restated in its home section and gated by the criteria block [CI-3](acceptance-criteria.md):
1. No code path outside `fenris-collect` interrogates the device (§2.1, §8.7).
2. Polkit authorizes exactly one binary, `fenris-monitor`, under `com.bongbetic.fenris.monitor` `auth_admin` (§8.5).
3. No `/run/fenris` coordination surface or export layer exists (§1.2).
4. No absent hour is ever interpolated, estimated, or fabricated (§5.3, §9.3).
5. No alerting, notification, or escalation machinery exists anywhere (§9.6–§9.7).
6. `/etc/fenris/fenris.conf` holds exactly one key — the device selector (§8.3).
7. No synthetic or capacity-derived baseline is ever created, including for legacy history (§6.1, §3.5).
8. Readers never partially interpret a newer-schema store (§3.6, §9.5).
9. Projections are never stored; always recomputed on read (§3.7, §6.10).
---
## 2. Collector acquisition
**ADR:** [0006](../adr/0006-collector-acquisition-path.md). **Criteria:** AC-1–AC-5; miss absorption per [0005](../adr/0005-failure-detection-and-recovery.md) §5.
### 2.1 Channels — the hard pin
Every collection run acquires exactly two ways:
- **Counters and thermal evidence** — solely from `smartctl -a -j <device>`: `data_units_written`, `data_units_read`, `percentage_used`, `available_spare`, `media_errors`, `power_on_hours`, `power_cycles`, `unsafe_shutdowns`, temperature, `critical_warning` — consumed as-is (smartmontools already trims the strings it copies).
- **Controller identity** — solely from sysfs (`/sys/class/nvme/<ctrl>/`): `subnqn`, `sn`, `mn`, `fr`, `transport`.
No other acquisition path exists anywhere in the codebase. There is no fallback: libnvme bindings and the `nvme` CLI JSON interface are excluded (ADR 0006, *Considered options*).
### 2.2 All-or-nothing runs
Any acquisition failure — missing `smartctl` binary, nonzero exit, malformed JSON, unreadable sysfs attribute — fails the **whole** collection run. A partial sample (identity without counters, or counters without identity) is never written: a transient read failure must never push a healthy drive down the degraded-identity path (§4). The miss surfaces through freshness grading (§8.9) and the flat retry cadence (§9.6), never as degraded identity.
### 2.3 Identity normalization — once, at write time
One collector-side function normalizes every identity field, applied exactly once at write time:
- strip trailing spaces and newlines;
- no case folding;
- empty-after-strip is stored blank.
Padded and unpadded renderings of the same field therefore yield byte-identical stored values — a collector implementation change can never split a drive's own history. A future acquisition-path change must deliver byte-identical normalized identity values, or the change itself forces a controller-segment boundary.
### 2.4 Segment metadata sourcing
`transport` comes from the NVMe class sysfs directory. `vid`/`ssvid` come from the PCI node (`/sys/class/nvme/<ctrl>/device/{vendor,subsystem_vendor}`) when present and are stored null otherwise. Both are segment **metadata only** (§4.2), never key components.
### 2.5 Prerequisites
`make install` verifies `smartctl` is present and fails cleanly otherwise (§10.1). The acquisition path adds no Python dependency and no OS package beyond smartmontools; the dependency lockfile (§10.5) is untouched by this section.
---
## 3. Observation store
**ADR:** [0001](../adr/0001-observation-store-sqlite.md) as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) and [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14). **Criteria:** ST-1–ST-12; FL-5.
### 3.1 Substrate and access
- One SQLite database in **WAL mode** at `/var/lib/fenris/observations.db`. An unprivileged reader querying during a collector write sees a consistent snapshot.
- The database is root-owned and group-readable through the `fenris` read group created by packaging; the TUI and `status` open it **read-only**. No `/run` snapshot, no export layer.
- `/var/lib/fenris` is created by the installer with root-written group-read permissions; the database file itself is created lazily by the first write, so "no observations yet" remains a real state the TUI can greet (§7.6, §10.1).
- Migration, schema changes, prune, and import are each single transactions — a killed timer run can never leave partial state.
### 3.2 Entities and column sets
The schema carries exactly six entities:
**`samples`** — recent raw samples (14-day retention, §3.4): timestamp (UTC); the normalized controller-identity fields captured at acquisition (§2.3); raw integer `data_units_written`, `data_units_read`; `percentage_used`; `available_spare`; `media_errors`; `power_on_hours`; `power_cycles`; `unsafe_shutdowns`; temperature; `critical_warning`.
**`hour_observations`** — one row per UTC hour: the usage-habit split `seconds_active`, `seconds_idle`, `seconds_powered_off`, `seconds_unknown` (summing to 3600, §5.1); DUW/DUR deltas; temperature min/avg/max; sample count; coverage flag. Classification thresholds belong to the projection model (§5.1), not the store.
**`day_aggregates`** — one row per UTC day, the habit-evidence grain: each day row carries, at minimum, the day's activity-split sums, write deltas, and coverage share — the inputs the evidence gates of §6.6 consume — derived monotonically from its hour rows.
**`monitoring_periods`** — `started_at`; `ended_at` (NULL = open); `end_cause` enum (`user_disabled`, `migrated`, …). Powered-off time stays inside a period; deliberately disabled time does not (§5.2, §8.6).
**`controller_segments`** — spans of unchanged controller identity and monotonic counters; write deltas are never computed across a segment boundary. Columns: the identity key (§4.1) and the frozen metadata snapshot of §4.2, plus the segment's span bounds.
**`endurance_baseline`** — one active row, replaced on edit (§6.2): the rated-TBW value in bytes (`E_rated = entered_TBW × 10¹²`); mandatory provenance — source URL, document revision, entry date, model string, nominal capacity; frozen validation facts — detected model, detected capacity bytes, `validated_by` (`machine`/`user`), `validated_at`.
### 3.3 Time model
Hours and days are UTC-bounded. Day derivation from hour rows is monotonic; DST-ambiguous 23- or 25-hour days never exist in the store.
### 3.4 Retention
Raw samples are pruned opportunistically by the collector to **14 days**. Hour observations and day aggregates are retained indefinitely.
### 3.5 Legacy migration
The migration procedure, invoked from the entry points below, is **idempotent and interruption-safe**:
1. If the store already carries the legacy-import marker, do nothing.
2. `history.jsonl` is the sole authority: import raw samples and derive hour observations and day aggregates from them.
3. `hourly.jsonl` is never trusted as input: mismatches against derived data are diffed and logged.
4. Open one implicit `monitoring_periods` row at the first legacy sample, closed `end_cause = migrated` at the migration moment. Pre-migration hours carry an unknown activity split except directly evidenced facts — a sample present means powered on; a DUW delta means writes occurred.
5. The import is a single transaction: a scripted kill mid-import leaves the store fully pre- or fully post-migration.
6. Only after commit are legacy files renamed `*.migrated` — never deleted.
7. Malformed legacy lines are quarantined with a logged count, never silently dropped.
No synthetic or capacity-derived baseline is ever created for legacy history (§1.4–7). Entry points: the installer's import detection at `./data/history.jsonl` (or an explicit path) (§10.1); `fenris import <path>` for later finds (§8.8); and the collector's first new-version run, which performs this same procedure (ADR 0001 §6).
### 3.6 Schema versioning
`PRAGMA user_version` plus ordered migration steps in code, each in its own transaction. The collector refuses to run against an unknown **newer** version; readers refuse symmetrically with the exact wording of §9.5 and never partially interpret. *Reconciliation note:* ADR 0004 §6 describes upgrade-time migrations as "governed by a `schema_version` table" — the operative mechanism is this section's `user_version` (ADR 0001 §8, criterion ST-12); there is one version authority, not two.
### 3.7 Nothing derived is stored
Projections are not stored; there is no separate latest-status table and no stored health flag. The freshest sample timestamp is the store's own staleness signal (§8.9). The baseline lives in the database (§6.2); `/etc/fenris/` holds only operational configuration (§8.3).
---
## 4. Controller identity and segmentation
**Decisions:** [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11), [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14), [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15). **ADRs:** [0001](../adr/0001-observation-store-sqlite.md) §3 (as amended), [0002](../adr/0002-projection-model-sustained-regime.md) §§8–9 (as amended). **Criteria:** ID-1–ID-4, PR-9, PR-15, PR-16.
### 4.1 Identity key ladder
The controller-segment identity key is the **normalized, kernel-exposed subsystem NQN** (`subnqn`), with fallbacks, in order:
1. kernel-exposed subsystem NQN;
2. the kernel composite;
3. `model|serial`.
`fr` (firmware revision) is metadata only — it may go stale after a mid-segment firmware update. Identity change and DUW decrease are **independent axes** (§4.3).
### 4.2 Frozen metadata snapshot
Each segment freezes, at open, a fully nullable metadata snapshot — immutable thereafter: normalized `subnqn`, `sn`, `mn`, `fr`, plus `vid`, `ssvid`, `transport`, and the `identity_degraded` flag. All columns are nullable so incompleteness stays explicit: legacy-imported segments carry `mn` with NULLs (§4.4); degraded segments carry whatever was observed. These are human diagnostics, never key components. `cntlid` is excluded — it distinguishes controllers within one subsystem, out of scope for a single-drive monitor.
### 4.3 Segmentation axes
- **DUW decrease, unchanged identity** — a segment boundary within the same drive. Prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.8).
- **Identity-key change** — quarantines prior history from projection entirely: it describes a different drive (§6.8).
- **Degraded identity** — a segment whose identity key is **blank** (every rung of the ladder empty). `identity_degraded` is set at segment open exactly when the key is blank; keys from the kernel-composite or `model|serial` rungs are not degraded. Blank-key semantics extend identity-change rules verbatim: any visible change of the recorded key — including to or from blank — is a controller-identity change and quarantines; equal blank keys continue the segment, segmented by DUW monotonicity alone. Even a degraded→healthy transition quarantines, so the projection window only ever spans segments sharing one key (§6.8).
- Ephemeral markers (model "Linux", non-pcie transport) are segment metadata, never confidence facts.
The confidence consequence of degraded identity — capped at Limited with its fixed contributing fact — is §6.7's rule.
### 4.4 Legacy identity
Legacy history imports under a labeled, model-scoped **legacy identity** (mn-only segments), so it never blends with the post-redesign identity of the same physical drive.
---
## 5. Hour and day derivation
**ADRs:** [0002](../adr/0002-projection-model-sustained-regime.md) §§4–6; [0001](../adr/0001-observation-store-sqlite.md) §3; [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) §2 (power-on-hours evidence). **Criteria:** PR-4–PR-6, ST-4, FL-3.
### 5.1 Hour classification
Each UTC hour is classified by named constants, in this order of evidence:
- **Powered-off** — the hour's power-on-hours delta is below **90 %** of its wall-clock span.
- **Active** — DUW delta ≥ **256 MiB** in the hour.
- **Idle** — powered on, sampled, below the active threshold.
- **Unknown** — everything else: unsampled without power-on-hours evidence (machine-off and collector failure are indistinguishable by design), or inconsistent counters.
There is no configuration surface for these thresholds; they are documented constants in one projection module.
### 5.2 Denominator and disabled time
The projection denominator is **wall-clock seconds inside monitoring periods**, including powered-off and unknown time. Disabled periods — wall-clock outside monitoring periods — are excluded from numerator and denominator. **Disabled time is not an hour state.**
### 5.3 Gaps and coverage — never backfill
No absent hour is ever interpolated, estimated, or fabricated. Unexplained gaps inside a period keep the aggregate counter delta, remain in the denominator as unknown seconds, and reduce coverage. Power-on-hours classification (§5.1) is the only inference admitted. **Coverage** is the share of wall-clock seconds inside monitoring periods whose classification is known rather than unknown — a first-class displayed fact (§6.10, §7.3).
### 5.4 Day aggregates
One row per UTC day, derived monotonically from hour rows (§3.2–§3.3) — the grain at which usage-habit evidence is judged (§6.6).
---
## 6. Projection and confidence
**ADRs:** [0002](../adr/0002-projection-model-sustained-regime.md) as amended by [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15); baseline per [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12). **Criteria:** PR-1–PR-17, CI-4.
### 6.1 One projection; baseline precedence
Exactly **one** usage-adjusted theoretical lifespan is computed, against the endurance baseline chosen by precedence:
1. **Verified override** — a rated-TBW override with complete provenance whose applicability to the detected drive was confirmed by machine match or explicit user attestation;
2. **Unverified override** — a rated-TBW override knowingly stored with incomplete provenance; always presented as user-supplied, never as verified;
3. **Implied baseline** — derived from vendor wear (§6.3), eligible only per §6.3's gate;
4. otherwise the projection is **Unavailable**.
Percentage Used is context, never a second projection: it renders as a vendor-wear context line, and when the wear it implies disagrees with the observed write rate by more than a factor of 2, a note says so. The legacy PU-slope regression and `capacity × 600` synthesis are gone; no synthetic or capacity-derived baseline is ever created (§1.4–7).
### 6.2 Endurance baseline: provenance and validation
The baseline lives in the observation store's `endurance_baseline` table (§3.2) and is edited via the CLI (§8.8) — `/etc/fenris/` holds no baseline.
- **Mandatory provenance:** source URL, document revision, entry date, model string, nominal capacity.
- **One active row**, replaced on edit.
- **Verification is derived at read** — complete provenance and a drive match (machine or attested) — never a stored boolean.
- **Unverified tier:** incomplete provenance stores only behind an explicit unverified acknowledgment, as NULL fields in that precedence tier.
- **Entry-time validation** (unprivileged, live sysfs read of the configured device): normalized model containment, with an interactive confirm recorded as `validated_by = user`; nominal capacity within ± 1 %.
- **Read-time applicability:** a model match against the current controller segment (§4). A mismatch is **retained — never auto-deleted** — and leaves the projection Unavailable.
- Persistence goes through the polkit-guarded `fenris-monitor` verb after CLI-side validation (§8.5).
### 6.3 Arithmetic
```text
rate = regime DUW delta bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
```
- `E_rated` is exact: rated TBW converts to bytes by × 10¹².
- `E_implied` is computed **only** for `1 ≤ p ≤ 254`; Percentage Used of 0 or saturated 255 implies no baseline — that precedence tier is unavailable. The implied baseline is labeled *implied from vendor wear estimate* and shown with few significant digits.
- **Implied-baseline eligibility:** the implied tier is used only after ≥ 2 Percentage-Used increments within the current controller segment; until then the projection is Unavailable with the fixed phrase *"vendor wear estimate too coarse to imply endurance"*.
### 6.4 Sustained regime and habit change
The headline rate is the **sustained-regime** rate: regime DUW bytes ÷ in-period wall-clock seconds. The default regime is the full observation history capped at **90 days**.
A **habit change** is declared when the trailing 7-day mean of daily written bytes stays ≥ 2× (or ≤ 0.5×) the mean of the preceding 28 days for **3 consecutive days**. The new regime starts at the **first day of divergence**, is adopted automatically, and is labeled *"usage habit changed N days ago"*; the scenario range keeps the longer horizons visible. A regime younger than **7 days** caps projection confidence at Limited evidence.
### 6.5 Scenario range
The 7-, 28-, and 90-day rates are computed **independently of the regime** and shown as the scenario range. Only horizons the history actually covers appear — no placeholders. The scenario range is the only spread shown anywhere (§6.9).
### 6.6 Minimum evidence
Warming up until **14 distinct UTC day aggregates** of which at most **2** fall below 50 % coverage. The projection still renders while warming up, labeled with its facts (e.g. *"warming up: N of 14 qualifying days"*). Every Unavailable condition renders **no lifespan number**.
### 6.7 Confidence rule table
Confidence renders as **state plus contributing facts, never a percentage**. Three states:
- **Unavailable** — no applicable baseline; DUW unsupported; zero rate over the regime; controller-identity change.
- **Supported** — verified baseline **and** ≥ 14 qualifying days **and** coverage ≥ 80 % **and** fresh (< 48 h) **and** 7/28/90 rates within a factor of 2 across existing horizons **and** no single day ≥ 50 % of trailing 28-day bytes **and** regime ≥ 7 days old **and** the current controller segment's identity key is not degraded.
- **Limited** — every other case with a baseline and a positive rate; the failing facts are shown.
**Staleness:** a newest day aggregate older than **48 hours** drops confidence one level (Supported → Limited) and is shown as a contributing fact.
**Degraded identity:** a controller segment whose identity key is blank (§4.3) caps confidence at **Limited evidence**, with the contributing fact *"controller identity unavailable — replacement detection relies on write-counter continuity only"* rendered in every state. Supported is unreachable while the current segment is degraded. The cap combines idempotently with the staleness drop (both land at Limited).
### 6.8 Segment-break effects
- **DUW decrease, unchanged identity:** prior day aggregates remain habit evidence; the projection is Unavailable only until the new segment re-warms (§6.6).
- **Controller-identity change** — including any to-or-from-blank key change (§4.3): prior history is quarantined from projection entirely.
- Since even degraded→healthy transitions quarantine, the projection window only ever spans segments sharing one key; no cross-segment propagation rule is needed.
### 6.9 Zero rate and uncertainty
Zero rate renders *"no finite projection from this history"* — never infinity, never zero. No statistical confidence interval appears anywhere; the scenario range is the only spread.
### 6.10 The projection contract
The projection function hands the TUI and `status` exactly: the confidence state; the contributing facts — including the degraded-identity fact when the current segment's key is blank; the headline remaining time when one exists; the scenario range; the Percentage-Used context line; the disclosure text (§6.11). Recomputed on read, never stored.
### 6.11 User-facing language
Adopted from the endurance research as fixed by ADR 0002 §12; rendered identically by TUI and `status`.
**Headline wording** (equivalent phrasing required):
> Estimated time until the selected host-write endurance baseline is consumed, if future write usage resembles the observed usage habit. This is not a predicted hardware-failure date.
**Fixed phrases** (exact): *no finite projection from this history* (zero rate); *vendor wear estimate too coarse to imply endurance* (§6.3); *usage habit changed N days ago* (§6.4); *controller identity unavailable — replacement detection relies on write-counter continuity only* (§6.7); *observation store unreadable* (§9.4); *observation store written by a newer Fenris — upgrade Fenris* (§9.5); *no observations yet* with an enable hint (§8.9); *configuration error: ⟨reason⟩* (§8.3).
**Confidence rendering:** state plus contributing facts, in the research's evidence style, e.g.
> Supported evidence · verified manufacturer TBW · 42 calendar days · 96 % interval coverage · 6 weekly cycles · recent and 28-day rates agree
Never "82 % confidence" or "95 % accurate".
**The six disclosures** (verbatim, always available — TUI disclosures view and `status`):
1. This is an endurance projection, not a predicted hardware-failure date.
2. Percentage Used is vendor-specific; 100 means estimated endurance consumed but may not mean failure, it can exceed 100, and 255 is saturated.
3. Rated TBW can be a warranty/endurance threshold with separate time and eligibility terms, not a failure threshold.
4. DUW is upward-rounded host writes excluding metadata and selected commands, not exact physical NAND writes.
5. Projection quality depends on baseline provenance, history duration and completeness, recentness, stability, and representative usage cycles; future workload and firmware behavior remain outside the observed evidence.
6. Gaps can preserve an aggregate counter delta without preserving hourly timing; unexplained and deliberately disabled periods must be distinguished.
---
## 7. Panes TUI
**Presentation amendment (2026-09-19):** The accepted
[Glint-inspired dashboard design](glint-dashboard-design.md) supersedes this
section's panel arrangement and graph appearance with an activity-first layout,
dotted volume plots, Live / Day / History tabs, and focused-panel zoom. Shared
evidence, projection, authentication, and monitoring-action contracts still apply.
**Decisions:** [Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3) (Variant A adopted), [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6) (Textual). **ADRs:** [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) §§8, 10; [0004](../adr/0004-install-upgrade-removal-lifecycle.md) §10. **Criteria:** TUI-1–TUI-4, CI-1, CI-2, CI-4. The [prototype](https://git.bongbetic.com/xavierk/Fenris/src/branch/prototype/tui-information-architecture/prototype/tui-ia) is visual reference only; this section is normative.
### 7.1 Framework and floor
The TUI is built on **Textual**. It runs on Python 3.9+, gated at install time (§10.1) — never a runtime crash. The tty-passthrough mechanism below was validated live under Textual on a real terminal (prototype decision).
### 7.2 Layout — one dense keyboard-first screen
Variant A **Panes**: everything on one screen, no page navigation. The screen is a grid of four regions:
1. **Headline band** — full width, top: the lifespan headline (or its no-projection wording) with its regime line (*"if current habits continue · sustained regime: N days at R GB/day"*); the confidence state with contributing facts; the scenario range.
2. **Usage-history pane** — left, wider column: the write-history sparkline with ▲ habit-change and ? unexplained-gap markers plus their legend; the habit-split bar with active/idle/powered-off/unknown shares.
3. **Drive-health and settings pane** — right, narrower column: health facts (model, temperature, spare, media errors, unsafe shutdowns, power-on hours, power cycles, capacity); the vendor-wear context line (Percentage Used · total written of rated — *context, not a second projection*); a read-only settings view (device selector, cadence with drop-in pointer, raw retention, endurance baseline value with its provenance label). Edits happen via CLI / drop-ins, not in the TUI.
4. **Service strip** — full width, bottom: the four separate service facts (§7.3), the monitoring-period line, the action legend.
Exact proportions, glyphs, and borders follow the prototype's validated arrangement as visual reference; the region arrangement, contents, and bindings above are normative.
### 7.3 Content contracts per region
- **Confidence is evidence:** state plus contributing facts, never a percentage (§6.7); the headline band renders the §6.10 contract in full, including the disclosure affordance (`d`).
- **Four separate service facts, always:** boot enablement (enabled/disabled) · runtime activity (timer active/inactive) · last collect outcome (ok/FAILED, age, reason) · freshness (fresh/missed/stale with newest-sample age, §8.9). They are never merged into one "service status".
- **Monitoring-period line:** open-since / closed with end cause; deliberate-disable count where nonzero.
- The scenario range shows only covered horizons (§6.5); the vendor-wear context line carries the >2× disagreement note when it applies (§6.1).
### 7.4 Keybindings and asymmetry
Production bindings:
| Key | Action |
|---|---|
| `p` | Pause — **asks for confirmation** (y pause · n cancel), stating that paused time is excluded from the usage habit while powered-off time would still count. |
| `r` | Resume — **no confirmation** (benign; friction invites raw-systemctl escapes). |
| `c` | Collect now — synchronous outcome (§8.7), no confirmation. |
| `d` | Disclosures — the six disclosures of §6.11. |
| `q` | Quit. |
No bare start/stop exists anywhere; no page navigation keys exist (variant switching was prototype-only). Framework defaults apply for focus and scrolling otherwise.
### 7.5 Privileged actions and tty passthrough
Pause, resume, collect-now, and baseline operations run through `fenris-monitor` as a **terminal-attached subprocess**: the TUI suspends, the platform polkit agent prompts on the real terminal, and control returns cleanly with the outcome reflected in the service facts. Where no polkit agent exists the operation fails cleanly with the printed root equivalent (§8.5).
### 7.6 State rendering obligations
- From a synthetic observation store, the TUI renders **every** realizable combination of confidence state × freshness grade × baseline tier exactly as the §6.7 rule table and §8.9 constants dictate — headline number only when the rules allow it, contributing facts always, never a percentage (criterion CI-1).
- Empty store: *"no observations yet"* with an enable hint; the first-run prompt is an opt-in that enables the timer and opens the first period in one step (§10.1, dormant install).
- A `configuration error: ⟨reason⟩` fact renders when the device selector is invalid (§8.3); a store fault suppresses everything store-dependent (§9.4); a newer schema renders its fixed phrase (§9.5).
- Every TUI action has a CLI twin with identical outcomes and wording (§8.8, CI-2).
---
## 8. Service lifecycle and sanctioned toggle
**ADR:** [0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) as amended by [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12). **Criteria:** LC-1–LC-10, CI-2, CI-3.
### 8.1 Units
Exactly two system units exist:
- `fenris-collect.timer` — `WantedBy=timers.target`.
- `fenris-collect.service` — `Type=oneshot`, root, `ExecStart=/usr/libexec/fenris/fenris-collect`; no listener, no UI code.
The TUI and CLI are ordinary unprivileged processes and never units.
### 8.2 Cadence
Shipped defaults: `OnBootSec=2min`, `OnUnitInactiveSec=5min` (measured from run completion; drift accepted because hours are the evidence grain), `AccuracySec=30s`, `Persistent=no` (no suspend catch-up — absent hours classify through power-on-hours evidence, §5.1), `TimeoutStartSec=90s` so a hung device interrogation fails visibly as a bounded failed run retried next interval. Cadence changes are documented drop-ins on the timer unit (`systemctl edit` + daemon-reload); **no interval key exists in configuration**.
### 8.3 Configuration
`/etc/fenris/fenris.conf` holds exactly one key: the **device selector**, a stable `/dev/disk/by-id/…` path (raw nodes accepted with an instability warning), validated at collection time. The oneshot re-reads it every run — there is no reload path. An invalid selector is a bounded failed run (journal + failed unit result, retried next interval); `status` and the TUI also read the world-readable file directly and surface `configuration error: ⟨reason⟩`.
### 8.4 Entry points
Two privileged binaries — `/usr/libexec/fenris/fenris-collect` (device interrogation and store writes; the unit's `ExecStart`) and `/usr/libexec/fenris/fenris-monitor` (fixed operations `enable`/`disable` with optional `--now`, the collect trigger, monitoring-period bookkeeping, and `baseline set`/`baseline clear` persistence for the CLI-validated baseline; the only binary the polkit policy authorizes). One unprivileged `fenris` wrapper (§1.2). Root invokes the helpers directly; unprivileged users go through polkit.
### 8.5 Sanctioned toggle and polkit
- Pause = `fenris-monitor disable --now`; Resume = `enable --now`. Both perform the systemctl operation **and** the monitoring-period bookkeeping in one step. The human-facing twins `fenris monitor pause` / `fenris monitor resume` map to these and always act immediately; pause asks for confirmation in both TUI and CLI, resume does not (§7.4).
- Polkit action `com.bongbetic.fenris.monitor` (`auth_admin`) covers the toggle **and** the collect trigger **and** baseline persistence — authorizing exactly the one binary `fenris-monitor`.
- Where no polkit agent exists the operation fails cleanly and prints the root equivalent.
- This is the **only** sanctioned control path: a raw `systemctl stop`/`disable` never records `user_disabled` — only the sanctioned path records intent (§8.6).
### 8.6 Period-row idempotent matrix
| Situation | Effect on `monitoring_periods` |
|---|---|
| First-ever enable | Opens a period at the enable moment (hours before the first successful sample are unknown-but-inside — correct when the device errors). |
| Resume with an open period (a raw `systemctl stop` intervened) | No row changes; the gap remains inside as unknown seconds. |
| Resume with no open period | Opens a new row at the resume moment. |
| Pause with an open period | Closes it `user_disabled` at the pause moment. |
| Pause otherwise | No-op. |
| Raw `systemctl stop`/`disable` outside the helper | An unexplained gap, never `user_disabled`. |
### 8.7 On-demand collection
`fenris sample` and the TUI's collect-now route through `fenris-monitor` → `systemctl start fenris-collect.service`, which blocks until the oneshot exits; the outcome (freshness line or journal hint) is reported synchronously. No confirmation is required. No code path outside `fenris-collect` touches the device; the TUI never samples in-process.
### 8.8 CLI surface
| Command | Behavior |
|---|---|
| `fenris` (no arguments) | Opens the TUI (§7). |
| `fenris status` | Read-only composition of the observation store and allow-listed `systemctl show` properties: projection facts, enabled/active, last collect outcome, and a `journalctl -u fenris-collect.service` hint on failure or staleness. Never auto-samples, never prompts. |
| `fenris sample` | On-demand collection via the helper path (§8.7). |
| `fenris monitor pause` / `resume` | The sanctioned toggle (§8.5), pause asking confirmation. |
| `fenris baseline set` / `clear` | CLI-side validation (§6.2), then polkit-guarded persistence. |
| `fenris import ⟨path⟩` | The idempotent single-transaction legacy import (§3.5). |
| `--device` | Rejected with a pointer to the configuration file. |
| `start`, `stop`, `run` | Rejected with one-line migration pointers — never aliased (an alias would silently change meaning). |
`fenris.sh` is retired: not shipped, removed from the repository; the README maps its five menu options to their successors. Headless administration has full parity: every TUI action has a CLI twin (pause, resume, collect-now, baseline set/clear, the status fact set) with identical outcomes and wording.
### 8.9 Freshness grading
Constants defined once, consumed by TUI and CLI alike; the grade derives from the **newest sample timestamp**, never a stored flag:
- **fresh** — newest sample within 2 × cadence + `AccuracySec` + 60 s (7.5 min at default cadence);
- **missed** — between that and 48 h (a contributing fact);
- **stale** — ≥ 48 h, matching the §6.7 evidence gate;
- **empty store** — *"no observations yet"* with an enable hint.
---
## 9. Failure and recovery
**ADR:** [0005](../adr/0005-failure-detection-and-recovery.md). **Criteria:** FL-1–FL-8.
The posture: **visible degradation, never fabrication.**
### 9.1 Write-boundary validation
The collector validates every row it would write against the store's domain invariants: hour seconds sum to 3600; non-negative DUW delta within a controller segment; coverage consistent with sample count. A violating run **writes nothing**, logs the refused row to the journal for post-mortem, and fails visibly — retried next interval. Store invariant: everything persisted is well-formed.
### 9.2 Reader defense
Readers (TUI, `status`) defensively exclude and count malformed rows as a contributing fact. Under a single trusted writer they should never see one.
### 9.3 No backfill, ever
Gaps remain unknown seconds; degradation flows exclusively through coverage, freshness facts, and confidence categories; recovery is the timer's next successful run. Power-on-hours classification (§5.1) is the only inference admitted.
### 9.4 Store faults — degrade, never recreate over
An unreadable or corrupt database is a store fault: readers surface *"observation store unreadable"* with the journal hint and show nothing else that depends on the store; the collector treats it as a bounded failed run and **never recreates or overwrites** an existing file. Recovery is human-sanctioned and documented: back up or move the corrupt file aside; the next run starts a fresh store; if the legacy import never completed, the still-present `history.jsonl` is re-imported (§3.5). No built-in destructive command exists.
### 9.5 Newer schema — symmetric refusal
The TUI and `status` detect a `user_version` newer than they understand and display *"observation store written by a newer Fenris — upgrade Fenris"* without partial interpretation, matching the collector's refusal (§3.6) and the forward-only upgrade rule (§10.2).
### 9.6 Repeated collector failures — flat cadence, no escalation
The timer's retry is the recovery path; the freshness grading walks fresh → missed → stale as failures persist, so degradation is visible without new state. No backoff, no notification machinery; a persistent failure reads as stale exactly like any other gap.
### 9.7 Drive-reported anomalies — facts, not alerts
`critical_warning`, media errors, and unsafe shutdowns surface as ordinary facts in the TUI and `status` (§7.2); no alerting or notification surface exists. The projection is unaffected: endurance math consumes writes, not warnings.
### 9.8 Orphaned samples — the collector re-anchors observed fact
When a collection run finds no open monitoring period (fresh store after a store fault, completed legacy re-import, or first-ever run), it opens one at the **run moment**, never backdated. This records observed fact, not intent — only the sanctioned path records a `user_disabled` close (§8.5). Coverage semantics stay intact without requiring a re-run of `fenris-monitor enable` after recovery.
---
## 10. Installation
**ADR:** [0004](../adr/0004-install-upgrade-removal-lifecycle.md). **Criteria:** IN-1–IN-10.
### 10.1 Install
`sudo make install`:
1. Builds a wheel from the checkout and installs it, with pinned dependencies (§10.5), into the dedicated Fenris-owned venv at `/opt/fenris`; a `/usr/local/bin/fenris` wrapper makes the unprivileged TUI/CLI a PATH command. The checkout is build-time input only — after install, nothing references it.
2. Verifies `python3 ≥ 3.9` and `smartctl` presence, failing cleanly otherwise (never a runtime crash).
3. Creates `/var/lib/fenris` with root-written group-read permissions and the `fenris` read group; the database file is created lazily by the first write (§3.1).
4. Places units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/` — recording **every** placed file in an explicit manifest consumed by upgrade and uninstall (§10.6).
5. **Never enables or starts units.** A fresh install is fully dormant: units present but disabled, nothing running, no monitoring period. The only opt-in is the sanctioned toggle — `fenris monitor resume` or the first-run TUI prompt — enabling the timer and opening the first period in one step.
6. Detects `./data/history.jsonl` beside the source (or accepts an explicit path), runs the idempotent single-transaction import (§3.5), and reports imported counts.
### 10.2 Upgrade
`sudo make upgrade` installs the new wheel into the same venv, syncs units and polkit against the manifest (`daemon-reload`; restart the timer only if unit contents changed **and** it is active — safe with `Persistent=no`), leaves timer state untouched, and **never kills an in-flight collection run**: a running oneshot finishes on its mapped interpreter; at worst one old-code run completes to the store. It then applies forward-only observation-store schema migrations (§3.6). `/var/lib/fenris` is never rebuilt.
### 10.3 Rollback
Best-effort by design: before migrations run, the installer snapshots `observations.db` to a one-generation `observations.db.bak`; rollback means reinstalling the previous version and restoring the backup. Automatic schema downgrade does not exist.
### 10.4 Uninstall and purge
- `make uninstall` first performs the sanctioned disable (`fenris-monitor disable --now`) so an open monitoring period closes `user_disabled` — removal is deliberate, and only the sanctioned path records intent — then stops and disables the units and removes the venv, helpers, units, polkit policy, and wrapper, **keeping** `/etc/fenris` and the observation store. Journal entries age out naturally.
- `make purge` additionally removes configuration and store.
- Reinstall after uninstall resumes from the preserved observation store; only purge erases history.
### 10.5 Dependencies
Exact pins in a committed lockfile; install and upgrade both install from it. Refreshing pins is an explicit developer step (`make update-deps`, committed), never a side effect of installing. The acquisition path adds no Python dependency and no OS package beyond smartmontools (§2.5).
### 10.6 Placement manifest
Installed artifacts sit only at their fixed locations — units in `/etc/systemd/system`, helpers in `/usr/libexec/fenris`, polkit policy under `/usr/share/polkit-1/actions/`, configuration at `/etc/fenris`, observation store under `/var/lib/fenris`, venv at `/opt/fenris`, wrapper at `/usr/local/bin/fenris` — and every placed file is recorded in the manifest (criterion IN-10).
---
## Appendix A: Traceability matrix
Built as assembly's first step (assembly decision, recommendation 5). Two-way: every ADR section maps to at least one criterion ID; every criterion cites its ADR or ticket.
### A.1 ADR section → criteria
| ADR section | Criteria |
|---|---|
| 0001 §1 Substrate | ST-1, ST-2 |
| 0001 §2 Access | ST-1, CI-3 (/run) |
| 0001 §3 Entities (incl. #12/#14 amendments) | ST-3, ST-5, PR-13, ID-2, CI-3 (no stored projections) |
| 0001 §4 Day boundary | ST-4 |
| 0001 §5 Retention | ST-5 |
| 0001 §6 Migration | ST-6, ST-7, ST-8, ST-9, ST-10, ST-11 |
| 0001 §7 Projection inputs | ST-3, PR-14, CI-3 (one-key config) |
| 0001 §8 Versioning | ST-12, FL-5 |
| 0001 §9 Collector health | LC-10 |
| 0002 §1 One projection | PR-1 |
| 0002 §2 Rate/formulas/regime/scenario | PR-2, PR-17 |
| 0002 §3 Habit change | PR-3 |
| 0002 §4 Hour classification | PR-4 |
| 0002 §5 Denominator | PR-5 |
| 0002 §6 Minimum evidence | PR-6 |
| 0002 §7 Staleness | PR-7 |
| 0002 §8 Confidence table (incl. #15 amendment) | PR-8, PR-15, CI-1, CI-4 |
| 0002 §9 Segment breaks (incl. #15 amendment) | PR-9, PR-16 |
| 0002 §10 Implied eligibility | PR-10 |
| 0002 §11 Uncertainty | PR-11 |
| 0002 §12 Language | CI-4 |
| 0002 §13 Contract | PR-12 |
| 0003 §1 Units | LC-1, CI-3 (/run) |
| 0003 §2 Cadence | LC-2, LC-3 |
| 0003 §3 Configuration | LC-4, CI-3 |
| 0003 §4 Entry points | LC-5, IN-10 |
| 0003 §5 Sanctioned toggle | LC-6, CI-2, CI-3 |
| 0003 §6 Period rows | LC-7 |
| 0003 §7 On-demand collection | LC-8, CI-2, CI-3 |
| 0003 §8 TUI controls | TUI-1, TUI-2, TUI-4 |
| 0003 §9 CLI compatibility | LC-9, CI-2 |
| 0003 §10 Freshness constants | LC-10, CI-1, CI-2 |
| 0004 §1 Delivery | IN-1 |
| 0004 §2 Layout and manifest | IN-2, IN-10 |
| 0004 §3 Privilege | IN-3, CI-3 |
| 0004 §4 Dormant install | IN-3 |
| 0004 §5 Legacy import | IN-4 |
| 0004 §6 Upgrade | IN-5 |
| 0004 §7 Rollback | IN-6 |
| 0004 §8 Removal | IN-7 |
| 0004 §9 Dependencies | IN-8, AC-5 |
| 0004 §10 Scaffolding and floor | IN-9, TUI-3 |
| 0005 §1 Malformed observations | FL-1, FL-2 |
| 0005 §2 Missed observations | FL-3, CI-3 |
| 0005 §3 Store faults | FL-4 |
| 0005 §4 Newer schema | FL-5 |
| 0005 §5 Repeated failures | FL-6 |
| 0005 §6 Drive-reported anomalies | FL-7, CI-3 |
| 0005 §7 Orphaned samples | FL-8 |
| 0006 §1 Pin | AC-1 |
| 0006 §2 Hard pin, no fallback | AC-3 |
| 0006 §3 Normalization | AC-2 |
| 0006 §4 Segment metadata sourcing | AC-4 |
| 0006 §5 Prerequisites | AC-5 |
### A.2 Ticket decisions → criteria
| Ticket | Criteria |
|---|---|
| [Evaluate Python TUI frameworks](https://git.bongbetic.com/xavierk/Fenris/issues/6) | TUI-3 |
| [Prototype the TUI information architecture](https://git.bongbetic.com/xavierk/Fenris/issues/3) | TUI-1, TUI-2, TUI-4 |
| [Verify the controller identity that segments observation history](https://git.bongbetic.com/xavierk/Fenris/issues/11) | ID-1, ID-3 |
| [Define endurance-baseline provenance and validation](https://git.bongbetic.com/xavierk/Fenris/issues/12) | PR-13, PR-14, ST-3 |
| [Decide controller-segment metadata columns](https://git.bongbetic.com/xavierk/Fenris/issues/14) | ID-2 |
| [Decide how degraded identity affects projection confidence](https://git.bongbetic.com/xavierk/Fenris/issues/15) | PR-15, PR-16, ID-4 |
| [Define cross-cutting acceptance criteria](https://git.bongbetic.com/xavierk/Fenris/issues/13) | the register itself |
| [Assemble the implementation-ready specification](https://git.bongbetic.com/xavierk/Fenris/issues/17) / [Write the Fenris redesign specification and close the map](https://git.bongbetic.com/xavierk/Fenris/issues/19) | this document and this appendix |
### A.3 Criterion → source
Every criterion carries its citation inline in the [register](acceptance-criteria.md): CI-1–CI-4 (ADR 0002 §§6–8, 0003 §§1–10, 0001 §§2–3/8, 0005 §§2/4–6); ST-1–ST-12 (ADR 0001, with ST-3 amended by tickets #12/#14); LC-1–LC-10 (ADR 0003); PR-1–PR-12 (ADR 0002), PR-13–PR-14 (ticket #12), PR-15–PR-16 (ticket #15, ADR 0002 §§8–9 as amended), PR-17 (ADR 0002 §2); ID-1–ID-4 (tickets #11/#14/#15, ADR 0001 §3 as amended); TUI-1–TUI-4 (tickets #3/#6, ADR 0003 §§8/10, ADR 0004 §10); FL-1–FL-8 (ADR 0005); IN-1–IN-10 (ADR 0004, with IN-10 also citing ADR 0003 §4); AC-1–AC-5 (ADR 0006).
### A.4 Assembly result
- **Every ADR 0001–0006 section maps to at least one criterion** — the A.1 table is complete; no orphan sections.
- **Every criterion cites its ADR or ticket** — verified in the register; no orphan criteria.
- **Decided-but-uncitered gaps found and filled inline in the register during assembly:** CI-3 bullet (no `/run` coordination surface; ADR 0003 §1, ADR 0001 §2), PR-17 (projection arithmetic; ADR 0002 §2), TUI-4 (normative Panes layout and bindings; ticket #3), IN-10 (fixed artifact placement; ADR 0004 §2, ADR 0003 §4).
- **No genuinely undecided behavior remained** — no blocking ticket was raised.
- **One reconciliation:** ADR 0004 §6's "`schema_version` table" wording resolves to ADR 0001 §8's `PRAGMA user_version` as the single version authority (§3.6); criterion ST-12 already fixed the mechanism.
+148
View File
@@ -0,0 +1,148 @@
# Glint-inspired Fenris dashboard
Status: accepted on 2026-09-19. The user confirmed the complete design and
additionally requested Glint-style plotted graphs in place of block bars.
## Purpose and reference
Adopt the visual approach of [Glint](https://github.com/ntrospect0/glint) for
Fenris's terminal dashboard. The selected visual reference is the third README
screenshot, using Chalktone, at upstream commit
`c1d73d3e8ead2f4069630b2a237306af8f6e69c8`:
[reference screenshot](https://github.com/ntrospect0/glint/blob/c1d73d3e8ead2f4069630b2a237306af8f6e69c8/docs/screenshots/glint-demo3.png).
The source and architectural assessment is recorded in
[the research note](../research/glint-dashboard-adoption.md).
## Confirmed user choices
- Adopt panel styling, keyboard focus, and panel zoom. A freely configurable
dashboard builder is outside the selected scope.
- Give the live read/write activity chart the largest opening-screen area,
with endurance outlook and monitoring state visible alongside it.
- Use the muted Chalktone appearance from Glint's third screenshot as the
visual direction.
- Put Live / Day / History views inside one activity panel, keeping both
read/write totals and the selected date visible.
- Expand a focused panel within the dashboard while retaining a fixed strip
for monitoring state, observation freshness, and essential controls.
- Replace block bars with thin dotted time-series plots resembling Glint's
chart. Label volume and time axes, retain exact selected-point readouts,
and leave explicit breaks for missing or incompatible evidence. Joining
adjacent measured points is visual guidance, not additional observations.
## Implementation recommendation
Implement the selected visual and interaction patterns independently in
Fenris's existing Python/Textual presentation layer. Retain the collector,
observation store, shared status/projection contracts, and authenticated control
path. Glint is a Rust/Ratatui application under GPL-3.0-or-later; copying or
porting its implementation into Fenris would require a separate licensing
decision under [ADR 0009](../adr/0009-mit-license.md). Referencing general panel
and navigation patterns does not require adopting its application architecture.
Use thin borders, compact titles, an explicit focused-panel indicator, and
restrained cream/earth-tone accents. Preserve readable contrast and semantic
state labels; the reference's dim secondary text is not a readability target.
Use Fenris content and identity, without Glint's unrelated clock, weather,
finance, email, or gallery features.
## Existing behavior to preserve
- Activity is measured read/write volume, with both totals accessible and
writes selected initially. A visual stock-chart reference must not turn
interval volume into speed or imply continuity through unknown evidence.
- Preserve the latest-three-hour opening view, date selection, local-day and
timezone labels, available history precision, and historical selection across
refresh under [live drive activity](live-drive-activity.md).
- Preserve the usage-adjusted theoretical lifespan and categorical projection
confidence with contributing facts. Navigation and graph selection cannot
change the projection evidence window or endurance accounting.
- Preserve explicit zero, gap, incomplete, unallocated, unavailable, paused,
stale, collection-failure, and store-fault states.
- Keep boot enablement, runtime activity, collection outcome, and observation
freshness distinct. Quit leaves background monitoring running; pause remains
a deliberate disable through the existing control path.
- Preserve keyboard and mouse access, constrained-terminal text/reflow,
high-contrast availability, reduced motion, help, and disclosures.
## Concrete layout and interaction proposal
- A compact identity header above one dashboard workspace.
- A narrow supporting column for the endurance outlook and drive facts;
a wide activity panel receives the remaining workspace. On the normal
dashboard, confidence and monitoring state remain visible beside activity.
- Live starts with the latest three hours; Day exposes the selected local
day's available detail; History exposes the retained longer-term evidence.
Reuse existing measurements and ranges rather than introducing a new data
model. Show selected date/timezone, read/write totals, measurement, and
evidence state wherever applicable.
- Use visible, clickable tabs; Tab/Shift+Tab and mouse clicks move focus.
Use `z` to toggle focused-panel zoom and Escape to restore the dashboard
when an input/dialog is not consuming Escape. Preserve selection, date,
range, measurement, and focus across zoom and refresh.
- Keep the bottom status/control area outside the expanding workspace. It
may wrap when required: compactness cannot merge the four service facts or
hide a fault/pause state. Keep quit distinct from pause in wording and
behavior without spending three full rows on a separate heavy quit box.
- Use existing date-navigation and read/write shortcuts, with visible help
updated for tabs, focus, and zoom. Text entry must consume its own keys.
- Use Chalktone-inspired styling as the new default while retaining existing
selectable themes, High Contrast, and reduced motion. Preserve an explicitly
saved theme preference during upgrade.
- Design for 80×24 and larger, with existing text/reflow behavior below that
size. Essential facts and actions must remain accessible; the large-screen
reference does not require squeezing its entire density into small terminals.
These defaults implement the confirmed choices. Date entry and inspection
retain existing evidence precision: hourly/daily UTC evidence is labelled UTC,
while local-day totals keep their recorded timezone. This visual redesign does
not manufacture finer or local-hour precision from coarse UTC evidence.
## Smallest sufficient implementation proof
Render normal and zoomed views at a representative large terminal and 80×24,
plus a constrained terminal. Check focus/tab/zoom/date/read-write interactions
and selection persistence through refresh. Exercise paused, stale, store-fault,
missing-baseline, and incomplete-evidence displays using existing synthetic
stores and headless TUI patterns. Verify quit never invokes monitoring control
and existing action tests still cover the sanctioned helper. UI fixtures are
not evidence of the user's deployed drive state.
## Specification reconciliation
This accepted redesign supersedes conflicting presentation requirements in
[the original Panes specification](fenris-redesign.md#7-panes-tui) and
[dashboard clarity](dashboard-clarity.md): the full-width headline becomes a
supporting endurance panel, the history/live plots share a tabbed activity
panel, and a compact fixed control row replaces the heavy standalone quit
rail. The block-bar requirement is superseded by dotted volume plots. Their
behavioral requirements, including clear quit-versus-pause semantics, remain.
The live-drive-activity specification excludes unrelated redesign from that
earlier task. This is a separate design request, not permission to undo its
accepted data, date-navigation, evidence, or forecast behavior. Existing design
documents and source may differ in implementation status; this assessment is
not proof that every earlier acceptance criterion has shipped.
## Documentation scope
No new domain term has been resolved: panel, focus, zoom, and theme are general
interface concepts and do not belong in the domain glossary. No new ADR is
needed for a reversible presentation change that retains the existing stack,
license, data model, and privilege boundary. Record any later durable
architectural trade-off separately if one emerges.
## Implementation validation
Implemented in Fenris's existing presentation layer with Textual 8.2.8.
The full suite passed 801 tests; 43 packaging/signing checks were skipped for
missing package artifacts or signing tools. After the final incomplete-evidence
fixes, all 81 focused dashboard, history, and plot tests passed. New renderer
and regression files pass configured Ruff checks; affected production files
pass correctness lint and the diff passes whitespace checks. Existing broader
lint warnings were not part of this redesign.
Rendered normal and zoomed dashboards at 140×44, plus normal 80×24 and
constrained 70×20 views. The [saved preview](../../assets/dashboard-chalktone.png)
uses synthetic observations. No installation or release was performed.
+117
View File
@@ -0,0 +1,117 @@
## Problem Statement
Fenris users cannot readily understand how much data their monitored drive reads and writes each day, inspect recent activity as it arrives, or select historical dates from the keyboard. Existing graph shortcuts depend on focus, date navigation is limited, collection and screen refresh default to five minutes, and the graph reads hourly/daily summaries rather than three-minute activity. Source inspection also identified derivation and rendering paths that can leave displayed history incomplete or stale despite newly stored readings.
Users want daily data volume as the primary measurement because writes contribute to drive endurance. They also want an understandable end-of-life outlook even with limited history, without presenting an unsupported hardware-failure prediction or a misleading numerical confidence score.
## Solution
Show both local-day read and write totals, updating from background collection targeted every three minutes. Open on a live last-three-hours graph of written data volume, provide a read/write toggle, and support selected-day and longer-history views. Make date navigation discoverable through visible keyboard hints and clickable equivalents.
Retain three-minute detail for 14 days, followed by durable hourly/daily summaries. Preserve the timezone and boundaries of historical local-day summaries, identify incomplete evidence, and never invent missing activity or divide midnight-spanning measurements by assumption.
Show the usage-adjusted theoretical lifespan after one full local calendar day of observations, provided an applicable endurance baseline and usable write rate exist. Explain that this estimates remaining write endurance if observed habits continue, not a physical failure date. Present categorical projection confidence with contributing facts, including limited-history reasons.
## User Stories
1. As a drive owner, I want to see the amount of data written each local day, so that I can understand the activity that contributes to write endurance.
2. As a drive owner, I want to see the amount of data read each local day, so that I can understand the drive's broader activity.
3. As a drive owner, I want read and write totals shown separately, so that reads are not mistaken for writes consuming the endurance allowance.
4. As a TUI user, I want data-volume units rather than transfer speed as the primary graph measurement, so that the graph answers how much data was transferred.
5. As a TUI user, I want today's totals labelled as totals so far, so that a partial day is not presented as a completed day's usage.
6. As a TUI user, I want the opening graph to show written volume over the last three hours, so that recent activity is visible immediately.
7. As a TUI user, I want to switch the graph between writes and reads while both daily totals remain visible, so that I can inspect either measurement without losing context.
8. As a TUI user, I want each new activity point to include transfers between readings, so that brief bursts between collection runs contribute to the graph.
9. As a TUI user, I want new measurements approximately every three minutes, so that the recent-activity view stays useful while I work.
10. As a user who closes the TUI, I want collection to continue in the background, so that reopening it shows activity gathered while I was away.
11. As a user who reboots, I want observation history and the established monitoring lifecycle preserved, so that restarting does not erase activity or confuse monitoring state.
12. As a keyboard user, I want `[` and `]` to select the previous and next day, so that browsing nearby dates takes one action.
13. As a keyboard user, I want `g` to open a date field, so that I can jump directly to a historical date.
14. As a keyboard user, I want `t` to return to today and the live view, so that I can leave historical browsing immediately.
15. As a keyboard user, I want arrows to inspect graph points, so that I can read the volume, time, and evidence state behind a point.
16. As a new TUI user, I want visible shortcut hints that work from normal launch, so that I do not have to discover an invisible graph-focus prerequisite.
17. As a mouse user, I want clickable equivalents for navigation and the read/write toggle, so that I can use the same features without memorizing shortcuts.
18. As a user entering a date, I want typing and cancelling to remain within the date-entry workflow, so that graph or monitoring shortcuts do not fire accidentally.
19. As a user browsing history, I want background refresh to preserve my selected date and measurement, so that new activity does not interrupt inspection.
20. As a user inspecting daily totals, I want to switch between recent detail, a selected day, and longer history, so that I can understand both short bursts and daily habits.
21. As a user outside UTC, I want calendar navigation and timestamps expressed in labelled local time, so that the displayed day matches my calendar.
22. As a user in a timezone with a half-hour offset, I want local-day totals based on the correct midnight boundary, so that relabelled UTC totals do not misrepresent my day.
23. As a user experiencing a daylight-saving transition, I want real local-day boundaries respected, so that a 23-hour or 25-hour day remains understandable.
24. As a user who changes system timezone, I want historical summaries to retain their recorded timezone and boundaries, so that old totals do not silently change meaning.
25. As a user reviewing recent history, I want three-minute detail available for 14 days, so that I can investigate recent usage.
26. As a user reviewing older history, I want durable hourly/daily summaries after detailed readings expire, so that long-term activity remains available.
27. As a user with legacy history, I want dates that cannot be reconstructed at local-day precision labelled incomplete or unavailable, so that old summaries are not presented with invented precision.
28. As a user whose readings straddle midnight, I want the measured volume preserved once and its uncertain day allocation explained, so that it is neither lost nor counted twice.
29. As a new user with only one reading, I want an awaiting-another-reading state, so that a missing interval is not shown as zero activity.
30. As a user with a measured zero-activity interval, I want zero distinguished from missing evidence, so that a quiet drive is not confused with a collection failure.
31. As a user with missed collection runs, I want gaps, actual timestamps, and freshness facts, so that the graph does not imply measurements that never occurred.
32. As a user who deliberately pauses monitoring, I want paused time excluded according to monitoring-period rules, so that the pause does not distort the observed usage habit.
33. As a user whose controller resets or changes, I want counter and identity boundaries respected, so that unrelated readings do not create invalid activity or lifespan estimates.
34. As a user with a store fault, I want a clear explanation instead of guessed history, so that I know which information cannot be trusted.
35. As a drive owner, I want an estimate of the time until remaining write endurance is consumed, so that I can plan around my observed usage habit.
36. As a new user, I want that estimate withheld until one full local calendar day has been observed, so that a few minutes of activity do not immediately produce a lifespan number.
37. As a user who starts monitoring at noon, I want the partial first day excluded from the full-day gate, so that 24 hours since launch is not mistaken for a complete calendar day.
38. As a user with one complete day but limited history, I want an estimate when the other inputs support it, so that I do not have to wait for high confidence before seeing an outlook.
39. As a user with limited evidence, I want a confidence category and concrete reasons, so that I understand why an estimate may change without being shown an unjustified percentage.
40. As a user without an applicable endurance baseline or usable write rate, I want the missing prerequisite explained, so that unavailable evidence is not disguised as merely low confidence.
41. As a drive owner, I want the endurance estimate distinguished from a physical failure date, so that I do not treat a write-endurance projection as a hardware guarantee.
42. As a CLI user, I want status and the TUI to agree on forecast eligibility and confidence, so that choosing a different interface does not change the facts.
43. As a user on a narrow terminal, I want daily totals, dates, evidence labels, and controls to remain accessible, so that terminal size does not hide the information I need.
44. As a user of existing themes and reduced-motion settings, I want those features preserved when navigation shortcuts change, so that better date controls do not remove accessibility or preferences.
45. As a user upgrading Fenris, I want migration, repair, refresh, and pruning to preserve measured evidence without duplication, so that the new graph can be trusted across restarts and upgrades.
## Implementation Decisions
- Extend the existing collector, observation store, derivation, shared status/projection, service scheduling, and TUI responsibilities. The privileged collector remains the single device-acquisition and store-write owner; the TUI remains an unprivileged reader. Do not add another sampler, service, export layer, or generic plotting abstraction.
- Set the default background collection target to three minutes on both systemd and runit. Update shared cadence/freshness semantics, scheduler definitions, user documentation, and relevant acceptance contracts together. Use actual observation timestamps: delayed runs must not be represented as perfectly spaced measurements or synthetic catch-up points.
- Derive interval read/write volumes from compatible cumulative counters. A first reading is an anchor, not an interval. Preserve controller-segment and monitoring-period boundaries, measured zero, unknown time, deliberate disables, and unsupported evidence. Do not compute deltas across a controller identity or counter discontinuity.
- Make successful collection publish consistent data for dependent read views, using the existing transactional and migration patterns. Ensure raw samples, affected summaries, and retained boundary evidence cannot disagree through partial publication. Validate source-level derivation gaps before implementing fixes; this spec does not claim a runtime diagnosis.
- Keep daily read and write totals separate. The write graph is the default; a read/write toggle changes only the displayed measurement. Each live point represents measured interval volume, not instantaneous speed. Daily totals remain visible independently of graph mode, and graph selection must not redefine the projection's evidence window.
- Open on the latest three hours. Offer selected-day and longer-history views at the precision supported by retained evidence. Selected-point readouts expose time, labelled timezone, data-volume units, and incomplete/gap state. Do not imply that a partial current day or an incompletely allocated day is a known full-day total.
- Use `[` / `]` for previous/next day, `g` for date entry, `t` for today/live, and arrows for point inspection, with visible hints and clickable equivalents. Date entry must handle invalid or unavailable dates visibly and must not leak keystrokes into unrelated actions. Cancelling returns to the prior selection; refresh preserves historical selection until the user changes it.
- Resolve the existing `t` theme binding in favor of the accepted today/live action. Preserve theme selection through a discoverable non-conflicting control and update help consistently. Preserve pause, resume, collect-now, disclosures, quit, and reduced-motion behavior; quitting still does not pause monitoring. No exact replacement theme shortcut was selected in the interview.
- Preserve UTC timestamps and the established UTC hour/day evidence used by endurance calculations. Local-day activity is a separate presentation aggregate using the actual local midnight boundaries, including non-whole-hour offsets and daylight-saving transitions. Timestamp relabelling alone is insufficient.
- Extend the existing versioned observation store to retain local-day read/write summaries with their controller identity/segment context, recorded timezone, local date, UTC boundaries, evidence state, and the information needed to explain unallocated boundary volume. These are required semantics, not prescribed table names or a finalized column layout. The collector derives durable summaries before source evidence is eligible for pruning.
- Retain three-minute detail for 14 days and hourly/daily summaries indefinitely. Keep the boundary evidence and anchors necessary for honest derivation. Do not introduce unbounded retention of all fine-grained intervals solely to support arbitrary historical timezone reinterpretation, and do not delete existing durable evidence merely to simplify migration.
- Preserve the timezone and boundaries recorded for historical summaries when the system timezone changes. New observations follow the current local timezone without rewriting old days; labels must make historical timezone context explicit. Do not blend differently bounded summaries into an apparently exact daily total.
- Allocate a measured interval to a local day only when evidence supports that allocation. Preserve an ambiguous midnight-spanning delta once as shared/unallocated boundary evidence; neither prorate it nor copy its full value into both days. If only coarse legacy UTC summaries survive, show them at their actual precision and mark unreconstructable local-day totals incomplete/unavailable. Repeated derivation and repair must be idempotent.
- Use the existing usage-adjusted theoretical lifespan model and categorical projection confidence. Forecast remaining write endurance from an applicable endurance baseline and a usable observed write rate. Lifetime-written counters consume the endurance allowance; observation-window deltas determine the usage rate. Read volume is not included in endurance consumption.
- Add the agreed complete-observation-day gate to the shared projection contract. A completed local midnight-to-midnight day within a monitoring period with usable evidence is required; a partial first day, deliberately disabled span, or day with no usable evidence cannot independently qualify. Monday-noon setup can first qualify at Wednesday 00:00 after observing Tuesday. Real daylight-saving day boundaries apply; do not replace this with a fixed 24-hour timer.
- Before that gate, explain that Fenris is waiting for a full local observation day. Afterward, allow an estimate with Limited confidence when the existing baseline, rate, identity, and evidence-validity rules permit it; preserve the longer warm-up and Supported-confidence requirements. Missing evidence is still evaluated under the existing validity/coverage rules: completing a date does not turn unknown measurements into known data.
- Show confidence as a category plus contributing facts, never a numeric confidence percentage. Explain missing baseline, unusable/zero rate, stale evidence, and controller changes through their applicable states. Do not fabricate a baseline, suppress a missing prerequisite behind a generic confidence warning, or label the output as a predicted hardware-failure date. Preserve the existing baseline setup path and zero-rate explanation.
- Preserve visible freshness, last-collection outcome, paused state, and store-fault behavior. On constrained terminals retain the existing textual fallback/reflow approach so both daily totals, the selected date, evidence state, confidence, and navigation remain accessible; colour alone must not carry meaning.
- Migrate through the existing ordered, versioned, transactional store mechanism. Preserve observation history, refuse unknown newer schemas, establish new durable evidence before pruning, and keep readers consistent while the collector writes. Migration cannot manufacture local precision from historical data that lacks it.
- This design amends ADR 0003's five-minute default on both native service backends under ADR 0008, adds a first-full-local-day eligibility gate to ADR 0002 while retaining its forecasting and confidence model, and extends ADR 0001 with the local-history decision recorded as ADR 0010. ADR 0005's no-fabrication/store-fault contract remains authoritative. Existing historical design documents are not evidence that the current runtime already implements these amendments.
## Testing Decisions
- Prefer one primary acceptance path: controlled acquisition fixtures and an injected clock enter the existing public collection operation; it writes a real temporary observation store; the normal read path supplies the TUI and CLI. Assert resulting displayed volumes, evidence states, selections, forecast eligibility, and public outcomes. Do not mock the internal derivation or query results under test. A good test would continue passing after an internal refactor and fail when a user's observed result becomes incorrect.
- Reuse existing collector tracer and collector-history tracer fixtures, including SMART JSON, a temporary sysfs tree, the injected clock, and a real SQLite store. Compose these with the existing status reader, projection fixtures, and headless TUI driver rather than creating a new production testing interface. Fix defects that this path reproduces; do not make tests reproduce incorrect implementation arithmetic.
- Exercise normal launch, previous/next date, direct date entry, invalid input, cancel, today/live, arrow inspection, read/write switching, refresh during historical browsing, and hourly/day navigation through user actions. Assert visible graph/readout results, not only private selection indices or row counts. In particular, do not manually force graph focus to make the advertised shortcuts work. Existing headless graph, help, theme, motion, and constrained-terminal tests provide prior art.
- Drive known successive read/write counter changes through collection and verify exact interval volumes and aggregate totals. Cover first reading, first compatible pair, measured zero, repeated collection, repeated read/refresh, restart, same-hour accumulation, and a day-boundary crossing. Preserve each measured byte once: attributable totals plus separately retained unallocated evidence must account for the measured delta without duplication.
- Use the same public path for local midnight, Asia/Kolkata's half-hour boundary, a 23-hour local day, a 25-hour local day, repeated local clock labels, and a system-timezone change. Verify historical labels and boundaries remain stable and ambiguous intervals are not silently divided. Add direct derivation cases only when they prove an invariant that the primary path cannot isolate clearly.
- Advance controlled time beyond the 14-day detail window, invoke normal pruning, restart readers, and verify recent detail expiry with durable summary/boundary preservation. Use existing pruning, repair, legacy-migration, and store-migration tests for focused persistence cases that the main path cannot prove, including interruption, reruns, concurrent readers, unknown newer schema, and inability to reconstruct legacy local-day precision.
- Test the forecast gate at concrete local times: Monday noon setup; Tuesday noon with 24 elapsed hours but no complete observed calendar day; and Wednesday 00:00 after a usable Tuesday. Include partial/disabled/unknown days and daylight-saving dates. After eligibility, verify Limited confidence with reasons and a correct estimate where inputs support one; no baseline, invalid counters, zero rate, stale evidence, and segment changes must preserve the relevant unavailable/degraded behavior. Reuse existing projection and status-parity tests.
- Independently calculate an expected endurance result from known lifetime-written counters, a known applicable baseline, and known observed-window deltas. This guards against confusing lifetime consumption with writes inside the rate window. Rendering, browsing another day, and switching to reads must not change that accounting.
- Add focused checks at the existing native scheduler boundary for the three-minute defaults, serialization, failures, and freshness calculations. Reuse systemd/runit control and packaging tests; use only necessary controlled platform probes to verify actual scheduled/background behavior. No live-system scheduling, pause, install, or reboot action is authorized merely by writing this spec. Report any unavailable platform verification during implementation.
- Verify at 80×24 and below that visible totals, date entry, navigation, confidence, and evidence states remain accessible and that theme/reduced-motion controls still work. Avoid pixel-perfect snapshots and assertions about private widget layout when visible text and action results express the contract.
- The affected modules are collection/derivation, observation-store migration and retention, shared projection/status, native service scheduling, and the TUI. Most acceptance coverage should remain at the single collection-to-visible-result path; targeted persistence and scheduler checks are supporting boundaries, not a new layer of test-only architecture.
## Out of Scope
- Adding SATA/HDD acquisition, another monitored device, or multi-drive management; the selected existing NVMe drive remains the source.
- Reporting occupied/free filesystem capacity, making transfer speed the primary measurement, or treating read bytes as write-endurance consumption.
- Forecasting today's final data volume or future daily activity; the selected forecast is the usage-adjusted theoretical lifespan.
- Predicting a physical hardware-failure date, adding numeric confidence percentages, fabricating endurance baselines, or introducing a new forecasting algorithm when the existing model fits.
- Interpolating missing samples, zero-filling unknown activity, prorating ambiguous midnight intervals, or retroactively rewriting historical timezone boundaries.
- Requiring all historical dates to retain three-minute detail forever or to support arbitrary later timezone regrouping.
- A new web UI, plotting service, generic graph framework, notifications, telemetry, theme redesign, or unrelated dashboard rework.
- Package publication, release creation, installation, host-monitoring changes, or runtime diagnosis of the user's deployed observation store as part of this specification task.
## Further Notes
- This spec synthesizes the completed design conversation. The user accepted daily volume, a live three-hour opening view, the specified keyboard date workflow, local dates, 14-day detail retention, a writes-first graph with read toggle and both totals, an endurance forecast, categorical confidence, one full local observation day before forecasting, and historically labelled timezone boundaries with honest incomplete evidence.
- Related prior issue: [Implement Fenris TUI polish and hourly history](https://git.bongbetic.com/xavierk/Fenris/issues/72). This newer accepted design supersedes its conflicting requirements for the opening graph, writes-only presentation, the `t` theme shortcut, unrestricted historical timezone regrouping, indefinite fine-grained interval retention, and the earlier projection-eligibility presentation. Preserve unrelated requirements for identity, status, themes, motion, accessibility, collector publication, evidence conservation, and CLI parity. Do not treat the older issue as a reason to undo these newly agreed choices.
- ADR 0010 records the deliberate trade-off: retain labelled local summaries and necessary boundary evidence before detailed readings expire, rather than retaining every fine-grained interval for arbitrary future timezone reinterpretation. The design adds durable local activity evidence while retaining UTC projection evidence and the existing collector/reader ownership boundary.
- Source inspection identified normal collection paths that do not consistently rebuild daily totals, existing-hour updates that omit read deltas, cross-day deltas added to both dates, and drill-down rendering overwritten by a loading placeholder. These are concrete investigation/regression targets, not claims of executed reproductions or deployed-store corruption.
- The user confirmed the existing collector-to-store-to-TUI/CLI test boundary, with focused migration/pruning and scheduler checks where needed. Publication is an implementation handoff; no application behavior or test result is claimed solely by creating this issue.
+66
View File
@@ -0,0 +1,66 @@
# Native Void Linux and XBPS distribution
Status: Approved and published as [Support native Void Linux and signed XBPS distribution through Gitea](https://git.bongbetic.com/xavierk/Fenris/issues/82).
## Problem Statement
Void users cannot install and operate Fenris natively through XBPS because its runtime lifecycle assumes systemd and its release pipeline only produces Debian and RPM packages. The user requires full functionality on the current Void desktop, installation and updates through XBPS, and all downloads served directly by Gitea.
## Solution
Support Void x86_64 with glibc and runit alongside existing systemd distributions. Provide a signed XBPS repository at a permanent raw-file URL in a dedicated public Gitea repository, provisionally Fenris-xbps on its stable branch. Publish versioned assets and notes in the application's Gitea release. Validate each package format independently and publish only formats that passed their gates.
## User Stories
1. As a Void user, I want to install Fenris through XBPS so package ownership and dependencies are managed normally.
2. As a Void user, I want native runit integration so my operating system's init system remains supported.
3. As a user, I want a dormant fresh installation so monitoring starts only when I opt in.
4. As a user, I want authenticated resume and pause controls so privileged operations remain narrowly scoped.
5. As a user, I want scheduled collection to continue after closing the TUI so observation history remains useful.
6. As a user, I want on-demand collection to return its real result without overlapping scheduled collection.
7. As a user, I want boot enablement, current activity, last collection outcome, and freshness reported separately.
8. As a user, I want actionable native diagnostics when collection fails.
9. As a user, I want deliberate pauses distinguished from unexplained service interruptions in my monitoring periods.
10. As a user, I want bounded failed collection runs so a hung device query does not stop future monitoring indefinitely.
11. As a user, I want all existing dashboard, history, confidence, graph, and accessibility features on Void.
12. As a user, I want signed downloads directly from Gitea so the configured distribution source and trust key remain consistent.
13. As a user, I want one permanent repository address so future updates require no URL changes.
14. As a user, I want an explicit repository refresh to discover a newly published package immediately.
15. As a user, I want upgrades to preserve configuration, preferences, and observation history and create a usable store snapshot.
16. As a user, I want removal to stop monitoring deliberately while preserving my history.
17. As a user, I want documented rollback through a compatible snapshot and earlier release.
18. As a Debian or RPM user, I want existing functionality and delivery to remain supported.
19. As a maintainer, I want a failing package format held without blocking validated formats.
20. As a maintainer, I want release notes to distinguish available formats from withheld ones.
21. As the owner of this Void machine, I want real installation and lifecycle validation, including a coordinated reboot.
22. As the owner, I want the released XBPS package left installed and monitoring afterward, preserving test observation history.
## Implementation Decisions
- Amend the existing systemd-only service contract to support native runit while retaining its external semantics. Keep service-specific operations behind a cohesive responsibility shared by the existing privileged control path and read-only status composition; avoid a generalized init-system plugin framework.
- Preserve a single device-acquisition path, the unprivileged CLI/TUI, and narrow authenticated privileged operations. Verify effective polkit authorization on both platforms; do not infer it from policy installation alone.
- Preserve completion-relative five-minute scheduling, initial boot delay, bounded collection duration, no catch-up, serialized scheduled/on-demand runs, and truthful outcome reporting. Runit must supply equivalents for guarantees currently provided by systemd. Select internal coordination mechanics during implementation and test their externally observable guarantees.
- Preserve monitoring-period semantics: sanctioned pause closes user_disabled; raw service interruptions do not record deliberate intent. Preserve separate boot-enabled and runtime-active facts.
- Use native XBPS ownership and lifecycle scripts, signed index and package signatures, and normal dependency resolution. Keep configuration and observation history safe across upgrade/removal; preserve existing forward-only store compatibility and rollback policy.
- Host ordinary Git blobs without LFS in the dedicated Gitea repository. Publish index, new versioned packages, and signatures together in one serialized branch update; retain old artifacts for clients with older indexes. Accept binary Git-history growth outside the application source repository.
- Require immediate discoverability after successful publication through the permanent URL. Test clients that fetched the previous index first. Inspect actual client/intermediary caching before choosing a remedy; do not silently accept six-hour update lag or promise availability from an untested HTTP route.
- Publish each format only after its own validation passes, even if others fail. Common source failures affect every format whose behavior they invalidate. Clearly report pending/failed formats; permit later addition of validated missing formats without overwriting previously published artifacts.
- Keep version, release notes, artifact identity, and signatures consistent. Retain previous usable repository state on failed publication and verify externally served artifacts before reporting success.
- First Void target is x86_64/glibc. Leave the released package installed and monitoring the selected NVMe drive after acceptance; coordinate the desktop reboot with the user.
## Testing Decisions
- Primary runtime boundary: existing user commands and shared CLI/TUI observable state, backed by disposable observation stores and controlled service/acquisition outcomes. Test behavior, not a particular helper layout.
- Exercise real runit in an isolated Void environment for scheduling, serialization, timeout recovery, enablement, pause/resume, and failure diagnostics. Preserve meaningful systemd regression coverage.
- Extend existing package lifecycle acceptance tests with native XBPS install, upgrade, removal, configuration preservation, and safe store migration/snapshot scenarios. Use disposable stores for destructive removal/rollback cases.
- Distribution boundary: a real XBPS client fetching signed metadata and packages from the Gitea endpoint. Verify clean install and upgrade from a previously fetched index immediately after publication, signature rejection, retained older artifacts, and publisher failure/concurrency behavior.
- Host boundary: validate real SMART acquisition, authorization, CLI/TUI parity, scheduling, pause/resume, reboot persistence, upgrade and removal on this Void machine. Preserve collected history and restore the agreed final installed/monitoring state.
- Record actual outcomes, skips, and environment failures. Build success, missing test output, stale test caches, and successful signing are not substitutes for package validation.
## Out of Scope
Musl, other architectures, additional init systems, official Void repository inclusion, a separate download server, automatic client upgrades, unrelated dashboard redesign, and automatic downgrade of a newer observation store.
## Further Notes
The design decisions are recorded in ADR 0008. The dedicated distribution repository and end-to-end XBPS proof are complete; the host acceptance record is [issue #87](https://git.bongbetic.com/xavierk/Fenris/issues/87). The accepted target is Void x86_64/glibc with runit, with the released package installed and monitoring the selected NVMe drive after the coordinated reboot and lifecycle checks.
+71
View File
@@ -0,0 +1,71 @@
# Native Void implementation tickets
Status: Approved and published as Gitea issues 83–87 with native blocking edges and ready-for-agent labels. Parent specification: https://git.bongbetic.com/xavierk/Fenris/issues/82.
Each issue references the published native Void specification, which includes ADR 0008. Existing Debian/RPM behavior, observation-history preservation, narrowly scoped privilege, and independent publication gates apply throughout.
Published tickets: [1](https://git.bongbetic.com/xavierk/Fenris/issues/83), [2](https://git.bongbetic.com/xavierk/Fenris/issues/84), [3](https://git.bongbetic.com/xavierk/Fenris/issues/85), [4](https://git.bongbetic.com/xavierk/Fenris/issues/86), [5](https://git.bongbetic.com/xavierk/Fenris/issues/87).
## 1. Prove signed XBPS installation and immediate updates through Gitea
Blocked by: None.
Deliver a dedicated public Gitea distribution repository and a repeatable, isolated XBPS-client proof using clearly identified test artifacts, without representing them as a validated Fenris release.
- Verify permanent raw URL delivery of index, package and signature bytes without LFS indirection.
- Establish signing-key handling and trust verification without exposing private keys.
- Demonstrate install and update with a client that fetched the old index immediately before publication.
- Resolve actual cache behavior and verify binary size limits; escalate any hosting change outside the agreed Gitea scope.
- Publish index/artifacts together with serialized updates, preserve older downloadable artifacts, and demonstrate safe failure recovery.
- Record results and the usable publication mechanism for subsequent tickets.
## 2. Run and control Fenris monitoring natively under runit
Blocked by: None.
Deliver an end-to-end native monitoring path with existing CLI/TUI controls and truthful status in an isolated Void environment.
- Scheduled and on-demand collection use the same acquisition path with no overlap and bounded execution.
- Preserve cadence, initial boot delay, recovery after failures, and no catch-up semantics.
- Resume/pause preserve monitoring-period bookkeeping and boot/runtime distinctions; raw service stops do not record deliberate disable.
- Verify effective authentication and actionable native diagnostics, including missing-agent failures.
- Preserve systemd behavior and application feature parity through existing public behavior tests.
## 3. Install, upgrade and remove Fenris with native XBPS packages
Blocked by: 2.
Deliver buildable x86_64/glibc XBPS artifacts with dependency resolution and tested package lifecycle in an isolated Void environment.
- Fresh install remains dormant; native controls enable monitoring afterward.
- Package ownership, configuration preservation, permissions, and group access support real CLI/TUI reads and collector writes.
- Upgrade snapshots and migrates observation history safely without recording a deliberate pause.
- Removal performs sanctioned pause and retains history; reinstall and documented snapshot rollback behave correctly.
- Installation over incompatible unmanaged remnants fails with a useful migration path.
- Capture explicit lifecycle test results and verify Debian/RPM regressions relevant to changed packaging.
## 4. Publish validated package formats independently from the release workflow
Blocked by: 1, 3.
Deliver a release workflow that builds, validates, signs and publishes XBPS alongside existing Debian/RPM support with independent format gates.
- Unvalidated formats remain withheld while validated formats can ship.
- Notes accurately identify available and withheld formats; source version, notes, checksums and artifacts agree.
- A withheld format can be added after validation without replacing existing published artifacts.
- Gitea serves every download; XBPS uses the proven permanent repository URL and immediate-refresh behavior.
- Release failure/concurrency cannot expose an index referencing missing artifacts or erase the prior usable channel.
- Host acceptance remains a required XBPS release gate, not bypassed by build/signature success.
## 5. Validate and release on the user's Void machine
Blocked by: 4.
Deliver recorded host acceptance and the validated XBPS release, ending with Fenris installed and monitoring.
- Recheck host state, identify/configure the intended NVMe drive, and preserve pre-existing data before package lifecycle operations.
- Verify real acquisition, authenticated controls, scheduling, failure reporting, full dashboard behavior, and pause/resume semantics.
- Coordinate and verify reboot persistence; absence of the reboot test leaves that gate pending.
- Verify native upgrade/removal/reinstall with history preservation and immediate update discovery through Gitea.
- Publish only after the XBPS gate passes, then verify downloads and released-package installation.
- Preserve acceptance-test observation history and leave the released package monitoring; document installation, trust setup, diagnostics and rollback for Void users.
+107
View File
@@ -0,0 +1,107 @@
# Fenris release and packaging specification
**Status: decision-complete.** Assembled by [Task: Compose release spec + ADR amending 0004](https://git.bongbetic.com/xavierk/Fenris/issues/42) from the closed tickets of the Wayfinder map [Fenris deb + rpm release plan](https://git.bongbetic.com/xavierk/Fenris/issues/33). This document is normative for the follow-up **execution effort** that builds and publishes packages; no packages are built here.
**Canonical roles.** [ADR 0007](../adr/0007-package-delivery-amends-0004.md) records the lifecycle rationale (amending [ADR 0004](../adr/0004-install-upgrade-removal-lifecycle.md)); this document restates the **operative contracts** — compat matrix, channel, toolchain, signing, release mechanics, package layout, maintainer-script behavior, migration — so the executing effort never needs Wayfinder-ticket access. Runtime semantics come from [ADRs 0001–0006](../adr/) and the [redesign specification](fenris-redesign.md) verbatim; nothing here overrides them. Terminology follows the glossary in [`CONTEXT.md`](../../CONTEXT.md), including *Release* and *Rollback*.
**Binding language.** *Must*, *exactly*, and *never* are normative.
## 1. Compatibility matrix
| Target | Version | Format | Registry placement |
|---|---|---|---|
| Debian 12 (bookworm) | — | deb | `debian/pool/bookworm/main` |
| Ubuntu 22.04 (jammy) | — | deb | `debian/pool/jammy/main` |
| Ubuntu 24.04 (noble) | — | deb | `debian/pool/noble/main` |
| Fedora 40+ | every release | rpm | `rpm/fenris` group |
| openSUSE Tumbleweed | rolling | rpm | `rpm/fenris` group |
- Architecture: **x86_64 only** (arm64 only if real ARM hardware appears — map fog).
- Dependencies are vendored as locked, pure-Python runtime packages for every target: Debian 12 and Ubuntu 22.04/24.04 ship `python3-textual` 0.1.13, far below the floor; Fedora 40+ ships ≥ 0.48 but below the pin ([toolchain research](../research/deb-rpm-toolchain.md)). No distro `python3-textual` dependency ever enters package metadata.
- Package metadata `depends:`/`Requires:` are exactly `python3 (>= 3.10)`, `smartmontools`, `systemd` — the Python floor is 3.10 (oldest supported distro interpreter, Ubuntu 22.04), bumping ADR 0004 §10's 3.9 gate for packages; `make install` keeps the checkout's floor.
- Runtime packages are staged with `python3 -m pip --target /opt/fenris/vendor`; entry points run the target system's `python3` with that directory on the import path. No package ships a copied Python interpreter, avoiding build-host ABI paths and rolling-distribution minor-version breakage.
## 2. Distribution channel
- **Channel:** the self-hosted Gitea 1.27.1 package registry at `git.bongbetic.com`, owner public for anonymous consumers ([registry research](../research/gitea-package-registry.md); [OBS rejected](../research/obs-route.md)).
- **Single channel.** No stable/testing split — deferred until external users ask to track pre-release builds (map fog). Every published version is retained indefinitely (registry has no REST cleanup; republishing a filename is a 409).
- **deb publication:** one deb artifact PUT to each codename pool — `PUT /api/packages/{owner}/debian/pool/{bookworm|jammy|noble}/main/upload`.
- **rpm publication:** one rpm artifact PUT to the single `fenris` group — `PUT /api/packages/{owner}/rpm/fenris/upload` — serving Fedora 40+ collectively.
- **Consumer setup (install docs, normative):**
- apt: keyring file from `…/debian/repository.key` via `signed-by`, one sources line per distribution, instance Debian Registry Key **fingerprint printed beside the curl one-liner** (TOFU hardening).
- dnf: `dnf config-manager --add-repo <raw-url of packaging/fenris.repo>` — the in-repo, Fenris-owned `.repo` with `gpgkey` pointing at the published packaging key and `repo_gpgcheck=0`. **Gitea's auto-generated `.repo` is never mentioned in docs**: it sets `gpgcheck=1` against the instance auto-key, which never signed our rpm payload — a trap that breaks installs.
## 3. Build toolchain
- **nfpm** for both formats from a single `packaging/nfpm.yaml` — one config, `overrides:` for per-format deltas, two invocations (`nfpm pkg -p deb`, `nfpm pkg -p rpm`). fpm is dropped entirely (CLI-flag config drifts); no hand rpm spec; dh-virtualenv is deb-only and dormant since 2020.
- **Single source of truth:** version injected from `pyproject.toml`; file lists generated by a staging script (locked runtime packages → `/opt/fenris/vendor`, plus wrapper, helpers, units, polkit policy, sysusers/tmpfiles fragments) referenced by `nfpm.yaml` as a `type: tree` content entry — no hand-maintained file lists.
- **Entry point:** `make package` → `dist/fenris_<v>_amd64.deb` + `dist/fenris-<v>-1.x86_64.rpm`.
- **Version scheme:** `<pyproject-version>-1` in both formats; a rebuild of the same upstream version bumps the revision (`-2`, `-3`, …) — the same filename is never re-PUT (registry 409s duplicates).
- **Authoring `nfpm.yaml`, the staging script, and the workflow file is execution** — deliberately not part of the decision map. The [toolchain research doc](../research/deb-rpm-toolchain.md) sketches the pipeline.
## 4. Signing and key policy
- **RPM payload: signed.** rpmsign with the dedicated packaging key, invoked by `make sign-rpm` after the package is built. This is required, not optional: it is the only working dnf-native verification path.
- **deb: unsigned.** apt never verifies payload signatures; trust = instance-signed `InRelease` (signed-by keyring) + TLS + Acquire-By-Hash. Manual-download integrity is covered by SHA256SUMS.
- **SHA256SUMS: clearsigned** with the packaging key — the trust anchor for manually downloaded release assets, independent of TLS.
- **Packaging key:** single dedicated key, RSA 3072, UID `Fenris Packaging <packaging@bongbetic.com>`, 2-year expiry, no master/subkey hierarchy. The private key is stored as the repository Actions secret `GPG_PRIVATE_KEY`. The release workflow imports it on the self-hosted runner, verifies it against the in-repo public key, signs the RPM and SHA256SUMS, then deletes the runner's keyring copy in an `always()` cleanup step. The full ceremony is documented in `docs/install/signing-key-ceremony.md`.
- **XBPS key:** separate RSA 3072 key stored as the repository Actions secret `XBPS_SIGNING_KEY`; its public key and fingerprint are published in `Fenris-xbps`. The release workflow uses it for the XBPS package and, when publication is explicitly requested, the repository index. A final `always()` cleanup deletes its runner copy after publication and release asset upload.
- **Public key publication:** in-repo `packaging/keys/fenris-packaging.asc` (raw URL doubles as the `.repo` gpgkey target), release notes, docs page. No keyservers — TOFU-over-TLS.
- **Rotation (outline):** new key published alongside old; rpm signed with the new key; `fenris.repo` gpgkey lists both URLs (dnf accepts multiple); old key dropped after one release cycle. Procedure details stay in map fog.
## 5. Release mechanics
- **A Release is:** a version tag, its packages in the channel, a Gitea release entry with notes, and a clearsigned SHA256SUMS — all together. **Bare tags are forbidden** (tag without packages + release entry is not a Release).
- **Cadence: on-demand.** Tag when user-visible changes or fixes accumulate; no calendar, no empty releases, no frequency SLA, no RC ceremony — fixes ship as a revision bump of the current version.
- **Versioning: plain semver.** Major = breaking CLI/config/unit change; store schema changes ride the natural bump (the forward-only refusal handles old-reader/new-store).
- **Promotion flow:** bump `pyproject.toml` and the matching dated `CHANGELOG.md` section, then push tag `v<version>`. The repository-scoped Gitea Actions workflow builds and validates the deb, rpm, checksums, and release entry. XBPS publication remains a manual dispatch option after host acceptance. `make release` is for local artifact preparation and does not replace the tag workflow as the supported publication path. The ceremony is documented in `docs/install/signing-key-ceremony.md`.
- **Rollback:** installing an older package over a newer store is **unsupported** — the store's forward-only version refusal fails it by design. Documented rollback = restore the observation-store snapshot, then install the old Release. No automatic downgrade machinery exists or will be built.
- **CI:** a repository-scoped self-hosted runner is registered and online (checked 2026-09-29). `.gitea/workflows/release.yml` is the tag-triggered release path; maintainers must confirm runner availability and required Gitea secrets before tagging. The workflow publishes deb/rpm packages and release assets; Void publication is withheld unless the signed XBPS host-release step is explicitly requested.
## 6. Package layout and ownership
Per [ADR 0007](../adr/0007-package-delivery-amends-0004.md) §2 — the dpkg/rpm database is the manifest; no `manifest.txt` ships:
| Artifact | Location | Ownership |
|---|---|---|
| Bundled runtime packages | `/opt/fenris/vendor` | package (tree) |
| Wrapper | `/usr/bin/fenris` | package |
| Helpers | `/usr/libexec/fenris/{fenris-monitor,fenris-collect}` | package — exactly these two, no new polkit-reachable binaries |
| Units | `/usr/lib/systemd/system/fenris-collect.{timer,service}` | package (vendor placement; `/etc/systemd/system` is admin-only) |
| Polkit policy | `/usr/share/polkit-1/actions/com.bongbetic.fenris.monitor.policy` | package |
| sysusers fragment | `/usr/lib/sysusers.d/fenris.conf` (`g fenris -`) | package |
| tmpfiles fragment | `/usr/lib/tmpfiles.d/fenris.conf` (`d /var/lib/fenris 2750 root fenris -`) | package |
| Configuration | `/etc/fenris/fenris.conf` | package as conffile / `%config(noreplace)` — placeholder-commented default, no active selector |
| Observation store | `/var/lib/fenris/observations.db` (+ WAL, `.bak`) | **never owned, never ghosted** — the package owns the directory only |
## 7. Maintainer-script contracts
- **preinst / %pre:** abort with a pointer to the migration runbook (§9) if `/var/lib/fenris/manifest.txt` **or** `/etc/systemd/system/fenris-collect.timer` exists (dual marker covers pre-manifest make installs). No auto-clean — scripts never delete files outside the package DB.
- **postinst / %post (install):** `systemd-sysusers`, `systemd-tmpfiles --create`, `systemctl daemon-reload`. Nothing else — no enable, no preset, no start; no preset file ships.
- **postinst / %post (upgrade):** snapshot `observations.db` → `.bak` (one generation) → forward-only schema migration via target `python3` with `/opt/fenris/vendor` on its import path → `daemon-reload` → restart `fenris-collect.timer` only if unit contents changed **and** it is active. `/var/lib/fenris` is never rebuilt; an in-flight oneshot finishes on its old interpreter.
- **prerm / %preun:** sanctioned disable (`fenris-monitor disable --now`, closing the monitoring period `user_disabled`) on remove/erase **only, never on upgrade** — deb prerm upgrade case is a no-op; rpm `%preun` gated on `$1 -eq 0`.
- **Removal mapping:** deb `remove` ≈ `make uninstall` (conffile + store survive); deb `purge` ≈ `make purge` (+ `.bak`, group cleanup); rpm erase ≈ `make uninstall` (unmodified config removed, modified survives as `.rpmsave`); rpm purge = documented manual command.
## 8. Initial configuration
- The device selector is **entered by hand**: root edits `/etc/fenris/fenris.conf` (world-readable, exactly one key per [ADR 0003](../adr/0003-service-lifecycle-and-sanctioned-toggle.md) §3). The shipped default is placeholder-commented and carries no active selector — a fresh install reads as a `configuration error`-free dormant system until the first `fenris monitor resume` + edit, exactly the dormant-install contract.
- No configuration verb is added to `fenris-monitor`; the polkit surface stays at one binary. (This resolves the open item from [Package ownership + ADR 0004 amendment](https://git.bongbetic.com/xavierk/Fenris/issues/40): nothing in any delivery writes the selector — `make install` never wrote `fenris.conf` either; hand-editing has been the model since ADR 0003.)
- On upgrade, local edits survive; a changed package default lands as `.dpkg-new` / `.rpmnew`.
## 9. Migration from make-install systems
- **Runbook only** — no migration script, no auto-clean. Population is author machines plus a few testers; store and config survive by path continuity.
- **Remove-then-install, mandatory:** `sudo make uninstall` (preserves store + `/etc/fenris`) → `apt install fenris` / `dnf install fenris`. **Over-install is forbidden:** stale `/etc/systemd/system/fenris-collect.*` silently shadows vendor units (systemd precedence), `/usr/local/bin/fenris` shadows `/usr/bin/fenris` on PATH.
- **No-move continuity:** `/var/lib/fenris` untouched; existing `fenris` group → sysusers no-op; existing dir → tmpfiles no-op; hand-written `fenris.conf` survives (dpkg ships the default as `.dpkg-new`; rpm as `.rpmnew`); store schema caught up by the upgrade-path migration.
- **Reset-to-dormant:** `make uninstall`'s sanctioned disable closes the open period `user_disabled`; after migration the user opts back in with `fenris monitor resume` — one ≤15-min sample gap, honest against the endurance timeline.
- **Mutual exclusion:** package and `make install` never on the same machine. The dev loop is checkout + `make test`; a dev-local install mode stays in map fog.
## 10. Out of scope
- Building and publishing packages (the follow-up execution effort).
- Snap/Flatpak/AppImage/Homebrew, PyPI as an install path.
- Stable/testing channel split, COPR/PPA fallback, arm64 — map fog until demand appears.
---
*Assembled from [Lock channel + toolchain](https://git.bongbetic.com/xavierk/Fenris/issues/38), [Signing + key policy](https://git.bongbetic.com/xavierk/Fenris/issues/39), [Package ownership + ADR 0004 amendment](https://git.bongbetic.com/xavierk/Fenris/issues/40), [Migration path from make-install systems to packages](https://git.bongbetic.com/xavierk/Fenris/issues/41), [Release cadence + stable/testing channel split](https://git.bongbetic.com/xavierk/Fenris/issues/43), [Actions runner availability](https://git.bongbetic.com/xavierk/Fenris/issues/37), and the research tickets [deb + rpm packaging toolchain](https://git.bongbetic.com/xavierk/Fenris/issues/34), [Gitea 1.27 package registry feasibility](https://git.bongbetic.com/xavierk/Fenris/issues/35), [OBS route](https://git.bongbetic.com/xavierk/Fenris/issues/36).*
+153
View File
@@ -0,0 +1,153 @@
# XBPS Proof of Concept Results (Issue #83)
Status: Complete. All acceptance criteria satisfied.
## Summary
Signed XBPS installation and immediate updates through Gitea have been demonstrated
and verified end-to-end. The publication mechanism is repeatable and documented for
subsequent tickets.
## Acceptance Criteria Results
### 1. Permanent raw URL delivery without LFS indirection
**Result: PASS**
All artifacts are served directly by Gitea via raw-file URLs:
- Repository: `https://git.bongbetic.com/xavierk/Fenris-xbps`
- Raw URL pattern: `https://git.bongbetic.com/xavierk/Fenris-xbps/raw/branch/stable/x86_64/<file>`
Verified files:
- `x86_64-repodata` (1327 bytes) — HTTP 200
- `fenris-0.3.5_1.x86_64.xbps` (731 bytes) — HTTP 200
- `fenris-0.3.5_1.x86_64.xbps.sig2` (384 bytes) — HTTP 200
- `fenris-0.3.6_1.x86_64.xbps` (752 bytes) — HTTP 200
- `fenris-0.3.6_1.x86_64.xbps.sig2` (384 bytes) — HTTP 200
No LFS indirection detected. Files served as ordinary Git blobs.
### 2. Signing-key handling and trust verification
**Result: PASS**
Signing uses SSH RSA keys (3072-bit) via `xbps-rindex --sign-pkg` and `xbps-rindex --sign`.
The public key is embedded in the repository metadata (`index-meta.plist`) within the
repodata archive. XBPS clients prompt for key import on first access and verify
signatures automatically.
Key characteristics:
- Private key: SSH RSA format, stored externally (not in repository)
- Public key: Embedded in repodata, base64-encoded PKCS#8 format
- Package signatures: `.sig2` files alongside each `.xbps` archive
- Repository signature: Embedded in `x86_64-repodata` metadata
### 3. Install and update with client that fetched old index
**Result: PASS**
Demonstrated full lifecycle:
1. Fresh install of v0.3.5 from repository with only v0.3.5 in index
2. Added v0.3.6 to index and committed to `stable` branch
3. Client performed `xbps-install -Syu` and upgraded from 0.3.5 → 0.3.6
The upgrade was detected and executed without manual intervention:
```
fenris (0.3.5_1 -> 0.3.6_1)
```
### 4. Cache behavior and binary size limits
**Result: PASS (with documented caveat)**
Cache behavior:
- Gitea raw endpoint: `cache-control: public, max-age=21600` (6-hour cache)
- ETag changes on each commit (different file hash)
- Conditional requests with old ETag return 200 (full content), not 304
xbps-install behavior:
- `-M -S` fetches fresh repodata while bypassing the on-disk cache
- The ordinary `-S` path can reuse a cached repodata archive
- No 6-hour delay observed in practice
Binary size limits:
- Test packages: 731–752 bytes (small test artifacts)
- Real packages expected to be <10MB (vendored pure-Python)
- Gitea serves any file size without LFS
**Caveat**: Clients using `xbps-install -Su` or `-Syu` without `-M` may use
cached repodata. Users should use `-M -Syu` for updates when immediate
publication visibility matters.
### 5. Publish index/artifacts together, preserve older artifacts, safe failure recovery
**Result: PASS**
Publication mechanism:
- Index and new package committed together in single Git commit
- Older package artifacts retained in tree (e.g., v0.3.5 alongside v0.3.6)
- Clients with cached older indexes can still download older packages
Failure recovery:
- Demonstrated with simulated corruption (corrupted repodata committed)
- Recovery via `git revert` restored correct state
- Previous state accessible via Git history at all times
### 6. Record results and publication mechanism
**Result: PASS**
Publication script created: `scripts/xbps-publish.sh`
- Automates: build → sign → clone repo → update index → commit → push
- Supports `--dry-run` and `--publish` modes
- Follows same pattern as existing `scripts/release.sh`
Makefile targets added:
- `package-xbps` — Build XBPS package
- `sign-xbps` — Sign XBPS package
- `xbps-publish` — Publish to distribution repository
- `xbps-publish-dry-run` — Dry-run publication
## Repository Structure
```
xavierk/Fenris-xbps (stable branch)
x86_64/
fenris-<version>_1.x86_64.xbps # Package archives
fenris-<version>_1.x86_64.xbps.sig2 # Package signatures
x86_64-repodata # Repository index (zstd-compressed tar)
keys/
fenris-xbps-signing.pub # Public signing key
README.md
```
## Client Configuration
Users add the repository to `/etc/xbps.d/xbps.conf`:
```
repository=https://git.bongbetic.com/xavierk/Fenris-xbps/raw/branch/stable/x86_64
```
Or make a one-off request without writing the configuration file:
```sh
sudo xbps-install -R https://git.bongbetic.com/xavierk/Fenris-xbps/raw/branch/stable/x86_64 -M -S fenris
```
## Publication Mechanism
1. Build: `make package-xbps`
2. Sign: `make sign-xbps` (or `scripts/xbps-publish.sh` handles this)
3. Publish: `make xbps-publish`
4. Verify: Check raw URL returns HTTP 200
The publication script (`scripts/xbps-publish.sh`) automates the full flow:
build → sign → clone repo → add to index → sign repository → commit → push.
## Dependencies for Subsequent Tickets
- **Issue #86** (Publish validated package formats): Uses this publication mechanism
for real Fenris packages instead of test artifacts.
- **Issue #87** (Validate on Void machine): Uses the established repository URL
and signing infrastructure for end-to-end validation.
-1052
View File
File diff suppressed because it is too large Load Diff
-128
View File
@@ -1,128 +0,0 @@
#!/usr/bin/env bash
# Fenris — interactive menu for the NVMe wear monitor & dashboard.
# Created by Bongbetic.
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
PY="$SCRIPT_DIR/fenris.py"
PORT_DEFAULT=8420
INTERVAL_DEFAULT=300
banner() {
cat <<'EOF'
_____ _
| __|___ ___ _| |___
| __| -_| | . | _|
|__| |___|_|_|_|___|_|
NVMe wear monitor & live dashboard
Created by Bongbetic
EOF
}
pause() { read -rp "Press Enter to continue..." _; }
detect_device() {
# Query Python's auto-detect for the default device.
python3 -c "import sys; sys.path.insert(0,'$SCRIPT_DIR'); from fenris import detect_device; print(detect_device())" 2>/dev/null || echo /dev/nvme0
}
menu() {
clear
banner
echo
echo " 1) Start monitoring (background daemon + dashboard)"
echo " 2) Stop monitoring"
echo " 3) Status / current wear stats"
echo " 4) Take one sample right now"
echo " 5) Open dashboard URL"
echo " ---"
echo " h) Help / how this works"
echo " q) Exit"
echo
read -rp "Choose an option: " choice
echo
case "$choice" in
1) start_flow ;;
2) python3 "$PY" stop; pause ;;
3) python3 "$PY" status; pause ;;
4) read -rp "Device [default: auto-detect]: " dev
if [ -z "$dev" ]; then python3 "$PY" sample; else python3 "$PY" sample --device "$dev"; fi
pause ;;
5) show_url; pause ;;
h|H) help_text; pause ;;
q|Q) echo "Bye. — Fenris, by Bongbetic"; exit 0 ;;
*) echo "Invalid choice."; pause ;;
esac
}
start_flow() {
read -rp "NVMe device [Enter = auto-detect]: " dev
read -rp "Sample interval in seconds [Enter = ${INTERVAL_DEFAULT}]: " interval
read -rp "Dashboard port [Enter = ${PORT_DEFAULT}]: " port
interval="${interval:-$INTERVAL_DEFAULT}"
port="${port:-$PORT_DEFAULT}"
args=(start --interval "$interval" --port "$port")
if [ -n "${dev:-}" ]; then args+=(--device "$dev"); fi
echo
echo "Note: reading NVMe SMART data needs root."
echo "Fenris runs 'sudo -n smartctl ...' (no-prompt sudo). If this fails,"
echo "either run this menu with sudo, or allow passwordless smartctl via:"
echo " sudo visudo -> youruser ALL=(root) NOPASSWD: /usr/sbin/smartctl"
echo
python3 "$PY" "${args[@]}"
pause
}
show_url() {
if [ -f "$SCRIPT_DIR/data/fenris.pid" ]; then
# Try to read actual port from running process cmdline, else guess default.
local pid port
pid=$(<"$SCRIPT_DIR/data/fenris.pid")
port=$(tr '\0' '\n' < /proc/"$pid"/cmdline 2>/dev/null | grep -A1 -- '--port' | tail -1 || true)
port="${port:-$PORT_DEFAULT}"
echo "Dashboard: http://localhost:${port}"
else
echo "Fenris is not currently running. Start it first (option 1)."
fi
}
help_text() {
cat <<EOF
What Fenris does:
- Periodically reads your NVMe drive's SMART health data (via smartctl),
including "percentage_used" (the drive's own wear indicator), total
bytes written/read, temperature, spare capacity, and error counts.
- Logs every sample to: $SCRIPT_DIR/data/history.jsonl
and per-hour aggregates to: $SCRIPT_DIR/data/hourly.jsonl (GB/hour,
rebuilt from history on restart). Used for the trailing-24h bar chart
and exact rolling-24h write volume.
- Serves a live HTML dashboard (dense layout, interval-synced polling)
with wear-over-time and trailing-24h hourly-write charts, plus a
projected life-remaining estimate in hours/days/years derived from
your actual rolling-24h write rate and implied TBW endurance.
Requirements:
- smartmontools (smartctl) installed.
- Root access to read NVMe SMART logs — either run Fenris via sudo,
or set up passwordless sudo for smartctl (see option 1).
CLI usage (equivalent to this menu):
python3 fenris.py start [--device /dev/nvme0] [--interval 300] [--port 8420]
python3 fenris.py stop
python3 fenris.py status
python3 fenris.py sample [--device /dev/nvme0]
Leave it running in the background (option 1) and check back after a
few days/weeks of normal use — more samples = a more accurate lifespan
estimate.
Fenris — created by Bongbetic.
EOF
}
# Loop instead of recurse to avoid stack overflow.
while true; do menu; done
+16
View File
@@ -0,0 +1,16 @@
# Fenris configuration
#
# This file is managed by the fenris package. Local edits are preserved
# across upgrades; changed defaults appear as .dpkg-new / .rpmnew.
#
# The device selector specifies which NVMe drive to monitor.
# Uncomment and set exactly one device path:
#
# device = /dev/disk/by-id/nvme-Samsung_SSD_980_PRO_500GB_S5PANS0T123456
#
# The observation store path is optional and defaults to
# /var/lib/fenris/observations.db when unset:
#
# store_path = /var/lib/fenris/observations.db
#
# See https://git.bongbetic.com/xavierk/Fenris for documentation.
+7
View File
@@ -0,0 +1,7 @@
[fenris]
name=Fenris NVMe Monitor
baseurl=https://git.bongbetic.com/api/packages/xavierk/rpm/fenris
enabled=1
gpgcheck=1
gpgkey=https://git.bongbetic.com/xavierk/Fenris/raw/branch/main/packaging/keys/fenris-packaging.asc
repo_gpgcheck=0
+40
View File
@@ -0,0 +1,40 @@
# Fenris Packaging Key
#
# Public half of dedicated RSA-3072 key used to sign RPM payloads and
# clearsign SHA256SUMS manifests.
#
# Fingerprint: CE4542E1E23EB50F09EDFFA5A5E8B22D1872FB07
# Algorithm: RSA 3072
# UID: Fenris Packaging <packaging@bongbetic.com>
# Expiry: 2 years from creation
#
# Raw URL:
# https://git.bongbetic.com/xavierk/Fenris/raw/branch/main/packaging/keys/fenris-packaging.asc
#
# The private half lives in approved secret storage only. Each release uses
# import -> sign -> delete. See docs/install/signing-key-ceremony.md.
#
-----BEGIN PGP PUBLIC KEY BLOCK-----
mQGNBGqZYxgBDADEyYQhndEEzoETD17vk8/4x2DoXQm9hxW7hiX3TQNSmXORpgjR
NNt0vV/rTptwjmxgkrlevjrYqiBuoXfKJ0WRC16e9+NRnCGwJX5F4sR7jfgS9XbH
pAbLbySll5LfrD6JcPcB4JsSishKkY6X0zHQD0/zrCaOsuNdLj+fLhWDoxjpLFGy
92U7KHwtt87vSmUM4FgAjUY4keVKqIP5pSWIcPEy7z025RytL1JP6z3jBJR7KKD/
MLXd2KTGGaxTIvzgimcvjQYqFxrT2YIRsmhVYzddRyUnYYWgOh9dp5xn9CRM48Lz
klxHI/jI4lRPCQJy0atBpGZk1bRIc9XsBXRWiR9Zwvpjah/r3whkknBFv7C4srMr
Fl0Ml897zwbCsCP/Ejs43SYt+BJ6B2z4lg6KVsG6lypqV8B+gT+E6rJZ/ML6UpS3
2XQBlWDAw3hf0abjZIOQegi0f5igc7TVyDzYYihLOjKRwZuGSHY40w4mohxESLy8
ogdMveFJKoRZvBsAEQEAAbQqRmVucmlzIFBhY2thZ2luZyA8cGFja2FnaW5nQGJv
bmdiZXRpYy5jb20+iQH0BBMBCABeFiEEzkVC4eI+tQ8J7f+lpeiyLRhy+wcFAmqZ
YxgbFIAAAAAABAAObWFudTIsMi41KzEuMTIsMiwyAxsvBAUJA8JnAAULCQgHAgIi
AgYVCgkICwIEFgIDAQIeBwIXgAAKCRCl6LItGHL7B7qZC/4yFm3JwuhXuaJ6JvsH
ZNVCAVOywktFbdcfKJYCXayaVsQ0Yc1w/gW6XhYCr4EECfWplnjtta9zPnN61ODD
B8ZIuM9VUOqxpwvBWJnHcnny1FjmbJ0r0NOwmqKMj54cFHEDbVzmVPoshQSukThj
Uz28XXw/JOkeQQaVl6OF3MoLvhLrLWvnqX310Z151dpl1lEA6gYWd1eKau2oIfU4
e6u1JnX6mKWb0WaaEqo1QARXloQTaKV+NiSUavckTn1LXXMxGCFkbtNWYZv7uf+U
WFB8KuaR3u8if7R8Bab7Y0lzmPCCeSkLHXDLq9FyfDdGOj5LyXxxzlVnTTUyRANh
+JwGiokDqUzh3yUdUYnx6pE6+3tcxP+Gp92K/GZXulmPQYhw0sSyqfnK8GAtUVaD
KAnrV7fKZ9jve87NWeb3G0xfQiH9mNSsEmnQVzd/DCuczOf5fMFfHlUNgkI4t4G7
+hTpKrOOubozZwfB23mdM+H9pxwWFN6To85Iy1ge9JKTFTY=
=V4/R
-----END PGP PUBLIC KEY BLOCK-----
+60
View File
@@ -0,0 +1,60 @@
name: fenris
arch: amd64
platform: linux
version: "${VERSION}"
maintainer: Fenris Maintainers <ops@bongbetic.com>
description: >
NVMe wear monitor with persistent TUI — observes real-world drive use and
translates it into an understandable endurance outlook.
homepage: https://git.bongbetic.com/xavierk/Fenris
license: MIT
depends:
- python3 (>= 3.10)
- smartmontools
- systemd
contents:
# Staged tree: runtime packages, wrapper, helpers, units, polkit, sysusers, tmpfiles
- src: build/stage/
dst: /
type: tree
# Configuration directory
- dst: /etc/fenris
type: dir
file_info:
mode: 0755
# Default placeholder-commented config (deb conffile / rpm %config(noreplace))
- src: packaging/fenris.conf
dst: /etc/fenris/fenris.conf
type: config|noreplace
file_info:
mode: 0644
# Observation store directory — owned by package, never packed.
# Store files (observations.db, WAL sidecars, .bak) are never owned.
- dst: /var/lib/fenris
type: dir
file_info:
mode: 02770
group: fenris
scripts:
preinstall: packaging/preinst.sh
postinstall: packaging/postinst.sh
preremove: packaging/prerm.sh
postremove: packaging/postrm.sh
overrides:
rpm:
depends:
- python3 >= 3.10
- smartmontools
- systemd
scripts:
preinstall: packaging/preinst.sh
postinstall: packaging/rpm/post.sh
preremove: packaging/rpm/preun.sh
postremove: packaging/rpm/postun.sh
+78
View File
@@ -0,0 +1,78 @@
#!/bin/sh
# postinst — deb install and upgrade paths (spec §7).
#
# dpkg calls: postinst configure [most-recently-configured-version]
# fresh install: $1 = "configure", $2 = ""
# upgrade: $1 = "configure", $2 = old version
set -eu
STORE_DIR="/var/lib/fenris"
STORE_DB="${STORE_DIR}/observations.db"
STORE_BAK="${STORE_DIR}/observations.db.bak"
RUNTIME_PYTHON="/usr/bin/python3"
VENDOR_DIR="/opt/fenris/vendor"
# Detect init system
if [ -d /run/systemd/system ] || [ "$(cat /proc/1/comm 2>/dev/null)" = "systemd" ]; then
INIT_SYSTEM="systemd"
else
INIT_SYSTEM="runit"
fi
case "${1:-}" in
configure)
if [ -n "${2:-}" ]; then
# Upgrade — snapshot, migration, init-system-aware reload
if [ -f "${STORE_DB}" ]; then
cp "${STORE_DB}" "${STORE_BAK}" 2>/dev/null || true
fi
if [ -d "${VENDOR_DIR}" ] && [ -f "${STORE_DB}" ]; then
PYTHONPATH="${VENDOR_DIR}" "${RUNTIME_PYTHON}" -c "
from fenris.store import migrate_to_latest
from pathlib import Path
n = migrate_to_latest(Path('${STORE_DB}'))
print(f'Fenris migration: {n} step(s) applied') if n else None
" 2>&1 || echo "Fenris: migration skipped (store not yet initialized)"
fi
if [ "${INIT_SYSTEM}" = "systemd" ]; then
# Capture running unit content BEFORE daemon-reload (spec §7)
RUNNING_UNITS=""
for unit in fenris-collect.timer; do
if systemctl is-active --quiet "${unit}" 2>/dev/null; then
RUNNING_UNITS="${RUNNING_UNITS} ${unit}"
fi
done
systemctl daemon-reload 2>/dev/null || true
# Restart timer only if unit contents changed AND active
for unit in ${RUNNING_UNITS}; do
OLD_CONTENT="$(mktemp)"
NEW_PATH="/usr/lib/systemd/system/${unit}"
systemctl cat "${unit}" > "${OLD_CONTENT}" 2>/dev/null || true
if ! diff -q "${OLD_CONTENT}" "${NEW_PATH}" > /dev/null 2>&1; then
systemctl restart "${unit}" 2>/dev/null || true
fi
rm -f "${OLD_CONTENT}"
done
fi
fi
# Fresh install: init-system-specific setup
if [ "${INIT_SYSTEM}" = "systemd" ]; then
systemd-sysusers || true
systemd-tmpfiles --create || true
systemctl daemon-reload || true
else
# runit: create group, set directory permissions
groupadd -f fenris
install -d -o root -g fenris -m 2770 "${STORE_DIR}"
chmod 02770 "${STORE_DIR}"
# Mark runit service as dormant (down) for fresh install
if [ -d /etc/sv/fenris-collect ] && [ ! -e /var/service/fenris-collect ]; then
touch /etc/sv/fenris-collect/down
fi
# Ensure log directory exists
install -d -o root -g fenris -m 2770 /var/log/fenris-collect 2>/dev/null || true
fi
;;
abort-upgrade|abort-install|disappear)
;;
esac
+32
View File
@@ -0,0 +1,32 @@
#!/bin/sh
# postrm — deb post-removal (spec §7, §9).
#
# dpkg calls: postrm remove (after package files removed)
# postrm purge (after conffiles and config removed)
set -eu
# Detect init system
if [ -d /run/systemd/system ] || [ "$(cat /proc/1/comm 2>/dev/null)" = "systemd" ]; then
INIT_SYSTEM="systemd"
else
INIT_SYSTEM="runit"
fi
case "${1:-}" in
purge)
rm -rf /etc/fenris
rm -rf /var/lib/fenris
if getent group fenris > /dev/null 2>&1; then
groupdel fenris 2>/dev/null || true
fi
;;
remove|upgrade|failed-upgrade|abort-install|abort-upgrade|disappear)
;;
esac
if [ "${INIT_SYSTEM}" = "systemd" ]; then
systemctl daemon-reload 2>/dev/null || true
else
# runit: clean up service directory and log
rm -rf /etc/sv/fenris-collect 2>/dev/null || true
rm -rf /var/log/fenris-collect 2>/dev/null || true
fi
+19
View File
@@ -0,0 +1,19 @@
#!/bin/sh
# preinst — abort if make-install remnants detected (spec §7, §9).
set -eu
MARKER1="/var/lib/fenris/manifest.txt"
MARKER2="/etc/systemd/system/fenris-collect.timer"
if [ -f "${MARKER1}" ] || [ -f "${MARKER2}" ]; then
echo >&2
echo >&2 "Fenris make-install remnants detected — refusing to install."
echo >&2
echo >&2 "Migrate to the package with:"
echo >&2 " sudo make uninstall # removes make-install files, preserves store + config"
echo >&2 " sudo apt install fenris # or: sudo dnf install fenris"
echo >&2
echo >&2 "See: https://git.bongbetic.com/xavierk/Fenris/blob/main/docs/spec/release-packaging.md#9-migration-from-make-install-systems"
echo >&2
exit 1
fi
+33
View File
@@ -0,0 +1,33 @@
#!/bin/sh
# prerm — deb pre-removal (spec §7).
#
# dpkg calls: prerm remove (package being removed)
# prerm upgrade (old version about to be replaced)
set -eu
# Detect init system
if [ -d /run/systemd/system ] || [ "$(cat /proc/1/comm 2>/dev/null)" = "systemd" ]; then
INIT_SYSTEM="systemd"
else
INIT_SYSTEM="runit"
fi
case "${1:-}" in
remove)
# Sanctioned disable — close monitoring period (spec §7)
if [ -x /usr/libexec/fenris/fenris-monitor ]; then
/usr/libexec/fenris/fenris-monitor disable --now 2>/dev/null || true
fi
if [ "${INIT_SYSTEM}" = "systemd" ]; then
systemctl stop fenris-collect.timer 2>/dev/null || true
systemctl disable fenris-collect.timer 2>/dev/null || true
else
# runit: remove the service symlink
rm -f /var/service/fenris-collect
touch /etc/sv/fenris-collect/down 2>/dev/null || true
fi
;;
upgrade)
# Never interrupt monitoring on upgrade
;;
esac
+30
View File
@@ -0,0 +1,30 @@
## Install
Install Fenris from its package channel after following the [package setup instructions](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/README.md#install-from-package-recommended):
```bash
sudo apt update && sudo apt install fenris # Debian / Ubuntu
sudo dnf install fenris # Fedora
sudo zypper install fenris # openSUSE Tumbleweed
sudo xbps-install fenris # Void Linux
```
For Void Linux, configure the XBPS repository first:
```bash
sudo install -d -m 0755 /etc/xbps.d
echo 'repository=https://git.bongbetic.com/xavierk/Fenris-xbps/raw/branch/stable/x86_64' \
| sudo tee /etc/xbps.d/fenris.conf
sudo xbps-install -M -S fenris
```
## Verify downloads
```bash
gpg --output SHA256SUMS --decrypt SHA256SUMS.asc
sha256sum -c SHA256SUMS
```
## Rollback
Installing an older package over a newer observation store is unsupported. Restore the observation-store snapshot, then install the earlier Release; see the [upgrade and rollback guidance](https://git.bongbetic.com/xavierk/Fenris/src/branch/main/README.md#upgrade).
+77
View File
@@ -0,0 +1,77 @@
#!/bin/sh
# RPM %post — post-install/upgrade scriptlet (spec §7).
set -eu
STORE_DIR="/var/lib/fenris"
STORE_DB="${STORE_DIR}/observations.db"
STORE_BAK="${STORE_DIR}/observations.db.bak"
RUNTIME_PYTHON="/usr/bin/python3"
VENDOR_DIR="/opt/fenris/vendor"
# Detect init system
if [ -d /run/systemd/system ] || [ "$(cat /proc/1/comm 2>/dev/null)" = "systemd" ]; then
INIT_SYSTEM="systemd"
else
INIT_SYSTEM="runit"
fi
if [ "$1" -eq 1 ]; then
# Fresh install
if [ "${INIT_SYSTEM}" = "systemd" ]; then
systemd-sysusers || true
systemd-tmpfiles --create || true
systemctl daemon-reload || true
else
# runit: create group, set directory permissions
groupadd -f fenris
install -d -o root -g fenris -m 2770 "${STORE_DIR}"
chmod 02770 "${STORE_DIR}"
# Mark runit service as dormant (down) for fresh install
if [ -d /etc/sv/fenris-collect ] && [ ! -e /var/service/fenris-collect ]; then
touch /etc/sv/fenris-collect/down
fi
# Ensure log directory exists
install -d -o root -g fenris -m 2770 /var/log/fenris-collect 2>/dev/null || true
fi
elif [ "$1" -ge 2 ]; then
# Upgrade — snapshot, migration, init-system-aware reload
if [ -f "${STORE_DB}" ]; then
cp "${STORE_DB}" "${STORE_BAK}" 2>/dev/null || true
fi
if [ -d "${VENDOR_DIR}" ] && [ -f "${STORE_DB}" ]; then
PYTHONPATH="${VENDOR_DIR}" "${RUNTIME_PYTHON}" -c "
from fenris.store import migrate_to_latest
from pathlib import Path
n = migrate_to_latest(Path('${STORE_DB}'))
print(f'Fenris migration: {n} step(s) applied') if n else None
" 2>&1 || echo "Fenris: migration skipped (store not yet initialized)"
fi
if [ "${INIT_SYSTEM}" = "systemd" ]; then
# Capture running unit content BEFORE daemon-reload (spec §7)
RUNNING_UNITS=""
for unit in fenris-collect.timer; do
if systemctl is-active --quiet "${unit}" 2>/dev/null; then
RUNNING_UNITS="${RUNNING_UNITS} ${unit}"
fi
done
systemctl daemon-reload 2>/dev/null || true
# Restart timer only if unit contents changed AND active
for unit in ${RUNNING_UNITS}; do
OLD_CONTENT="$(mktemp)"
NEW_PATH="/usr/lib/systemd/system/${unit}"
systemctl cat "${unit}" > "${OLD_CONTENT}" 2>/dev/null || true
if ! diff -q "${OLD_CONTENT}" "${NEW_PATH}" > /dev/null 2>&1; then
systemctl restart "${unit}" 2>/dev/null || true
fi
rm -f "${OLD_CONTENT}"
done
# Re-apply placement modes (store dir group access, issue #54)
systemd-tmpfiles --create || true
else
# runit: repair log access when upgrading from an older package
groupadd -f fenris
install -d -o root -g fenris -m 2770 "${STORE_DIR}"
chmod 02770 "${STORE_DIR}"
install -d -o root -g fenris -m 2770 /var/log/fenris-collect 2>/dev/null || true
fi
fi
+26
View File
@@ -0,0 +1,26 @@
#!/bin/sh
# RPM %postun — post-uninstall scriptlet (spec §7).
set -eu
# Detect init system
if [ -d /run/systemd/system ] || [ "$(cat /proc/1/comm 2>/dev/null)" = "systemd" ]; then
INIT_SYSTEM="systemd"
else
INIT_SYSTEM="runit"
fi
if [ "$1" -eq 0 ]; then
# Package fully erased — remove config, store, group
rm -rf /etc/fenris
rm -rf /var/lib/fenris
if getent group fenris > /dev/null 2>&1; then
groupdel fenris 2>/dev/null || true
fi
fi
if [ "${INIT_SYSTEM}" = "systemd" ]; then
systemctl daemon-reload 2>/dev/null || true
else
# runit: clean up service directory and log
rm -rf /etc/sv/fenris-collect 2>/dev/null || true
rm -rf /var/log/fenris-collect 2>/dev/null || true
fi
+26
View File
@@ -0,0 +1,26 @@
#!/bin/sh
# RPM %preun — pre-uninstall scriptlet (spec §7).
set -eu
# Detect init system
if [ -d /run/systemd/system ] || [ "$(cat /proc/1/comm 2>/dev/null)" = "systemd" ]; then
INIT_SYSTEM="systemd"
else
INIT_SYSTEM="runit"
fi
if [ "$1" -eq 0 ]; then
# Package is being erased — sanctioned disable (spec §7)
if [ -x /usr/libexec/fenris/fenris-monitor ]; then
/usr/libexec/fenris/fenris-monitor disable --now 2>/dev/null || true
fi
if [ "${INIT_SYSTEM}" = "systemd" ]; then
systemctl stop fenris-collect.timer 2>/dev/null || true
systemctl disable fenris-collect.timer 2>/dev/null || true
else
# runit: remove the service symlink
rm -f /var/service/fenris-collect
touch /etc/sv/fenris-collect/down 2>/dev/null || true
fi
fi
# On upgrade ($1 -ge 1): do nothing
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env bash
# Stage a packaging tree at build/stage/ for nfpm consumption.
#
# Usage: packaging/stage.sh [VERSION]
#
# VERSION defaults to the version in pyproject.toml.
# The staged tree contains:
# /opt/fenris/vendor/ — bundled pure-Python application dependencies
# /usr/bin/fenris — unprivileged wrapper
# /usr/libexec/fenris/ — fenris-monitor, fenris-collect
# /usr/lib/systemd/system/ — fenris-collect.{timer,service}
# /usr/share/polkit-1/actions/ — polkit policy
# /usr/lib/sysusers.d/fenris.conf
# /usr/lib/tmpfiles.d/fenris.conf
set -euo pipefail
# Package directories must be traversable by every runtime user and agree
# with distribution-owned directories, regardless of the builder's umask.
umask 022
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
STAGE_DIR="${REPO_ROOT}/build/stage"
# --- Resolve version ---
if [ -n "${1:-}" ]; then
VERSION="$1"
else
VERSION="$(sed -n 's/^version = "\(.*\)"/\1/p' "${REPO_ROOT}/pyproject.toml")"
fi
if [ -z "${VERSION}" ]; then
echo "Error: could not determine version" >&2
exit 1
fi
echo "Staging fenris ${VERSION} ..."
# --- Clean previous stage ---
rm -rf "${STAGE_DIR}"
mkdir -p "${STAGE_DIR}"
# --- Use pre-built wheel from dist/ ---
WHEEL=$(ls "${REPO_ROOT}"/dist/fenris-"${VERSION}"-*.whl 2>/dev/null | head -1)
if [ -z "${WHEEL}" ]; then
echo "Error: no wheel found in dist/ — run 'make dist/fenris-*.whl' first" >&2
exit 1
fi
echo " Using wheel: $(basename "${WHEEL}")"
# --- Vendor runtime packages without an interpreter ---
# A copied Python binary contains an ABI and build-host dynamic-library path.
# It fails after rolling-distribution Python upgrades (for example Tumbleweed
# 3.12 -> 3.13). Fenris and its locked dependencies are pure Python, so place
# them in a version-neutral directory and execute with the target's python3.
echo " Installing version-neutral runtime packages ..."
VENDOR_DIR="${STAGE_DIR}/opt/fenris/vendor"
mkdir -p "${VENDOR_DIR}"
python3 -m pip install --disable-pip-version-check --no-compile \
--target "${VENDOR_DIR}" -r "${REPO_ROOT}/requirements.txt" "${WHEEL}"
# Package files must be importable by unprivileged Fenris users regardless of
# the builder's umask. pip otherwise preserves a restrictive umask in the
# vendored runtime, which makes the installed CLI fail before it can read the
# observation store.
chmod -R a+rX "${VENDOR_DIR}"
# Ship the application license in every native package.
install -D -m 0644 "${REPO_ROOT}/LICENSE" "${STAGE_DIR}/usr/share/licenses/fenris/LICENSE"
# --- Inject version into wrapper from pyproject.toml ---
# The wrapper has a hardcoded version string; patch it for packaging.
WRAPPER_SRC="${REPO_ROOT}/scripts/fenris"
WRAPPER_DST="${STAGE_DIR}/usr/bin/fenris"
mkdir -p "$(dirname "${WRAPPER_DST}")"
sed "s|version=\"%(prog)s [0-9.]*\"|version=\"%(prog)s ${VERSION}\"|g" \
"${WRAPPER_SRC}" > "${WRAPPER_DST}"
chmod 0755 "${WRAPPER_DST}"
# --- Privileged helpers ---
echo " Installing helpers ..."
mkdir -p "${STAGE_DIR}/usr/libexec/fenris"
install -m 0755 "${REPO_ROOT}/src/fenris/monitor.py" "${STAGE_DIR}/usr/libexec/fenris/fenris-monitor"
install -m 0755 "${REPO_ROOT}/src/fenris/collect.py" "${STAGE_DIR}/usr/libexec/fenris/fenris-collect"
# --- systemd units (vendor placement) ---
echo " Installing systemd units ..."
mkdir -p "${STAGE_DIR}/usr/lib/systemd/system"
install -m 0644 "${REPO_ROOT}/units/fenris-collect.timer" "${STAGE_DIR}/usr/lib/systemd/system/"
install -m 0644 "${REPO_ROOT}/units/fenris-collect.service" "${STAGE_DIR}/usr/lib/systemd/system/"
# --- runit service files ---
echo " Installing runit service files ..."
mkdir -p "${STAGE_DIR}/etc/sv/fenris-collect/log"
install -m 0755 "${REPO_ROOT}/units/runit/fenris-collect/run" "${STAGE_DIR}/etc/sv/fenris-collect/run"
install -m 0755 "${REPO_ROOT}/units/runit/fenris-collect/log/run" "${STAGE_DIR}/etc/sv/fenris-collect/log/run"
# --- polkit policy ---
echo " Installing polkit policy ..."
mkdir -p "${STAGE_DIR}/usr/share/polkit-1/actions"
install -m 0644 "${REPO_ROOT}/polkit/com.bongbetic.fenris.monitor.policy" \
"${STAGE_DIR}/usr/share/polkit-1/actions/"
# --- sysusers and tmpfiles fragments ---
echo " Installing sysusers/tmpfiles fragments ..."
mkdir -p "${STAGE_DIR}/usr/lib/sysusers.d"
install -m 0644 "${REPO_ROOT}/packaging/sysusers.d/fenris.conf" "${STAGE_DIR}/usr/lib/sysusers.d/"
mkdir -p "${STAGE_DIR}/usr/lib/tmpfiles.d"
install -m 0644 "${REPO_ROOT}/packaging/tmpfiles.d/fenris.conf" "${STAGE_DIR}/usr/lib/tmpfiles.d/"
echo "Stage complete: ${STAGE_DIR}"
+3
View File
@@ -0,0 +1,3 @@
# System user/group for Fenris observation store access
# Created by systemd-sysusers during package install
g fenris -
+2
View File
@@ -0,0 +1,2 @@
# Type Path Mode User Group Age Argument
d /var/lib/fenris 2770 root fenris - -
+70
View File
@@ -0,0 +1,70 @@
#!/bin/sh
# XBPS INSTALL script — post-install and post-upgrade paths.
#
# Arguments: $1=ACTION $2=PKGNAME $3=VERSION $4=UPDATE $5=CONF_FILE $6=ARCH
#
# Actions: pre (before files extracted), post (after files extracted)
# UPDATE: "yes" on upgrade, "no" on fresh install
set -eu
STORE_DIR="/var/lib/fenris"
STORE_DB="${STORE_DIR}/observations.db"
STORE_BAK="${STORE_DIR}/observations.db.bak"
RUNTIME_PYTHON="/usr/bin/python3"
VENDOR_DIR="/opt/fenris/vendor"
ACTION="$1"
UPDATE="$4"
case "${ACTION}" in
pre)
# Migration guard — abort if make-install remnants detected
MARKER="/var/lib/fenris/manifest.txt"
if [ -f "${MARKER}" ]; then
echo >&2
echo >&2 "Fenris make-install remnants detected — refusing to install."
echo >&2
echo >&2 "Migrate to the package with:"
echo >&2 " sudo make uninstall # removes make-install files, preserves store + config"
echo >&2 " sudo xbps-install fenris"
echo >&2
echo >&2 "See: https://git.bongbetic.com/xavierk/Fenris/blob/main/docs/spec/release-packaging.md#9-migration-from-make-install-systems"
echo >&2
exit 1
fi
;;
post)
# Keep observation-store access consistent across fresh installs and upgrades.
groupadd -f fenris
install -d -o root -g fenris -m 2770 "${STORE_DIR}"
chmod 02770 "${STORE_DIR}"
if [ "${UPDATE}" = "yes" ]; then
# Upgrade — snapshot, migration, runit-aware reload
if [ -f "${STORE_DB}" ]; then
cp "${STORE_DB}" "${STORE_BAK}" 2>/dev/null || true
fi
if [ -d "${VENDOR_DIR}" ] && [ -f "${STORE_DB}" ]; then
PYTHONPATH="${VENDOR_DIR}" "${RUNTIME_PYTHON}" -c "
from fenris.store import migrate_to_latest
from pathlib import Path
n = migrate_to_latest(Path('${STORE_DB}'))
print(f'Fenris migration: {n} step(s) applied') if n else None
" 2>&1 || echo "Fenris: migration skipped (store not yet initialized)"
fi
# Ensure log directory exists on upgrade
groupadd -f fenris
install -d -o root -g fenris -m 2770 /var/log/fenris-collect 2>/dev/null || true
else
# Fresh install — runit service setup
groupadd -f fenris
# Mark runit service as dormant (down) for fresh install
if [ -d /etc/sv/fenris-collect ] && [ ! -e /var/service/fenris-collect ]; then
touch /etc/sv/fenris-collect/down
fi
# Ensure log directory exists
install -d -o root -g fenris -m 2770 /var/log/fenris-collect 2>/dev/null || true
fi
;;
esac
exit 0
+30
View File
@@ -0,0 +1,30 @@
#!/bin/sh
# XBPS REMOVE script — pre-remove path.
#
# Arguments: $1=ACTION $2=PKGNAME $3=VERSION $4=UPDATE $5=CONF_FILE $6=ARCH
#
# Actions: pre (before files removed)
set -eu
ACTION="$1"
UPDATE="$4"
case "${ACTION}" in
pre)
# Only a real erase is a sanctioned disable. An upgrade must preserve
# both monitoring intent and the active runit service.
if [ "${UPDATE}" = "no" ]; then
# Close the monitoring period while the store is still available.
if [ -x /usr/libexec/fenris/fenris-monitor ]; then
/usr/libexec/fenris/fenris-monitor disable --now 2>/dev/null || true
fi
# Configuration and the observation store are deliberately not
# package-owned. Leave them in place: XBPS does not guarantee a
# post-remove callback before its cleanup action.
# runit: remove the service symlink and mark dormant
rm -f /var/service/fenris-collect
touch /etc/sv/fenris-collect/down 2>/dev/null || true
fi
;;
esac
exit 0
@@ -0,0 +1,21 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE policyconfig PUBLIC
"-//freedesktop//DTD PolicyKit Policy Configuration 1.0//EN"
"http://www.freedesktop.org/standards/PolicyKit/1/policyconfig.dtd">
<policyconfig>
<vendor>bongbetic</vendor>
<vendor_url>https://bongbetic.com</vendor_url>
<action id="com.bongbetic.fenris.monitor">
<description>Fenris Monitor Helper</description>
<message>Authentication is required to manage Fenris monitoring.</message>
<defaults>
<allow_any>no</allow_any>
<allow_inactive>no</allow_inactive>
<allow_active>auth_admin</allow_active>
</defaults>
<annotate key="org.freedesktop.policykit.imply">org.freedesktop.systemd1.manage-units</annotate>
</action>
</policyconfig>
+24
View File
@@ -0,0 +1,24 @@
[project]
name = "fenris"
version = "0.6.0"
description = "NVMe wear monitor with persistent TUI"
requires-python = ">=3.10"
license = {file = "LICENSE"}
dependencies = [
"textual>=0.40.0",
]
[project.optional-dependencies]
dev = [
"pytest>=7.0.0",
"pytest-cov>=4.0.0",
"pytest-asyncio>=0.20.0",
]
[tool.pytest.ini_options]
testpaths = ["tests"]
python_files = ["test_*.py"]
python_functions = ["test_*"]
markers = [
"slow: marks tests as slow",
]
+12
View File
@@ -0,0 +1,12 @@
# Fenris dependency lockfile
# Exact pins for reproducible installs (IN-8)
# Refresh with: make update-deps
textual==8.2.8
rich==15.0.0
markdown-it-py==4.2.0
mdit-py-plugins==0.6.1
mdurl==0.1.2
platformdirs==4.11.7
Pygments==2.21.0
linkify-it-py==2.2.0
typing-extensions==4.16.0
+129
View File
@@ -0,0 +1,129 @@
#!/usr/bin/env python3
"""Extract one validated Keep a Changelog version section."""
from __future__ import annotations
import argparse
from datetime import date
from pathlib import Path
import re
import sys
class ChangelogError(ValueError):
"""A release cannot safely use the supplied changelog."""
_SEMVER = r"(?:0|[1-9]\d*)\.(?:0|[1-9]\d*)\.(?:0|[1-9]\d*)"
_VERSION_HEADING = re.compile(
rf"^## \[(?P<version>{_SEMVER})\] - (?P<date>.+)$", re.MULTILINE
)
def extract_version_section(changelog: str, version: str) -> str:
"""Return *version*'s changelog section without altering its bytes.
The section ends immediately before the next level-two heading. A release
cannot use an absent, empty, or malformed version section.
"""
if not re.fullmatch(_SEMVER, version):
raise ChangelogError(f"requested version is not bare semver: {version!r}")
heading = next(
(match for match in _VERSION_HEADING.finditer(changelog)
if match.group("version") == version),
None,
)
if heading is None:
if re.search(rf"^## \[{re.escape(version)}\].*$", changelog, re.MULTILINE):
raise ChangelogError(f"version {version} has a malformed heading or date")
raise ChangelogError(f"version {version} is missing from the changelog")
heading_date = heading.group("date")
if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", heading_date):
raise ChangelogError(f"version {version} has a malformed release date")
try:
date.fromisoformat(heading_date)
except ValueError as error:
raise ChangelogError(f"version {version} has a malformed release date") from error
next_heading = re.search(r"^## ", changelog[heading.end():], re.MULTILINE)
section_end = heading.end() + next_heading.start() if next_heading else len(changelog)
section = changelog[heading.start():section_end]
if not re.search(r"^- \S", section[heading.end() - heading.start():], re.MULTILINE):
raise ChangelogError(f"version {version} has an empty changelog section")
return section
def extract_changelog(path: Path, version: str) -> str:
"""Read and extract a requested version from a changelog file."""
try:
return extract_version_section(path.read_text(encoding="utf-8"), version)
except OSError as error:
raise ChangelogError(f"cannot read changelog {path}: {error.strerror}") from error
def assemble_release_body(section: str, footer: str) -> str:
"""Append standing guidance while preserving the extracted section verbatim."""
separator = "\n" if section.endswith("\n") else "\n\n"
return f"{section}{separator}{footer}"
def format_availability_section(
available: list[str], withheld: list[str]
) -> str:
"""Generate a package formats section for release notes."""
if not available and not withheld:
return ""
lines = ["\n## Package formats\n"]
if available:
lines.append(f"Available: {', '.join(available)}")
if withheld:
lines.append(f"Withheld: {', '.join(withheld)}")
return "\n".join(lines)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("changelog", type=Path)
parser.add_argument("version")
parser.add_argument(
"--footer",
type=Path,
help="append this standing release guidance after the extracted section",
)
parser.add_argument(
"--available",
action="append",
default=[],
help="format available for this release (can be repeated)",
)
parser.add_argument(
"--withheld",
action="append",
default=[],
help="format withheld from this release (can be repeated)",
)
args = parser.parse_args(argv)
try:
section = extract_changelog(args.changelog, args.version)
if args.footer:
try:
footer = args.footer.read_text(encoding="utf-8")
except OSError as error:
raise ChangelogError(
f"cannot read release footer {args.footer}: {error.strerror}"
) from error
section = assemble_release_body(section, footer)
if args.available or args.withheld:
formats = format_availability_section(args.available, args.withheld)
if formats:
section = f"{section}\n{formats}"
sys.stdout.write(section)
except ChangelogError as error:
print(f"::error::{error}", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())
Executable
+206
View File
@@ -0,0 +1,206 @@
#!/usr/bin/env python3
"""fenris: unprivileged entry point for the Fenris TUI and CLI.
With no arguments, opens the TUI.
Subcommands route through fenris-monitor for privileged operations.
Spec: §1.2, §8.4
"""
import argparse
import sys
from pathlib import Path
def add_runtime_packages() -> None:
"""Make the package-owned, pure-Python dependencies importable.
RPM and deb installations deliberately use the target system's Python.
Their dependencies are vendored without a copied interpreter so a distro
Python minor-version update cannot leave Fenris linked to a removed ABI.
The legacy development install keeps its venv fallback.
"""
runtime_dir = Path("/opt/fenris")
vendor_dir = runtime_dir / "vendor"
if vendor_dir.is_dir():
sys.path.insert(0, str(vendor_dir))
return
site_packages = next((runtime_dir / "lib").glob("python*/site-packages"), None)
if site_packages:
sys.path.insert(0, str(site_packages))
add_runtime_packages()
def run_monitor(*args: str) -> None:
"""Render the shared privileged-action outcome for the CLI."""
from fenris.control import MonitorError, run_monitor as invoke_monitor
try:
invoke_monitor(*args)
except MonitorError as exc:
print(str(exc), file=sys.stderr)
sys.exit(exc.exit_code)
sys.exit(0)
def cmd_tui(args: argparse.Namespace) -> None:
"""Open the TUI."""
from fenris.tui import run_tui
run_tui()
def cmd_status(args: argparse.Namespace) -> None:
"""Show status."""
from fenris.status import render_status
print(render_status())
def cmd_sample(args: argparse.Namespace) -> None:
"""Trigger on-demand collection."""
run_monitor("collect")
def cmd_monitor_pause(args: argparse.Namespace) -> None:
"""Pause monitoring."""
# Pause asks confirmation (§7.4)
if not args.yes:
response = input("Pause monitoring? [y/N] ")
if response.lower() not in ("y", "yes"):
print("Aborted.")
return
run_monitor("disable", "--now")
def cmd_monitor_resume(args: argparse.Namespace) -> None:
"""Resume monitoring."""
# Resume does not ask confirmation (§7.4)
run_monitor("enable", "--now")
def cmd_baseline_set(args: argparse.Namespace) -> None:
"""Set baseline."""
run_monitor("baseline", "set", args.baseline_json)
def cmd_baseline_clear(args: argparse.Namespace) -> None:
"""Clear baseline."""
run_monitor("baseline", "clear")
def cmd_import(args: argparse.Namespace) -> None:
"""Import legacy history."""
# This is a one-off migration, not a privileged operation
print("Legacy import: use fenris-import directly")
def cmd_migrate(args: argparse.Namespace) -> None:
"""Apply forward-only schema migrations (IN-5, IN-6).
Called by 'sudo make upgrade'. Raises on newer-schema store.
"""
from fenris.store import migrate_to_latest
from pathlib import Path
store_path = Path("/var/lib/fenris/observations.db")
if not store_path.exists():
print("No observation store found — nothing to migrate.")
return
steps = migrate_to_latest(store_path)
if steps:
print(f"Migration complete: {steps} step(s) applied.")
else:
print("Schema already current.")
def main() -> None:
parser = argparse.ArgumentParser(
prog="fenris",
description="Fenris NVMe endurance monitor",
)
parser.add_argument(
"--version", action="version", version="%(prog)s 0.3.6"
)
subparsers = parser.add_subparsers(dest="command")
# Default: TUI (no subcommand)
subparsers.add_parser("tui", help="Open the TUI (default)")
# Status
subparsers.add_parser("status", help="Show status")
# Sample (on-demand collection)
subparsers.add_parser("sample", help="Trigger on-demand collection")
# Monitor subcommand
monitor_parser = subparsers.add_parser("monitor", help="Monitor control")
monitor_sub = monitor_parser.add_subparsers(dest="monitor_action")
# monitor pause
pause_parser = monitor_sub.add_parser("pause", help="Pause monitoring")
pause_parser.add_argument(
"-y", "--yes", action="store_true", help="Skip confirmation"
)
pause_parser.set_defaults(func=cmd_monitor_pause)
# monitor resume
resume_parser = monitor_sub.add_parser("resume", help="Resume monitoring")
resume_parser.set_defaults(func=cmd_monitor_resume)
# Baseline subcommand
baseline_parser = subparsers.add_parser("baseline", help="Baseline operations")
baseline_sub = baseline_parser.add_subparsers(dest="baseline_action")
baseline_set = baseline_sub.add_parser("set", help="Set baseline")
baseline_set.add_argument("baseline_json", help="Baseline JSON data")
baseline_set.set_defaults(func=cmd_baseline_set)
baseline_clear = baseline_sub.add_parser("clear", help="Clear baseline")
baseline_clear.set_defaults(func=cmd_baseline_clear)
# Import
import_parser = subparsers.add_parser("import", help="Import legacy history")
import_parser.add_argument("path", help="Path to history.jsonl")
import_parser.set_defaults(func=cmd_import)
# Migrate (IN-5, IN-6) — called by upgrade, not for human use
migrate_parser = subparsers.add_parser("migrate", help=argparse.SUPPRESS)
migrate_parser.set_defaults(func=cmd_migrate)
# Rejected commands
for cmd in ["start", "stop", "run"]:
reject_parser = subparsers.add_parser(cmd, help=argparse.SUPPRESS)
reject_parser.set_defaults(func=lambda a: print(
f"'{cmd}' is not a valid command. "
f"Use 'fenris monitor resume' instead.",
file=sys.stderr,
))
args = parser.parse_args()
if args.command is None or args.command == "tui":
cmd_tui(args)
elif args.command == "status":
cmd_status(args)
elif args.command == "sample":
cmd_sample(args)
elif args.command == "monitor":
if args.monitor_action is None:
monitor_parser.error("a subcommand is required")
args.func(args)
elif args.command == "baseline":
if args.baseline_action is None:
baseline_parser.error("a subcommand is required")
args.func(args)
elif args.command == "import":
cmd_import(args)
elif args.command == "migrate":
cmd_migrate(args)
if __name__ == "__main__":
main()
+206
View File
@@ -0,0 +1,206 @@
#!/usr/bin/env bash
set -euo pipefail
# Fenris one-command release flow (issue #52).
# Builds both packages, signs, uploads to registry, creates release entry,
# and attaches artifacts — or in dry-run mode, prints every command.
#
# Usage:
# scripts/release.sh --dry-run # Print commands without executing
# scripts/release.sh --publish # Execute the full release flow
#
# Environment:
# GITEA_TOKEN - API token for Gitea registry and release API
# PACKAGING_KEY - GPG key UID (default: packaging@bongbetic.com)
#
# Spec: release-packaging.md §5
# ── Defaults ─────────────────────────────────────────────────────────────
DRY_RUN=false
PUBLISH=false
GITEA_URL="https://git.bongbetic.com"
GITEA_OWNER="xavierk"
GITEA_REPO="Fenris"
PACKAGING_KEY="${PACKAGING_KEY:-packaging@bongbetic.com}"
CODENAMES=(bookworm jammy noble)
RPM_GROUP="fenris"
# ── Parse arguments ──────────────────────────────────────────────────────
for arg in "$@"; do
case "$arg" in
--dry-run) DRY_RUN=true ;;
--publish) PUBLISH=true ;;
--help|-h)
echo "Usage: $0 [--dry-run | --publish]"
echo ""
echo "Modes:"
echo " --dry-run Print commands without executing (default)"
echo " --publish Execute the full release flow"
echo ""
echo "Environment:"
echo " GITEA_TOKEN API token for Gitea registry and release API"
echo " PACKAGING_KEY GPG key UID (default: packaging@bongbetic.com)"
exit 0
;;
*)
echo "Unknown argument: $arg" >&2
echo "Usage: $0 [--dry-run | --publish]" >&2
exit 1
;;
esac
done
if ! $DRY_RUN && ! $PUBLISH; then
DRY_RUN=true
fi
# ── Helpers ──────────────────────────────────────────────────────────────
_version() {
sed -n 's/^version = "\(.*\)"/\1/p' pyproject.toml
}
_deb_name() {
local ver="$1"
echo "fenris_${ver}_amd64.deb"
}
_rpm_name() {
local ver="$1" rel="$2"
echo "fenris-${ver}-${rel}.x86_64.rpm"
}
_run() {
if $DRY_RUN; then
echo " $*"
else
eval "$@"
fi
}
# ── Main ─────────────────────────────────────────────────────────────────
VERSION=$(_version)
REVISION=1
DEB=$(_deb_name "$VERSION")
RPM=$(_rpm_name "$VERSION" "$REVISION")
echo "=== Fenris Release v${VERSION} ==="
echo ""
if $DRY_RUN; then
echo "[dry-run] Commands below will be executed in --publish mode."
echo ""
fi
# ── Step 1: Build both formats ──────────────────────────────────────────
echo "--- Build packages ---"
_run "make package"
echo ""
# ── Step 2: Sign RPM payload ────────────────────────────────────────────
echo "--- Sign RPM payload ---"
_run "rpmsign --addsign --define '_gpg_name ${PACKAGING_KEY}' dist/${RPM}"
echo ""
# ── Step 3: Generate and clearsign SHA256SUMS ────────────────────────────
echo "--- Generate SHA256SUMS ---"
_run "cd dist && sha256sum ${DEB} ${RPM} > SHA256SUMS"
echo ""
echo "--- Clearsign SHA256SUMS ---"
_run "gpg --batch --yes --clearsign --local-user ${PACKAGING_KEY} dist/SHA256SUMS"
echo ""
# ── Step 4: Upload to Gitea package registry ─────────────────────────────
echo "--- Upload packages to registry ---"
for codename in "${CODENAMES[@]}"; do
_run "curl --fail -X PUT -u ${GITEA_OWNER}:\$GITEA_TOKEN -T dist/${DEB} '${GITEA_URL}/api/packages/${GITEA_OWNER}/debian/pool/${codename}/main/upload'"
done
_run "curl --fail -X PUT -u ${GITEA_OWNER}:\$GITEA_TOKEN -T dist/${RPM} '${GITEA_URL}/api/packages/${GITEA_OWNER}/rpm/${RPM_GROUP}/upload'"
echo ""
# ── Step 5: Create Gitea release with notes ─────────────────────────────
echo "--- Create Gitea release ---"
_release_notes="Release v${VERSION}
## Packages
Install via apt (Debian/Ubuntu):
\`\`\`bash
curl --fail -fsSL https://git.bongbetic.com/${GITEA_OWNER}/${GITEA_REPO}/raw/branch/main/packaging/keys/fenris-packaging.asc | sudo gpg --dearmor -o /etc/apt/keyrings/fenris.asc
echo \"deb [signed-by=/etc/apt/keyrings/fenris.asc] https://git.bongbetic.com/api/packages/${GITEA_OWNER}/debian bookworm main\" | sudo tee /etc/apt/sources.list.d/fenris.list
sudo apt update && sudo apt install fenris
\`\`\`
Install via dnf (Fedora):
\`\`\`bash
sudo dnf config-manager --add-repo https://git.bongbetic.com/${GITEA_OWNER}/${GITEA_REPO}/raw/branch/main/packaging/fenris.repo
sudo dnf install fenris
\`\`\`
## Verification
\`\`\`bash
rpm -Kv fenris-${VERSION}-1.x86_64.rpm
gpg --verify SHA256SUMS.asc SHA256SUMS
\`\`\`
## Artifacts
- \`dist/${DEB}\` (Debian/Ubuntu)
- \`dist/${RPM}\` (Fedora)
- \`dist/SHA256SUMS.asc\` (clearsigned checksums)
See [docs/install/signing-key-ceremony.md](docs/install/signing-key-ceremony.md) for key ceremony details.
See [docs/install/migrate-from-makeinstall.md](docs/install/migrate-from-makeinstall.md) for migration from make install."
if $DRY_RUN; then
_run "curl --fail -X POST -u ${GITEA_OWNER}:\$GITEA_TOKEN -H 'Content-Type: application/json' -d '{\"tag_name\":\"v${VERSION}\",\"name\":\"v${VERSION}\",\"body\":\"...\"}' '${GITEA_URL}/api/v1/repos/${GITEA_OWNER}/${GITEA_REPO}/releases'"
else
# Create release via Gitea API (creates the tag atomically — no bare tag)
RELEASE_RESPONSE=$(curl --fail -s -X POST \
-u "${GITEA_OWNER}:${GITEA_TOKEN}" \
-H "Content-Type: application/json" \
-d "$(jq -n \
--arg tag "v${VERSION}" \
--arg name "v${VERSION}" \
--arg body "$_release_notes" \
'{tag_name: $tag, name: $name, body: $body}')" \
"${GITEA_URL}/api/v1/repos/${GITEA_OWNER}/${GITEA_REPO}/releases")
RELEASE_ID=$(echo "$RELEASE_RESPONSE" | jq -r '.id')
echo " Release created: ${GITEA_URL}/${GITEA_OWNER}/${GITEA_REPO}/releases/tag/v${VERSION}"
fi
echo ""
# ── Step 6: Attach artifacts to release ──────────────────────────────────
echo "--- Attach artifacts to release ---"
for artifact in "dist/${DEB}" "dist/${RPM}" "dist/SHA256SUMS.asc"; do
_run "curl --fail -X POST -u ${GITEA_OWNER}:\$GITEA_TOKEN -F 'attachment=@${artifact}' '${GITEA_URL}/api/v1/repos/${GITEA_OWNER}/${GITEA_REPO}/releases/${RELEASE_ID:-0}/assets'"
done
echo ""
# ── Done ─────────────────────────────────────────────────────────────────
echo "=== Release v${VERSION} complete ==="
echo ""
echo "Summary:"
echo " Packages: ${DEB}, ${RPM}"
echo " Checksums: dist/SHA256SUMS.asc"
echo " Registry: deb → bookworm, jammy, noble; rpm → ${RPM_GROUP}"
echo " Release: ${GITEA_URL}/${GITEA_OWNER}/${GITEA_REPO}/releases/tag/v${VERSION}"
echo ""
echo "Key ceremony: delete the private key after release."
echo " See docs/install/signing-key-ceremony.md"
+70
View File
@@ -0,0 +1,70 @@
#!/usr/bin/env python3
"""Describe the Gitea request that creates or resynchronizes a release."""
from __future__ import annotations
import argparse
import json
from pathlib import Path
import re
import sys
from typing import Any
_SEMVER = r"(?:0|[1-9]\d*)\.(?:0|[1-9]\d*)\.(?:0|[1-9]\d*)"
class ReleaseRequestError(ValueError):
"""A release request could not be prepared safely."""
def build_release_request(
version: str, body: str, existing_release: dict[str, Any] | None
) -> dict[str, Any]:
"""Return the observable POST or PATCH request for a Gitea release."""
if not re.fullmatch(_SEMVER, version):
raise ReleaseRequestError(f"version is not bare semver: {version!r}")
if existing_release is None:
return {
"method": "POST",
"path": "/releases",
"payload": {"tag_name": f"v{version}", "name": f"v{version}", "body": body},
}
release_id = existing_release.get("id")
if not isinstance(release_id, int):
raise ReleaseRequestError("existing release does not contain an integer id")
return {
"method": "PATCH",
"path": f"/releases/{release_id}",
"payload": {"body": body},
}
def _read_json(path: Path) -> dict[str, Any]:
try:
value = json.loads(path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError) as error:
raise ReleaseRequestError(f"cannot read existing release {path}: {error}") from error
if not isinstance(value, dict):
raise ReleaseRequestError("existing release must be a JSON object")
return value
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--version", required=True)
parser.add_argument("--body-file", type=Path, required=True)
parser.add_argument("--existing-release", type=Path)
args = parser.parse_args(argv)
try:
body = args.body_file.read_text(encoding="utf-8")
existing = _read_json(args.existing_release) if args.existing_release else None
print(json.dumps(build_release_request(args.version, body, existing)))
except (OSError, ReleaseRequestError) as error:
print(f"::error::{error}", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())
+208
View File
@@ -0,0 +1,208 @@
#!/usr/bin/env bash
set -euo pipefail
# Fenris XBPS publication script (issue #83).
# Builds, signs, and publishes XBPS packages to the Fenris-xbps repository.
#
# Usage:
# scripts/xbps-publish.sh --dry-run # Print commands without executing
# scripts/xbps-publish.sh --publish # Execute the full publication flow
#
# Environment:
# XBPS_SIGNING_KEY - Path to SSH RSA private key for XBPS signing
# (default: ~/.ssh/id_xbps)
# SIGNED_BY - Signature identity string
# (default: "Fenris Packaging <packaging@bongbetic.com>")
#
# Spec: native-void-support.md, ADR 0008
# ── Defaults ─────────────────────────────────────────────────────────────
DRY_RUN=false
PUBLISH=false
GITEA_URL="https://git.bongbetic.com"
GITEA_OWNER="xavierk"
GITEA_REPO="Fenris-xbps"
GITEA_BRANCH="stable"
ARCH="x86_64"
XBPS_SIGNING_KEY="${XBPS_SIGNING_KEY:-$HOME/.ssh/id_xbps}"
SIGNED_BY="${SIGNED_BY:-Fenris Packaging <packaging@bongbetic.com>}"
# ── Parse arguments ──────────────────────────────────────────────────────
for arg in "$@"; do
case "$arg" in
--dry-run) DRY_RUN=true ;;
--publish) PUBLISH=true ;;
--help|-h)
echo "Usage: $0 [--dry-run | --publish]"
echo ""
echo "Modes:"
echo " --dry-run Print commands without executing (default)"
echo " --publish Execute the full publication flow"
echo ""
echo "Environment:"
echo " XBPS_SIGNING_KEY Path to SSH RSA private key (default: ~/.ssh/id_xbps)"
echo " SIGNED_BY Signature identity (default: Fenris Packaging <packaging@bongbetic.com>)"
exit 0
;;
*)
echo "Unknown argument: $arg" >&2
echo "Usage: $0 [--dry-run | --publish]" >&2
exit 1
;;
esac
done
if ! $DRY_RUN && ! $PUBLISH; then
DRY_RUN=true
fi
# ── Helpers ──────────────────────────────────────────────────────────────
_version() {
sed -n 's/^version = "\(.*\)"/\1/p' pyproject.toml
}
_run() {
if $DRY_RUN; then
echo " $*"
else
eval "$@"
fi
}
# ── Pre-flight checks ───────────────────────────────────────────────────
if [[ ! -f "$XBPS_SIGNING_KEY" ]]; then
echo "ERROR: Signing key not found at $XBPS_SIGNING_KEY" >&2
echo "Generate one with: ssh-keygen -t rsa -b 3072 -f $XBPS_SIGNING_KEY" >&2
exit 1
fi
for tool in xbps-create xbps-rindex git; do
if ! command -v "$tool" &>/dev/null; then
echo "ERROR: Required tool not found: $tool" >&2
exit 1
fi
done
# ── Main ─────────────────────────────────────────────────────────────────
VERSION=$(_version)
REVISION="${XBPS_REVISION:-1}"
PKGVER="fenris-${VERSION}_${REVISION}"
XBPS_FILE="${PKGVER}.${ARCH}.xbps"
SOURCE_DIR=$(pwd)
XBPS_PATH="${SOURCE_DIR}/${XBPS_FILE}"
GIT_USER_NAME=$(git config user.name || true)
GIT_USER_EMAIL=$(git config user.email || true)
if [[ -z "${GIT_USER_NAME}" || -z "${GIT_USER_EMAIL}" ]]; then
echo "ERROR: Configure git user.name and user.email in the source repository before publishing" >&2
exit 1
fi
echo "=== Fenris XBPS Publication v${VERSION} ==="
echo ""
if $DRY_RUN; then
echo "[dry-run] Commands below will be executed in --publish mode."
echo ""
fi
# ── Step 1: Build XBPS package ───────────────────────────────────────────
echo "--- Build XBPS package ---"
_run "make XBPS_REVISION=${REVISION} package-xbps"
echo ""
# ── Step 2: Sign package ─────────────────────────────────────────────────
echo "--- Sign XBPS package ---"
_run "rm -f ${XBPS_PATH}.sig2"
_run "xbps-rindex --sign-pkg --privkey ${XBPS_SIGNING_KEY} ${XBPS_PATH}"
echo ""
# ── Step 3: Clone/update distribution repository ─────────────────────────
echo "--- Prepare distribution repository ---"
WORK_DIR=$(mktemp -d)
_run "git clone ${GITEA_URL}/${GITEA_OWNER}/${GITEA_REPO}.git ${WORK_DIR}"
_run "git -C ${WORK_DIR} config user.name '${GIT_USER_NAME}'"
_run "git -C ${WORK_DIR} config user.email '${GIT_USER_EMAIL}'"
_run "git -C ${WORK_DIR} checkout ${GITEA_BRANCH}"
_run "mkdir -p ${WORK_DIR}/${ARCH}"
echo ""
# ── Step 4: Copy artifacts and update index ──────────────────────────────
echo "--- Update repository index ---"
_run "cp ${XBPS_PATH} ${XBPS_PATH}.sig2 ${WORK_DIR}/${ARCH}/"
_run "(cd ${WORK_DIR} && xbps-rindex --add ${ARCH}/${XBPS_FILE})"
_run "(cd ${WORK_DIR} && xbps-rindex --sign --privkey ${XBPS_SIGNING_KEY} --signedby '${SIGNED_BY}' ${ARCH})"
echo ""
# ── Step 5: Commit and push ──────────────────────────────────────────────
echo "--- Commit and push ---"
_run "git -C ${WORK_DIR} add -A"
_run "git -C ${WORK_DIR} commit -m 'Release fenris ${VERSION}'"
_run "git -C ${WORK_DIR} push origin ${GITEA_BRANCH}"
echo ""
# ── Step 6: Verify publication ───────────────────────────────────────────
echo "--- Verify publication ---"
RAW_BASE="${GITEA_URL}/${GITEA_OWNER}/${GITEA_REPO}/raw/branch/${GITEA_BRANCH}/${ARCH}"
if $PUBLISH; then
# Compare every served byte with the artifact that was indexed and pushed.
# A successful HEAD request alone can still hide a stale or incomplete
# publication behind the raw endpoint's cache.
# Repository signatures are embedded in x86_64-repodata by xbps-rindex;
# only package signatures are separate .sig2 files.
for artifact in "${XBPS_FILE}" "${XBPS_FILE}.sig2" "x86_64-repodata"; do
case "${artifact}" in
"${XBPS_FILE}") local_path="${XBPS_PATH}" ;;
"${XBPS_FILE}.sig2") local_path="${XBPS_PATH}.sig2" ;;
*) local_path="${WORK_DIR}/${ARCH}/${artifact}" ;;
esac
downloaded="${WORK_DIR}/.${artifact}.download"
curl --fail --silent --show-error --location \
--output "${downloaded}" "${RAW_BASE}/${artifact}"
cmp -- "${local_path}" "${downloaded}"
rm -f "${downloaded}"
done
else
_run "curl --fail --silent --show-error --location --output ${WORK_DIR}/.${XBPS_FILE}.download ${RAW_BASE}/${XBPS_FILE}"
_run "cmp -- ${XBPS_PATH} ${WORK_DIR}/.${XBPS_FILE}.download"
_run "curl --fail --silent --show-error --location --output ${WORK_DIR}/.${XBPS_FILE}.sig2.download ${RAW_BASE}/${XBPS_FILE}.sig2"
_run "cmp -- ${XBPS_PATH}.sig2 ${WORK_DIR}/.${XBPS_FILE}.sig2.download"
_run "curl --fail --silent --show-error --location --output ${WORK_DIR}/.x86_64-repodata.download ${RAW_BASE}/x86_64-repodata"
_run "cmp -- ${WORK_DIR}/${ARCH}/x86_64-repodata ${WORK_DIR}/.x86_64-repodata.download"
_run "rm -f ${WORK_DIR}/.${XBPS_FILE}.download ${WORK_DIR}/.${XBPS_FILE}.sig2.download ${WORK_DIR}/.x86_64-repodata.download"
fi
echo ""
# ── Cleanup ──────────────────────────────────────────────────────────────
if $PUBLISH; then
rm -rf "${WORK_DIR}"
fi
# ── Done ─────────────────────────────────────────────────────────────────
echo "=== XBPS Publication v${VERSION} complete ==="
echo ""
echo "Summary:"
echo " Package: ${XBPS_FILE}"
echo " Repository: ${GITEA_URL}/${GITEA_OWNER}/${GITEA_REPO}"
echo " Branch: ${GITEA_BRANCH}"
echo " Repository URL: https://git.bongbetic.com/${GITEA_OWNER}/${GITEA_REPO}/raw/branch/${GITEA_BRANCH}/${ARCH}"
echo ""
echo "Client installation:"
echo " echo 'repository=https://git.bongbetic.com/${GITEA_OWNER}/${GITEA_REPO}/raw/branch/${GITEA_BRANCH}/${ARCH}' | sudo tee /etc/xbps.d/fenris.conf"
echo " sudo xbps-install -M -S fenris"
echo ""
echo "Key ceremony: delete the private key after publication."
echo " See docs/install/signing-key-ceremony.md"
+2
View File
@@ -0,0 +1,2 @@
"""Fenris: NVMe wear monitor with persistent TUI."""
__version__ = "0.5.0"
+104
View File
@@ -0,0 +1,104 @@
"""A terminal volume plot shared by Fenris's activity views.
Braille provides two by four dots per terminal cell. Only adjacent, complete
measurements are connected; gaps and partial evidence never imply continuity.
The returned columns also place mouse inspection on the plotted time axis.
"""
from collections.abc import Mapping, Sequence
from dataclasses import dataclass
from rich.text import Text
@dataclass(frozen=True)
class VolumePoint:
time: float
volume: int | None
label: str
state: str = "measured"
def volume_plot(
points: Sequence[VolumePoint], width: int, height: int,
selected: int, colors: Mapping[str, str],
) -> tuple[Text, list[int], str]:
"""Render bounded axes and a dotted trace, without resampling evidence."""
width, height = max(12, width), max(4, height)
axis_width = 8
columns, rows = width - axis_width, height - 2
pixel_width, pixel_height = columns * 2, rows * 4
maximum = max((p.volume or 0 for p in points), default=0)
scale, unit = (1e12, "TB") if maximum >= 1e12 else (
(1e9, "GB") if maximum >= 1e9 else
(1e6, "MB") if maximum >= 1e6 else
(1e3, "KB") if maximum >= 1e3 else (1, "B")
)
ceiling = maximum or scale
start = points[0].time if points else 0
span = (points[-1].time - start) if len(points) > 1 else 0
xs = [round((p.time - start) / span * (pixel_width - 1)) if span else 0
for p in points]
cells = [[0] * columns for _ in range(rows)]
# Unicode braille dot numbering, indexed by y within cell and then x.
bits = ((1, 8), (2, 16), (4, 32), (64, 128))
def dot(x: int, y: int) -> None:
cells[y // 4][x // 2] |= bits[y % 4][x % 2]
previous = None
markers = {}
for index, (point, x) in enumerate(zip(points, xs)):
if point.volume is None:
previous = None
if point.state != "future":
markers[x // 2] = "?"
continue
y = round((1 - point.volume / ceiling) * (pixel_height - 1))
y = min(pixel_height - 1, max(0, y))
if previous is not None and point.state == "measured":
px, py = previous
steps = max(abs(x - px), abs(y - py), 1)
for step in range(steps + 1):
dot(round(px + (x - px) * step / steps),
round(py + (y - py) * step / steps))
dot(x, y)
previous = (x, y) if point.state == "measured" else None
if point.state != "measured":
markers[x // 2] = "~" if point.state == "partial" else "u"
elif point.volume == 0:
markers.setdefault(x // 2, "·")
selected_column = xs[selected] // 2 if 0 <= selected < len(xs) else -1
result = Text(no_wrap=True, overflow="crop")
ticks = {0, rows // 2, rows - 1}
for row, values in enumerate(cells):
value = ceiling * (rows - 1 - row) / max(1, rows - 1) / scale
label = f"{value:6.2f}"[-6:] if row in ticks else " "
result.append(label + " │", style=colors["muted"])
for col, value in enumerate(values):
char = chr(0x2800 + value) if value else " "
style = colors["allocated"]
if col == selected_column:
style = "bold " + colors["selection"]
if not value:
char = "┊"
style = colors["muted"]
result.append(char, style=style)
result.append("\n")
result.append(" └", style=colors["muted"])
for col in range(columns):
result.append("▼" if col == selected_column else markers.get(col, "─"),
style=colors["selection"] if col == selected_column else colors["muted"])
result.append("\n" + " " * axis_width)
labels = [" "] * columns
end_of_label = -1
for index in sorted({0, len(points) // 2, len(points) - 1}):
if index < 0 or not points:
continue
label = points[index].label[:columns]
column = max(0, min(columns - len(label), xs[index] // 2 - len(label) // 2))
if column > end_of_label:
labels[column:column + len(label)] = label
end_of_label = column + len(label)
result.append("".join(labels), style=colors["muted"])
return result, [axis_width + x // 2 for x in xs], unit
+132
View File
@@ -0,0 +1,132 @@
"""User choices for live and historical activity views."""
from collections.abc import Sequence
from dataclasses import dataclass
HISTORY_RANGE_DEFAULT = 14
@dataclass(frozen=True)
class IntervalIdentity:
"""Stable identity for one measured interval."""
start_ts: str
end_ts: str
@dataclass(frozen=True)
class HistoryDayIdentity:
"""Stable identity for one recorded local-day summary."""
local_date: str
timezone: str
utc_start: str
utc_end: str
class ActivitySelection:
"""Keep activity navigation state separate from its Textual rendering."""
def __init__(self) -> None:
self.view = "live"
self.browse_date: str | None = None
self.selected_history_day: HistoryDayIdentity | None = None
self.selected_history_hour: str | None = None
self.history_range_days = HISTORY_RANGE_DEFAULT
self.measure = "written"
self.following_live = True
self.selected_live_interval: IntervalIdentity | None = None
self.live_interval_expired = False
def set_view(self, view: str) -> None:
"""Select live, day, or history view."""
if view not in ("live", "day", "history"):
raise ValueError(f"unknown activity view: {view}")
self.view = view
if view == "live":
self.follow_live()
def select_history_date(
self,
date: str | None,
identity: HistoryDayIdentity | None = None,
) -> None:
"""Keep requested date and evidence identity across refresh."""
if date != self.browse_date or (
identity is not None and identity != self.selected_history_day
):
self.selected_history_hour = None
if date != self.browse_date:
self.selected_history_day = None
self.browse_date = date
if identity is not None:
self.selected_history_day = identity
elif date is None:
self.selected_history_day = None
def select_history_hour(self, hour: str | None) -> None:
"""Keep the chosen historical hour across refresh."""
self.selected_history_hour = hour
def toggle_measure(self) -> str:
"""Switch read/write presentation without changing point selection."""
self.measure = "read" if self.measure == "written" else "written"
return self.measure
def follow_live(self) -> None:
"""Resume following the newest interval and clear stale-pin notices."""
self.view = "live"
self.browse_date = None
self.selected_history_day = None
self.selected_history_hour = None
self.following_live = True
self.selected_live_interval = None
self.live_interval_expired = False
def update_live(self, intervals: Sequence[IntervalIdentity]) -> None:
"""Reconcile the selected interval with the rolling live window."""
if self.following_live:
self.selected_live_interval = intervals[-1] if intervals else None
return
if self.selected_live_interval in intervals:
return
self.following_live = True
self.selected_live_interval = intervals[-1] if intervals else None
self.live_interval_expired = True
def move_live(
self, intervals: Sequence[IntervalIdentity], offset: int,
) -> IntervalIdentity | None:
"""Inspect an adjacent interval by stable identity."""
if not intervals:
return None
if self.selected_live_interval in intervals:
current = intervals.index(self.selected_live_interval)
else:
current = len(intervals) - 1
selected = intervals[max(0, min(len(intervals) - 1, current + offset))]
self._pin_live_interval(selected)
return selected
def inspect_live(
self,
intervals: Sequence[IntervalIdentity],
interval: IntervalIdentity,
) -> None:
"""Pin the interval chosen by mouse inspection."""
if interval in intervals:
self._pin_live_interval(interval)
def selected_live_index(self, intervals: Sequence[IntervalIdentity]) -> int:
"""Return the selected point's current render index, or -1."""
if self.selected_live_interval is None:
return -1
try:
return intervals.index(self.selected_live_interval)
except ValueError:
return -1
def _pin_live_interval(self, interval: IntervalIdentity) -> None:
self.selected_live_interval = interval
self.following_live = False
self.live_interval_expired = False
+141
View File
@@ -0,0 +1,141 @@
#!/usr/bin/env python3
"""fenris-collect: device interrogation and store writes.
This is the root oneshot unit's ExecStart. It reads the device selector
from /etc/fenris/fenris.conf, interrogates the drive via smartctl and sysfs,
and writes the sample to the observation store.
Spec: §8.4, §8.5
When run as a script, uses the fenris package from the installed wheel.
"""
import json
import subprocess
import sys
from datetime import datetime, timezone
from pathlib import Path
# Package dependencies are vendored independently of the host Python minor
# version. Keep the venv fallback for the legacy development install.
VENV_DIR = Path("/opt/fenris")
if VENV_DIR.exists():
vendor_dir = VENV_DIR / "vendor"
if vendor_dir.is_dir():
sys.path.insert(0, str(vendor_dir))
else:
site_packages = next((VENV_DIR / "lib").glob("python*/site-packages"), None)
if site_packages:
sys.path.insert(0, str(site_packages))
from fenris.collector import run_collection
CONFIG_PATH = Path("/etc/fenris/fenris.conf")
def load_config() -> dict:
"""Load configuration from /etc/fenris/fenris.conf.
The file holds exactly one key: the device selector.
Spec §8.3: re-read every run; no reload path.
"""
if not CONFIG_PATH.exists():
raise RuntimeError(f"Configuration file not found: {CONFIG_PATH}")
config = {}
try:
with open(CONFIG_PATH, "r") as f:
for line in f:
line = line.strip()
if not line or line.startswith("#"):
continue
if "=" in line:
key, value = line.split("=", 1)
config[key.strip()] = value.strip()
except Exception as e:
raise RuntimeError(f"Failed to read configuration: {e}")
if "device" not in config:
raise RuntimeError("Configuration error: missing 'device' key")
return config
def interrogate_drive(device: str) -> dict:
"""Interrogate the drive via smartctl.
Returns the smartctl JSON output.
Raises RuntimeError on failure.
"""
result = subprocess.run(
["smartctl", "-a", "-j", device],
capture_output=True,
text=True,
)
if result.returncode != 0:
raise RuntimeError(
f"smartctl failed for {device}: {result.stderr}"
)
try:
return json.loads(result.stdout)
except json.JSONDecodeError as e:
raise RuntimeError(f"Failed to parse smartctl output: {e}")
def find_nvme_sysfs() -> Path | None:
"""Find the NVMe controller sysfs path."""
nvme_ctrl = Path("/sys/class/nvme")
if not nvme_ctrl.exists():
return None
for ctrl in sorted(nvme_ctrl.iterdir()):
if ctrl.name.startswith("nvme"):
return ctrl
return None
def main() -> None:
"""Run one collection cycle."""
try:
config = load_config()
device = config["device"]
# Inject a simple clock
class SimpleClock:
def utcnow(self):
return datetime.now(timezone.utc)
clock = SimpleClock()
def acquire():
"""Interrogate the drive after pending-work preflight."""
smartctl_data = interrogate_drive(device)
sysfs_path = find_nvme_sysfs()
if sysfs_path is None:
raise RuntimeError("No NVMe controller found in sysfs")
return smartctl_data, sysfs_path
# Run collection
result = run_collection(config=config, clock=clock, acquire=acquire)
if result["ok"]:
print(f"Collection successful: {result['sample_count']} sample(s)")
if result.get("retention_error"):
print(
f"Retention deferred: {result['retention_error']}",
file=sys.stderr,
)
sys.exit(0)
else:
print(f"Collection failed: {result['error']}", file=sys.stderr)
sys.exit(1)
except Exception as e:
print(f"Collection error: {e}", file=sys.stderr)
sys.exit(1)
if __name__ == "__main__":
main()
+583
View File
@@ -0,0 +1,583 @@
"""Collector: acquires counters and identity, writes to observation store.
This module implements the thinnest complete write path:
- Acquire counters and thermal evidence from smartctl -a -j
- Acquire controller identity from sysfs
- Normalize identity exactly once at write time
- Validate every row against store invariants
- Stage valid observations privately before derivation
- Publish the sample and derived evidence in one collection-owned transaction
No code path outside the collector interrogates the device.
"""
import json
import sqlite3
from collections.abc import Callable
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, Optional
from .derive import derive_hours_from_interval, find_previous_sample
from .monitoring_periods import ensure_period_open
from .store import init_store, get_store_path
class AcquisitionError(Exception):
"""Raised when acquisition fails - whole run is refused."""
pass
class InvariantViolationError(Exception):
"""Raised when a row would violate store invariants - writes nothing."""
pass
class _PendingDerivationError(Exception):
"""A valid staged observation failed while deriving dependent evidence."""
def __init__(self, cause: Exception):
super().__init__(str(cause))
self.cause = cause
class _PendingRecoveryFailure(Exception):
"""A publication attempt failed after earlier pending rows committed."""
def __init__(self, error: Exception, published_count: int):
derivation_error = isinstance(error, _PendingDerivationError)
cause = error.cause if derivation_error else error
super().__init__(str(cause))
self.cause = cause
self.published_count = published_count
self.derivation_error = derivation_error
PENDING_PUBLICATION_LIMIT = 6720
def acquire_from_smartctl(smartctl_data: Dict[str, Any]) -> Dict[str, Any]:
"""Acquire counters and thermal evidence from smartctl -a -j data.
Validates that all required fields are present.
Raises AcquisitionError on any failure.
"""
required_fields = [
"nvme_smart_health_information_log",
"user_capacity",
"model_name",
"serial_number",
"firmware_version",
]
for field in required_fields:
if field not in smartctl_data:
raise AcquisitionError(f"Missing required field in smartctl data: {field}")
log = smartctl_data["nvme_smart_health_information_log"]
required_log_fields = [
"data_units_written",
"data_units_read",
"percentage_used",
"power_on_hours",
"temperature",
]
for field in required_log_fields:
if field not in log:
raise AcquisitionError(f"Missing required field in SMART log: {field}")
return {
"model": smartctl_data["model_name"],
"serial": smartctl_data["serial_number"],
"firmware_rev": smartctl_data["firmware_version"],
"capacity_bytes": smartctl_data["user_capacity"]["bytes"],
"percentage_used": log["percentage_used"],
"available_spare": log.get("available_spare"),
"media_errors": log.get("media_errors", 0),
"power_on_hours": log["power_on_hours"],
"power_cycles": log.get("power_cycles"),
"unsafe_shutdowns": log.get("unsafe_shutdowns"),
"temperature_c": log["temperature"],
"data_units_written": log["data_units_written"],
"data_units_read": log["data_units_read"],
"bytes_written": log["data_units_written"] * 512000,
"bytes_read": log["data_units_read"] * 512000,
"critical_warning": log.get("critical_warning", 0),
}
def acquire_from_sysfs(sysfs_path: Path) -> Dict[str, Any]:
"""Acquire controller identity from sysfs.
Reads identity from:
- /sys/class/nvme/<ctrl>/subsysnqn (primary)
- /sys/class/nvme/<ctrl>/model
- /sys/class/nvme/<ctrl>/serial
- /sys/class/nvme/<ctrl>/firmware_rev
- /sys/class/nvme/<ctrl>/transport/ (optional)
Raises AcquisitionError on any failure.
"""
identity_files = {
"subnqn": "subsysnqn",
"mn": "model",
"sn": "serial",
"fr": "firmware_rev",
}
identity = {}
for key, filename in identity_files.items():
filepath = sysfs_path / filename
if not filepath.exists():
raise AcquisitionError(f"Missing sysfs file: {filepath}")
try:
value = filepath.read_text().strip()
identity[key] = value if value else ""
except Exception as e:
raise AcquisitionError(f"Failed to read {filepath}: {e}")
# Transport info (optional)
transport_dir = sysfs_path / "transport"
if transport_dir.exists():
try:
transport_file = transport_dir / "trstring"
if transport_file.exists():
identity["transport"] = transport_file.read_text().strip()
else:
identity["transport"] = None
except Exception:
identity["transport"] = None
else:
identity["transport"] = None
# vid/ssvid from PCI node (optional, metadata only - never key components)
# PCI device directory is the sysfs_path itself (the controller dir is a symlink to PCI)
pci_device = sysfs_path
for attr, key in [("vendor", "vid"), ("subsystem_vendor", "ssvid")]:
filepath = pci_device / attr
if filepath.exists():
try:
value = filepath.read_text().strip()
identity[key] = value if value else None
except Exception:
identity[key] = None
else:
identity[key] = None
return identity
def normalize_identity(identity: Dict[str, Any]) -> str:
"""Normalize identity exactly once at write time.
Rules:
- Strip trailing spaces and newlines
- No case folding
- Empty-after-strip stored blank
Returns normalized identity key.
"""
# Primary key: normalized kernel-exposed subsystem NQN
key = identity.get("subnqn", "")
if key:
key = key.rstrip()
return key
# Fallback 1: kernel composite (not implemented yet)
# Fallback 2: model|serial
mn = identity.get("mn", "").rstrip()
sn = identity.get("sn", "").rstrip()
if mn or sn:
return f"{mn}|{sn}"
# All keys blank - degraded identity
return ""
def compute_identity_degraded(identity: Dict[str, Any]) -> bool:
"""Check if identity is degraded (all key rungs empty)."""
key = normalize_identity(identity)
return key == ""
def validate_sample_invariants(sample: Dict[str, Any], conn: sqlite3.Connection) -> None:
"""Validate sample against store invariants.
Raises InvariantViolationError if any invariant is violated.
"""
# TODO: Implement more complex invariants as needed
# For now, just check basic constraints
if sample.get("bytes_written", 0) < 0:
raise InvariantViolationError("Negative bytes_written")
if sample.get("bytes_read", 0) < 0:
raise InvariantViolationError("Negative bytes_read")
def write_sample(
sample: Dict[str, Any],
identity: Dict[str, Any],
conn: sqlite3.Connection,
observed_at: datetime,
) -> Dict[str, Any]:
"""Write one sample to the observation store.
Identity normalization happens exactly once here.
Returns segment info for the caller. Caller owns the transaction.
"""
from .segment import find_current_segment, should_open_new_segment, open_segment
# Normalize identity exactly once at write time
identity_key = normalize_identity(identity)
identity_degraded = compute_identity_degraded(identity)
# Find current segment
current_segment = find_current_segment(conn)
# Determine if we need a new segment
should_open, reason = should_open_new_segment(
current_segment, identity_key, sample["bytes_written"], conn
)
# Open new segment if needed
segment_opened = False
if should_open:
open_segment(conn, observed_at, identity, identity_key, identity_degraded)
segment_opened = True
# Get current segment_id for provenance
current_segment = find_current_segment(conn)
segment_id = current_segment["id"] if current_segment else None
# Insert sample with segment_id
cursor = conn.execute(
"""
INSERT INTO samples (
ts, device, subnqn, sn, mn, fr, capacity_bytes,
percentage_used, available_spare, media_errors, power_on_hours,
power_cycles, unsafe_shutdowns, temperature_c,
data_units_written, data_units_read, bytes_written, bytes_read,
critical_warning, segment_id, local_tz
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
sample["ts"],
sample["device"],
identity.get("subnqn", ""),
identity.get("sn", ""),
identity.get("mn", ""),
identity.get("fr", ""),
sample["capacity_bytes"],
sample["percentage_used"],
sample["available_spare"],
sample["media_errors"],
sample["power_on_hours"],
sample["power_cycles"],
sample["unsafe_shutdowns"],
sample["temperature_c"],
sample["data_units_written"],
sample["data_units_read"],
sample["bytes_written"],
sample["bytes_read"],
sample["critical_warning"],
segment_id,
sample.get("local_tz"),
),
)
return {
"segment_opened": segment_opened,
"segment_reason": reason,
"identity_key": identity_key,
"identity_degraded": identity_degraded,
"segment_id": segment_id,
"sample_id": cursor.lastrowid,
}
def _observation_time(sample: Dict[str, Any]) -> datetime:
"""Read an observation's original timestamp for ordered recovery."""
observed_at = datetime.fromisoformat(sample["ts"])
if observed_at.tzinfo is None:
return observed_at.replace(tzinfo=timezone.utc)
return observed_at.astimezone(timezone.utc)
def _publish_observation(
conn: sqlite3.Connection,
sample: Dict[str, Any],
identity: Dict[str, Any],
tz_name: str,
) -> None:
"""Publish one staged sample and all dependent evidence in caller transaction."""
observed_at = _observation_time(sample)
sample["local_tz"] = tz_name
validate_sample_invariants(sample, conn)
ensure_period_open(conn, observed_at)
seg_info = write_sample(sample, identity, conn, observed_at)
current_id = seg_info["sample_id"]
prev = find_previous_sample(conn, seg_info.get("segment_id"), current_id)
try:
if prev is not None:
current = {
"id": current_id,
"ts": sample["ts"],
"bytes_written": sample["bytes_written"],
"bytes_read": sample["bytes_read"],
"power_on_hours": sample["power_on_hours"],
"temperature_c": sample["temperature_c"],
"data_units_written": sample["data_units_written"],
"data_units_read": sample["data_units_read"],
"local_tz": tz_name,
}
derive_hours_from_interval(conn, prev, current)
from .local_day import record_local_activity_interval
record_local_activity_interval(
conn,
prev,
current,
start_sample_id=prev["id"],
end_sample_id=current_id,
segment_id=seg_info.get("segment_id"),
)
else:
previous_any_segment = find_previous_sample(conn, None, current_id)
if previous_any_segment is not None:
from .local_day import mark_local_activity_gap
mark_local_activity_gap(
conn,
sample["ts"],
tz_name,
previous_any_segment,
)
from .day_aggregate import derive_all_days, persist_day_aggregate
for aggregate in derive_all_days(conn):
persist_day_aggregate(conn, aggregate)
from .local_day import derive_local_day_summary, persist_local_day
local_summary = derive_local_day_summary(conn, tz_name, observed_at)
if local_summary is not None:
persist_local_day(conn, local_summary)
except Exception as exc: # noqa: BLE001 - preserve valid staged evidence for retry.
raise _PendingDerivationError(exc) from exc
def _pending_count(conn: sqlite3.Connection) -> int:
return conn.execute("SELECT COUNT(*) FROM pending_publications").fetchone()[0]
def _stage_observation(
conn: sqlite3.Connection,
sample: Dict[str, Any],
identity: Dict[str, Any],
tz_name: str,
) -> int:
"""Durably stage valid acquired evidence before attempting publication."""
payload = json.dumps(
{"sample": sample, "identity": identity, "tz_name": tz_name},
separators=(",", ":"),
sort_keys=True,
)
cursor = conn.execute(
"INSERT INTO pending_publications (sample_ts, payload) VALUES (?, ?)",
(sample["ts"], payload),
)
return cursor.lastrowid
def _recover_pending(conn: sqlite3.Connection) -> int:
"""Publish pending observations oldest first; stop at first failure."""
recovered = 0
while True:
conn.execute("BEGIN IMMEDIATE")
try:
row = conn.execute(
"SELECT id, payload FROM pending_publications ORDER BY id LIMIT 1"
).fetchone()
if row is None:
conn.rollback()
return recovered
pending_id, payload = row
observation = json.loads(payload)
try:
_publish_observation(
conn,
observation["sample"],
observation["identity"],
observation["tz_name"],
)
except Exception as exc:
raise _PendingRecoveryFailure(exc, recovered) from exc
conn.execute("DELETE FROM pending_publications WHERE id = ?", (pending_id,))
conn.commit()
recovered += 1
except Exception:
conn.rollback()
raise
def _is_store_or_invariant_failure(exc: Exception) -> bool:
return (
isinstance(exc, sqlite3.Error)
and not isinstance(exc, sqlite3.IntegrityError)
) or isinstance(exc, InvariantViolationError)
def _pending_capacity_error(
reason: str,
*,
acquisition_skipped: bool = False,
) -> RuntimeError:
outcome = "; no new observation acquired" if acquisition_skipped else ""
return RuntimeError(
"pending publication capacity full "
f"({PENDING_PUBLICATION_LIMIT} observations){outcome}; "
f"{reason}"
)
def _recover_pending_for_collection(conn: sqlite3.Connection) -> int:
"""Retry queued work and report exhausted capacity without masking store faults."""
try:
return _recover_pending(conn)
except Exception as exc: # noqa: BLE001 - preserve pending evidence on recovery errors.
cause = exc.cause if isinstance(exc, _PendingRecoveryFailure) else exc
if _is_store_or_invariant_failure(cause):
raise cause
try:
capacity_full = _pending_count(conn) >= PENDING_PUBLICATION_LIMIT
except sqlite3.Error:
raise cause
if capacity_full:
raise _pending_capacity_error(f"recovery failed: {cause}") from cause
raise cause
def _recover_pending_for_admission(conn: sqlite3.Connection) -> int:
"""Recover in order, then admit one observation if bounded space remains."""
try:
recovered = _recover_pending(conn)
except _PendingRecoveryFailure as failure:
if (
not failure.derivation_error
or _is_store_or_invariant_failure(failure.cause)
):
raise failure.cause
try:
pending_count = _pending_count(conn)
except sqlite3.Error:
raise failure.cause
if pending_count >= PENDING_PUBLICATION_LIMIT:
raise _pending_capacity_error(
f"recovery failed: {failure.cause}",
acquisition_skipped=True,
) from failure.cause
return failure.published_count
if _pending_count(conn) >= PENDING_PUBLICATION_LIMIT:
raise _pending_capacity_error(
"recovery left the queue full",
acquisition_skipped=True,
)
return recovered
def run_collection(
smartctl_data: Optional[Dict[str, Any]] = None,
sysfs_path: Optional[Path] = None,
config: Optional[Dict[str, Any]] = None,
clock=None,
*,
acquire: Optional[Callable[[], tuple[Dict[str, Any], Path]]] = None,
) -> Dict[str, Any]:
"""Recover old work, acquire one observation, then publish it atomically.
Production callers pass ``acquire`` so recovery and capacity checks run
before device interrogation. Direct sample arguments remain useful for
deterministic collector tests.
"""
conn = None
try:
if config is None or clock is None:
raise ValueError("config and clock are required")
store_path = get_store_path(config)
conn = init_store(store_path)
from .legacy import import_legacy_history
history_path = Path(config.get("data_dir", ".")) / "history.jsonl"
if history_path.exists():
import_legacy_history(conn, history_path, clock=clock)
published_count = _recover_pending_for_admission(conn)
checked_pending_count = _pending_count(conn)
# Hold the writer reservation across the capacity check and acquisition.
# Recheck recovery if another collector appended work during preflight.
while True:
conn.execute("BEGIN IMMEDIATE")
waiting = _pending_count(conn)
if (
waiting >= PENDING_PUBLICATION_LIMIT
or waiting > checked_pending_count
):
conn.rollback()
published_count += _recover_pending_for_admission(conn)
checked_pending_count = _pending_count(conn)
continue
from .tz_util import detect_system_tz
tz_name = detect_system_tz()
if acquire is not None:
smartctl_data, sysfs_path = acquire()
if smartctl_data is None or sysfs_path is None:
raise AcquisitionError("No acquired SMART data or sysfs identity")
counters = acquire_from_smartctl(smartctl_data)
identity = acquire_from_sysfs(sysfs_path)
sample = {
"ts": clock.utcnow().isoformat(),
"device": config["device"],
**counters,
**identity,
}
validate_sample_invariants(sample, conn)
_stage_observation(conn, sample, identity, tz_name)
conn.commit()
break
published_count += _recover_pending_for_collection(conn)
retention_error = None
try:
if _pending_count(conn) == 0:
from .pruning import prune_old_samples
prune_old_samples(conn, clock.utcnow())
except Exception as exc: # noqa: BLE001 - publication already committed.
# Publication already committed. Keep collection successful and
# retry atomic retention on the next normal collection.
conn.rollback()
retention_error = str(exc)
return {
"ok": True,
"sample_count": published_count,
"retention_error": retention_error,
"store_path": str(store_path),
}
except Exception as exc:
if conn is not None:
conn.rollback()
return {
"ok": False,
"error": str(exc),
"error_type": type(exc).__name__,
}
finally:
if conn is not None:
conn.close()
+56
View File
@@ -0,0 +1,56 @@
"""Terminal-attached invocation of Fenris's fixed privileged operations.
Both human entry points use this module. Authentication has no frontend
deadline; collection runtime is bounded by the native scheduler (ADR 0003).
"""
import os
import shlex
import subprocess
MONITOR_HELPER = "/usr/libexec/fenris/fenris-monitor"
class MonitorError(Exception):
"""An action failed, with a message and exit status for either renderer."""
def __init__(self, message: str, exit_code: int = 1):
super().__init__(message)
self.exit_code = exit_code
def run_monitor(*args: str, helper_path: str = MONITOR_HELPER) -> None:
"""Run one helper operation, inheriting the terminal for authentication.
Never invoke a shell, retry an action, or fall back to sudo automatically.
The helper owns the operation allow-list and privileged state changes.
"""
helper_command = [helper_path, *args]
needs_auth = os.geteuid() != 0
command = ["pkexec", *helper_command] if needs_auth else helper_command
root_hint = (
" If authentication is unavailable, run in your terminal: "
+ shlex.join(["sudo", *helper_command])
) if needs_auth else ""
try:
# The collector owns its 90-second runtime limit. A frontend timeout
# would also count time spent authenticating or waiting for a run.
result = subprocess.run(command)
except FileNotFoundError as exc:
raise MonitorError(
"Command not found: %s.%s" % (exc.filename or command[0], root_hint), 127,
) from exc
except OSError as exc:
raise MonitorError("Cannot run monitoring action: %s.%s" % (exc, root_hint)) from exc
except KeyboardInterrupt as exc:
raise MonitorError(
"Action interrupted. Check fenris status before retrying.", 130,
) from exc
if result.returncode:
exit_code = result.returncode if result.returncode > 0 else 128 - result.returncode
raise MonitorError(
"Action failed (exit %d). Check fenris status before retrying.%s"
% (exit_code, root_hint), exit_code,
)
+203
View File
@@ -0,0 +1,203 @@
"""Day aggregate derivation per spec §5.4, §3.3.
One row per UTC day, derived monotonically from hour rows — the grain at
which usage-habit evidence is judged. No absent hour is ever interpolated,
estimated, or fabricated (§5.3, FL-3).
Coverage: the share of wall-clock seconds inside monitoring periods whose
usage-habit classification is known rather than unknown (§5.3).
"""
import sqlite3
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
@dataclass(frozen=True)
class DayAggregate:
"""One UTC day's aggregated stats."""
day: str # ISO 8601 UTC date, e.g. "2026-09-01"
seconds_active: int
seconds_idle: int
seconds_powered_off: int
seconds_unknown: int
bytes_written_delta: int
bytes_read_delta: int
sample_count: int
coverage: float
def _period_wall_clock_for_day(conn: sqlite3.Connection, day: str) -> int:
"""Total wall-clock seconds inside monitoring periods for a UTC day.
Clamps each period to the day boundary [dayT00:00, dayT24:00).
"""
day_start = datetime.fromisoformat(f"{day}T00:00:00+00:00")
day_end = day_start + timedelta(days=1)
day_start_str = day_start.isoformat()
day_end_str = day_end.isoformat()
cursor = conn.execute(
"SELECT started_at, ended_at FROM monitoring_periods "
"WHERE (ended_at IS NULL OR ended_at > ?) AND started_at < ? "
"ORDER BY started_at",
(day_start_str, day_end_str),
)
total = 0
for row in cursor.fetchall():
period_start = row[0]
period_end = row[1]
effective_start = max(period_start, day_start_str)
if period_end is not None:
effective_end = min(period_end, day_end_str)
else:
effective_end = day_end_str
if effective_start < effective_end:
s = datetime.fromisoformat(effective_start)
e = datetime.fromisoformat(effective_end)
total += int((e - s).total_seconds())
return total
def _hour_overlaps_period(conn: sqlite3.Connection, hour_iso: str) -> bool:
"""Check if an hour's wall-clock span overlaps any monitoring period."""
hour_start = datetime.fromisoformat(hour_iso)
hour_end = hour_start + timedelta(hours=1)
hs = hour_start.isoformat()
he = hour_end.isoformat()
cursor = conn.execute(
"SELECT 1 FROM monitoring_periods "
"WHERE started_at < ? AND (ended_at IS NULL OR ended_at > ?) "
"LIMIT 1",
(he, hs),
)
return cursor.fetchone() is not None
def derive_day(conn: sqlite3.Connection, day: str) -> DayAggregate | None:
"""Derive a single day aggregate from its hour rows + monitoring periods.
Only hours overlapping a monitoring period contribute to the aggregate.
Gap hours inside periods contribute unknown seconds. Hours outside all
monitoring periods are excluded entirely (§5.2).
Returns None if no hours exist for the day.
"""
cursor = conn.execute(
"SELECT hour, active_seconds, idle_seconds, powered_off_seconds, unknown_seconds, "
" bytes_written_delta, bytes_read_delta, sample_count "
"FROM hour_observations "
"WHERE hour LIKE ? "
"ORDER BY hour",
(day + "T%",),
)
rows = cursor.fetchall()
if not rows:
return None
total_active = 0
total_idle = 0
total_powered_off = 0
total_unknown_from_hours = 0
total_bw = 0
total_br = 0
total_samples = 0
total_hour_wall_clock = 0
for row in rows:
# Only count hours overlapping a monitoring period
if not _hour_overlaps_period(conn, row[0]):
continue
total_active += row[1]
total_idle += row[2]
total_powered_off += row[3]
total_unknown_from_hours += row[4]
total_bw += row[5]
total_br += row[6]
total_samples += row[7]
total_hour_wall_clock += row[1] + row[2] + row[3] + row[4]
# Wall-clock seconds inside monitoring periods for this day
period_wc = _period_wall_clock_for_day(conn, day)
# Gap seconds = period wall-clock - sum of existing hour wall-clock
gap_seconds = max(0, period_wc - total_hour_wall_clock)
total_unknown = total_unknown_from_hours + gap_seconds
# Coverage: known seconds / period wall-clock (§5.2, §5.3)
known_seconds = total_active + total_idle + total_powered_off
coverage = known_seconds / period_wc if period_wc > 0 else 0.0
return DayAggregate(
day=day,
seconds_active=total_active,
seconds_idle=total_idle,
seconds_powered_off=total_powered_off,
seconds_unknown=total_unknown,
bytes_written_delta=total_bw,
bytes_read_delta=total_br,
sample_count=total_samples,
coverage=coverage,
)
def derive_all_days(conn: sqlite3.Connection) -> list[DayAggregate]:
"""Derive day aggregates for all days that have hour rows.
Returns days sorted by date.
"""
cursor = conn.execute(
"SELECT DISTINCT substr(hour, 1, 10) as day FROM hour_observations ORDER BY day"
)
days = [row[0] for row in cursor.fetchall()]
results = []
for day in days:
agg = derive_day(conn, day)
if agg is not None:
results.append(agg)
return results
def persist_day_aggregate(conn: sqlite3.Connection, agg: DayAggregate) -> None:
"""Upsert a derived day aggregate into the day_aggregates table.
Merges attributed bytes from hour observations with any existing
unattributed cross-hour evidence already stored for this day.
Caller must manage transactions and commits.
"""
existing = conn.execute(
"SELECT id FROM day_aggregates WHERE day = ?",
(agg.day,),
).fetchone()
if existing is None:
conn.execute(
"INSERT INTO day_aggregates "
"(day, active_seconds, idle_seconds, powered_off_seconds, unknown_seconds, "
" bytes_written_delta, bytes_read_delta, sample_count, coverage) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
(agg.day, agg.seconds_active, agg.seconds_idle,
agg.seconds_powered_off, agg.seconds_unknown,
agg.bytes_written_delta, agg.bytes_read_delta,
agg.sample_count, agg.coverage),
)
else:
conn.execute(
"UPDATE day_aggregates "
"SET active_seconds = ?, idle_seconds = ?, powered_off_seconds = ?, "
" unknown_seconds = ?, bytes_written_delta = ?, bytes_read_delta = ?, "
" sample_count = ?, coverage = ? "
"WHERE id = ?",
(agg.seconds_active, agg.seconds_idle,
agg.seconds_powered_off, agg.seconds_unknown,
agg.bytes_written_delta, agg.bytes_read_delta,
agg.sample_count, agg.coverage,
existing[0]),
)
+241
View File
@@ -0,0 +1,241 @@
"""Interval derivation: samples → hour observations → day aggregates.
After each collection run, the collector calls into this module to:
1. Find the previous sample in the same segment
2. Compute deltas (bytes, POH, temperature)
3. Classify the hour(s) the interval spans
4. Write/update hour_observations for each affected hour
5. Update day_aggregates with unattributed cross-hour bytes
Cross-hour deltas are retained once with unknown shares explicit (issue #73 AC4).
No proportional allocation, endpoint assignment, or double counting.
"""
import sqlite3
from datetime import datetime, timedelta, timezone
from typing import Any, Dict, List, Optional, Tuple
from .hour_classify import classify_hour, HourSplit
def find_previous_sample(
conn: sqlite3.Connection,
segment_id: Optional[int],
current_sample_id: int,
) -> Optional[Dict[str, Any]]:
"""Find the most recent sample before current_sample_id in the same segment.
Returns None if no previous sample exists (first sample in segment).
"""
if segment_id is not None:
cursor = conn.execute(
"SELECT id, ts, bytes_written, bytes_read, power_on_hours, "
" temperature_c, data_units_written, data_units_read, local_tz "
"FROM samples WHERE id < ? AND segment_id = ? "
"ORDER BY id DESC LIMIT 1",
(current_sample_id, segment_id),
)
else:
cursor = conn.execute(
"SELECT id, ts, bytes_written, bytes_read, power_on_hours, "
" temperature_c, data_units_written, data_units_read, local_tz "
"FROM samples WHERE id < ? "
"ORDER BY id DESC LIMIT 1",
(current_sample_id,),
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0], "ts": row[1], "bytes_written": row[2],
"bytes_read": row[3], "power_on_hours": row[4],
"temperature_c": row[5], "data_units_written": row[6],
"data_units_read": row[7], "local_tz": row[8],
}
def _parse_ts(ts: str) -> datetime:
"""Parse ISO timestamp to datetime with UTC."""
dt = datetime.fromisoformat(ts)
if dt.tzinfo is None:
dt = dt.replace(tzinfo=timezone.utc)
return dt
def _hour_floor(dt: datetime) -> datetime:
"""Floor a datetime to its UTC hour boundary."""
return dt.replace(minute=0, second=0, microsecond=0)
def _hours_spanned(start: datetime, end: datetime) -> List[datetime]:
"""Return list of UTC hour boundaries spanned by [start, end)."""
hours = []
h = _hour_floor(start)
while h < end:
hours.append(h)
h += timedelta(hours=1)
return hours
def _compute_sampled_seconds_in_hour(
start: datetime, end: datetime, hour_start: datetime
) -> int:
"""How many seconds of the sample interval fall within this hour."""
hour_end = hour_start + timedelta(hours=1)
effective_start = max(start, hour_start)
effective_end = min(end, hour_end)
if effective_start >= effective_end:
return 0
return int((effective_end - effective_start).total_seconds())
def derive_hours_from_interval(
conn: sqlite3.Connection,
prev_sample: Dict[str, Any],
next_sample: Dict[str, Any],
) -> List[Dict[str, Any]]:
"""Derive hour observations from a sample pair interval.
Returns list of hour observation dicts that were written/updated.
Caller owns the transaction.
"""
prev_ts = _parse_ts(prev_sample["ts"])
next_ts = _parse_ts(next_sample["ts"])
# Deltas
bw_delta = max(0, next_sample["bytes_written"] - prev_sample["bytes_written"])
br_delta = max(0, next_sample["bytes_read"] - prev_sample["bytes_read"])
poh_delta_s = max(0, (next_sample["power_on_hours"] - prev_sample["power_on_hours"])) * 3600
hours = _hours_spanned(prev_ts, next_ts)
total_span_s = int((next_ts - prev_ts).total_seconds())
results = []
if len(hours) == 1:
# Same-hour interval: fully attributed to this hour
hour_key = hours[0].strftime("%Y-%m-%dT%H:00:00+00:00")
sampled_s = total_span_s
# Classify hour
split = classify_hour(
wall_clock_seconds=3600,
poh_delta=poh_delta_s,
duw_delta=bw_delta,
dur_delta=br_delta,
sampled_seconds=sampled_s,
)
_upsert_hour_observation(
conn, hour_key, split,
bw_delta, br_delta,
prev_sample.get("temperature_c"), next_sample.get("temperature_c"),
2, # 2 samples contributed (prev + next)
)
results.append({"hour": hour_key, "bytes_written": bw_delta, "attributed": True})
elif len(hours) >= 2:
# Cross-hour interval: split wall-clock time, bytes unattributed
for h in hours:
hour_key = h.strftime("%Y-%m-%dT%H:00:00+00:00")
sampled_s = _compute_sampled_seconds_in_hour(prev_ts, next_ts, h)
# For cross-hour, we classify based on time only (no byte attribution)
# The hour gets its time split but NOT the byte delta
split = classify_hour(
wall_clock_seconds=3600,
poh_delta=0, # POH attribution unknown for cross-hour
duw_delta=0, # Bytes unattributed
dur_delta=0,
sampled_seconds=sampled_s,
)
_upsert_hour_observation(
conn, hour_key, split,
0, 0, # No byte attribution for cross-hour
None, None,
0, # No sample falls IN this hour
)
results.append({"hour": hour_key, "bytes_written": 0, "attributed": False})
# Track unattributed bytes at day level
_add_unattributed_bytes(conn, prev_ts, next_ts, bw_delta, br_delta)
return results
def _upsert_hour_observation(
conn: sqlite3.Connection,
hour_key: str,
split: HourSplit,
bw_delta: int,
br_delta: int,
temp_min: Optional[int],
temp_max: Optional[int],
sample_count: int,
) -> None:
"""Insert or update an hour observation."""
# Check if hour exists
existing = conn.execute(
"SELECT id, bytes_written_delta, bytes_read_delta, sample_count "
"FROM hour_observations WHERE hour = ?",
(hour_key,),
).fetchone()
if existing is None:
temp_avg = ((temp_min or 0) + (temp_max or 0)) / 2 if temp_min is not None else None
coverage = (split.seconds_active + split.seconds_idle + split.seconds_powered_off) / 3600.0
conn.execute(
"""INSERT INTO hour_observations
(hour, active_seconds, idle_seconds, powered_off_seconds, unknown_seconds,
bytes_written_delta, bytes_read_delta,
temperature_min, temperature_avg, temperature_max,
sample_count, coverage)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)""",
(hour_key, split.seconds_active, split.seconds_idle,
split.seconds_powered_off, split.seconds_unknown,
bw_delta, br_delta,
temp_min, temp_avg, temp_max,
sample_count, coverage),
)
else:
# Merge: accumulate bytes and sample count
new_bw = existing[1] + bw_delta
new_br = existing[2] + br_delta
new_samples = existing[3] + sample_count
conn.execute(
"UPDATE hour_observations "
"SET bytes_written_delta = ?, bytes_read_delta = ?, sample_count = ? "
"WHERE id = ?",
(new_bw, new_br, new_samples, existing[0]),
)
def _add_unattributed_bytes(
conn: sqlite3.Connection,
prev_ts: datetime,
next_ts: datetime,
bw_delta: int,
br_delta: int,
) -> None:
"""Store unattributed byte deltas once as shared boundary evidence.
A cross-midnight interval's delta is preserved on the day where it
STARTS (the earlier day). It is not duplicated into both days;
the spec requires preserving the measured volume once as shared
unallocated boundary evidence (issue #88).
"""
day = prev_ts.strftime("%Y-%m-%d")
existing = conn.execute(
"SELECT id FROM day_aggregates WHERE day = ?", (day,)
).fetchone()
if existing is None:
conn.execute(
"INSERT INTO day_aggregates (day, unattributed_bytes_written, unattributed_bytes_read) "
"VALUES (?, ?, ?)",
(day, bw_delta, br_delta),
)
else:
conn.execute(
"UPDATE day_aggregates SET unattributed_bytes_written = unattributed_bytes_written + ?, "
"unattributed_bytes_read = unattributed_bytes_read + ? WHERE day = ?",
(bw_delta, br_delta, day),
)
+94
View File
@@ -0,0 +1,94 @@
"""Hour classification per spec §5.1.
Each UTC hour is classified by named constants, in this order of evidence:
- Powered-off: power-on-hours delta < 90% of wall-clock span
- Active: DUW delta >= 256 MiB in the hour
- Idle: powered on, sampled, below active threshold
- Unknown: everything else (unsampled without POH evidence)
Four splits sum to exactly wall_clock_seconds. Disabled time is never an
hour state — it is wall-clock outside monitoring periods (§5.2).
"""
from dataclasses import dataclass
# Spec §5.1: Active hour threshold — 256 MiB DUW delta
ACTIVE_THRESHOLD_BYTES = 256 * 1024 * 1024 # 256 MiB
# Spec §5.1: Powered-off threshold — 90% of wall-clock span
POWERED_OFF_THRESHOLD_PERCENT = 0.90
@dataclass(frozen=True)
class HourSplit:
"""Usage-habit split for one UTC hour. Fields sum to wall_clock_seconds."""
seconds_active: int
seconds_idle: int
seconds_powered_off: int
seconds_unknown: int
def classify_hour(
wall_clock_seconds: int,
poh_delta: int,
duw_delta: int,
dur_delta: int,
sampled_seconds: int | None = None,
) -> HourSplit:
"""Classify a UTC hour into the four usage-habit states.
Args:
wall_clock_seconds: Total seconds in this hour boundary (3600 for a
full hour, less for partial-hours at period edges).
poh_delta: Power-on-hours delta since previous sample (in seconds).
duw_delta: Data-units-written delta since previous sample (in bytes).
dur_delta: Data-units-read delta since previous sample (in bytes).
sampled_seconds: Seconds within this hour covered by a sample.
None or 0 means no sample fell in this hour.
Returns:
HourSplit whose four fields sum to wall_clock_seconds.
"""
if sampled_seconds is None:
sampled_seconds = 0
# Clamp sampled_seconds to wall_clock_seconds
sampled_seconds = min(sampled_seconds, wall_clock_seconds)
# --- Decision order per spec §5.1 ---
# 1. Powered-off: POH delta < 90% of wall-clock span
powered_off_threshold = wall_clock_seconds * POWERED_OFF_THRESHOLD_PERCENT
if poh_delta < powered_off_threshold:
return HourSplit(
seconds_active=0,
seconds_idle=0,
seconds_powered_off=wall_clock_seconds,
seconds_unknown=0,
)
# 2. Active: DUW delta >= 256 MiB
if duw_delta >= ACTIVE_THRESHOLD_BYTES:
return HourSplit(
seconds_active=wall_clock_seconds,
seconds_idle=0,
seconds_powered_off=0,
seconds_unknown=0,
)
# 3. Idle: powered on, sampled, below active threshold
# Unsampled portion within the hour is unknown
if sampled_seconds > 0:
return HourSplit(
seconds_active=0,
seconds_idle=sampled_seconds,
seconds_powered_off=0,
seconds_unknown=wall_clock_seconds - sampled_seconds,
)
# 4. Unknown: unsampled without POH evidence
return HourSplit(
seconds_active=0,
seconds_idle=0,
seconds_powered_off=0,
seconds_unknown=wall_clock_seconds,
)
+500
View File
@@ -0,0 +1,500 @@
"""Init system abstraction for Fenris monitoring (issue #84).
Detects the active init system (systemd or runit) and provides a unified
interface for timer/collection control and service-state queries. This keeps
the privileged helper and status composition init-system agnostic without
introducing a generalized plugin framework.
Design: ADR 0008 — keep service-specific operations behind a cohesive
responsibility shared by the privileged control path and read-only status
composition.
Runit service layout:
/etc/sv/fenris-collect/run — scheduler (sleep 120; loop { collect; sleep 180 })
/etc/sv/fenris-collect/log/run — logger to /var/log/fenris-collect/
/var/service/fenris-collect — symlink to enable
/etc/sv/fenris-collect/down — marker for dormant install
Runit guarantees:
- Completion-relative 3-minute cadence (sleep 180 after each collect)
- Initial 2-minute boot delay (sleep 120 before first collect)
- Bounded execution (90s timeout via timeout(1))
- No catch-up (service sleeps fixed interval, no Persistent= flag)
- Serialized scheduled runs (runsv does not restart until exit)
- Serialized on-demand (flock serializes fenris-collect execution)
Spec: §8.4, §8.5, §8.6, §8.7, §8.8, ADR 0008
"""
import os
import shutil
import subprocess
import sys
from enum import Enum
from pathlib import Path
from typing import Any, Dict, Optional
# ---------------------------------------------------------------------------
# Init system detection
# ---------------------------------------------------------------------------
class InitSystem(Enum):
SYSTEMD = "systemd"
RUNIT = "runit"
FENRIS_SV_DIR = Path("/etc/sv/fenris-collect")
FENRIS_SERVICE_LINK = Path("/var/service/fenris-collect")
FENRIS_LOG_DIR = Path("/var/log/fenris-collect")
COLLECT_TIMEOUT_S = 90
def detect_init_system() -> InitSystem:
"""Detect the active init system.
Checks for systemd first (PID 1 is systemd or /run/systemd/system exists),
then falls back to runit (PID 1 is runsv or /etc/sv exists).
"""
# systemd detection: /run/systemd/system exists when systemd is PID 1
if Path("/run/systemd/system").exists():
return InitSystem.SYSTEMD
# Check PID 1 name
try:
pid1_comm = Path("/proc/1/comm").read_text().strip()
if pid1_comm == "systemd":
return InitSystem.SYSTEMD
if pid1_comm in ("runsv", "runsvdir"):
return InitSystem.RUNIT
except OSError:
pass
# Fallback: check for /etc/sv (Void Linux default)
if Path("/etc/sv").is_dir():
return InitSystem.RUNIT
# Default to systemd (existing behavior)
return InitSystem.SYSTEMD
def get_init_system() -> InitSystem:
"""Get the detected init system (cached)."""
if not hasattr(get_init_system, "_cached"):
get_init_system._cached = detect_init_system()
return get_init_system._cached
def reset_init_system_cache() -> None:
"""Reset the cached init system detection (for testing)."""
if hasattr(get_init_system, "_cached"):
delattr(get_init_system, "_cached")
# ---------------------------------------------------------------------------
# systemd backend
# ---------------------------------------------------------------------------
def _systemd_enable(now: bool) -> None:
"""Enable and optionally start the systemd timer."""
cmd = ["systemctl", "enable"]
if now:
cmd.append("--now")
cmd.append("fenris-collect.timer")
result = subprocess.run(cmd, capture_output=True, text=True)
if result.returncode != 0:
print("Error enabling timer:", result.stderr, file=sys.stderr)
sys.exit(1)
print("Timer enabled" + (" and started" if now else ""))
def _systemd_disable(now: bool) -> None:
"""Disable and optionally stop the systemd timer."""
cmd = ["systemctl", "disable"]
if now:
cmd.append("--now")
cmd.append("fenris-collect.timer")
result = subprocess.run(cmd, capture_output=True, text=True)
if result.returncode != 0:
print("Error disabling timer:", result.stderr, file=sys.stderr)
sys.exit(1)
print("Timer disabled" + (" and stopped" if now else ""))
def _systemd_collect() -> None:
"""Trigger on-demand collection via systemd (blocking)."""
result = subprocess.run(
["systemctl", "start", "fenris-collect.service"],
capture_output=True,
text=True,
)
if result.returncode == 0:
print("Collection completed successfully")
else:
print("Collection failed:", result.stderr, file=sys.stderr)
sys.exit(1)
def _systemctl_show(unit: str, *properties: str) -> Dict[str, str]:
"""Query systemctl show for specific properties."""
try:
result = subprocess.run(
["systemctl", "show", unit, "--property=" + ",".join(properties)],
capture_output=True, text=True, timeout=5,
)
if result.returncode != 0:
return {}
out = {}
for line in result.stdout.splitlines():
if "=" in line:
key, _, value = line.partition("=")
out[key.strip()] = value.strip()
return out
except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
return {}
def _systemd_query_state() -> Dict[str, Any]:
"""Query systemd for the four separate service facts."""
from datetime import datetime, timezone
timer_props = _systemctl_show(
"fenris-collect.timer",
"UnitFileState", "ActiveState", "LastTriggerUSec",
)
service_props = _systemctl_show(
"fenris-collect.service",
"ActiveState", "ExecMainStatus", "ExecMainExitTimestamp",
)
boot_enabled_str = timer_props.get("UnitFileState")
boot_enabled = boot_enabled_str == "enabled" if boot_enabled_str else None
active_state = timer_props.get("ActiveState")
timer_active = active_state == "active" if active_state else None
last_collect_ok = None
last_collect_age_s = None
last_collect_reason = None
last_trigger = timer_props.get("LastTriggerUSec", "")
if last_trigger and last_trigger != "n/a":
try:
trigger_dt = datetime.fromisoformat(last_trigger.replace("Z", "+00:00"))
now = datetime.now(timezone.utc)
last_collect_age_s = int((now - trigger_dt).total_seconds())
except (ValueError, TypeError):
pass
exec_status = service_props.get("ExecMainStatus", "")
if exec_status:
try:
exit_code = int(exec_status)
last_collect_ok = exit_code == 0
if exit_code != 0:
last_collect_reason = "exit code %d" % exit_code
except (ValueError, TypeError):
pass
return {
"boot_enabled": boot_enabled,
"timer_active": timer_active,
"last_collect_ok": last_collect_ok,
"last_collect_age_s": last_collect_age_s,
"last_collect_reason": last_collect_reason,
}
def _systemd_journal_hint(lines: int = 5) -> Optional[str]:
"""Get the last N journal lines for fenris-collect.service."""
try:
result = subprocess.run(
["journalctl", "-u", "fenris-collect.service",
"--no-pager", "-n", str(lines), "--output=short-iso"],
capture_output=True, text=True, timeout=5,
)
if result.returncode != 0 or not result.stdout.strip():
return None
return result.stdout.strip()
except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
return None
# ---------------------------------------------------------------------------
# runit backend
# ---------------------------------------------------------------------------
def _runit_enable(_now: bool) -> None:
"""Enable the runit service by creating a symlink.
runit activates the service immediately when the symlink appears.
If the service is already enabled but stopped (e.g., via `sv stop`),
restart it when `now=True`.
"""
if FENRIS_SERVICE_LINK.exists():
# Already enabled — check if we need to restart
if _now and not _runit_is_running():
# Service is stopped but enabled — restart via sv
try:
subprocess.run(
["sv", "restart", "fenris-collect"],
capture_output=True, text=True, timeout=5,
)
except (FileNotFoundError, subprocess.TimeoutExpired):
pass
print("Service already enabled (idempotent)")
return
# Remove the 'down' file if present (dormant install marker)
down_file = FENRIS_SV_DIR / "down"
if down_file.exists():
down_file.unlink()
FENRIS_SERVICE_LINK.symlink_to(FENRIS_SV_DIR)
print("Service enabled")
def _runit_disable(_now: bool) -> None:
"""Disable the runit service by removing the symlink.
runit stops the service immediately when the symlink is removed.
"""
if not FENRIS_SERVICE_LINK.exists():
print("Service already disabled (idempotent)")
return
FENRIS_SERVICE_LINK.unlink()
# Place 'down' file to mark as intentionally disabled
(FENRIS_SV_DIR / "down").touch()
print("Service disabled")
def _runit_collect() -> None:
"""Trigger on-demand collection via direct execution with flock.
Serializes against the scheduler using the same lock file.
The timeout(1) command enforces bounded execution.
"""
lock_path = Path("/var/lib/fenris/fenris-collect.lock")
collect_script = Path("/usr/libexec/fenris/fenris-collect")
fallback_script = Path(__file__).parent.parent.parent / "src" / "fenris" / "collect.py"
if collect_script.exists():
script = str(collect_script)
elif fallback_script.exists():
script = str(fallback_script)
else:
print("Error: fenris-collect script not found", file=sys.stderr)
sys.exit(1)
try:
result = subprocess.run(
["flock", "--nonblock", str(lock_path),
"timeout", str(COLLECT_TIMEOUT_S), "nice", "ionice", "-c3",
sys.executable, script],
capture_output=True,
text=True,
)
if result.returncode == 0:
print("Collection completed successfully")
elif result.returncode == 124:
print("Collection timed out after %ds" % COLLECT_TIMEOUT_S, file=sys.stderr)
sys.exit(1)
else:
# exit code 1 from flock means lock is held (scheduled run in progress)
if result.returncode == 1 and "Resource temporarily unavailable" in result.stderr:
print("Collection already in progress (serialized)", file=sys.stderr)
sys.exit(1)
print("Collection failed:", result.stderr, file=sys.stderr)
sys.exit(1)
except FileNotFoundError:
print("Error: flock/timeout not found", file=sys.stderr)
sys.exit(1)
def _runit_is_enabled() -> bool:
"""Check if the runit service is enabled (symlink exists)."""
return FENRIS_SERVICE_LINK.exists()
def _runit_is_running() -> bool:
"""Check if the runit service is currently running.
Looks for a 'supervise/pid' file in the service directory. Void creates
that directory root-only, so unprivileged dashboard reads fall back to
the public process table when they cannot traverse it.
"""
def runsv_process_exists() -> bool:
try:
result = subprocess.run(
["pgrep", "-f", "^runsv fenris-collect$"],
capture_output=True,
text=True,
timeout=5,
)
return result.returncode == 0
except (FileNotFoundError, subprocess.TimeoutExpired, OSError):
return False
pid_file = FENRIS_SV_DIR / "supervise" / "pid"
# On Void, Path.exists() is false for an unprivileged process when it
# cannot traverse runit's root-only supervise directory.
if not pid_file.exists():
return FENRIS_SERVICE_LINK.exists() and runsv_process_exists()
try:
pid = int(pid_file.read_text().strip())
# Check if the process is alive
os.kill(pid, 0)
return True
except PermissionError:
# runsv's supervisor state is root-only on Void. Its process command
# is still observable, which gives the read-only UI the same runtime
# fact without granting it service-control permissions.
return FENRIS_SERVICE_LINK.exists() and runsv_process_exists()
except (ValueError, OSError):
return False
def _runit_query_state() -> Dict[str, Any]:
"""Query runit for the four separate service facts.
Checks: boot_enabled (symlink), timer_active (running), last_collect
(store-based), freshness (store-based).
"""
from datetime import datetime, timezone
from pathlib import Path
boot_enabled = _runit_is_enabled()
timer_active = _runit_is_running()
last_collect_ok = None
last_collect_age_s = None
last_collect_reason = None
# Try to get last collection info from the store
store_path = Path("/var/lib/fenris/observations.db")
if store_path.exists():
try:
import sqlite3
conn = sqlite3.connect("file:%s?mode=ro" % store_path, uri=True)
conn.row_factory = sqlite3.Row
# Get newest sample
cursor = conn.execute(
"SELECT ts FROM samples ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
if row and row[0]:
try:
ts = datetime.fromisoformat(row[0])
if ts.tzinfo is None:
ts = ts.replace(tzinfo=timezone.utc)
now = datetime.now(timezone.utc)
last_collect_age_s = int((now - ts).total_seconds())
last_collect_ok = True
except (ValueError, TypeError):
pass
# Check for failed collections via monitoring_periods
cursor = conn.execute(
"SELECT end_cause FROM monitoring_periods "
"WHERE ended_at IS NOT NULL ORDER BY ended_at DESC LIMIT 1"
)
row = cursor.fetchone()
if row and row[0] and row[0] not in ("user_disabled", None):
last_collect_reason = row[0]
conn.close()
except Exception:
pass
# Check for timeout/exit failures via supervise exit status
if timer_active:
exit_file = FENRIS_SV_DIR / "supervise" / "exit"
if exit_file.exists():
try:
exit_code = int(exit_file.read_text().strip())
if exit_code != 0:
last_collect_ok = False
last_collect_reason = "exit code %d" % exit_code
except (ValueError, OSError):
pass
return {
"boot_enabled": boot_enabled,
"timer_active": timer_active,
"last_collect_ok": last_collect_ok,
"last_collect_age_s": last_collect_age_s,
"last_collect_reason": last_collect_reason,
}
def _runit_journal_hint(lines: int = 5) -> Optional[str]:
"""Get the last N log lines for fenris-collect from runit logging."""
log_current = FENRIS_LOG_DIR / "current"
if not log_current.exists():
return None
try:
# Use tail to get the last N lines
result = subprocess.run(
["tail", "-n", str(lines), str(log_current)],
capture_output=True, text=True, timeout=5,
)
if result.returncode != 0 or not result.stdout.strip():
return None
return result.stdout.strip()
except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
return None
# ---------------------------------------------------------------------------
# Public API — unified interface
# ---------------------------------------------------------------------------
def enable_timer(now: bool) -> None:
"""Enable the collection timer/service."""
init = get_init_system()
if init == InitSystem.SYSTEMD:
_systemd_enable(now)
else:
_runit_enable(now)
def disable_timer(now: bool) -> None:
"""Disable the collection timer/service."""
init = get_init_system()
if init == InitSystem.SYSTEMD:
_systemd_disable(now)
else:
_runit_disable(now)
def collect_now() -> None:
"""Trigger on-demand collection (blocking)."""
init = get_init_system()
if init == InitSystem.SYSTEMD:
_systemd_collect()
else:
_runit_collect()
def query_service_state() -> Dict[str, Any]:
"""Query the four separate service facts."""
init = get_init_system()
if init == InitSystem.SYSTEMD:
return _systemd_query_state()
else:
return _runit_query_state()
def journal_hint(lines: int = 5, unit: str = "fenris-collect.service") -> Optional[str]:
"""Get journal/log hints for diagnostics.
The unit parameter is accepted for backward compatibility with callers
that pass 'fenris-collect.service'. For runit, the unit parameter is
ignored since the log directory is always /var/log/fenris-collect/.
"""
init = get_init_system()
if init == InitSystem.SYSTEMD:
return _systemd_journal_hint(lines)
else:
return _runit_journal_hint(lines)
+440
View File
@@ -0,0 +1,440 @@
"""Legacy migration: import history.jsonl into the observation store.
Spec §3.5, ADR 0001 §6. Idempotent and interruption-safe.
Entry points:
- Installer import detection at ./data/history.jsonl (§10.1)
- fenris import <path> (§8.8)
- Collector's first new-version run (ADR 0001 §6)
The import is a single transaction — a scripted kill mid-import leaves the
store fully pre- or fully post-migration. history.jsonl is the sole authority;
hourly.jsonl is diffed and logged but never trusted. Malformed lines are
quarantined with a logged count, never silently dropped. Legacy files are
renamed *.migrated only after commit and never deleted.
"""
import json
import logging
import sqlite3
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, List, Optional, Tuple
from .hour_classify import classify_hour
from .segment import open_segment, normalize_identity
from .monitoring_periods import close_period
logger = logging.getLogger(__name__)
# Legacy import marker — stored in a metadata table
LEGACY_IMPORT_MARKER = "legacy_imported"
def _ensure_metadata_table(conn: sqlite3.Connection) -> None:
"""Create the metadata table if it doesn't exist."""
conn.execute("""
CREATE TABLE IF NOT EXISTS store_metadata (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
)
""")
def is_legacy_imported(conn: sqlite3.Connection) -> bool:
"""Check if legacy history has already been imported.
Spec §3.5.1: If the store already carries the legacy-import marker, do nothing.
"""
_ensure_metadata_table(conn)
cursor = conn.execute(
"SELECT value FROM store_metadata WHERE key = ?",
(LEGACY_IMPORT_MARKER,),
)
row = cursor.fetchone()
return row is not None and row[0] == "true"
def _parse_history_line(line: str, line_num: int) -> Optional[Dict[str, Any]]:
"""Parse a single line from history.jsonl.
Returns None for malformed lines (quarantined, not silently dropped).
"""
line = line.strip()
if not line:
return None
try:
record = json.loads(line)
except json.JSONDecodeError as e:
logger.warning("Malformed JSON at line %d: %s", line_num, e)
return None
# Validate required fields
required_fields = [
"timestamp", "model", "serial", "firmware_version",
"data_units_written", "data_units_read", "percentage_used",
"power_on_hours", "temperature",
]
for field in required_fields:
if field not in record:
logger.warning("Missing field '%s' at line %d", field, line_num)
return None
return record
def _record_to_sample(record: Dict[str, Any]) -> Dict[str, Any]:
"""Convert a legacy history.jsonl record to a sample dict."""
# Legacy records use different field names
duw = record["data_units_written"]
dur = record["data_units_read"]
return {
"ts": record["timestamp"],
"device": record.get("device", "/dev/nvme0"),
"subnqn": record.get("subsystem_nqn", ""),
"sn": record["serial"],
"mn": record["model"],
"fr": record["firmware_version"],
"capacity_bytes": record.get("capacity_bytes", 0),
"percentage_used": record["percentage_used"],
"available_spare": record.get("available_spare"),
"media_errors": record.get("media_errors", 0),
"power_on_hours": record["power_on_hours"],
"power_cycles": record.get("power_cycles"),
"unsafe_shutdowns": record.get("unsafe_shutdowns"),
"temperature_c": record["temperature"],
"data_units_written": duw,
"data_units_read": dur,
"bytes_written": duw * 512000,
"bytes_read": dur * 512000,
"critical_warning": record.get("critical_warning", 0),
}
def _derive_hour_observation(
samples: List[Dict[str, Any]],
hour_start: datetime,
) -> Dict[str, Any]:
"""Derive a single hour observation from samples in that hour.
Pre-migration hours carry an unknown activity split except directly
evidenced facts — a sample present means powered on; a DUW delta means
writes occurred (§3.5.4).
"""
hour_end = hour_start.replace(hour=hour_start.hour + 1) if hour_start.hour < 23 else hour_start.replace(hour=0, day=hour_start.day + 1)
# Filter samples in this hour
hour_samples = []
for s in samples:
ts = datetime.fromisoformat(s["ts"])
if hour_start <= ts < hour_end:
hour_samples.append(s)
if not hour_samples:
return None
# Sort by timestamp
hour_samples.sort(key=lambda x: x["ts"])
# Compute deltas from first to last sample in the hour
first = hour_samples[0]
last = hour_samples[-1]
duw_delta = last["bytes_written"] - first["bytes_written"]
dur_delta = last["bytes_read"] - first["bytes_read"]
poh_delta = (last["power_on_hours"] - first["power_on_hours"]) * 3600
# Temperature stats
temps = [s["temperature_c"] for s in hour_samples]
# Classify the hour
split = classify_hour(
wall_clock_seconds=3600,
poh_delta=poh_delta,
duw_delta=duw_delta,
dur_delta=dur_delta,
sampled_seconds=3600, # Legacy samples cover the full hour
)
return {
"hour": hour_start.strftime("%Y-%m-%dT%H:00:00Z"),
"active_seconds": split.seconds_active,
"idle_seconds": split.seconds_idle,
"powered_off_seconds": split.seconds_powered_off,
"unknown_seconds": split.seconds_unknown,
"bytes_written_delta": duw_delta,
"bytes_read_delta": dur_delta,
"temperature_min": min(temps),
"temperature_avg": sum(temps) / len(temps),
"temperature_max": max(temps),
"sample_count": len(hour_samples),
"coverage": 1.0 if split.seconds_unknown == 0 else (3600 - split.seconds_unknown) / 3600,
}
def _diff_hourly_jsonl(
hourly_path: Path,
derived_hours: Dict[str, Dict[str, Any]],
) -> None:
"""Diff hourly.jsonl against derived data and log mismatches.
Spec §3.5.3: hourly.jsonl is never trusted; mismatches are diffed and logged.
"""
if not hourly_path.exists():
logger.info("No hourly.jsonl found for diffing")
return
try:
with open(hourly_path, "r") as f:
for line_num, line in enumerate(f, 1):
line = line.strip()
if not line:
continue
try:
record = json.loads(line)
except json.JSONDecodeError:
logger.warning("Malformed hourly.jsonl at line %d", line_num)
continue
hour_key = record.get("hour")
if hour_key not in derived_hours:
logger.info("Hourly.jsonl has hour %s not in derived data", hour_key)
continue
derived = derived_hours[hour_key]
mismatches = []
for field in ["bytes_written_delta", "bytes_read_delta", "sample_count"]:
if field in record and record[field] != derived.get(field):
mismatches.append(
f"{field}: hourly={record[field]} derived={derived.get(field)}"
)
if mismatches:
logger.info(
"Hourly.jsonl mismatch for %s: %s",
hour_key,
"; ".join(mismatches),
)
except Exception as e:
logger.warning("Failed to diff hourly.jsonl: %s", e)
def import_legacy_history(
conn: sqlite3.Connection,
history_path: Path,
hourly_path: Optional[Path] = None,
clock=None,
) -> Dict[str, Any]:
"""Import legacy history.jsonl into the observation store.
This is the main entry point for legacy migration. It is:
- Idempotent: second run no-ops on the legacy-import marker
- Interruption-safe: single transaction
- Never creates synthetic baselines
Args:
conn: Connection to the observation store
history_path: Path to history.jsonl
hourly_path: Optional path to hourly.jsonl for diffing
clock: Injected clock (for testing)
Returns:
Dict with migration outcome
"""
# Check idempotency
if is_legacy_imported(conn):
return {"ok": True, "skipped": True, "reason": "already_imported"}
# Read and parse history.jsonl
samples = []
malformed_count = 0
if not history_path.exists():
return {"ok": False, "error": f"History file not found: {history_path}"}
with open(history_path, "r") as f:
for line_num, line in enumerate(f, 1):
record = _parse_history_line(line, line_num)
if record is None:
malformed_count += 1
continue
samples.append(_record_to_sample(record))
if malformed_count > 0:
logger.warning("Quarantined %d malformed lines from history.jsonl", malformed_count)
if not samples:
return {"ok": False, "error": "No valid samples found in history.jsonl"}
# Sort samples by timestamp
samples.sort(key=lambda x: x["ts"])
# Derive hour observations
hour_observations = {}
for sample in samples:
ts = datetime.fromisoformat(sample["ts"])
hour_start = ts.replace(minute=0, second=0, microsecond=0)
hour_key = hour_start.strftime("%Y-%m-%dT%H:00:00Z")
if hour_key not in hour_observations:
hour_observations[hour_key] = {
"hour": hour_start,
"samples": [],
}
hour_observations[hour_key]["samples"].append(sample)
derived_hours = {}
for hour_key, hour_data in hour_observations.items():
obs = _derive_hour_observation(hour_data["samples"], hour_data["hour"])
if obs is not None:
derived_hours[hour_key] = obs
# Diff against hourly.jsonl if provided
if hourly_path:
_diff_hourly_jsonl(hourly_path, derived_hours)
# Single transaction for the entire import
try:
# Begin transaction
conn.execute("BEGIN IMMEDIATE")
# 1. Insert raw samples
for sample in samples:
conn.execute(
"""
INSERT INTO samples (
ts, device, subnqn, sn, mn, fr, capacity_bytes,
percentage_used, available_spare, media_errors, power_on_hours,
power_cycles, unsafe_shutdowns, temperature_c,
data_units_written, data_units_read, bytes_written, bytes_read,
critical_warning
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
sample["ts"],
sample["device"],
sample["subnqn"],
sample["sn"],
sample["mn"],
sample["fr"],
sample["capacity_bytes"],
sample["percentage_used"],
sample["available_spare"],
sample["media_errors"],
sample["power_on_hours"],
sample["power_cycles"],
sample["unsafe_shutdowns"],
sample["temperature_c"],
sample["data_units_written"],
sample["data_units_read"],
sample["bytes_written"],
sample["bytes_read"],
sample["critical_warning"],
),
)
# 2. Insert derived hour observations
for hour_key, obs in sorted(derived_hours.items()):
conn.execute(
"""
INSERT OR REPLACE INTO hour_observations (
hour, active_seconds, idle_seconds, powered_off_seconds,
unknown_seconds, bytes_written_delta, bytes_read_delta,
temperature_min, temperature_avg, temperature_max,
sample_count, coverage
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
obs["hour"],
obs["active_seconds"],
obs["idle_seconds"],
obs["powered_off_seconds"],
obs["unknown_seconds"],
obs["bytes_written_delta"],
obs["bytes_read_delta"],
obs["temperature_min"],
obs["temperature_avg"],
obs["temperature_max"],
obs["sample_count"],
obs["coverage"],
),
)
# 3. Open implicit monitoring period at first legacy sample
first_sample_ts = datetime.fromisoformat(samples[0]["ts"])
conn.execute(
"INSERT INTO monitoring_periods (started_at) VALUES (?)",
(first_sample_ts.isoformat(),),
)
# 4. Close period with end_cause = migrated at migration moment
if clock:
migration_time = clock.utcnow()
else:
migration_time = datetime.now(timezone.utc)
conn.execute(
"UPDATE monitoring_periods SET ended_at = ?, end_cause = ? WHERE ended_at IS NULL",
(migration_time.isoformat(), "migrated"),
)
# 5. Open legacy controller segment (mn-only)
# Legacy identity is model-scoped only (§4.4)
first_sample = samples[0]
legacy_identity = {
"mn": first_sample["mn"],
"sn": "", # Legacy segments are mn-only
"subnqn": "",
"fr": "",
}
legacy_identity_key = f"legacy|{first_sample['mn']}"
open_segment(
conn,
migration_time,
legacy_identity,
legacy_identity_key,
identity_degraded=False,
)
# 6. Set legacy import marker
_ensure_metadata_table(conn)
conn.execute(
"INSERT OR REPLACE INTO store_metadata (key, value) VALUES (?, ?)",
(LEGACY_IMPORT_MARKER, "true"),
)
# Commit
conn.commit()
except Exception as e:
conn.rollback()
raise RuntimeError(f"Migration failed: {e}") from e
# 7. Rename legacy files to *.migrated (only after commit)
try:
migrated_path = history_path.with_suffix(history_path.suffix + ".migrated")
history_path.rename(migrated_path)
logger.info("Renamed %s to %s", history_path, migrated_path)
if hourly_path and hourly_path.exists():
hourly_migrated = hourly_path.with_suffix(hourly_path.suffix + ".migrated")
hourly_path.rename(hourly_migrated)
logger.info("Renamed %s to %s", hourly_path, hourly_migrated)
except Exception as e:
# Non-fatal: files weren't renamed but migration succeeded
logger.warning("Failed to rename legacy files: %s", e)
return {
"ok": True,
"skipped": False,
"samples_imported": len(samples),
"hours_imported": len(derived_hours),
"malformed_lines": malformed_count,
"first_sample": samples[0]["ts"],
"last_sample": samples[-1]["ts"],
"legacy_identity_key": legacy_identity_key,
}
File diff suppressed because it is too large Load Diff
+243
View File
@@ -0,0 +1,243 @@
#!/usr/bin/env python3
"""fenris-monitor: privileged helper for toggle, collect, and baseline operations.
This binary is the ONLY sanctioned control path for:
- enable/disable (toggle) with monitoring-period bookkeeping
- on-demand collection trigger
- baseline set/clear persistence
Polkit authorizes this binary under com.bongbetic.fenris.monitor (auth_admin).
Spec: §8.4, §8.5, §8.6, §8.7
When run as a script, uses the fenris package from the installed wheel.
"""
import argparse
import json
import os
import sys
from datetime import datetime, timezone
from pathlib import Path
# Package dependencies are vendored independently of the host Python minor
# version. Keep the venv fallback for the legacy development install.
VENV_DIR = Path("/opt/fenris")
if VENV_DIR.exists():
vendor_dir = VENV_DIR / "vendor"
if vendor_dir.is_dir():
sys.path.insert(0, str(vendor_dir))
else:
site_packages = next((VENV_DIR / "lib").glob("python*/site-packages"), None)
if site_packages:
sys.path.insert(0, str(site_packages))
from fenris.store import init_store, get_store_path
from fenris.monitoring_periods import (
ensure_period_open,
close_period,
get_open_period,
)
from fenris.init_system import (
enable_timer,
disable_timer,
collect_now,
)
DEFAULT_STORE_PATH = Path("/var/lib/fenris/observations.db")
def is_root() -> bool:
"""Check if running as root."""
return os.geteuid() == 0
def cmd_enable(args: argparse.Namespace) -> None:
"""Enable monitoring: enable timer + open monitoring period.
Idempotent matrix (§8.6):
- First-ever enable: opens a period at the enable moment
- Resume with open period: no-op (gap stays inside as unknown)
- Resume with no open period: opens a new row
"""
store_path = getattr(args, 'store_path', DEFAULT_STORE_PATH)
conn = init_store(store_path)
now = datetime.now(timezone.utc)
try:
# Open monitoring period if none exists (§8.6)
open_period = get_open_period(conn)
if open_period is None:
ensure_period_open(conn, now)
conn.commit()
print("Monitoring period opened at", now.isoformat())
else:
print("Monitoring period already open (id=%d)" % open_period["id"])
# Enable the timer (init-system aware)
enable_timer(args.now)
finally:
conn.close()
def cmd_disable(args: argparse.Namespace) -> None:
"""Disable monitoring: disable timer + close monitoring period.
Idempotent matrix (§8.6):
- Pause with open period: closes it user_disabled
- Pause otherwise: no-op
"""
store_path = getattr(args, 'store_path', DEFAULT_STORE_PATH)
if not store_path.exists():
print("Error: Observation store not found at", store_path, file=sys.stderr)
sys.exit(1)
conn = init_store(store_path)
now = datetime.now(timezone.utc)
try:
# Close monitoring period if open (§8.6)
open_period = get_open_period(conn)
if open_period is not None:
close_period(conn, now, "user_disabled")
print("Monitoring period closed (id=%d)" % open_period["id"])
else:
print("No open monitoring period (no-op)")
# Disable the timer (init-system aware)
disable_timer(args.now)
finally:
conn.close()
def cmd_collect(args: argparse.Namespace) -> None:
"""Trigger on-demand collection.
Starts fenris-collect, blocks until exit, reports outcome.
Spec §8.7: fenris sample routes through fenris-monitor → collect,
which blocks until the collection exits; outcome reported synchronously.
"""
collect_now()
def cmd_baseline_set(args: argparse.Namespace) -> None:
"""Persist a baseline row after CLI-side validation.
Spec §8.4: fenris-monitor persists CLI-validated baseline rows.
Spec PR-14: baseline set persists through polkit-guarded helper.
"""
store_path = getattr(args, 'store_path', DEFAULT_STORE_PATH)
if not store_path.exists():
print("Error: Observation store not found at", store_path, file=sys.stderr)
sys.exit(1)
conn = init_store(store_path)
now = datetime.now(timezone.utc)
try:
# Parse and validate baseline data
data = json.loads(args.baseline_json)
required_fields = [
"tbw_terabytes",
"source_url",
"document_revision",
"entry_date",
"model_string",
"nominal_capacity_bytes",
]
for field in required_fields:
if field not in data:
print(f"Error: Missing required field: {field}", file=sys.stderr)
sys.exit(1)
# One active row replaced on edit (§6.2)
conn.execute("DELETE FROM endurance_baseline")
conn.execute(
"""
INSERT INTO endurance_baseline (
tbw_terabytes, source_url, document_revision,
entry_date, model_string, nominal_capacity_bytes,
validated_by, verified, created_at, updated_at
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
data["tbw_terabytes"],
data["source_url"],
data["document_revision"],
data["entry_date"],
data["model_string"],
data["nominal_capacity_bytes"],
data.get("validated_by", "user"),
data.get("verified", False),
now.isoformat(),
now.isoformat(),
),
)
conn.commit()
print("Baseline persisted")
finally:
conn.close()
def cmd_baseline_clear(args: argparse.Namespace) -> None:
"""Clear the endurance baseline.
Spec PR-14: baseline clear persists through polkit-guarded helper.
"""
store_path = getattr(args, 'store_path', DEFAULT_STORE_PATH)
if not store_path.exists():
print("Error: Observation store not found at", store_path, file=sys.stderr)
sys.exit(1)
conn = init_store(store_path)
try:
conn.execute("DELETE FROM endurance_baseline")
conn.commit()
print("Baseline cleared")
finally:
conn.close()
def main() -> None:
parser = argparse.ArgumentParser(
prog="fenris-monitor",
description="Fenris privileged helper for toggle, collect, and baseline operations.",
)
subparsers = parser.add_subparsers(dest="command", required=True)
# enable/disable
enable_parser = subparsers.add_parser("enable", help="Enable monitoring")
enable_parser.add_argument(
"--now", action="store_true", help="Also start the timer immediately"
)
enable_parser.set_defaults(func=cmd_enable)
disable_parser = subparsers.add_parser("disable", help="Disable monitoring")
disable_parser.add_argument(
"--now", action="store_true", help="Also stop the timer immediately"
)
disable_parser.set_defaults(func=cmd_disable)
# collect
collect_parser = subparsers.add_parser("collect", help="Trigger on-demand collection")
collect_parser.set_defaults(func=cmd_collect)
# baseline
baseline_parser = subparsers.add_parser("baseline", help="Baseline operations")
baseline_sub = baseline_parser.add_subparsers(dest="baseline_action", required=True)
baseline_set = baseline_sub.add_parser("set", help="Persist baseline")
baseline_set.add_argument("baseline_json", help="Baseline JSON data")
baseline_set.set_defaults(func=cmd_baseline_set)
baseline_clear = baseline_sub.add_parser("clear", help="Clear baseline")
baseline_clear.set_defaults(func=cmd_baseline_clear)
args = parser.parse_args()
args.func(args)
if __name__ == "__main__":
main()
+163
View File
@@ -0,0 +1,163 @@
"""Monitoring period bookkeeping per spec §5.2, §8.6, §9.8.
A monitoring period is a span during which Fenris monitoring is enabled.
Powered-off time stays inside a period; deliberately disabled time does not.
Key contracts:
- Run finding no open period opens one at the run moment, never backdated (§9.8)
- Wall-clock outside periods excluded from numerator and denominator (§5.2)
- End causes: user_disabled, migrated, unknown_gap
"""
import sqlite3
from datetime import datetime, timezone
def ensure_period_open(conn: sqlite3.Connection, run_time: datetime) -> None:
"""Ensure a monitoring period is open. If none exists, open one at run_time.
Spec §9.8: A collection run finding no open monitoring period opens one
at the run moment, never backdated. Caller owns the transaction.
"""
if get_open_period(conn) is not None:
return # Already open — no-op
ts = run_time.isoformat()
conn.execute(
"INSERT INTO monitoring_periods (started_at) VALUES (?)",
(ts,),
)
def close_period(
conn: sqlite3.Connection,
closed_at: datetime,
end_cause: str,
) -> None:
"""Close the current open monitoring period.
Spec §8.6: Pause with an open period closes it user_disabled.
If no period is open, this is a no-op (pause otherwise).
"""
open_period = get_open_period(conn)
if open_period is None:
return # No-op
ts = closed_at.isoformat()
conn.execute(
"UPDATE monitoring_periods SET ended_at = ?, end_cause = ? WHERE id = ?",
(ts, end_cause, open_period["id"]),
)
conn.commit()
def get_open_period(conn: sqlite3.Connection) -> dict | None:
"""Return the currently open monitoring period, or None."""
cursor = conn.execute(
"SELECT id, started_at, ended_at, end_cause "
"FROM monitoring_periods WHERE ended_at IS NULL LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0],
"started_at": row[1],
"ended_at": row[2],
"end_cause": row[3],
}
def is_inside_period(conn: sqlite3.Connection, ts: datetime) -> bool:
"""Check if a timestamp falls inside any monitoring period.
Spec §5.2: Wall-clock outside periods is excluded from numerator/denominator.
"""
ts_str = ts.isoformat()
cursor = conn.execute(
"SELECT 1 FROM monitoring_periods "
"WHERE started_at <= ? AND (ended_at IS NULL OR ended_at > ?) "
"LIMIT 1",
(ts_str, ts_str),
)
return cursor.fetchone() is not None
def interval_within_one_monitoring_period(
conn: sqlite3.Connection,
start: str | datetime,
end: str | datetime,
) -> bool:
"""Return whether one monitoring period contains the full interval.
Compare timestamps as UTC instants. Malformed timestamps fail closed.
"""
try:
interval_start = as_utc_datetime(start)
interval_end = as_utc_datetime(end)
except (TypeError, ValueError, OverflowError):
return False
if interval_end <= interval_start:
return False
for started_at, ended_at in conn.execute(
"SELECT started_at, ended_at FROM monitoring_periods"
):
try:
period_start = as_utc_datetime(started_at)
period_end = as_utc_datetime(ended_at) if ended_at is not None else None
except (TypeError, ValueError, OverflowError):
continue
if period_start <= interval_start and (
period_end is None or interval_end <= period_end
):
return True
return False
def as_utc_datetime(value: str | datetime) -> datetime:
"""Parse a timestamp and normalize it to an aware UTC datetime."""
parsed = value if isinstance(value, datetime) else datetime.fromisoformat(value)
if parsed.tzinfo is None:
parsed = parsed.replace(tzinfo=timezone.utc)
return parsed.astimezone(timezone.utc)
def wall_clock_in_periods(
conn: sqlite3.Connection,
start: datetime,
end: datetime,
) -> int:
"""Compute total wall-clock seconds between start and end that fall inside
any monitoring period.
Used for denominator computation (§5.2).
"""
start_str = start.isoformat()
end_str = end.isoformat()
cursor = conn.execute(
"SELECT started_at, ended_at FROM monitoring_periods "
"WHERE ended_at IS NULL OR ended_at > ? "
"ORDER BY started_at",
(start_str,),
)
total = 0
for row in cursor.fetchall():
period_start = row[0]
period_end = row[1] # None if open
# Clip period to [start, end]
effective_start = max(period_start, start_str)
if period_end is not None:
effective_end = min(period_end, end_str)
else:
effective_end = end_str
if effective_start < effective_end:
# Parse for arithmetic
s = datetime.fromisoformat(effective_start)
e = datetime.fromisoformat(effective_end)
total += int((e - s).total_seconds())
return total
+98
View File
@@ -0,0 +1,98 @@
"""User-scoped TUI preferences (issue #80).
Persists theme preset and reduced-motion choice per unprivileged user.
Preferences live at XDG_CONFIG_HOME/fenris/preferences.json and must not
affect collection, projection, history evidence, helper state, package
config, or CLI status.
Safe failures: invalid/unreadable/unwritable data never crashes the
dashboard, corrupts previous preferences, or affects monitoring.
Failures are understandable rather than silently implying persistence
succeeded.
Criteria: TPH-10, AC80-2, AC80-3.
"""
import json
import os
from pathlib import Path
from typing import Any, Dict
PREFERENCE_FILE_NAME = "preferences.json"
VALID_THEMES = {"chalktone", "amber", "nord", "high_contrast"}
DEFAULT_THEME = "chalktone"
DEFAULT_REDUCED_MOTION = False
def get_preference_path() -> Path:
"""Return the user-scoped preference file path.
Uses XDG_CONFIG_HOME/fenris/preferences.json.
Falls back to ~/.config/fenris/preferences.json if unset.
"""
xdg = os.environ.get("XDG_CONFIG_HOME")
if xdg:
base = Path(xdg)
else:
base = Path.home() / ".config"
return base / "fenris" / PREFERENCE_FILE_NAME
def load_preferences() -> Dict[str, Any]:
"""Load user preferences with safe defaults.
Returns a dict with keys:
theme: str (one of VALID_THEMES)
reduced_motion: bool
If the file is missing, corrupt, unreadable, or contains invalid
values, returns safe defaults (Chalktone theme, normal motion).
"""
path = get_preference_path()
try:
text = path.read_text()
except (OSError, FileNotFoundError):
return _defaults()
try:
data = json.loads(text)
except (json.JSONDecodeError, ValueError):
return _defaults()
if not isinstance(data, dict):
return _defaults()
theme = data.get("theme", DEFAULT_THEME)
if theme not in VALID_THEMES:
theme = DEFAULT_THEME
reduced_motion = data.get("reduced_motion", DEFAULT_REDUCED_MOTION)
if not isinstance(reduced_motion, bool):
reduced_motion = DEFAULT_REDUCED_MOTION
return {"theme": theme, "reduced_motion": reduced_motion}
def save_preferences(theme: str = DEFAULT_THEME,
reduced_motion: bool = DEFAULT_REDUCED_MOTION) -> None:
"""Save user preferences.
Creates the config directory if needed. If the write fails
(read-only filesystem, permissions), the failure is swallowed —
the TUI continues with whatever was loaded, and the user sees
no crash or error.
"""
path = get_preference_path()
try:
path.parent.mkdir(parents=True, exist_ok=True)
payload = json.dumps({"theme": theme, "reduced_motion": reduced_motion}, indent=2)
path.write_text(payload + "\n")
except (OSError, PermissionError):
# Best-effort persistence — failure must not crash the TUI
pass
def _defaults() -> Dict[str, Any]:
return {"theme": DEFAULT_THEME, "reduced_motion": DEFAULT_REDUCED_MOTION}
+668
View File
@@ -0,0 +1,668 @@
"""Projection core: the pure-function read path (spec §6).
Recomputes the complete projection contract on every read, never stores
anything derived. Takes a read-only observation store connection and an
injected clock; returns a ProjectionResult with confidence state,
contributing facts, headline remaining time (when one exists), scenario
range, Percentage-Used context line, and disclosure text.
Baseline precedence (§6.1):
verified override → unverified override → implied → unavailable
Confidence rule table (§6.7):
Supported — all conjuncts satisfied
Limited — baseline + positive rate, failing facts shown
Unavailable — no applicable baseline / zero rate / identity change
Arithmetic (§6.3):
rate = regime DUW bytes / in-period wall-clock seconds
projected = max(E_baseline − W_t, 0) / rate (rate > 0)
E_rated = entered_TBW × 10¹² bytes
E_implied = 100 · W_t / p (1 ≤ p ≤ 254)
Criteria: PR-1–PR-17, CI-4.
"""
import sqlite3
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
from enum import Enum
from typing import Any, Dict, List, Optional, Tuple
from .monitoring_periods import interval_within_one_monitoring_period
# ---------------------------------------------------------------------------
# Constants (spec §6)
# ---------------------------------------------------------------------------
HORIZON_DAYS = (7, 28, 90)
TBW_TO_BYTES = 10 ** 12
IMPLIED_P_MIN = 1
IMPLIED_P_MAX = 254
IMPLIED_MIN_PU_INCREMENTS = 2
WARMING_MIN_DAYS = 14
WARMING_MAX_LOW_COVERAGE = 2
WARMING_COVERAGE_FLOOR = 0.50
SUPPORTED_COVERAGE_FLOOR = 0.80
HORIZON_AGREEMENT_FACTOR = 2
BURST_GUARD_FRACTION = 0.50
BURST_GUARD_LOOKBACK = 28
YOUNG_REGIME_DAYS = 7
HABIT_CHANGE_SHORT_WINDOW = 7
HABIT_CHANGE_LONG_WINDOW = 28
HABIT_CHANGE_UPPER_FACTOR = 2
HABIT_CHANGE_LOWER_FACTOR = 0.5
HABIT_CHANGE_CONSECUTIVE_DAYS = 3
STALENESS_HOURS = 48
WEAR_DISAGREEMENT_FACTOR = 2
class ConfidenceState(Enum):
UNSUPPORTED = "Unavailable"
LIMITED = "Limited"
SUPPORTED = "Supported"
class BaselineTier(Enum):
VERIFIED = "verified_override"
UNVERIFIED = "unverified_override"
IMPLIED = "implied"
NONE = "none"
@dataclass(frozen=True)
class ScenarioRange:
rates: Dict[int, float]
min_days: int
max_days: int
horizon_reasons: Dict[int, str] = field(default_factory=dict)
@dataclass(frozen=True)
class ProjectionResult:
confidence_state: ConfidenceState
contributing_facts: List[str]
headline_remaining_seconds: Optional[float]
scenario_range: Optional[ScenarioRange]
pu_context_line: str
disclosure_text: List[str]
baseline_tier: BaselineTier
baseline_label: str
regime_days: Optional[int]
habit_change_fact: Optional[str]
warming_fact: Optional[str]
staleness_fact: Optional[str]
degraded_identity_fact: Optional[str]
zero_rate_fact: Optional[str]
qualifying_days_progress: Optional[str] = None # Issue #77: honest qualifying-day progress
# ---------------------------------------------------------------------------
# Store queries
# ---------------------------------------------------------------------------
def _get_baseline(conn: sqlite3.Connection) -> Optional[Dict[str, Any]]:
cursor = conn.execute(
"SELECT id, tbw_terabytes, source_url, document_revision, entry_date, "
" model_string, nominal_capacity_bytes, validated_by, verified "
"FROM endurance_baseline LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0], "tbw_terabytes": row[1], "source_url": row[2],
"document_revision": row[3], "entry_date": row[4],
"model_string": row[5], "nominal_capacity_bytes": row[6],
"validated_by": row[7], "verified": bool(row[8]),
}
def _get_current_segment(conn: sqlite3.Connection) -> Optional[Dict[str, Any]]:
cursor = conn.execute(
"SELECT id, opened_at, identity_key, identity_degraded, mn "
"FROM controller_segments ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0], "opened_at": row[1], "identity_key": row[2],
"identity_degraded": bool(row[3]), "mn": row[4],
}
def _get_days_in_segment(conn, segment_opened_at):
cursor = conn.execute(
"SELECT day, bytes_written_delta, coverage, sample_count "
"FROM day_aggregates WHERE day >= ? ORDER BY day",
(segment_opened_at[:10],),
)
return [{"day": r[0], "bytes_written": r[1], "coverage": r[2], "sample_count": r[3]}
for r in cursor.fetchall()]
def _get_all_days(conn):
cursor = conn.execute(
"SELECT day, bytes_written_delta, coverage, sample_count "
"FROM day_aggregates ORDER BY day"
)
return [{"day": r[0], "bytes_written": r[1], "coverage": r[2], "sample_count": r[3]}
for r in cursor.fetchall()]
def _get_latest_pu(conn):
cursor = conn.execute("SELECT percentage_used FROM samples ORDER BY id DESC LIMIT 1")
row = cursor.fetchone()
return row[0] if row else None
def _get_pu_increments_in_segment(conn, segment_opened_at):
cursor = conn.execute(
"SELECT COUNT(DISTINCT percentage_used) FROM samples WHERE ts >= ?",
(segment_opened_at,),
)
row = cursor.fetchone()
return max(0, (row[0] if row else 0) - 1)
def _wall_clock_in_range(conn, start, end):
start_str = start.isoformat()
end_str = end.isoformat()
cursor = conn.execute(
"SELECT started_at, ended_at FROM monitoring_periods "
"WHERE (ended_at IS NULL OR ended_at > ?) AND started_at < ? "
"ORDER BY started_at", (start_str, end_str),
)
total = 0
for row in cursor.fetchall():
eff_start = max(row[0], start_str)
eff_end = min(row[1], end_str) if row[1] is not None else end_str
if eff_start < eff_end:
total += int((datetime.fromisoformat(eff_end) - datetime.fromisoformat(eff_start)).total_seconds())
return total
# ---------------------------------------------------------------------------
# Baseline resolution (§6.1, §6.2)
# ---------------------------------------------------------------------------
def _resolve_baseline(conn, current_segment):
baseline = _get_baseline(conn)
facts: list[str] = []
if baseline is None:
return BaselineTier.NONE, None, "no baseline", facts
mandatory = [baseline["source_url"], baseline["document_revision"],
baseline["entry_date"], baseline["model_string"],
baseline["nominal_capacity_bytes"]]
provenance_complete = all(f is not None and f != "" for f in mandatory)
model_matches = True
if current_segment is not None and baseline["model_string"] is not None:
seg_mn = (current_segment.get("mn") or "").lower()
bl_model = (baseline["model_string"] or "").lower()
model_matches = bl_model in seg_mn or seg_mn in bl_model
if provenance_complete and model_matches and baseline["verified"]:
label = "verified manufacturer TBW (%.1f TB)" % baseline["tbw_terabytes"]
return BaselineTier.VERIFIED, baseline, label, facts
if not model_matches:
facts.append(
"baseline model '%s' does not match current drive '%s'"
" — baseline retained but not applicable"
% (baseline.get("model_string", ""),
current_segment.get("mn", "") if current_segment else "")
)
return BaselineTier.NONE, baseline, "baseline model mismatch", facts
if not provenance_complete:
label = "unverified TBW (%.1f TB) — user-supplied" % baseline["tbw_terabytes"]
return BaselineTier.UNVERIFIED, baseline, label, facts
label = "verified manufacturer TBW (%.1f TB)" % baseline["tbw_terabytes"]
return BaselineTier.VERIFIED, baseline, label, facts
# ---------------------------------------------------------------------------
# Rate computation
# ---------------------------------------------------------------------------
def _compute_regime_rate(days, conn, regime_start_day, clock_now):
regime_bytes = sum(d["bytes_written"] for d in days if d["day"] >= regime_start_day)
regime_start_dt = datetime.fromisoformat(regime_start_day + "T00:00:00+00:00")
regime_wc = _wall_clock_in_range(conn, regime_start_dt, clock_now)
if regime_wc <= 0:
return None, regime_bytes, 0
return regime_bytes / regime_wc, regime_bytes, regime_wc
def _compute_horizon_rate(days, conn, horizon_days, evidence_endpoint):
"""Compute horizon rate anchored at the latest evidence endpoint T.
Uses exact trailing horizon_days × 86400 seconds from T, not clock_now.
Reader refresh alone never moves T or dilutes rates.
Returns (rate, reason) where reason is None on success or a string
describing why the rate is unavailable.
"""
if not days:
return None, "no observation history"
# T is the latest evidence endpoint — the end of the last day aggregate
T = datetime.fromisoformat(days[-1]["day"] + "T23:59:59+00:00")
# Exact trailing start: T minus horizon_days × 86400 seconds
h_start = T - timedelta(days=horizon_days)
cutoff = h_start.strftime("%Y-%m-%d")
# History must span the full horizon — no placeholders
if days[0]["day"] > cutoff:
return None, "%d-day window starts before earliest data" % horizon_days
h_bytes = sum(d["bytes_written"] for d in days if d["day"] >= cutoff)
covered = sum(1 for d in days if d["day"] >= cutoff)
if covered == 0:
return None, "%d-day window has no data" % horizon_days
h_wc = _wall_clock_in_range(conn, h_start, T)
if h_wc <= 0:
return None, "%d-day window has no monitored wall-clock time" % horizon_days
return h_bytes / h_wc, None
# ---------------------------------------------------------------------------
# Habit change detection (§6.4)
# ---------------------------------------------------------------------------
def _detect_habit_change(days):
"""Detect habit change per spec §6.4.
Trailing 7-day mean >= 2x (or <= 0.5x) the preceding 28-day mean
for 3 consecutive days. Returns (change_day, days_since) or None.
The first divergence day is the earliest day in the consecutive run.
"""
need = HABIT_CHANGE_SHORT_WINDOW + HABIT_CHANGE_LONG_WINDOW
if len(days) < need:
return None
def _ratio_at(end_idx):
"""Compute 7-day / preceding-28-day mean ratio ending at end_idx."""
if end_idx < HABIT_CHANGE_SHORT_WINDOW - 1:
return None
se = end_idx + 1
ss = se - HABIT_CHANGE_SHORT_WINDOW
s_bytes = sum(d["bytes_written"] for d in days[ss:se])
s_mean = s_bytes / HABIT_CHANGE_SHORT_WINDOW
le = ss
ls = le - HABIT_CHANGE_LONG_WINDOW
if ls < 0:
return None
l_bytes = sum(d["bytes_written"] for d in days[ls:le])
l_mean = l_bytes / HABIT_CHANGE_LONG_WINDOW
if l_mean == 0:
return None
return s_mean / l_mean
# Scan backwards from the most recent day
for i in range(len(days) - 1, HABIT_CHANGE_LONG_WINDOW + HABIT_CHANGE_SHORT_WINDOW - 2, -1):
ratio = _ratio_at(i)
if ratio is None:
continue
is_upper = ratio >= HABIT_CHANGE_UPPER_FACTOR
is_lower = ratio <= HABIT_CHANGE_LOWER_FACTOR
if not (is_upper or is_lower):
continue
# Count consecutive days going backwards from i
consecutive = 1
for j in range(i - 1, HABIT_CHANGE_LONG_WINDOW + HABIT_CHANGE_SHORT_WINDOW - 3, -1):
r = _ratio_at(j)
if r is None:
break
if (is_upper and r >= HABIT_CHANGE_UPPER_FACTOR) or \
(is_lower and r <= HABIT_CHANGE_LOWER_FACTOR):
consecutive += 1
else:
break
if consecutive >= HABIT_CHANGE_CONSECUTIVE_DAYS:
change_idx = i - consecutive + 1
change_day = days[change_idx]["day"]
days_since = (datetime.fromisoformat(days[-1]["day"]) - datetime.fromisoformat(change_day)).days
return change_day, days_since
return None
# ---------------------------------------------------------------------------
# Confidence rule table (§6.7)
# ---------------------------------------------------------------------------
def _evaluate_confidence(tier, rate, regime_days, days, current_segment,
clock_now, warming_days, warming_low_coverage,
habit_change, staleness_hours, scenario_range):
facts = []
if tier == BaselineTier.NONE:
facts.append("no applicable endurance baseline")
return ConfidenceState.UNSUPPORTED, facts
if rate is None or rate <= 0:
facts.append("no finite projection from this history")
return ConfidenceState.UNSUPPORTED, facts
supported_facts = []
failing = False
# 1. Verified baseline
if tier != BaselineTier.VERIFIED:
failing = True
else:
supported_facts.append("verified manufacturer TBW")
# 2. >= 14 qualifying days
qualifying = sum(1 for d in days if d["coverage"] >= WARMING_COVERAGE_FLOOR and d["sample_count"] > 0)
if qualifying < WARMING_MIN_DAYS:
failing = True
else:
supported_facts.append("%d calendar days" % qualifying)
# 3. Coverage >= 80%
total_wc = len(days) * 86400
total_known = sum(int(d["coverage"] * 86400) for d in days)
avg_cov = total_known / total_wc if total_wc > 0 else 0.0
if avg_cov < SUPPORTED_COVERAGE_FLOOR:
failing = True
else:
supported_facts.append("%d%% interval coverage" % int(avg_cov * 100))
# 4. Fresh (< 48h)
if staleness_hours is not None and staleness_hours > STALENESS_HOURS:
failing = True
elif staleness_hours is not None:
supported_facts.append("recent data")
# 5. Horizon agreement
if scenario_range is not None and len(scenario_range.rates) >= 2:
rl = list(scenario_range.rates.values())
if min(rl) > 0 and max(rl) / min(rl) > HORIZON_AGREEMENT_FACTOR:
failing = True
else:
supported_facts.append("%d weekly cycles" % len(scenario_range.rates))
else:
failing = True
# 6. Burst guard
if not failing and len(days) >= BURST_GUARD_LOOKBACK:
t28 = sum(d["bytes_written"] for d in days[-BURST_GUARD_LOOKBACK:])
for d in days[-BURST_GUARD_LOOKBACK:]:
if t28 > 0 and d["bytes_written"] >= BURST_GUARD_FRACTION * t28:
failing = True
break
if not failing:
supported_facts.append("no burst days")
# 7. Regime >= 7 days
if regime_days < YOUNG_REGIME_DAYS:
failing = True
# 8. Degraded identity
if current_segment and current_segment.get("identity_degraded"):
failing = True
facts.append("controller identity unavailable — replacement detection relies on write-counter continuity only")
if not failing:
return ConfidenceState.SUPPORTED, supported_facts
# Limited
limited_facts = list(supported_facts)
if staleness_hours is not None and staleness_hours > STALENESS_HOURS:
limited_facts.append("newest data %dh old (≥48h)" % staleness_hours)
if habit_change is not None:
limited_facts.append("usage habit changed %d days ago" % habit_change[1])
if regime_days < YOUNG_REGIME_DAYS:
limited_facts.append("regime only %d days old (≥7 required)" % regime_days)
if current_segment and current_segment.get("identity_degraded"):
degraded_fact = "controller identity unavailable — replacement detection relies on write-counter continuity only"
if degraded_fact not in limited_facts:
limited_facts.append(degraded_fact)
return ConfidenceState.LIMITED, limited_facts
# ---------------------------------------------------------------------------
# Disclosure text (§6.11)
# ---------------------------------------------------------------------------
DISCLOSURES = [
"This is an endurance projection, not a predicted hardware-failure date.",
("Percentage Used is vendor-specific; 100 means estimated endurance consumed "
"but may not mean failure, it can exceed 100, and 255 is saturated."),
("Rated TBW can be a warranty/endurance threshold with separate time and "
"eligibility terms, not a failure threshold."),
("DUW is upward-rounded host writes excluding metadata and selected commands, "
"not exact physical NAND writes."),
("Projection quality depends on baseline provenance, history duration and "
"completeness, recentness, stability, and representative usage cycles; "
"future workload and firmware behavior remain outside the observed evidence."),
("Gaps can preserve an aggregate counter delta without preserving hourly "
"timing; unexplained and deliberately disabled periods must be distinguished."),
]
# ---------------------------------------------------------------------------
# PU context line (§6.1)
# ---------------------------------------------------------------------------
def _build_pu_context_line(conn, rate, days, clock_now):
pu = _get_latest_pu(conn)
if pu is None:
return "Percentage Used: unknown"
if rate is None or rate <= 0 or not days:
return "Percentage Used: %d%%" % pu
total_bytes = sum(d["bytes_written"] for d in days)
if total_bytes <= 0:
return "Percentage Used: %d%%" % pu
total_days_count = len(days)
if total_days_count == 0:
return "Percentage Used: %d%%" % pu
pu_daily = total_bytes / total_days_count
obs_daily = rate * 86400
if pu_daily > 0:
ratio = obs_daily / pu_daily
if ratio > WEAR_DISAGREEMENT_FACTOR or ratio < 1.0 / WEAR_DISAGREEMENT_FACTOR:
return ("Percentage Used: %d%% — vendor wear estimate disagrees "
"with observed write rate (>2× difference)") % pu
return "Percentage Used: %d%%" % pu
# ---------------------------------------------------------------------------
# Complete observation day gate (issue #94)
# ---------------------------------------------------------------------------
def _has_complete_local_day(conn):
"""Check if at least one complete local observation day exists.
A complete local day is a full midnight-to-midnight calendar day
within a monitoring period that has usable observation evidence.
This is the prerequisite for showing an endurance outlook.
"""
local_days = conn.execute(
"SELECT utc_start, utc_end FROM local_days "
"WHERE complete = 1 "
"AND activity_precision IN ('measured', 'coarse') "
"AND activity_intervals > 0 ORDER BY id"
).fetchall()
return any(
interval_within_one_monitoring_period(conn, utc_start, utc_end)
for utc_start, utc_end in local_days
)
# ---------------------------------------------------------------------------
# Main projection function
# ---------------------------------------------------------------------------
def compute_projection(conn, clock_now):
facts = []
habit_change_fact = None
warming_fact = None
staleness_fact = None
degraded_identity_fact = None
zero_rate_fact = None
current_segment = _get_current_segment(conn)
tier, baseline, baseline_label, baseline_facts = _resolve_baseline(conn, current_segment)
facts.extend(baseline_facts)
# --- Complete observation day gate (issue #94) ---
# An endurance outlook requires at least one complete local
# midnight-to-midnight calendar day with usable observation evidence.
has_complete_day = _has_complete_local_day(conn)
if not has_complete_day:
facts.append("waiting for a full local observation day")
return ProjectionResult(
confidence_state=ConfidenceState.UNSUPPORTED,
contributing_facts=facts,
headline_remaining_seconds=None,
scenario_range=None,
pu_context_line="Percentage Used: unknown" if _get_latest_pu(conn) is None else "Percentage Used: %d%%" % (_get_latest_pu(conn) or 0),
disclosure_text=list(DISCLOSURES),
baseline_tier=tier,
baseline_label=baseline_label,
regime_days=None,
habit_change_fact=None,
warming_fact=None,
staleness_fact=None,
degraded_identity_fact=None,
zero_rate_fact=None,
qualifying_days_progress=None,
)
segment_days = _get_days_in_segment(conn, current_segment["opened_at"]) if current_segment else _get_all_days(conn)
all_days = _get_all_days(conn)
regime_start_day = None
habit_change = None
if segment_days:
earliest = segment_days[0]["day"]
cutoff_90 = (clock_now - timedelta(days=90)).strftime("%Y-%m-%d")
regime_start_day = max(earliest, cutoff_90)
habit_change = _detect_habit_change(segment_days)
if habit_change is not None:
regime_start_day = habit_change[0]
habit_change_fact = "usage habit changed %d days ago" % habit_change[1]
facts.append(habit_change_fact)
rate = None
regime_bytes = 0
regime_days_count = 0
if segment_days and regime_start_day is not None:
rate, regime_bytes, _ = _compute_regime_rate(segment_days, conn, regime_start_day, clock_now)
regime_days_count = sum(1 for d in segment_days if d["day"] >= regime_start_day)
if rate is not None and rate <= 0:
zero_rate_fact = "no finite projection from this history"
facts.append(zero_rate_fact)
scenario = None
horizon_rates = {}
horizon_reasons = {}
for h in HORIZON_DAYS:
hr, reason = _compute_horizon_rate(all_days, conn, h, clock_now)
if hr is not None:
horizon_rates[h] = hr
else:
horizon_reasons[h] = reason
if horizon_rates:
scenario = ScenarioRange(
rates=horizon_rates,
min_days=min(horizon_rates),
max_days=max(horizon_rates),
horizon_reasons=horizon_reasons,
)
total_days_count = len(segment_days)
days_below_coverage = sum(1 for d in segment_days
if d["coverage"] < WARMING_COVERAGE_FLOOR or d["sample_count"] == 0)
qualifying = total_days_count - days_below_coverage
# Issue #77: Show honest qualifying-day progress
if total_days_count >= WARMING_MIN_DAYS:
# After warm-up, show qualifying day details for transparency
qualifying_days_progress = "%d of %d qualifying days" % (qualifying, total_days_count)
if days_below_coverage > 0:
qualifying_days_progress += " (%d below coverage)" % days_below_coverage
else:
qualifying_days_progress = None
if total_days_count < WARMING_MIN_DAYS or days_below_coverage > WARMING_MAX_LOW_COVERAGE:
warming_fact = "warming up: %d of %d qualifying days" % (qualifying, WARMING_MIN_DAYS)
facts.append(warming_fact)
staleness_hours = None
if segment_days:
newest_dt = datetime.fromisoformat(segment_days[-1]["day"] + "T12:00:00+00:00")
staleness_hours = int((clock_now - newest_dt).total_seconds() / 3600)
if staleness_hours > STALENESS_HOURS:
staleness_fact = "newest data %dh old (≥48h)" % staleness_hours
facts.append(staleness_fact)
if current_segment and current_segment.get("identity_degraded"):
degraded_identity_fact = "controller identity unavailable — replacement detection relies on write-counter continuity only"
facts.append(degraded_identity_fact)
state, conf_facts = _evaluate_confidence(
tier, rate, regime_days_count, segment_days, current_segment, clock_now,
qualifying, 0, habit_change, staleness_hours, scenario,
)
all_facts = list(facts)
for cf in conf_facts:
if cf not in all_facts:
all_facts.append(cf)
headline_seconds = None
if state != ConfidenceState.UNSUPPORTED and rate is not None and rate > 0 and baseline is not None:
if tier in (BaselineTier.VERIFIED, BaselineTier.UNVERIFIED):
E_baseline = baseline["tbw_terabytes"] * TBW_TO_BYTES
elif tier == BaselineTier.IMPLIED:
p = _get_latest_pu(conn)
if p is not None and IMPLIED_P_MIN <= p <= IMPLIED_P_MAX:
E_baseline = 100 * regime_bytes / p
else:
E_baseline = None
else:
E_baseline = None
if E_baseline is not None:
headline_seconds = max(E_baseline - regime_bytes, 0) / rate
pu_line = _build_pu_context_line(conn, rate, segment_days, clock_now)
return ProjectionResult(
confidence_state=state,
contributing_facts=all_facts,
headline_remaining_seconds=headline_seconds,
scenario_range=scenario,
pu_context_line=pu_line,
disclosure_text=list(DISCLOSURES),
baseline_tier=tier,
baseline_label=baseline_label,
regime_days=regime_days_count,
habit_change_fact=habit_change_fact,
warming_fact=warming_fact,
staleness_fact=staleness_fact,
degraded_identity_fact=degraded_identity_fact,
zero_rate_fact=zero_rate_fact,
qualifying_days_progress=qualifying_days_progress,
)
+372
View File
@@ -0,0 +1,372 @@
"""Expire detail only after UTC and local_days replacement evidence is durable.
The 14-day cutoff never overrides local-day preservation, shared boundary
evidence, publication recovery, or the newest sample's successor-anchor role.
"""
import sqlite3
from datetime import date, datetime, timedelta
from .monitoring_periods import (
as_utc_datetime,
interval_within_one_monitoring_period,
)
RAW_SAMPLE_RETENTION_DAYS = 14
def _parse_sample_time(value: str) -> datetime | None:
"""Parse old timestamps conservatively; malformed legacy values stay."""
try:
parsed = datetime.fromisoformat(value)
except (TypeError, ValueError):
return None
return as_utc_datetime(parsed)
def _sample_rows(conn: sqlite3.Connection) -> list[tuple]:
return conn.execute(
"SELECT id, ts, segment_id, bytes_written, bytes_read, local_tz "
"FROM samples ORDER BY id"
).fetchall()
def _required_hour_window(
start: datetime,
end: datetime,
) -> tuple[datetime, datetime, int] | None:
"""Return first hour, exclusive end and count for [start, end)."""
if end <= start:
return None
first_hour = start.replace(minute=0, second=0, microsecond=0)
last_hour = end.replace(minute=0, second=0, microsecond=0)
try:
exclusive_end = last_hour if end == last_hour else last_hour + timedelta(hours=1)
except OverflowError:
return None
count = int((exclusive_end - first_hour).total_seconds() // 3600)
return first_hour, exclusive_end, count
def _has_utc_replacement(
conn: sqlite3.Connection,
previous: tuple,
current: tuple,
start: datetime,
end: datetime,
) -> bool:
window = _required_hour_window(start, end)
if window is None:
return False
first_hour, exclusive_end, expected_hour_count = window
actual_hour_count = conn.execute(
"SELECT COUNT(*) FROM hour_observations WHERE hour >= ? AND hour < ?",
(
first_hour.strftime("%Y-%m-%dT%H:00:00+00:00"),
exclusive_end.strftime("%Y-%m-%dT%H:00:00+00:00"),
),
).fetchone()[0]
if actual_hour_count < expected_hour_count:
return False
start_written, end_written = previous[3], current[3]
start_read, end_read = previous[4], current[4]
if None in (start_written, end_written, start_read, end_read):
return False
if end_written < start_written or end_read < start_read:
return False
bytes_written = end_written - start_written
bytes_read = end_read - start_read
if expected_hour_count == 1:
totals = conn.execute(
"SELECT bytes_written_delta, bytes_read_delta FROM hour_observations "
"WHERE hour = ?",
(first_hour.strftime("%Y-%m-%dT%H:00:00+00:00"),),
).fetchone()
else:
totals = conn.execute(
"SELECT unattributed_bytes_written, unattributed_bytes_read "
"FROM day_aggregates WHERE day = ?",
(start.date().isoformat(),),
).fetchone()
if totals is None or totals[0] < bytes_written or totals[1] < bytes_read:
return False
last_day = (exclusive_end - timedelta(microseconds=1)).date()
expected_day_count = (last_day - first_hour.date()).days + 1
actual_day_count = conn.execute(
"SELECT COUNT(*) FROM day_aggregates WHERE day >= ? AND day <= ?",
(first_hour.date().isoformat(), last_day.isoformat()),
).fetchone()[0]
return actual_day_count >= expected_day_count
def _local_day_row_exists(
conn: sqlite3.Connection,
local_date: str,
tz_name: str,
) -> bool:
return conn.execute(
"SELECT 1 FROM local_days WHERE local_date = ? AND tz_name = ?",
(local_date, tz_name),
).fetchone() is not None
def _local_day_bounds(
conn: sqlite3.Connection,
local_date: str,
tz_name: str,
) -> tuple[datetime, datetime] | None:
row = conn.execute(
"SELECT utc_start, utc_end FROM local_days "
"WHERE local_date = ? AND tz_name = ?",
(local_date, tz_name),
).fetchone()
if row is None:
return None
start = _parse_sample_time(row[0])
end = _parse_sample_time(row[1])
if start is None or end is None or end <= start:
return None
return start, end
def _contains_instant(
conn: sqlite3.Connection,
local_date: str,
tz_name: str,
instant: datetime,
) -> bool:
bounds = _local_day_bounds(conn, local_date, tz_name)
return bounds is not None and bounds[0] <= instant < bounds[1]
def _monitoring_period_covers(
conn: sqlite3.Connection,
start: datetime,
end: datetime,
) -> bool:
"""Require continuous monitoring before a local-day total replaces detail."""
return interval_within_one_monitoring_period(conn, start, end)
def _has_local_replacement(
conn: sqlite3.Connection,
previous: tuple,
current: tuple,
start: datetime,
end: datetime,
) -> bool:
start_tz = previous[5]
end_tz = current[5]
if not end_tz:
return False
start_id, end_id = previous[0], current[0]
segment_id = current[2]
if (
start_tz
and start_tz == end_tz
and segment_id is not None
and _monitoring_period_covers(conn, start, end)
):
known_days = conn.execute(
"SELECT local_days.utc_start, local_days.utc_end "
"FROM local_days JOIN local_day_segment_totals "
" ON local_day_segment_totals.local_day_id = local_days.id "
"WHERE local_days.tz_name = ? "
" AND local_days.activity_precision IN ('measured', 'coarse') "
" AND local_days.last_sample_id >= ? "
" AND local_days.activity_intervals > 0 "
" AND local_day_segment_totals.segment_id = ? "
" AND local_day_segment_totals.activity_intervals > 0",
(end_tz, end_id, segment_id),
).fetchall()
for row in known_days:
bounds_start = _parse_sample_time(row[0])
bounds_end = _parse_sample_time(row[1])
if (
bounds_start is not None
and bounds_end is not None
and bounds_start <= start
and end < bounds_end
):
return True
evidence = conn.execute(
"SELECT bytes_written, bytes_read, reason, start_local_date, end_local_date, "
" start_tz_name, end_tz_name, started_at, ended_at "
"FROM local_day_unallocated_evidence "
"WHERE start_sample_id = ? AND end_sample_id = ?",
(start_id, end_id),
).fetchone()
if evidence is None:
return False
start_written, end_written = previous[3], current[3]
start_read, end_read = previous[4], current[4]
if None in (start_written, end_written, start_read, end_read):
return False
if end_written < start_written or end_read < start_read:
return False
# The source IDs uniquely identify this interval. Confirm its durable
# deltas and retain every recorded local-day boundary it touches.
if (end_written - start_written, end_read - start_read) != evidence[:2]:
return False
if evidence[5] != start_tz or evidence[6] != end_tz:
return False
if evidence[2] not in {
"counter_discontinuity",
"legacy_timezone_unknown",
"timezone_change",
"monitoring_period",
"local_midnight",
"segment_unknown",
}:
return False
if _parse_sample_time(evidence[7]) != start or _parse_sample_time(evidence[8]) != end:
return False
affected_days: set[tuple[str, str]] = set()
if start_tz:
affected_days.add((evidence[3], start_tz))
affected_days.add((evidence[4], end_tz))
if start_tz and start_tz == end_tz and evidence[3] < evidence[4]:
day = date.fromisoformat(evidence[3]) + timedelta(days=1)
last_day = date.fromisoformat(evidence[4])
while day < last_day:
affected_days.add((day.isoformat(), end_tz))
day += timedelta(days=1)
if not all(
_local_day_row_exists(conn, local_date, zone)
for local_date, zone in affected_days
):
return False
if start_tz and not _contains_instant(conn, evidence[3], start_tz, start):
return False
return _contains_instant(conn, evidence[4], end_tz, end)
def _interval_has_replacement(
conn: sqlite3.Connection,
previous: tuple,
current: tuple,
) -> bool:
start = _parse_sample_time(previous[1])
end = _parse_sample_time(current[1])
if start is None or end is None or end <= start:
return False
try:
return (
_has_utc_replacement(conn, previous, current, start, end)
and _has_local_replacement(conn, previous, current, start, end)
)
except (sqlite3.Error, ValueError, OverflowError, TypeError):
# Missing or unreadable derived evidence never authorizes deletion.
return False
def _sample_needs_preservation(
conn: sqlite3.Connection,
rows: list[tuple],
index: int,
newest_id: int,
) -> bool:
sample = rows[index]
segment_id = sample[2]
if segment_id is None or conn.execute(
"SELECT 1 FROM controller_segments WHERE id = ?",
(segment_id,),
).fetchone() is None:
return True
if sample[0] == newest_id:
# Keep newest sample as the source anchor for the next collection.
return True
pairs = []
if index > 0 and rows[index - 1][2] == segment_id:
pairs.append((rows[index - 1], sample))
if index + 1 < len(rows) and rows[index + 1][2] == segment_id:
pairs.append((sample, rows[index + 1]))
return any(
not _interval_has_replacement(conn, previous, current)
for previous, current in pairs
)
def needs_boundary_anchor(
conn: sqlite3.Connection,
sample_ts: str,
now: datetime,
) -> bool:
"""Return whether an old sample still carries unreplaced evidence."""
sample_time = _parse_sample_time(sample_ts)
cutoff = as_utc_datetime(now) - timedelta(days=RAW_SAMPLE_RETENTION_DAYS)
if sample_time is None:
return True
if sample_time >= cutoff:
return False
rows = _sample_rows(conn)
matching = [index for index, row in enumerate(rows) if row[1] == sample_ts]
if not matching:
return False
newest_id = rows[-1][0] if rows else -1
return any(
_sample_needs_preservation(conn, rows, index, newest_id)
for index in matching
)
def prune_old_samples(
conn: sqlite3.Connection,
now: datetime,
retention_days: int = RAW_SAMPLE_RETENTION_DAYS,
) -> int:
"""Delete expired samples only after every dependent fact is durable.
Pending publications are stored separately and never selected here. A
sample stays when it is the newest successor anchor, has an unpublishable
neighbour interval, lacks UTC or local-day replacement evidence, or has a
timestamp that cannot be safely interpreted.
"""
cutoff = as_utc_datetime(now) - timedelta(days=retention_days)
owns_transaction = not conn.in_transaction
savepoint = "fenris_sample_retention"
if owns_transaction:
conn.execute("BEGIN IMMEDIATE")
else:
conn.execute(f"SAVEPOINT {savepoint}")
try:
rows = _sample_rows(conn)
if not rows:
if owns_transaction:
conn.commit()
else:
conn.execute(f"RELEASE SAVEPOINT {savepoint}")
return 0
newest_id = rows[-1][0]
expired = []
for index, row in enumerate(rows):
sample_time = _parse_sample_time(row[1])
if sample_time is None or sample_time >= cutoff:
continue
if _sample_needs_preservation(conn, rows, index, newest_id):
continue
expired.append(row[0])
conn.executemany("DELETE FROM samples WHERE id = ?", ((sample_id,) for sample_id in expired))
if owns_transaction:
conn.commit()
else:
conn.execute(f"RELEASE SAVEPOINT {savepoint}")
return len(expired)
except Exception:
if owns_transaction:
conn.rollback()
else:
conn.execute(f"ROLLBACK TO SAVEPOINT {savepoint}")
conn.execute(f"RELEASE SAVEPOINT {savepoint}")
raise
+448
View File
@@ -0,0 +1,448 @@
"""Historical repair and retention (issue #74, #99).
Re-derives hour observations and day aggregates from surviving raw samples,
rebuilds legacy local-day evidence from surviving samples or trustworthy UTC
hours, and preserves import markers, boundary anchors, and valid summaries.
Repair is idempotent and interruption-safe.
Contracts:
- Idempotent: running repair multiple times produces no duplicates
- Interruption-safe: partial repair preserves prior valid history
- Preserves existing valid data: never overwrites valid derived data
- Surfaces failures explicitly for retry
"""
import logging
import sqlite3
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
from typing import Optional, List, Tuple
from .derive import find_previous_sample, derive_hours_from_interval, _parse_ts
logger = logging.getLogger(__name__)
@dataclass
class RepairResult:
"""Result of a repair operation."""
ok: bool
hours_created: int = 0
hours_updated: int = 0
days_created: int = 0
days_updated: int = 0
intervals_derived: int = 0
boundary_anchors_retained: int = 0
legacy_summaries_preserved: int = 0
error: Optional[str] = None
@dataclass
class RepairStatus:
"""Current repair status for read-only views."""
last_repair: Optional[str] = None # ISO timestamp of last successful repair
repair_in_progress: bool = False
hours_derived: int = 0
days_derived: int = 0
def _ensure_repair_metadata(conn: sqlite3.Connection) -> None:
"""Ensure metadata table exists for tracking repair state."""
conn.execute("""
CREATE TABLE IF NOT EXISTS store_metadata (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
)
""")
def is_repair_in_progress(conn: sqlite3.Connection) -> bool:
"""Check if a repair operation is currently in progress."""
_ensure_repair_metadata(conn)
cursor = conn.execute(
"SELECT value FROM store_metadata WHERE key = 'repair_in_progress'"
)
row = cursor.fetchone()
return row is not None and row[0] == "true"
def _set_repair_in_progress(conn: sqlite3.Connection, in_progress: bool) -> None:
"""Mark repair as in progress or complete."""
_ensure_repair_metadata(conn)
conn.execute(
"INSERT OR REPLACE INTO store_metadata (key, value) VALUES (?, ?)",
("repair_in_progress", "true" if in_progress else "false"),
)
def _update_repair_status(conn: sqlite3.Connection, result: RepairResult) -> None:
"""Update repair status after successful completion."""
_ensure_repair_metadata(conn)
now = datetime.now(timezone.utc).isoformat()
# Update last repair timestamp
conn.execute(
"INSERT OR REPLACE INTO store_metadata (key, value) VALUES (?, ?)",
("last_repair", now),
)
# Update derived counts
cursor = conn.execute("SELECT COUNT(*) FROM hour_observations")
hours = cursor.fetchone()[0]
cursor = conn.execute("SELECT COUNT(*) FROM day_aggregates")
days = cursor.fetchone()[0]
conn.execute(
"INSERT OR REPLACE INTO store_metadata (key, value) VALUES (?, ?)",
("hours_derived", str(hours)),
)
conn.execute(
"INSERT OR REPLACE INTO store_metadata (key, value) VALUES (?, ?)",
("days_derived", str(days)),
)
def get_repair_status(conn: sqlite3.Connection) -> RepairStatus:
"""Get current repair status for read-only views."""
_ensure_repair_metadata(conn)
last_repair = None
cursor = conn.execute(
"SELECT value FROM store_metadata WHERE key = 'last_repair'"
)
row = cursor.fetchone()
if row:
last_repair = row[0]
in_progress = is_repair_in_progress(conn)
hours_derived = 0
cursor = conn.execute(
"SELECT value FROM store_metadata WHERE key = 'hours_derived'"
)
row = cursor.fetchone()
if row:
hours_derived = int(row[0])
days_derived = 0
cursor = conn.execute(
"SELECT value FROM store_metadata WHERE key = 'days_derived'"
)
row = cursor.fetchone()
if row:
days_derived = int(row[0])
return RepairStatus(
last_repair=last_repair,
repair_in_progress=in_progress,
hours_derived=hours_derived,
days_derived=days_derived,
)
def needs_boundary_anchor(
conn: sqlite3.Connection,
sample_ts: str,
now: datetime,
) -> bool:
"""Check if a sample is needed as a boundary anchor for derivation.
A sample is a boundary anchor if:
1. It's older than retention_days
2. It has no derived hour observation for its hour
3. It's the last sample before a gap that needs derivation (gap > 24 hours)
"""
from .pruning import RAW_SAMPLE_RETENTION_DAYS
sample_dt = _parse_ts(sample_ts)
retention_cutoff = now - timedelta(days=RAW_SAMPLE_RETENTION_DAYS)
# If sample is within retention, not an anchor (will be kept anyway)
if sample_dt >= retention_cutoff:
return False
# Check if this sample's hour already has a derived observation
hour_start = sample_dt.replace(minute=0, second=0, microsecond=0).isoformat()
hour_end = (sample_dt + timedelta(hours=1)).replace(minute=0, second=0, microsecond=0).isoformat()
cursor = conn.execute(
"""SELECT COUNT(*) FROM hour_observations
WHERE hour >= ? AND hour < ?""",
(hour_start, hour_end),
)
# If there's an hour observation in this sample's hour, it's been derived
if cursor.fetchone()[0] > 0:
return False
# Check if this sample is the last sample before a gap
# (i.e., the next sample is significantly later)
cursor = conn.execute(
"""SELECT ts FROM samples WHERE ts > ? ORDER BY ts LIMIT 1""",
(sample_ts,),
)
next_row = cursor.fetchone()
if next_row is None:
# No next sample - this is the last sample, might be needed
# But if it's old and fully derived, it's not needed
return False
next_ts = _parse_ts(next_row[0])
gap = (next_ts - sample_dt).total_seconds()
# If gap > 24 hours, this sample is a boundary anchor
# (needed to derive the interval spanning the gap)
return gap > 24 * 3600
def _get_unlinked_intervals(conn: sqlite3.Connection) -> List[Tuple[dict, dict]]:
"""Find sample pairs that form intervals but have no hour observations."""
cursor = conn.execute(
"""SELECT id, ts, bytes_written, bytes_read, power_on_hours,
temperature_c, data_units_written, data_units_read, segment_id
FROM samples ORDER BY ts"""
)
all_samples = []
for row in cursor.fetchall():
all_samples.append({
"id": row[0], "ts": row[1], "bytes_written": row[2],
"bytes_read": row[3], "power_on_hours": row[4],
"temperature_c": row[5], "data_units_written": row[6],
"data_units_read": row[7], "segment_id": row[8],
})
intervals = []
for i in range(len(all_samples) - 1):
prev = all_samples[i]
next_s = all_samples[i + 1]
# Skip if different segments
if prev["segment_id"] != next_s["segment_id"]:
continue
# Check if the interval spans hours that need derivation
prev_dt = _parse_ts(prev["ts"])
next_dt = _parse_ts(next_s["ts"])
# Check if any hour in the span lacks an observation
current = prev_dt.replace(minute=0, second=0, microsecond=0)
end = next_dt.replace(minute=0, second=0, microsecond=0)
needs_derivation = False
while current <= end:
cursor2 = conn.execute(
"SELECT id FROM hour_observations WHERE hour = ?",
(current.isoformat(),),
)
if cursor2.fetchone() is None:
needs_derivation = True
break
current += timedelta(hours=1)
if needs_derivation:
intervals.append((prev, next_s))
return intervals
def _derive_day_aggregate_from_hours(
conn: sqlite3.Connection,
day: str,
) -> Optional[dict]:
"""Derive a day aggregate from its hour observations."""
cursor = conn.execute(
"""SELECT SUM(active_seconds), SUM(idle_seconds),
SUM(powered_off_seconds), SUM(unknown_seconds),
SUM(bytes_written_delta), SUM(bytes_read_delta),
SUM(sample_count)
FROM hour_observations WHERE hour LIKE ?""",
(day + "T%",),
)
row = cursor.fetchone()
if row is None or row[0] is None:
return None
return {
"day": day,
"active_seconds": row[0] or 0,
"idle_seconds": row[1] or 0,
"powered_off_seconds": row[2] or 0,
"unknown_seconds": row[3] or 0,
"bytes_written_delta": row[4] or 0,
"bytes_read_delta": row[5] or 0,
"sample_count": row[6] or 0,
}
def _upsert_day_aggregate(conn: sqlite3.Connection, day_data: dict) -> bool:
"""Insert or update a day aggregate. Returns True if created."""
existing = conn.execute(
"SELECT id FROM day_aggregates WHERE day = ?",
(day_data["day"],),
).fetchone()
if existing is None:
# Calculate coverage
total_seconds = (day_data["active_seconds"] + day_data["idle_seconds"] +
day_data["powered_off_seconds"] + day_data["unknown_seconds"])
coverage = (day_data["active_seconds"] + day_data["idle_seconds"] +
day_data["powered_off_seconds"]) / total_seconds if total_seconds > 0 else 0.0
conn.execute(
"""INSERT INTO day_aggregates
(day, active_seconds, idle_seconds, powered_off_seconds, unknown_seconds,
bytes_written_delta, bytes_read_delta, sample_count, coverage)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)""",
(day_data["day"], day_data["active_seconds"], day_data["idle_seconds"],
day_data["powered_off_seconds"], day_data["unknown_seconds"],
day_data["bytes_written_delta"], day_data["bytes_read_delta"],
day_data["sample_count"], coverage),
)
return True
else:
# Update existing (but only if new data is more complete)
# This implements the "don't overwrite valid older history" rule
cursor = conn.execute(
"""SELECT bytes_written_delta, sample_count
FROM day_aggregates WHERE day = ?""",
(day_data["day"],),
)
existing_data = cursor.fetchone()
# Only update if new data has more samples or more bytes
if (day_data["sample_count"] > existing_data[1] or
day_data["bytes_written_delta"] > existing_data[0]):
total_seconds = (day_data["active_seconds"] + day_data["idle_seconds"] +
day_data["powered_off_seconds"] + day_data["unknown_seconds"])
coverage = (day_data["active_seconds"] + day_data["idle_seconds"] +
day_data["powered_off_seconds"]) / total_seconds if total_seconds > 0 else 0.0
conn.execute(
"""UPDATE day_aggregates SET
active_seconds = ?, idle_seconds = ?, powered_off_seconds = ?,
unknown_seconds = ?, bytes_written_delta = ?, bytes_read_delta = ?,
sample_count = ?, coverage = ?
WHERE day = ?""",
(day_data["active_seconds"], day_data["idle_seconds"],
day_data["powered_off_seconds"], day_data["unknown_seconds"],
day_data["bytes_written_delta"], day_data["bytes_read_delta"],
day_data["sample_count"], coverage, day_data["day"]),
)
return False # Updated, not created
else:
return False # No update needed
def repair_derivation(
conn: sqlite3.Connection,
clock=None,
) -> RepairResult:
"""Repair hour observations and day aggregates from surviving samples.
This is the main entry point for historical repair. It:
1. Finds sample pairs that need interval derivation
2. Derives hour observations from those intervals
3. Updates day aggregates from the hour observations
4. Preserves existing valid data
5. Is idempotent and interruption-safe
Args:
conn: Connection to the observation store
clock: Injected clock (for testing)
Returns:
RepairResult with operation details
"""
if clock is None:
clock = datetime.now(timezone.utc)
elif hasattr(clock, 'utcnow'):
clock = clock.utcnow()
result = RepairResult(ok=True)
try:
# A committed in-progress marker may be stale after process death.
# SQLite's write lock serializes active repairs; retry idempotently.
if conn.in_transaction:
conn.commit()
conn.execute("BEGIN IMMEDIATE")
_set_repair_in_progress(conn, True)
conn.commit()
conn.execute("BEGIN IMMEDIATE")
# 1. Find and derive intervals from sample pairs
intervals = _get_unlinked_intervals(conn)
for prev, next_s in intervals:
try:
derived_hours = derive_hours_from_interval(conn, prev, next_s)
result.intervals_derived += 1
result.hours_created += len([h for h in derived_hours if h.get("attributed", True)])
except Exception as e:
logger.warning("Failed to derive interval %s -> %s: %s",
prev["ts"], next_s["ts"], e)
# Continue with other intervals (resilient)
# 2. Update day aggregates from hour observations
cursor = conn.execute(
"SELECT DISTINCT substr(hour, 1, 10) as day FROM hour_observations ORDER BY day"
)
days = [row[0] for row in cursor.fetchall()]
for day in days:
day_data = _derive_day_aggregate_from_hours(conn, day)
if day_data is not None:
created = _upsert_day_aggregate(conn, day_data)
if created:
result.days_created += 1
else:
result.days_updated += 1
# Rebuild only legacy or unavailable local-day activity. This shares
# the migration derivation and leaves trustworthy published totals intact.
from .local_day import repair_legacy_local_day_evidence
repair_legacy_local_day_evidence(conn)
# 3. Count boundary anchors retained
now = clock if isinstance(clock, datetime) else datetime.now(timezone.utc)
cursor = conn.execute("SELECT ts FROM samples ORDER BY ts")
anchor_count = 0
for row in cursor.fetchall():
if needs_boundary_anchor(conn, row[0], now):
anchor_count += 1
result.boundary_anchors_retained = anchor_count
# 4. Count preserved legacy summaries
# Legacy summaries are day aggregates without corresponding hour observations
cursor = conn.execute(
"""SELECT COUNT(*) FROM day_aggregates d
WHERE NOT EXISTS (
SELECT 1 FROM hour_observations h
WHERE h.hour LIKE d.day || 'T%'
)"""
)
result.legacy_summaries_preserved = cursor.fetchone()[0]
# Publish derived rows and completion status together. A process
# interruption rolls back the in-progress marker with the repair.
_update_repair_status(conn, result)
_set_repair_in_progress(conn, False)
conn.commit()
except Exception as e:
conn.rollback()
try:
_set_repair_in_progress(conn, False)
conn.commit()
except sqlite3.Error:
conn.rollback()
logger.error("Repair failed: %s", e)
return RepairResult(
ok=False,
error=str(e),
)
return result
+147
View File
@@ -0,0 +1,147 @@
"""Controller segment management.
Handles identity-based segmentation of observation history:
- Find current (most recent) segment
- Determine if a new segment should open
- Open new segments with frozen metadata snapshot
Segmentation axes (independent):
- Identity key change → quarantines prior history
- DUW decrease with unchanged identity → new segment, prior history stays as habit evidence
Blank-key semantics (PR-16):
- To/from blank is an identity change → quarantines
- Equal blanks continue the segment, segmented by DUW monotonicity alone
"""
import sqlite3
from datetime import datetime
from typing import Any, Dict, Optional, Tuple
from .collector import normalize_identity
def find_current_segment(conn: sqlite3.Connection) -> Optional[Dict[str, Any]]:
"""Find the most recent (open) controller segment.
Returns the segment dict or None if no segments exist.
"""
cursor = conn.execute(
"SELECT id, opened_at, identity_key, identity_degraded, "
"subnqn, sn, mn, fr, vid, ssvid, transport "
"FROM controller_segments ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return None
return {
"id": row[0],
"opened_at": row[1],
"identity_key": row[2],
"identity_degraded": bool(row[3]),
"subnqn": row[4],
"sn": row[5],
"mn": row[6],
"fr": row[7],
"vid": row[8],
"ssvid": row[9],
"transport": row[10],
}
def get_last_duw(conn: sqlite3.Connection, segment_id: int) -> Optional[int]:
"""Get the bytes_written from the most recent sample in a segment.
Returns None if no samples exist in the segment.
"""
# Samples don't have a segment_id FK yet, so we need to find the
# latest sample before the segment's opened_at, or the latest sample
# if this is the first segment.
#
# For now, we'll use a simpler approach: get the latest sample's bytes_written.
# TODO: Add segment_id FK to samples table in next schema migration
cursor = conn.execute(
"SELECT bytes_written FROM samples ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
return row[0] if row else None
def should_open_new_segment(
current_segment: Optional[Dict[str, Any]],
new_identity_key: str,
new_bytes_written: int,
conn: sqlite3.Connection,
) -> Tuple[bool, Optional[str]]:
"""Determine if a new segment should open.
Returns (should_open, reason).
reason is None if no new segment, or a string describing why.
"""
# No current segment → must open first segment
if current_segment is None:
return True, "first_segment"
old_key = current_segment["identity_key"] or ""
# Identity key change (including to/from blank)
if old_key != new_identity_key:
return True, "identity_change"
# DUW decrease (counter reset or controller replacement with same identity)
last_duw = get_last_duw(conn, current_segment["id"])
if last_duw is not None and new_bytes_written < last_duw:
return True, "duw_decrease"
# Same identity, DUW non-decreasing → continue segment
return False, None
def open_segment(
conn: sqlite3.Connection,
now: datetime,
identity: Dict[str, Any],
identity_key: str,
identity_degraded: bool,
) -> Dict[str, Any]:
"""Open a new controller segment with frozen metadata snapshot.
The metadata is immutable once frozen.
"""
cursor = conn.execute(
"""
INSERT INTO controller_segments (
opened_at, identity_key, identity_degraded,
subnqn, sn, mn, fr, vid, ssvid, transport
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
now.isoformat(),
identity_key if identity_key else None,
identity_degraded,
identity.get("subnqn") or None,
identity.get("sn") or None,
identity.get("mn") or None,
identity.get("fr") or None,
identity.get("vid") or None,
identity.get("ssvid") or None,
identity.get("transport") or None,
),
)
segment_id = cursor.lastrowid
return {
"id": segment_id,
"opened_at": now.isoformat(),
"identity_key": identity_key if identity_key else None,
"identity_degraded": identity_degraded,
"subnqn": identity.get("subnqn") or None,
"sn": identity.get("sn") or None,
"mn": identity.get("mn") or None,
"fr": identity.get("fr") or None,
"vid": identity.get("vid") or None,
"ssvid": identity.get("ssvid") or None,
"transport": identity.get("transport") or None,
}
+554
View File
@@ -0,0 +1,554 @@
"""Read-only CLI status command: the CLI twin of the TUI (spec §8.8, LC-9, CI-2).
Composes from the observation store (read-only) and allow-listed service
properties: projection facts, four separate service facts (boot enablement,
runtime activity, last collect outcome, freshness), and a journal/log hint on
failure or staleness. Never auto-samples, never prompts.
Freshness constants are defined once here and shared with the TUI (§8.9):
fresh — newest sample within 2 × cadence + AccuracySec + 60 s
missed — between fresh and 48 h
stale — ≥ 48 h
empty — no observations yet
Criteria: LC-9, CI-2, CI-4, FL-4, FL-5, FL-7.
"""
import sqlite3
from contextlib import contextmanager
import sys
from datetime import datetime, timedelta, timezone
from pathlib import Path
from typing import Any, Dict, Iterator, List, Optional, Tuple, TYPE_CHECKING
if TYPE_CHECKING:
from .status_composition import StatusComposition
from .projection import compute_projection, ConfidenceState, DISCLOSURES
from .store import SCHEMA_VERSION
from .init_system import (
query_service_state as _init_query_service_state,
journal_hint as _init_journal_hint,
)
# ---------------------------------------------------------------------------
# Freshness constants (§8.9, §8.2)
# ---------------------------------------------------------------------------
CADENCE_DEFAULT_S = 180 # 3 min
ACCURACY_SEC = 30
FRESH_THRESHOLD_S = 2 * CADENCE_DEFAULT_S + ACCURACY_SEC + 60 # 450 s
STALENESS_THRESHOLD_S = 48 * 3600 # 48 h
# ---------------------------------------------------------------------------
# Configuration reading (§8.3)
# ---------------------------------------------------------------------------
CONFIG_PATH = Path("/etc/fenris/fenris.conf")
def read_config() -> Dict[str, Any]:
"""Read the world-readable configuration file.
Returns a dict with at least 'device'.
Raises ConfigError with a reason string on any failure.
"""
if not CONFIG_PATH.exists():
raise ConfigError("configuration file not found at %s" % CONFIG_PATH)
try:
text = CONFIG_PATH.read_text()
except OSError as e:
raise ConfigError("cannot read %s: %s" % (CONFIG_PATH, e))
device = None
for line in text.splitlines():
line = line.strip()
if not line or line.startswith("#"):
continue
if "=" in line:
key, _, value = line.partition("=")
key = key.strip()
value = value.strip().strip('"').strip("'")
if key == "device":
device = value
break
if not device:
raise ConfigError("no device selector in %s" % CONFIG_PATH)
return {"device": device}
class ConfigError(Exception):
"""Configuration is invalid — surfaced in status as a fact (§8.3)."""
pass
# ---------------------------------------------------------------------------
# Store opening (read-only, §3, §9.4, §9.5)
# ---------------------------------------------------------------------------
def open_store_readonly(store_path: Path) -> sqlite3.Connection:
"""Open the observation store read-only.
Raises StoreFault if unreadable, NewerSchema if user_version > SCHEMA_VERSION.
"""
try:
exists = store_path.exists()
except OSError as e:
# A non-group user stat()ing a 2750 store directory gets
# PermissionError before any StoreFault can be raised (issue #54).
raise StoreFault("observation store not readable: %s" % e)
if not exists:
raise MissingStore("observation store not found at %s" % store_path)
try:
conn = sqlite3.connect("file:%s?mode=ro" % store_path, uri=True)
conn.row_factory = sqlite3.Row
except sqlite3.Error as e:
raise StoreFault("observation store unreadable: %s" % e)
try:
cursor = conn.execute("PRAGMA user_version")
version = cursor.fetchone()[0]
except sqlite3.Error as e:
conn.close()
raise StoreFault("observation store unreadable: %s" % e)
if version > SCHEMA_VERSION:
conn.close()
raise NewerSchema(version)
return conn
class StoreFault(Exception):
"""Store is present but unreadable or corrupt (§9.4)."""
pass
class MissingStore(StoreFault):
"""No observation history has been created yet."""
class NewerSchema(Exception):
"""Store has a newer user_version (§9.5)."""
def __init__(self, version: int):
self.version = version
super().__init__("schema version %d" % version)
# ---------------------------------------------------------------------------
# Service state queries (§8.8 — init-system agnostic)
# ---------------------------------------------------------------------------
# Re-export for backward compatibility with tests that import directly
query_service_state = _init_query_service_state
_journalctl_hint = _init_journal_hint
# ---------------------------------------------------------------------------
# Freshness grading (§8.9)
# ---------------------------------------------------------------------------
def grade_freshness(newest_sample_ts: Optional[str], clock_now: datetime) -> str:
"""Grade freshness from the newest sample timestamp (never a stored flag).
Returns 'fresh', 'missed', 'stale', or 'empty'.
"""
if newest_sample_ts is None:
return "empty"
try:
ts = datetime.fromisoformat(newest_sample_ts)
if ts.tzinfo is None:
ts = ts.replace(tzinfo=timezone.utc)
else:
ts = ts.astimezone(timezone.utc)
except (ValueError, TypeError):
return "empty"
age_s = (clock_now - ts).total_seconds()
if age_s <= FRESH_THRESHOLD_S:
return "fresh"
elif age_s < STALENESS_THRESHOLD_S:
return "missed"
else:
return "stale"
def freshness_age_human(age_s: Optional[int]) -> str:
"""Human-readable age string for freshness fact."""
if age_s is None:
return "unknown age"
if age_s < 60:
return "%ds ago" % age_s
if age_s < 3600:
return "%dm ago" % (age_s // 60)
if age_s < 86400:
return "%dh %dm ago" % (age_s // 3600, (age_s % 3600) // 60)
return "%dd ago" % (age_s // 86400)
# ---------------------------------------------------------------------------
# Drive anomalies (§9.7 — FL-7)
# ---------------------------------------------------------------------------
def _query_drive_facts(conn: sqlite3.Connection) -> List[str]:
"""Query drive-reported anomalies from the latest sample (§9.7, FL-7).
critical_warning, media errors, and unsafe shutdowns render as ordinary
facts and never affect the projection.
"""
cursor = conn.execute(
"SELECT critical_warning, media_errors, unsafe_shutdowns, "
"temperature_c, available_spare "
"FROM samples ORDER BY id DESC LIMIT 1"
)
row = cursor.fetchone()
if row is None:
return []
facts = []
cw = row[0]
if cw and cw != 0:
facts.append("critical warning: %s" % hex(cw) if isinstance(cw, int) else str(cw))
me = row[1]
if me and me > 0:
facts.append("media errors: %d" % me)
us = row[2]
if us and us > 0:
facts.append("unsafe shutdowns: %d" % us)
return facts
# ---------------------------------------------------------------------------
# Retired command rejection (§8.8)
# ---------------------------------------------------------------------------
RETIRED_COMMANDS = {
"start": "use 'fenris monitor resume' to enable monitoring",
"stop": "use 'fenris monitor pause' to disable monitoring",
"run": "use 'fenris monitor resume' to enable monitoring; the timer runs in the background",
}
MIGRATION_POINTERS = {
"--device": "device is configured in /etc/fenris/fenris.conf",
}
def check_retired_command(cmd: str) -> Optional[str]:
"""Check if a command is retired and return the migration pointer, or None."""
return RETIRED_COMMANDS.get(cmd)
def check_retired_flag(flag: str) -> Optional[str]:
"""Check if a flag is retired and return the migration pointer, or None."""
return MIGRATION_POINTERS.get(flag)
# ---------------------------------------------------------------------------
# Formatting
# ---------------------------------------------------------------------------
def _format_projection(proj, freshness: str, drive_facts: List[str],
config_error: Optional[str],
sample_count: int = 0, day_count: int = 0) -> str:
"""Format projection details; monitoring status has its own renderer."""
lines = []
# --- Configuration error (§8.3) ---
if config_error:
lines.append("configuration error: %s" % config_error)
lines.append("")
# --- Empty store (§8.9) ---
if freshness == "empty":
lines.append("no observations yet")
lines.append("")
lines.append("Enable monitoring: fenris monitor resume")
return "\n".join(lines)
# --- Single sample: awaiting another sample (issue #73 AC3) ---
# Only show awaiting state when there are no day aggregates (e.g., legacy import
# or hand-crafted stores can have 1 sample but sufficient day data for projection)
if sample_count <= 1 and day_count == 0:
lines.append("awaiting another sample")
lines.append("")
lines.append("Collecting usage data — the first projection requires at least two samples.")
return "\n".join(lines)
if proj is None:
lines.append("no projection available")
return "\n".join(lines)
# --- Projection headline ---
headline = _format_headline(proj)
lines.append(headline)
lines.append("")
# --- Confidence state + contributing facts (§6.7, §6.11) ---
state_label = proj.confidence_state.value
facts_list = list(proj.contributing_facts) if proj.contributing_facts else []
# Issue #77: Show qualifying day progress for honesty
if proj.qualifying_days_progress:
facts_list.insert(0, proj.qualifying_days_progress)
if facts_list:
facts_str = " · ".join(facts_list)
lines.append("%s evidence · %s" % (state_label, facts_str))
else:
lines.append("%s evidence" % state_label)
lines.append("")
# --- Scenario range (§6.5) with horizon reasons ---
if proj.scenario_range:
parts = []
for horizon in sorted(proj.scenario_range.rates.keys()):
rate_gb_day = proj.scenario_range.rates[horizon] * 86400 / 1e9
parts.append("%dd: %.2f GB/day" % (horizon, rate_gb_day))
for horizon, reason in sorted(proj.scenario_range.horizon_reasons.items()):
parts.append("%dd: %s" % (horizon, reason))
if parts:
lines.append("scenario range: %s" % " · ".join(parts))
lines.append("")
# --- PU context line (§6.1) ---
lines.append(proj.pu_context_line)
lines.append("")
# --- Drive anomalies (§9.7, FL-7) ---
if drive_facts:
for fact in drive_facts:
lines.append(fact)
lines.append("")
return "\n".join(lines)
def _format_headline(proj) -> str:
"""Format the lifespan headline or its no-projection wording (§6.11)."""
if proj.headline_remaining_seconds is None:
if proj.zero_rate_fact:
return "no finite projection from this history"
if proj.warming_fact:
return proj.warming_fact
return "no projection available"
secs = proj.headline_remaining_seconds
if secs <= 0:
return "endurance exhausted"
# Human-readable time
years = int(secs // 31557600)
rem = secs % 31557600
days = int(rem // 86400)
rem %= 86400
hours = int(rem // 3600)
parts = []
if years:
parts.append("%d yr" % years)
if days or years:
parts.append("%d d" % days)
parts.append("%d h" % hours)
remaining_human = " ".join(parts)
# Regime line
regime_parts = []
if proj.regime_days:
regime_parts.append("sustained regime: %d days" % proj.regime_days)
headline = "%s remaining" % remaining_human
if regime_parts:
headline += " · %s" % " · ".join(regime_parts)
return headline
# ---------------------------------------------------------------------------
# Dashboard clarity parity wording (DC-2, DC-3)
# ---------------------------------------------------------------------------
_CONTINUITY_ACTIVE = "monitoring: active in background · persists across reboots"
_CONTINUITY_DISABLED = "monitoring: does not start on next boot"
_PAUSED_TITLE = "monitoring: paused — deliberate disable"
_PAUSED_CONSEQUENCE = (
"paused time is excluded from your usage habit · resume: fenris monitor resume"
)
def monitoring_continuity(service: Dict[str, Any]) -> str:
"""Return the boot-persistence wording, independent of timer runtime."""
if service.get("boot_enabled") is None:
return "monitoring: boot persistence unknown"
return _CONTINUITY_ACTIVE if service["boot_enabled"] else _CONTINUITY_DISABLED
def deliberate_pause_lines() -> List[str]:
"""Return the exact CLI/TUI presentation for a sanctioned pause."""
return [_PAUSED_TITLE, _PAUSED_CONSEQUENCE]
def is_deliberately_paused(conn: sqlite3.Connection, service: Dict[str, Any]) -> bool:
"""Whether the latest closed period was ended by Fenris's own pause path.
Raw systemd operations have no `user_disabled` row, so they must never be
presented as a Deliberate disable. A live enabled timer also wins over a
stale period marker, keeping the presentation consistent with service facts.
"""
if service.get("boot_enabled") or service.get("timer_active"):
return False
open_period = conn.execute(
"SELECT 1 FROM monitoring_periods WHERE ended_at IS NULL LIMIT 1"
).fetchone()
if open_period is not None:
return False
row = conn.execute(
"SELECT end_cause FROM monitoring_periods "
"WHERE ended_at IS NOT NULL "
"ORDER BY ended_at DESC, id DESC LIMIT 1"
).fetchone()
return row is not None and row[0] == "user_disabled"
def format_disclosures() -> str:
"""Format the six disclosures (§6.11, CI-4)."""
lines = []
lines.append("Disclosures")
lines.append("")
for i, disc in enumerate(DISCLOSURES, 1):
lines.append("%d. %s" % (i, disc))
return "\n".join(lines)
# ---------------------------------------------------------------------------
# Main status entry point
# ---------------------------------------------------------------------------
@contextmanager
def read_status(
store_path: Optional[Path] = None,
clock_now: Optional[datetime] = None,
query_services: bool = True,
collecting: bool = False,
reduced_motion: bool = False,
) -> Iterator[Tuple[Optional[sqlite3.Connection], "StatusComposition"]]:
"""Yield a read-only store snapshot and its composed monitoring status.
Own acquisition, fault classification, and connection lifetime for both
renderers. An absent store is empty; an unreadable or newer store exposes
no connection. Unknown monitoring facts are never coerced to disabled.
"""
from .status_composition import compose_status
clock_now = clock_now or datetime.now(timezone.utc)
service = None
if query_services:
try:
service = query_service_state()
except (OSError, RuntimeError):
pass
conn = None
store_fault = newer_schema = None
try:
try:
conn = open_store_readonly(store_path or Path("/var/lib/fenris/observations.db"))
conn.execute("BEGIN")
except MissingStore:
pass
except (StoreFault, sqlite3.Error) as exc:
store_fault = str(exc)
except NewerSchema as exc:
newer_schema = str(exc)
comp = compose_status(
conn, service, clock_now, store_fault=store_fault,
newer_schema=newer_schema, collecting=collecting,
reduced_motion=reduced_motion,
)
if conn is not None and (comp.store_fault or comp.newer_schema):
conn.close()
conn = None
yield conn, comp
finally:
if conn is not None:
conn.close()
def get_status(store_path: Optional[Path] = None, clock_now: Optional[datetime] = None,
query_services: bool = True, query_journal: bool = True) -> str:
"""Render CLI status through the shared read-only acquisition path."""
from .status_composition import render_status_cli
clock_now = clock_now or datetime.now(timezone.utc)
config_error = None
try:
read_config()
except ConfigError as exc:
config_error = str(exc)
with read_status(store_path, clock_now, query_services) as (conn, comp):
parts = [render_status_cli(comp)]
if not comp.store_fault and not comp.newer_schema:
drive_facts = []
proj = None
if conn is not None:
try:
drive_facts = _query_drive_facts(conn)
proj = compute_projection(conn, clock_now)
except (sqlite3.Error, ValueError, TypeError):
pass
parts.append(_format_projection(
proj, comp.freshness, drive_facts, config_error,
comp.sample_count, comp.day_count,
))
if query_journal and (
comp.store_fault or comp.last_collect_ok is False
or comp.freshness in ("missed", "stale")
):
hint = _journalctl_hint()
if hint:
parts.append("Recent collector logs:\n" + hint)
return "\n\n".join(part for part in parts if part)
def render_status(store_path: Optional[Path] = None, clock_now: Optional[datetime] = None,
query_services: bool = True, query_journal: bool = True,
show_disclosures: bool = False) -> str:
"""High-level status renderer: status + optional disclosures.
Used by the CLI entry point.
"""
parts = [get_status(store_path, clock_now, query_services, query_journal)]
if show_disclosures:
parts.append("")
parts.append(format_disclosures())
return "\n\n".join(parts)
# ---------------------------------------------------------------------------
# Shared status composition integration (issue #78)
# ---------------------------------------------------------------------------
def get_status_composition(
store_path: Optional[Path] = None,
clock_now: Optional[datetime] = None,
query_services: bool = True,
collecting: bool = False,
reduced_motion: bool = False,
) -> 'StatusComposition':
"""Return monitoring status without retaining the read-only snapshot."""
with read_status(
store_path, clock_now, query_services, collecting, reduced_motion,
) as (_, comp):
return comp
+558
View File
@@ -0,0 +1,558 @@
"""Shared status composition for TUI and CLI (issue #78, TPH-2/3).
Centralizes the monitoring status lattice consumed by both surfaces.
States: Monitoring, Collecting, Paused, Waiting, Interrupted, Error, Stale, Unknown.
Precedence: Error > Interrupted > Paused > Stale > Waiting > Monitoring > Unknown.
Collecting overlays every base except store fault.
Freshness grading uses shared constants from status.py.
"""
import enum
import sqlite3
from dataclasses import dataclass, field
from datetime import datetime, timedelta, timezone
from typing import Any, Dict, List, Optional
# Status poll interval (AC78-6): lightweight 5s systemctl show poll
STATUS_POLL_INTERVAL_S = 5
from .status import (
FRESH_THRESHOLD_S,
STALENESS_THRESHOLD_S,
grade_freshness,
freshness_age_human,
is_deliberately_paused,
monitoring_continuity,
deliberate_pause_lines,
open_store_readonly,
StoreFault,
NewerSchema,
)
# ---------------------------------------------------------------------------
# Status state enum with glyph and label
# ---------------------------------------------------------------------------
class StatusState(enum.Enum):
"""The eight monitoring status states (TPH-2)."""
MONITORING = "monitoring"
COLLECTING = "collecting"
PAUSED = "paused"
WAITING = "waiting"
INTERRUPTED = "interrupted"
ERROR = "error"
STALE = "stale"
UNKNOWN = "unknown"
@property
def glyph(self) -> str:
_glyphs = {
"monitoring": "●",
"collecting": "◐",
"paused": "‖",
"waiting": "○",
"interrupted": "⊘",
"error": "✖",
"stale": "◌",
"unknown": "?",
}
return _glyphs[self.value]
@property
def label(self) -> str:
return self.value.capitalize()
# Precedence order: higher index = higher precedence
_PRECEDENCE = [
StatusState.UNKNOWN,
StatusState.MONITORING,
StatusState.WAITING,
StatusState.STALE,
StatusState.PAUSED,
StatusState.INTERRUPTED,
StatusState.ERROR,
]
_PRECEDENCE_RANK = {s: i for i, s in enumerate(_PRECEDENCE)}
# ---------------------------------------------------------------------------
# Status composition result
# ---------------------------------------------------------------------------
@dataclass
class StatusComposition:
"""The composed status result shared between TUI and CLI."""
state: StatusState
glyph: str
label: str
explanation: str
# Separate facts (never folded into the status word)
freshness: str = "unknown"
freshness_age_s: Optional[int] = None
last_collect_ok: Optional[bool] = None
last_collect_age_s: Optional[int] = None
last_collect_reason: Optional[str] = None
boot_enabled: Optional[bool] = None
timer_active: Optional[bool] = None
deliberately_paused: bool = False
pause_age_s: Optional[int] = None
external_stop_reason: Optional[str] = None
# Store fault / newer schema (suppress store-dependent views)
store_fault: Optional[str] = None
newer_schema: Optional[str] = None
# Collecting overlay
overlay_base: Optional[StatusState] = None
# Sample counts for waiting explanations
sample_count: int = 0
day_count: int = 0
pending_publication_count: int = 0
# Whether the status dot should blink (only Monitoring)
should_blink: bool = False
# Continuity line
continuity: str = ""
# Paused lines (for TUI banner)
paused_lines: List[str] = field(default_factory=list)
# ---------------------------------------------------------------------------
# Status composition logic
# ---------------------------------------------------------------------------
def _determine_base_state(
freshness: str,
sample_count: int,
day_count: int,
boot_enabled: Optional[bool],
timer_active: Optional[bool],
last_collect_ok: Optional[bool],
deliberately_paused: bool,
external_stop_reason: Optional[str],
service_available: bool,
) -> StatusState:
"""Determine the base status state from facts.
Precedence: Error > Interrupted > Paused > Stale > Waiting > Monitoring > Unknown.
"""
# Unknown: service query failed and no store-derived fact places us higher
if not service_available:
return StatusState.UNKNOWN
# Error: last collect failed
if last_collect_ok is False:
return StatusState.ERROR
# Interrupted: external stop (not user_disabled)
if external_stop_reason is not None:
return StatusState.INTERRUPTED
# Paused: deliberate disable
if deliberately_paused:
return StatusState.PAUSED
# Stale: timer active, no failure, but data ≥ 48h old
if freshness == "stale" and timer_active and last_collect_ok is not False:
return StatusState.STALE
# Waiting: empty store, single sample, or fresh data but not yet enough evidence
if sample_count == 0:
return StatusState.WAITING
if sample_count <= 1 and day_count == 0:
return StatusState.WAITING
# Monitoring: everything is fine
return StatusState.MONITORING
def _determine_explanation(
state: StatusState,
freshness: str,
freshness_age_s: Optional[int],
last_collect_ok: Optional[bool],
last_collect_reason: Optional[str],
sample_count: int,
deliberately_paused: bool,
store_fault: Optional[str],
newer_schema: Optional[str],
) -> str:
"""Determine the explanation line for the status state."""
if store_fault:
return "observation store unreadable — see collector logs"
if newer_schema:
return "observation store written by a newer Fenris — upgrade Fenris"
if state == StatusState.ERROR:
parts = []
if last_collect_ok is False:
parts.append("last run failed")
if last_collect_reason:
parts.append("(%s)" % last_collect_reason)
if freshness_age_s is not None and freshness != "empty":
parts.append("· last good sample %s" % freshness_age_human(freshness_age_s))
return " ".join(parts) if parts else "last run failed"
if state == StatusState.INTERRUPTED:
return "collection stopped outside Fenris — monitoring period still open"
if state == StatusState.PAUSED:
return "monitoring paused — paused time excluded from your usage habit"
if state == StatusState.STALE:
if freshness_age_s is not None:
return "last sample %s" % freshness_age_human(freshness_age_s)
return "data is stale"
if state == StatusState.WAITING:
if sample_count == 0:
return "awaiting first sample"
if sample_count <= 1:
return "awaiting another sample"
return "waiting for data"
if state == StatusState.MONITORING:
if freshness_age_s is not None:
return "last sample %s" % freshness_age_human(freshness_age_s)
return "monitoring active"
if state == StatusState.UNKNOWN:
return "service state unavailable"
return ""
def _determine_collecting_overlay(
base_state: StatusState,
last_collect_ok: Optional[bool],
deliberately_paused: bool,
store_fault: Optional[str],
newer_schema: Optional[str],
) -> str:
"""Determine the explanation line when Collecting overlays a base state."""
if store_fault or newer_schema:
return "run in flight — store fault"
if base_state == StatusState.PAUSED:
return "run in flight — paused"
if base_state == StatusState.INTERRUPTED:
return "run in flight — interrupted"
if base_state == StatusState.ERROR:
return "run in flight — retry"
if base_state == StatusState.STALE:
return "run in flight — stale data"
if base_state == StatusState.WAITING:
return "run in flight"
return "run in flight"
def compose_status(
conn: Optional[sqlite3.Connection],
service: Optional[Dict[str, Any]],
clock_now: datetime,
store_fault: Optional[str] = None,
newer_schema: Optional[str] = None,
collecting: bool = False,
reduced_motion: bool = False,
) -> StatusComposition:
"""Compose the shared status from service state and store data.
This is the single entry point consumed by both TUI and CLI.
"""
service_available = bool(service) and any(
service.get(key) is not None
for key in ("boot_enabled", "timer_active")
)
# --- Separate facts from service ---
boot_enabled = service.get("boot_enabled") if service else None
timer_active = service.get("timer_active") if service else None
last_collect_ok = service.get("last_collect_ok") if service else None
last_collect_age_s = service.get("last_collect_age_s") if service else None
last_collect_reason = service.get("last_collect_reason") if service else None
# --- Freshness from store ---
freshness = "unknown"
freshness_age_s = None
sample_count = 0
day_count = 0
pending_publication_count = 0
deliberately_paused = False
if conn is None and store_fault is None and newer_schema is None:
freshness = "empty"
if conn is not None and store_fault is None and newer_schema is None:
try:
cursor = conn.execute("SELECT ts FROM samples ORDER BY id DESC LIMIT 1")
row = cursor.fetchone()
newest_ts = row[0] if row else None
freshness = grade_freshness(newest_ts, clock_now)
if newest_ts:
try:
ts = datetime.fromisoformat(newest_ts)
if ts.tzinfo is None:
ts = ts.replace(tzinfo=timezone.utc)
else:
ts = ts.astimezone(timezone.utc)
freshness_age_s = int((clock_now - ts).total_seconds())
except (ValueError, TypeError):
pass
except sqlite3.Error as exc:
store_fault = str(exc)
try:
cursor = conn.execute("SELECT COUNT(*) FROM samples")
sample_count = cursor.fetchone()[0]
cursor = conn.execute("SELECT COUNT(*) FROM day_aggregates")
day_count = cursor.fetchone()[0]
schema_version = conn.execute("PRAGMA user_version").fetchone()[0]
if schema_version >= 4:
cursor = conn.execute("SELECT COUNT(*) FROM pending_publications")
pending_publication_count = cursor.fetchone()[0]
except sqlite3.Error as exc:
store_fault = str(exc)
try:
svc_for_pause = service if service_available else {}
deliberately_paused = is_deliberately_paused(conn, svc_for_pause)
except sqlite3.Error as exc:
store_fault = str(exc)
if store_fault or newer_schema:
freshness = "unknown"
freshness_age_s = None
sample_count = day_count = 0
pending_publication_count = 0
deliberately_paused = False
# --- External stop detection ---
# External stop = timer inactive + boot disabled + NOT deliberately paused
# + period still open (the timer was stopped but Fenris didn't close the period)
external_stop_reason = None
if (conn is not None and not store_fault and not newer_schema
and not deliberately_paused and timer_active is False and boot_enabled is False):
# Check if there's an open monitoring period (external stop left it open)
try:
open_period = conn.execute(
"SELECT 1 FROM monitoring_periods WHERE ended_at IS NULL LIMIT 1"
).fetchone()
if open_period is not None:
external_stop_reason = "external_stop"
except sqlite3.Error:
pass
# --- Determine base state ---
# Store fault and newer schema always override to ERROR
if store_fault is not None or newer_schema is not None:
base_state = StatusState.ERROR
else:
base_state = _determine_base_state(
freshness=freshness,
sample_count=sample_count,
day_count=day_count,
boot_enabled=boot_enabled,
timer_active=timer_active,
last_collect_ok=last_collect_ok,
deliberately_paused=deliberately_paused,
external_stop_reason=external_stop_reason,
service_available=service_available,
)
# --- Apply Collecting overlay ---
state = base_state
overlay_base = None
explanation = ""
if collecting and store_fault is None and newer_schema is None:
# Collecting overlays every base except store fault
overlay_base = base_state
state = StatusState.COLLECTING
explanation = _determine_collecting_overlay(
base_state, last_collect_ok, deliberately_paused,
store_fault, newer_schema,
)
else:
explanation = _determine_explanation(
base_state, freshness, freshness_age_s,
last_collect_ok, last_collect_reason,
sample_count, deliberately_paused,
store_fault, newer_schema,
)
# --- Continuity line ---
continuity = ""
if service_available:
continuity = monitoring_continuity(service)
# --- Paused lines ---
paused_lines = []
if deliberately_paused:
paused_lines = deliberate_pause_lines()
# Determine blink flag: only Monitoring dot blinks, never text or other states
should_blink = (state == StatusState.MONITORING and not reduced_motion)
return StatusComposition(
state=state,
glyph=state.glyph,
label=state.label,
explanation=explanation,
should_blink=should_blink,
freshness=freshness,
freshness_age_s=freshness_age_s,
last_collect_ok=last_collect_ok,
last_collect_age_s=last_collect_age_s,
last_collect_reason=last_collect_reason,
boot_enabled=boot_enabled,
timer_active=timer_active,
deliberately_paused=deliberately_paused,
external_stop_reason=external_stop_reason,
store_fault=store_fault,
newer_schema=newer_schema,
overlay_base=overlay_base,
sample_count=sample_count,
day_count=day_count,
pending_publication_count=pending_publication_count,
continuity=continuity,
paused_lines=paused_lines,
)
# ---------------------------------------------------------------------------
# CLI rendering (static, no styling)
# ---------------------------------------------------------------------------
def render_status_cli(comp: StatusComposition) -> str:
"""Render the status composition as static CLI text."""
lines = []
# Status line
lines.append("%s %s" % (comp.glyph, comp.label))
# Explanation
if comp.explanation:
lines.append(comp.explanation)
lines.append("")
# Separate facts
facts = []
facts.append("freshness: %s" % comp.freshness)
if comp.freshness_age_s is not None and facts:
facts[-1] += " (%s)" % freshness_age_human(comp.freshness_age_s)
if comp.last_collect_ok is True:
facts.append("last collect: ok")
elif comp.last_collect_ok is False:
collect_str = "last collect: FAILED"
if comp.last_collect_reason:
collect_str += " (%s)" % comp.last_collect_reason
facts.append(collect_str)
else:
facts.append("last collect: unknown")
facts.append("boot: %s" % (
"unknown" if comp.boot_enabled is None else "enabled" if comp.boot_enabled else "disabled"
))
facts.append("timer: %s" % (
"unknown" if comp.timer_active is None else "active" if comp.timer_active else "inactive"
))
if comp.pending_publication_count:
facts.append(_pending_publication_text(comp.pending_publication_count))
if facts:
lines.append(" · ".join(facts))
# Continuity
if comp.continuity:
lines.append("")
lines.append("CONTINUITY: %s" % comp.continuity)
# Deliberate pause
if comp.paused_lines:
for pl in comp.paused_lines:
lines.append(pl)
return "\n".join(lines)
# ---------------------------------------------------------------------------
# TUI rendering (with styling tokens)
# ---------------------------------------------------------------------------
def render_status_tui(comp: StatusComposition) -> str:
"""Render the status composition as TUI text with Textual markup."""
lines = []
# Status line with color
color = {
StatusState.MONITORING: "green",
StatusState.COLLECTING: "green",
StatusState.PAUSED: "yellow",
StatusState.WAITING: "yellow",
StatusState.INTERRUPTED: "red",
StatusState.ERROR: "red",
StatusState.STALE: "red",
StatusState.UNKNOWN: "dim",
}[comp.state]
lines.append("[%s]%s %s[/%s]" % (color, comp.glyph, comp.label, color))
# Explanation
if comp.explanation:
lines.append(comp.explanation[:1].upper() + comp.explanation[1:])
lines.append("")
# Separate facts
facts = []
facts.append("Freshness: %s" % comp.freshness)
if comp.freshness_age_s is not None and facts:
facts[-1] += " (%s)" % freshness_age_human(comp.freshness_age_s)
if comp.last_collect_ok is True:
facts.append("Last collect: ok")
elif comp.last_collect_ok is False:
collect_str = "Last collect: failed"
if comp.last_collect_reason:
collect_str += " (%s)" % comp.last_collect_reason
facts.append(collect_str)
else:
facts.append("Last collect: unknown")
facts.append("Boot: %s" % (
"unknown" if comp.boot_enabled is None else "enabled" if comp.boot_enabled else "disabled"
))
facts.append("Timer: %s" % (
"unknown" if comp.timer_active is None else "active" if comp.timer_active else "inactive"
))
if comp.pending_publication_count:
facts.append(_pending_publication_text(comp.pending_publication_count).capitalize())
if facts:
lines.append(" · ".join(facts))
# Continuity
if comp.continuity:
lines.append("")
lines.append("[bold]Continuity[/bold] %s" % comp.continuity)
# Deliberate pause
if comp.paused_lines:
for pl in comp.paused_lines:
lines.append(pl[:1].upper() + pl[1:])
return "\n".join(lines)
def _pending_publication_text(count: int) -> str:
noun = "observation" if count == 1 else "observations"
return f"pending publication: {count} {noun} retained for retry"
+465
View File
@@ -0,0 +1,465 @@
"""Observation store: SQLite database for persisting observation history.
This module handles:
- Store initialization with WAL mode
- Schema versioning with PRAGMA user_version
- Observation records, derived activity, publication state, and metadata
"""
import sqlite3
from pathlib import Path
from typing import Optional
# Schema version - increment on each migration
SCHEMA_VERSION = 6
# Packaged default placement (spec §8.3). The config may override it, but a
# fresh install that sets only the device selector must collect cleanly.
DEFAULT_STORE_PATH = Path("/var/lib/fenris/observations.db")
def get_store_path(config: dict) -> Path:
"""Get the store path from config.
Falls back to the packaged default when the config does not pin one,
so a fresh install whose config holds only the device selector works
instead of crashing with KeyError 'store_path' (issue #53).
"""
return Path(config.get("store_path", DEFAULT_STORE_PATH))
def init_store(store_path: Path) -> sqlite3.Connection:
"""Initialize the observation store if not present.
Creates the observation schema, including private pending-publication storage.
Returns a connection to the store.
"""
conn = sqlite3.connect(str(store_path))
# Enable WAL mode for concurrent reads during writes
conn.execute("PRAGMA journal_mode=WAL")
# Group members (fenris group) read the live store read-only, but SQLite
# in WAL mode needs write access to the db and its -wal/-shm sidecars even
# for readers. Best effort: root-created stores stay group-accessible
# without relying on the creating process's umask (issue #54).
import os as _os
for sidecar in (store_path,
store_path.with_name(store_path.name + "-wal"),
store_path.with_name(store_path.name + "-shm")):
try:
mode = _os.stat(sidecar).st_mode & 0o777
_os.chmod(sidecar, mode | 0o060)
except OSError:
pass
# Check if this is a new database
cursor = conn.execute("PRAGMA user_version")
current_version = cursor.fetchone()[0]
if current_version == 0:
# New database - create schema
_create_schema(conn)
conn.execute(f"PRAGMA user_version={SCHEMA_VERSION}")
conn.commit()
elif current_version > SCHEMA_VERSION:
# Unknown newer version - refuse
conn.close()
raise ValueError(
f"Observation store written by a newer Fenris (version {current_version}) "
f"— upgrade Fenris"
)
elif current_version < SCHEMA_VERSION:
# Older version - apply migrations
try:
_apply_migrations(conn, current_version)
except Exception:
conn.close()
raise
return conn
def _create_schema(conn: sqlite3.Connection):
"""Create the initial schema with all six entities."""
# Samples: raw collection runs (14-day retention)
conn.execute("""
CREATE TABLE IF NOT EXISTS samples (
id INTEGER PRIMARY KEY AUTOINCREMENT,
ts TEXT NOT NULL, -- ISO 8601 UTC timestamp
device TEXT NOT NULL,
-- Normalized controller-identity fields captured at acquisition
subnqn TEXT,
sn TEXT,
mn TEXT,
fr TEXT,
capacity_bytes INTEGER,
percentage_used INTEGER,
available_spare INTEGER,
media_errors INTEGER,
power_on_hours INTEGER,
power_cycles INTEGER,
unsafe_shutdowns INTEGER,
temperature_c INTEGER,
data_units_written INTEGER,
data_units_read INTEGER,
bytes_written INTEGER,
bytes_read INTEGER,
critical_warning INTEGER,
segment_id INTEGER,
local_tz TEXT
)
""")
# Hour observations: UTC-hour usage-habit split
conn.execute("""
CREATE TABLE IF NOT EXISTS hour_observations (
id INTEGER PRIMARY KEY AUTOINCREMENT,
hour TEXT NOT NULL UNIQUE, -- ISO 8601 UTC hour (e.g., "2026-09-01T12:00:00Z")
active_seconds INTEGER DEFAULT 0,
idle_seconds INTEGER DEFAULT 0,
powered_off_seconds INTEGER DEFAULT 0,
unknown_seconds INTEGER DEFAULT 0,
bytes_written_delta INTEGER DEFAULT 0,
bytes_read_delta INTEGER DEFAULT 0,
temperature_min INTEGER,
temperature_avg REAL,
temperature_max INTEGER,
sample_count INTEGER DEFAULT 0,
coverage REAL DEFAULT 0.0
)
""")
# Day aggregates: derived from hour observations
conn.execute("""
CREATE TABLE IF NOT EXISTS day_aggregates (
id INTEGER PRIMARY KEY AUTOINCREMENT,
day TEXT NOT NULL UNIQUE, -- ISO 8601 UTC day (e.g., "2026-09-01")
active_seconds INTEGER DEFAULT 0,
idle_seconds INTEGER DEFAULT 0,
powered_off_seconds INTEGER DEFAULT 0,
unknown_seconds INTEGER DEFAULT 0,
bytes_written_delta INTEGER DEFAULT 0,
bytes_read_delta INTEGER DEFAULT 0,
sample_count INTEGER DEFAULT 0,
coverage REAL DEFAULT 0.0,
unattributed_bytes_written INTEGER DEFAULT 0,
unattributed_bytes_read INTEGER DEFAULT 0
)
""")
# Monitoring periods: tracking when monitoring was enabled/disabled
conn.execute("""
CREATE TABLE IF NOT EXISTS monitoring_periods (
id INTEGER PRIMARY KEY AUTOINCREMENT,
started_at TEXT NOT NULL, -- ISO 8601 UTC timestamp
ended_at TEXT, -- NULL if currently active
end_cause TEXT CHECK(end_cause IN ('user_disabled', 'migrated', 'unknown_gap'))
)
""")
# Controller segments: identity key plus metadata snapshot
conn.execute("""
CREATE TABLE IF NOT EXISTS controller_segments (
id INTEGER PRIMARY KEY AUTOINCREMENT,
opened_at TEXT NOT NULL, -- ISO 8601 UTC timestamp
identity_key TEXT, -- Normalized identity key (NULL if degraded)
identity_degraded BOOLEAN DEFAULT 0,
subnqn TEXT,
sn TEXT,
mn TEXT,
fr TEXT,
vid TEXT,
ssvid TEXT,
transport TEXT
)
""")
# Endurance baseline: one active row, replaced on edit
conn.execute("""
CREATE TABLE IF NOT EXISTS endurance_baseline (
id INTEGER PRIMARY KEY AUTOINCREMENT,
tbw_terabytes REAL NOT NULL,
source_url TEXT,
document_revision TEXT,
entry_date TEXT,
model_string TEXT,
nominal_capacity_bytes INTEGER,
validated_by TEXT, -- 'user' or 'machine_match'
verified BOOLEAN DEFAULT 0,
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL
)
""")
# Local-day activity summaries derived from UTC hour observations.
# Each row retains its recorded timezone and UTC boundaries so that
# historical summaries survive a system-timezone change (ADR 0010).
conn.execute("""
CREATE TABLE IF NOT EXISTS local_days (
id INTEGER PRIMARY KEY AUTOINCREMENT,
local_date TEXT NOT NULL, -- e.g. "2026-09-01" in the recorded tz
tz_name TEXT NOT NULL, -- POSIX tz name, e.g. "Asia/Kolkata"
tz_offset TEXT NOT NULL, -- e.g. "+05:30"
utc_start TEXT NOT NULL, -- ISO 8601 UTC: local midnight start
utc_end TEXT NOT NULL, -- ISO 8601 UTC: local midnight end
bytes_written INTEGER DEFAULT 0,
bytes_read INTEGER DEFAULT 0,
coverage REAL DEFAULT 0.0,
sample_count INTEGER DEFAULT 0,
complete BOOLEAN DEFAULT 0,
activity_seconds INTEGER NOT NULL DEFAULT 0,
activity_intervals INTEGER NOT NULL DEFAULT 0,
activity_incomplete BOOLEAN NOT NULL DEFAULT 0,
activity_precision TEXT NOT NULL DEFAULT 'measured',
last_sample_id INTEGER,
UNIQUE(local_date, tz_name)
)
""")
_create_local_day_shared_evidence(conn)
_create_local_day_segment_totals(conn)
# Metadata table for store state (e.g., legacy import marker)
conn.execute("""
CREATE TABLE IF NOT EXISTS store_metadata (
key TEXT PRIMARY KEY,
value TEXT NOT NULL
)
""")
_create_pending_publications(conn)
def _create_pending_publications(conn: sqlite3.Connection) -> None:
"""Create private staging for valid observations awaiting derivation."""
conn.execute("""
CREATE TABLE IF NOT EXISTS pending_publications (
id INTEGER PRIMARY KEY AUTOINCREMENT,
sample_ts TEXT NOT NULL,
payload TEXT NOT NULL
)
""")
def _create_local_day_shared_evidence(conn: sqlite3.Connection) -> None:
"""Create once-only local activity evidence that cannot be day-allocated."""
conn.execute("""
CREATE TABLE IF NOT EXISTS local_day_unallocated_evidence (
id INTEGER PRIMARY KEY AUTOINCREMENT,
start_sample_id INTEGER NOT NULL,
end_sample_id INTEGER NOT NULL,
start_local_date TEXT NOT NULL,
end_local_date TEXT NOT NULL,
start_tz_name TEXT,
end_tz_name TEXT,
started_at TEXT NOT NULL,
ended_at TEXT NOT NULL,
bytes_written INTEGER NOT NULL DEFAULT 0,
bytes_read INTEGER NOT NULL DEFAULT 0,
reason TEXT NOT NULL,
segment_id INTEGER,
UNIQUE(start_sample_id, end_sample_id)
)
""")
conn.execute(
"CREATE INDEX IF NOT EXISTS local_day_evidence_start "
"ON local_day_unallocated_evidence(start_local_date, start_tz_name)"
)
conn.execute(
"CREATE INDEX IF NOT EXISTS local_day_evidence_end "
"ON local_day_unallocated_evidence(end_local_date, end_tz_name)"
)
def _create_local_day_segment_totals(conn: sqlite3.Connection) -> None:
"""Retain the controller-segment provenance behind known day totals."""
conn.execute("""
CREATE TABLE IF NOT EXISTS local_day_segment_totals (
id INTEGER PRIMARY KEY AUTOINCREMENT,
local_day_id INTEGER NOT NULL,
segment_id INTEGER NOT NULL,
bytes_written INTEGER NOT NULL DEFAULT 0,
bytes_read INTEGER NOT NULL DEFAULT 0,
activity_seconds INTEGER NOT NULL DEFAULT 0,
activity_intervals INTEGER NOT NULL DEFAULT 0,
UNIQUE(local_day_id, segment_id)
)
""")
def _apply_migrations(conn: sqlite3.Connection, current_version: int):
"""Apply forward-only migrations from current_version to SCHEMA_VERSION.
Commit each version transition independently. A failed step rolls back in
full while earlier successful steps remain versioned and retryable.
"""
migrations = {
2: _migrate_1_to_2,
3: _migrate_2_to_3,
4: _migrate_3_to_4,
5: _migrate_4_to_5,
6: _migrate_5_to_6,
}
while current_version < SCHEMA_VERSION:
target_version = current_version + 1
migration = migrations.get(target_version)
if migration is None:
raise ValueError(f"No migration registered for schema {target_version}")
conn.execute("BEGIN IMMEDIATE")
try:
migration(conn)
conn.execute(f"PRAGMA user_version={target_version}")
conn.commit()
except Exception:
conn.rollback()
raise
current_version = target_version
def _migrate_1_to_2(conn: sqlite3.Connection) -> None:
"""Add segment provenance and unattributed UTC byte tracking."""
tables = {row[0] for row in conn.execute(
"SELECT name FROM sqlite_master WHERE type='table'"
).fetchall()}
if "samples" in tables:
cols = {row[1] for row in conn.execute(
"PRAGMA table_info(samples)"
).fetchall()}
if "segment_id" not in cols:
conn.execute("ALTER TABLE samples ADD COLUMN segment_id INTEGER")
if "day_aggregates" in tables:
cols = {row[1] for row in conn.execute(
"PRAGMA table_info(day_aggregates)"
).fetchall()}
for column in ("unattributed_bytes_written", "unattributed_bytes_read"):
if column not in cols:
conn.execute(
f"ALTER TABLE day_aggregates ADD COLUMN {column} INTEGER DEFAULT 0"
)
def _migrate_2_to_3(conn: sqlite3.Connection) -> None:
"""Add local-day activity summaries."""
tables = {row[0] for row in conn.execute(
"SELECT name FROM sqlite_master WHERE type='table'"
).fetchall()}
if "local_days" not in tables:
conn.execute("""
CREATE TABLE local_days (
id INTEGER PRIMARY KEY AUTOINCREMENT,
local_date TEXT NOT NULL,
tz_name TEXT NOT NULL,
tz_offset TEXT NOT NULL,
utc_start TEXT NOT NULL,
utc_end TEXT NOT NULL,
bytes_written INTEGER DEFAULT 0,
bytes_read INTEGER DEFAULT 0,
coverage REAL DEFAULT 0.0,
sample_count INTEGER DEFAULT 0,
complete BOOLEAN DEFAULT 0,
UNIQUE(local_date, tz_name)
)
""")
def _migrate_3_to_4(conn: sqlite3.Connection) -> None:
"""Add private publication staging for acquired observations."""
_create_pending_publications(conn)
def _migrate_4_to_5(conn: sqlite3.Connection) -> None:
"""Add measured local-day evidence storage and mark old totals legacy."""
tables = {row[0] for row in conn.execute(
"SELECT name FROM sqlite_master WHERE type='table'"
).fetchall()}
if "samples" in tables:
sample_cols = {row[1] for row in conn.execute(
"PRAGMA table_info(samples)"
).fetchall()}
if "local_tz" not in sample_cols:
conn.execute("ALTER TABLE samples ADD COLUMN local_tz TEXT")
if "local_days" in tables:
local_cols = {row[1] for row in conn.execute(
"PRAGMA table_info(local_days)"
).fetchall()}
for column, declaration in (
("activity_seconds", "INTEGER NOT NULL DEFAULT 0"),
("activity_intervals", "INTEGER NOT NULL DEFAULT 0"),
("activity_incomplete", "BOOLEAN NOT NULL DEFAULT 0"),
("activity_precision", "TEXT NOT NULL DEFAULT 'legacy'"),
("last_sample_id", "INTEGER"),
):
if column not in local_cols:
conn.execute(
f"ALTER TABLE local_days ADD COLUMN {column} {declaration}"
)
_create_local_day_shared_evidence(conn)
_create_local_day_segment_totals(conn)
def _migrate_5_to_6(conn: sqlite3.Connection) -> None:
"""Rebuild local-day summaries from surviving trustworthy evidence."""
from .local_day import repair_legacy_local_day_evidence
repair_legacy_local_day_evidence(conn)
def migrate_to_latest(store_path: Path) -> int:
"""Apply forward-only migrations to bring the store to SCHEMA_VERSION.
Called by the upgrade target (§10.2). Returns the number of migration
steps applied. Raises ValueError on newer-schema store (§3.6, §9.5).
Spec: §3.6, §10.2, §10.3
"""
conn = sqlite3.connect(str(store_path))
conn.execute("PRAGMA journal_mode=WAL")
cursor = conn.execute("PRAGMA user_version")
current_version = cursor.fetchone()[0]
if current_version > SCHEMA_VERSION:
conn.close()
raise ValueError(
f"Observation store written by a newer Fenris (version {current_version}) "
f"— upgrade Fenris"
)
if current_version == SCHEMA_VERSION:
conn.close()
return 0 # Already up to date
# Version 0 means no schema — create fresh (issue #73)
if current_version == 0:
_create_schema(conn)
conn.execute(f"PRAGMA user_version={SCHEMA_VERSION}")
conn.commit()
conn.close()
return SCHEMA_VERSION
steps = SCHEMA_VERSION - current_version
try:
_apply_migrations(conn, current_version)
except Exception:
conn.close()
raise
conn.close()
return steps
def is_store_faulty(store_path: Path) -> bool:
"""Check if the store is present but cannot be read or trusted."""
if not store_path.exists():
return False
try:
conn = sqlite3.connect(f"file:{store_path}?mode=ro", uri=True)
conn.execute("PRAGMA user_version")
conn.close()
return False
except sqlite3.Error:
return True
+177
View File
@@ -0,0 +1,177 @@
"""Fenris theme presets (issue #80).
Chalktone-inspired default plus Amber, Nord, and High Contrast.
Themes style chrome, borders, accents, muted text, and graph roles;
status semantic colours/glyphs/text always win.
Theme roles for graph rendering expose distinct colours per preset so
the bar graph can reflect the user's visual preference without depending
on graph-ticket completion.
Criteria: TPH-10, AC80-1, AC80-5.
"""
from typing import Dict
from textual.theme import Theme
# ---------------------------------------------------------------------------
# Status semantic colours — always win, never themed (AC80-1)
# ---------------------------------------------------------------------------
STATUS_COLORS = {
"monitoring": "green",
"collecting": "green",
"paused": "yellow",
"waiting": "yellow",
"interrupted": "red",
"error": "red",
"stale": "red",
"unknown": "dim",
}
# ---------------------------------------------------------------------------
# Amber theme — warm golden tones (default, amber graph role)
# ---------------------------------------------------------------------------
_AMBER = Theme(
name="fenris-amber",
primary="#d4a017", # warm amber
secondary="#c49b0a", # darker amber
accent="#ffd54f", # light amber highlight
warning="#e6a817", # amber warning
error="#e74c3c", # red error
success="#27ae60", # green success
foreground="#e8e0d0", # warm light
background="#1a1510", # warm dark
surface="#241f16", # warm surface
panel="#2a2318", # warm panel
boost="#332a1c", # warm boost
dark=True,
variables={
"graph-allocated": "#d4a017",
"graph-unallocated": "#8b6914",
"graph-gap": "#554422",
"graph-zero": "#665533",
"graph-partial": "#aa8822",
"graph-selection": "#ffd54f",
"border-default": "#554422",
"muted-text": "#887755",
},
)
# ---------------------------------------------------------------------------
# Nord theme — cool blue-gray polar night palette
# ---------------------------------------------------------------------------
_NORD = Theme(
name="fenris-nord",
primary="#88c0d0", # nord8 frost
secondary="#81a1c1", # nord9
accent="#8fbcbb", # nord7
warning="#ebcb8b", # nord13
error="#bf616a", # nord11
success="#a3be8c", # nord14
foreground="#eceff4", # nord6
background="#2e3440", # nord0
surface="#3b4252", # nord1
panel="#434c5e", # nord2
boost="#4c566a", # nord3
dark=True,
variables={
"graph-allocated": "#88c0d0",
"graph-unallocated": "#5e81ac",
"graph-gap": "#4c566a",
"graph-zero": "#616e88",
"graph-partial": "#81a1c1",
"graph-selection": "#8fbcbb",
"border-default": "#4c566a",
"muted-text": "#7b88a1",
},
)
# ---------------------------------------------------------------------------
# High Contrast — maximum readability, pure black and white
# ---------------------------------------------------------------------------
_HIGH_CONTRAST = Theme(
name="fenris-high-contrast",
primary="#ffffff", # pure white
secondary="#dddddd", # light gray
accent="#ffff00", # bright yellow
warning="#ff8800", # bright orange
error="#ff0000", # pure red
success="#00ff00", # pure green
foreground="#ffffff", # pure white
background="#000000", # pure black
surface="#111111", # near-black surface
panel="#1a1a1a", # near-black panel
boost="#222222", # near-black boost
dark=True,
variables={
"graph-allocated": "#ffffff",
"graph-unallocated": "#aaaaaa",
"graph-gap": "#555555",
"graph-zero": "#666666",
"graph-partial": "#cccccc",
"graph-selection": "#ffff00",
"border-default": "#ffffff",
"muted-text": "#aaaaaa",
},
)
# ---------------------------------------------------------------------------
# Theme registry
# ---------------------------------------------------------------------------
_CHALKTONE = Theme(
name="fenris-chalktone",
primary="#abc4b3", secondary="#c9b69a", accent="#dfc49a",
warning="#e9bc79", error="#e58d89", success="#acd29c",
foreground="#e0d8c5", background="#202426", surface="#202426",
panel="#252a2c", boost="#333b3d", dark=True,
variables={
"graph-allocated": "#abc4b3", "graph-unallocated": "#dfc49a",
"graph-gap": "#e58d89", "graph-zero": "#a7b1a9",
"graph-partial": "#e9bc79", "graph-selection": "#f3dbb0",
"border-default": "#606d6a", "muted-text": "#a7b1a9",
},
)
THEMES = {
"chalktone": _CHALKTONE,
"amber": _AMBER,
"nord": _NORD,
"high_contrast": _HIGH_CONTRAST,
}
THEME_NAMES = set(THEMES.keys())
def get_theme(name: str) -> Theme:
"""Return a registered theme by preset name.
Unknown names fall back to Chalktone.
"""
return THEMES.get(name, _CHALKTONE)
def get_graph_colors(theme_name: str) -> Dict[str, str]:
"""Return the graph colour roles for a theme preset.
Returns a dict with keys: allocated, unallocated, gap, zero, partial,
selection and muted text. Falls back to Chalktone for unknown names.
"""
theme = get_theme(theme_name)
variables = theme.variables or {}
return {
"allocated": variables.get("graph-allocated", "#d4a017"),
"unallocated": variables.get("graph-unallocated", "#8b6914"),
"gap": variables.get("graph-gap", "#554422"),
"zero": variables.get("graph-zero", "#665533"),
"partial": variables.get("graph-partial", "#aa8822"),
"selection": variables.get("graph-selection", "#ffd54f"),
"muted": variables.get("muted-text", "#a7b1a9"),
}
+2451
View File
File diff suppressed because it is too large Load Diff
+52
View File
@@ -0,0 +1,52 @@
"""System timezone detection for local-day activity totals.
Provides timezone detection (TZ env, /etc/localtime) and offset
computation used by local-day derivation. All functions are
stateless and safe to call from the collector and TUI reader.
"""
import os
from datetime import datetime, timezone
from pathlib import Path
def detect_system_tz() -> str:
"""Detect the system timezone name.
Resolution order:
1. ``TZ`` environment variable (may be empty or ":UTC")
2. Symlink target of ``/etc/localtime``
3. Falls back to ``"UTC"``
Returns a POSIX timezone name like ``"Asia/Kolkata"`` or ``"UTC"``.
"""
tz = os.environ.get("TZ", "").strip()
if tz:
return tz.lstrip(":")
localtime = Path("/etc/localtime")
if localtime.is_symlink():
target_path = Path(os.readlink(str(localtime)))
if not target_path.is_absolute():
target_path = localtime.parent / target_path
target = target_path.resolve().as_posix()
marker = "/zoneinfo/"
marker_index = target.find(marker)
if marker_index >= 0:
return target[marker_index + len(marker):]
return target
return "UTC"
def get_tz_offset_str(dt: datetime, tz_name: str) -> str:
"""UTC offset as ``+HH:MM`` or ``-HH:MM`` string for *dt* in *tz_name*."""
from zoneinfo import ZoneInfo
local_tz = ZoneInfo(tz_name)
local_dt = dt.astimezone(local_tz)
offset = local_dt.utcoffset()
total_seconds = int(offset.total_seconds())
sign = "+" if total_seconds >= 0 else "-"
total_seconds = abs(total_seconds)
hours = total_seconds // 3600
minutes = (total_seconds % 3600) // 60
return "%s%02d:%02d" % (sign, hours, minutes)
+20
View File
@@ -0,0 +1,20 @@
"""Shared test helpers for Fenris test suite."""
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).resolve().parent.parent
VERSION_FILE = REPO_ROOT / "pyproject.toml"
def get_version() -> str:
"""Extract version from pyproject.toml."""
for line in VERSION_FILE.read_text().splitlines():
if line.startswith("version"):
return line.split("=")[1].strip().strip('"')
raise RuntimeError("Could not determine version from pyproject.toml")
def read(path: str | Path) -> str:
"""Read a file relative to the repository root."""
return (REPO_ROOT / path).read_text()
+755
View File
@@ -0,0 +1,755 @@
"""Cross-cutting acceptance sweep (issue #32).
Systematic verification of every acceptance criterion that spans multiple
subsystems. Grouped by criterion ID; each test cites its clause.
CI-1 Exhaustive state matrix: confidence × freshness × baseline tier
CI-2 TUI/CLI parity: identical outcomes and wording
CI-3 Prohibition set: automated structural checks
CI-4 Required wording and six disclosures in both views
"""
import re
import sqlite3
from datetime import datetime, timedelta, timezone
from pathlib import Path
from unittest.mock import patch
import pytest
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store, SCHEMA_VERSION
from fenris.monitoring_periods import ensure_period_open
from fenris.projection import (
compute_projection,
ConfidenceState,
BaselineTier,
DISCLOSURES,
STALENESS_HOURS,
WARMING_MIN_DAYS,
YOUNG_REGIME_DAYS,
)
from fenris.status import (
grade_freshness,
get_status,
render_status,
format_disclosures,
FRESH_THRESHOLD_S,
STALENESS_THRESHOLD_S,
CADENCE_DEFAULT_S,
ACCURACY_SEC,
)
from fenris.tui import (
FenrisTuiApp,
_format_remaining,
)
SRC_DIR = Path(__file__).parent.parent / "src"
FENRIS_PKG = SRC_DIR / "fenris"
def _clock(year=2026, month=9, day=30, hour=12):
return datetime(year, month, day, hour, 0, 0, tzinfo=timezone.utc)
def _insert_baseline(conn, tbw_tb=1.0, verified=True,
model="Samsung SSD 970 EVO Plus 1TB",
source_url="https://example.com/spec",
doc_rev="v1.0", entry_date="2026-01-01",
nominal_cap=1024000000000):
conn.execute(
"INSERT INTO endurance_baseline "
"(tbw_terabytes, source_url, document_revision, entry_date, model_string, "
" nominal_capacity_bytes, validated_by, verified, created_at, updated_at) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(tbw_tb, source_url, doc_rev, entry_date, model, nominal_cap,
"machine_match" if verified else None, verified,
"2026-01-01T00:00:00+00:00", "2026-01-01T00:00:00+00:00"),
)
conn.commit()
def _insert_segment(conn, opened_at="2026-09-01T00:00:00+00:00",
identity_key="nqn.test", degraded=False,
mn="Samsung SSD 970 EVO Plus 1TB"):
conn.execute(
"INSERT INTO controller_segments "
"(opened_at, identity_key, identity_degraded, subnqn, sn, mn, fr, vid, ssvid, transport) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(opened_at, identity_key, degraded, "nqn.test", "SN123", mn, "FW1",
"0x144d", "0x144d", "pcie"),
)
conn.commit()
def _insert_day(conn, day, bw=1024*1024*100, coverage=0.95, samples=24):
conn.execute(
"INSERT INTO day_aggregates (day, active_seconds, idle_seconds, "
"powered_off_seconds, unknown_seconds, bytes_written_delta, "
"bytes_read_delta, sample_count, coverage) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
(day, 3600, 0, 0, 0, bw, 0, samples, coverage),
)
conn.commit()
def _insert_sample(conn, ts, pu=5):
conn.execute(
"INSERT INTO samples (ts, device, data_units_written, data_units_read, "
"percentage_used, bytes_written, bytes_read, power_on_hours) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?)",
(ts, "/dev/nvme0n1", 1000000, 500000, pu, 512000000000, 256000000000, 8765),
)
conn.commit()
def _open_period(conn, start="2026-09-01T00:00:00+00:00"):
ensure_period_open(conn, datetime.fromisoformat(start))
def _insert_local_day(conn, local_date, tz_name="UTC", tz_offset="+00:00",
utc_start=None, utc_end=None, bw=1024*1024*100,
br=0, coverage=0.95, samples=24, complete=True):
"""Insert a local_days row (issue #94 gate prerequisite)."""
if utc_start is None:
utc_start = local_date + "T00:00:00+00:00"
if utc_end is None:
dt = datetime.strptime(local_date, "%Y-%m-%d") + timedelta(days=1)
utc_end = dt.strftime("%Y-%m-%dT00:00:00+00:00")
conn.execute(
"INSERT INTO local_days "
"(local_date, tz_name, tz_offset, utc_start, utc_end, "
" bytes_written, bytes_read, coverage, sample_count, complete, activity_intervals) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, 1)",
(local_date, tz_name, tz_offset, utc_start, utc_end,
bw, br, coverage, samples, complete),
)
conn.commit()
def _insert_complete_local_days(conn, start_date, count, bw=1024*1024*100):
"""Insert multiple complete local days to satisfy the issue #94 gate."""
for i in range(count):
d = (datetime.strptime(start_date, "%Y-%m-%d") + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_local_day(conn, d, bw=bw)
def _setup_full_store(conn, *, baseline=True, segment=True, days=30,
bw=1024*1024*100, coverage=0.95, samples_per_day=24,
sample_ts="2026-09-30T10:00:00+00:00",
period_start="2026-09-01T00:00:00+00:00",
segment_opened="2026-09-01T00:00:00+00:00",
baseline_kw=None, segment_kw=None,
local_days=True):
if baseline:
_insert_baseline(conn, **(baseline_kw or {}))
if segment:
_insert_segment(conn, opened_at=segment_opened, **(segment_kw or {}))
_open_period(conn, start=period_start)
for i in range(days):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=bw, coverage=coverage, samples=samples_per_day)
if sample_ts:
_insert_sample(conn, sample_ts)
# Issue #94: satisfy the complete-observation-day gate
if local_days and days > 0:
_insert_complete_local_days(conn, "2026-09-29", 1, bw=bw)
# ===================================================================
# CI-1: Exhaustive state matrix
# ===================================================================
class TestCI1StateMatrix:
"""Systematic walk of confidence x freshness x baseline tier combinations."""
def test_no_baseline_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert proj.headline_remaining_seconds is None
assert proj.baseline_tier == BaselineTier.NONE
conn.close()
def test_verified_baseline_possible_supported(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.SUPPORTED
assert proj.baseline_tier == BaselineTier.VERIFIED
assert proj.headline_remaining_seconds is not None
conn.close()
def test_unverified_baseline_possible_limited(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(
tbw_tb=10.0, verified=False, source_url=None))
proj = compute_projection(conn, _clock())
assert proj.baseline_tier == BaselineTier.UNVERIFIED
assert proj.confidence_state != ConfidenceState.SUPPORTED
conn.close()
def test_model_mismatch_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(model="Different Model"))
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert proj.baseline_tier == BaselineTier.NONE
conn.close()
def test_fresh_sample_grades_fresh(self, tmp_path):
now = _clock()
ts = (now - timedelta(seconds=FRESH_THRESHOLD_S - 10)).isoformat()
assert grade_freshness(ts, now) == "fresh"
def test_missed_sample_grades_missed(self, tmp_path):
now = _clock()
ts = (now - timedelta(hours=2)).isoformat()
assert grade_freshness(ts, now) == "missed"
def test_stale_sample_grades_stale(self, tmp_path):
now = _clock()
ts = (now - timedelta(hours=49)).isoformat()
assert grade_freshness(ts, now) == "stale"
def test_empty_store_grades_empty(self, tmp_path):
now = _clock()
assert grade_freshness(None, now) == "empty"
def test_unsupported_fresh(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100)
fresh_ts = (_clock() - timedelta(seconds=60)).isoformat()
_insert_sample(conn, fresh_ts)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert grade_freshness(fresh_ts, _clock()) == "fresh"
conn.close()
def test_limited_young_regime(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(5):
d = (datetime(2026, 9, 25) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
_insert_complete_local_days(conn, "2026-09-29", 1)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
conn.close()
def test_limited_warming(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(10):
d = (datetime(2026, 9, 20) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
_insert_complete_local_days(conn, "2026-09-29", 1)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
assert proj.warming_fact is not None
conn.close()
def test_limited_stale_data(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn, opened_at="2026-08-01T00:00:00+00:00")
_open_period(conn, start="2026-08-01T00:00:00+00:00")
for i in range(30):
d = (datetime(2026, 8, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
stale_ts = (_clock() - timedelta(days=5)).isoformat()
_insert_sample(conn, stale_ts)
_insert_complete_local_days(conn, "2026-08-30", 1)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
assert proj.staleness_fact is not None
conn.close()
def test_limited_degraded_identity(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn, identity_key=None, degraded=True)
_open_period(conn)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100, coverage=0.95, samples=24)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
_insert_complete_local_days(conn, "2026-09-29", 1)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.LIMITED
assert proj.degraded_identity_fact is not None
conn.close()
def test_unsupported_zero_rate(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=10.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(30):
d = (datetime(2026, 9, 1) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=0)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
_insert_complete_local_days(conn, "2026-09-29", 1)
proj = compute_projection(conn, _clock())
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
assert proj.zero_rate_fact is not None
conn.close()
def test_headline_present_when_projection_exists(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
proj = compute_projection(conn, _clock())
assert proj.headline_remaining_seconds is not None
assert proj.headline_remaining_seconds > 0
conn.close()
def test_headline_absent_when_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
proj = compute_projection(conn, _clock())
assert proj.headline_remaining_seconds is None
conn.close()
def test_headline_absent_when_zero_rate(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_baseline(conn, tbw_tb=1.0, verified=True)
_insert_segment(conn)
_open_period(conn)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=0)
proj = compute_projection(conn, _clock())
assert proj.headline_remaining_seconds is None
conn.close()
def test_facts_always_list(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
proj = compute_projection(conn, _clock())
assert isinstance(proj.contributing_facts, list)
conn.close()
def test_facts_never_empty_for_unavailable(self, tmp_path):
conn = init_store(tmp_path / "db")
_insert_segment(conn)
_open_period(conn)
proj = compute_projection(conn, _clock())
assert len(proj.contributing_facts) > 0
conn.close()
def test_confidence_never_percentage(self, tmp_path):
conn = init_store(tmp_path / "db")
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
proj = compute_projection(conn, _clock())
assert proj.confidence_state in (
ConfidenceState.UNSUPPORTED, ConfidenceState.LIMITED, ConfidenceState.SUPPORTED)
for f in proj.contributing_facts:
if re.match(r"^\\d+%$", f.strip()):
pytest.fail("Bare percentage in facts: %r" % f)
conn.close()
def test_status_renders_same_state_as_projection(self, tmp_path):
db = tmp_path / "observations.db"
conn = init_store(db)
_setup_full_store(conn, baseline_kw=dict(tbw_tb=10.0, verified=True))
conn.close()
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": True, "timer_active": True,
"last_collect_ok": True, "last_collect_age_s": 60,
"last_collect_reason": None,
}):
status = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "Supported" in status or "supported" in status.lower()
assert "remaining" in status.lower()
# ===================================================================
# CI-2: TUI/CLI parity
# ===================================================================
class TestCI2Parity:
"""Verify TUI and CLI share the same constants, formatting, and wording."""
def test_freshness_constants_shared(self):
from fenris import tui as tui_mod
from fenris import status as status_mod
assert tui_mod.FRESH_THRESHOLD_S == status_mod.FRESH_THRESHOLD_S
assert tui_mod.STALENESS_THRESHOLD_S == status_mod.STALENESS_THRESHOLD_S
def test_grade_freshness_shared(self):
from fenris.tui import grade_freshness as tui_gf
from fenris.status import grade_freshness as status_gf
assert tui_gf is status_gf
def test_disclosures_shared(self):
from fenris.projection import DISCLOSURES as proj_disc
from fenris.status import format_disclosures
output = format_disclosures()
for d in proj_disc:
assert d in output
def test_status_four_facts_match_tui_strip(self, tmp_path):
db = tmp_path / "observations.db"
conn = init_store(db)
_insert_segment(conn)
_open_period(conn)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(conn, d, bw=1024*1024*100)
_insert_sample(conn, "2026-09-30T10:00:00+00:00")
conn.close()
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": True, "timer_active": True,
"last_collect_ok": True, "last_collect_age_s": 120,
"last_collect_reason": None,
}):
status = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "boot:" in status
assert "timer:" in status
assert "last collect:" in status
assert "freshness:" in status
def test_dashboard_clarity_parity_strings_have_one_status_source(self):
"""DC-2/DC-3 wording originates in status and the TUI imports it."""
status_src = (FENRIS_PKG / "status.py").read_text()
tui_src = (FENRIS_PKG / "tui.py").read_text()
for wording in (
"monitoring: active in background · persists across reboots",
"monitoring: does not start on next boot",
"monitoring: paused — deliberate disable",
"paused time is excluded from your usage habit · resume: fenris monitor resume",
):
assert status_src.count(wording) == 1
assert wording not in tui_src
@pytest.mark.asyncio
@pytest.mark.parametrize(
("state", "service", "expected_lines"),
[
(
"active_enabled",
{"boot_enabled": True, "timer_active": True},
["monitoring: active in background · persists across reboots"],
),
(
"boot_disabled",
{"boot_enabled": False, "timer_active": False},
["monitoring: does not start on next boot"],
),
(
"deliberately_paused",
{"boot_enabled": False, "timer_active": False},
[
"monitoring: does not start on next boot",
"monitoring: paused — deliberate disable",
"paused time is excluded from your usage habit · resume: fenris monitor resume",
],
),
],
)
async def test_dashboard_clarity_monitoring_lines_match_both_views(
self, tmp_path, state, service, expected_lines
):
"""CI-2 synthetic-store sweep covers active, disabled, and paused states."""
db = tmp_path / (state + ".db")
conn = init_store(db)
if state == "active_enabled":
ensure_period_open(conn, _clock())
elif state == "deliberately_paused":
conn.execute(
"INSERT INTO monitoring_periods (started_at, ended_at, end_cause) "
"VALUES (?, ?, ?)",
("2026-09-30T09:00:00+00:00", "2026-09-30T10:00:00+00:00", "user_disabled"),
)
conn.commit()
conn.close()
service_state = {
**service,
"last_collect_ok": None,
"last_collect_age_s": None,
"last_collect_reason": None,
}
with patch("fenris.status.query_service_state", return_value=service_state), patch(
"fenris.status.query_service_state", return_value=service_state
):
status = get_status(
store_path=db, clock_now=_clock(), query_services=True, query_journal=False
).lower()
app = FenrisTuiApp(store_path=db)
async with app.run_test(size=(100, 40)):
tui_text = "\n".join(
(
str(app.query_one("#service-strip").render()),
str(app.query_one("#paused-banner").render()),
)
).lower()
for expected in expected_lines:
assert expected in status
assert expected in tui_text
if state != "deliberately_paused":
assert "monitoring: paused — deliberate disable" not in status
assert "monitoring: paused — deliberate disable" not in tui_text
def test_pause_resume_action_names(self):
tui_keys = {b.key for b in FenrisTuiApp.BINDINGS}
assert "p" in tui_keys
assert "r" in tui_keys
assert "c" in tui_keys
assert "q" in tui_keys
def test_empty_store_greeting_both_views(self, tmp_path):
db = tmp_path / "observations.db"
init_store(db)
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
status = get_status(store_path=db, clock_now=now,
query_services=True, query_journal=False)
assert "no observations yet" in status.lower()
def test_status_never_prompts(self):
status_src = (FENRIS_PKG / "status.py").read_text()
assert "input(" not in status_src
# ===================================================================
# CI-3: Prohibition set
# ===================================================================
class TestCI3ProhibitionSet:
"""Structural codebase checks for every prohibition clause."""
def _read_all_sources(self):
files = {}
for py in FENRIS_PKG.glob("*.py"):
files[py.name] = py.read_text()
return files
def test_single_acquisition_path(self):
"""Only fenris-collect may interrogate the device. [2.1, 8.7]
collector.py contains the acquisition functions; collect.py is the
fenris-collect entry point that invokes them. No other module may
reference smartctl.
"""
sources = self._read_all_sources()
allowed = {"collector.py", "collect.py"}
for name, text in sources.items():
if name in allowed:
continue
assert "smartctl" not in text, (
"%s must not contain smartctl" % name
)
def test_no_run_surface(self):
"""No /run/fenris coordination surface. [1.2, 3]"""
sources = self._read_all_sources()
for name, text in sources.items():
assert "/run/fenris" not in text, (
"%s references /run/fenris" % name
)
def test_single_config_key(self):
"""Config holds exactly one key: device. [8.3]"""
status_src = (FENRIS_PKG / "status.py").read_text()
in_read_config = False
config_keys = []
for line in status_src.split("\n"):
if "def read_config" in line:
in_read_config = True
elif in_read_config and line.strip().startswith("def "):
break
elif in_read_config and "key ==" in line:
match = re.search(r'key\s*==\s*["\']([^"\']+)["\']', line)
if match:
config_keys.append(match.group(1))
assert "device" in config_keys
assert len(config_keys) == 1, "Found keys: %s" % config_keys
def test_no_alerting_machinery(self):
"""No alerting, notification, or escalation. [9.6]"""
sources = self._read_all_sources()
alert_keywords = ["send_email", "smtp", "webhook", "push_notification"]
for name, text in sources.items():
for kw in alert_keywords:
for line in text.split("\n"):
stripped = line.strip()
if kw in stripped and not stripped.startswith("#"):
pytest.fail(
"%s contains alerting keyword '%s': %s" % (name, kw, stripped)
)
def test_no_synthetic_baselines(self):
"""No synthetic or capacity-derived baseline. [6.1]"""
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert "synthetic" not in proj_src.lower()
def test_no_stored_projections(self):
"""Projections never stored; recomputed on read. [3.7, 6.10]"""
store_src = (FENRIS_PKG / "store.py").read_text()
create_tables = re.findall(r"CREATE TABLE.*?(?=\n\n|$)", store_src, re.DOTALL)
table_names = []
for ct in create_tables:
m = re.search(r"IF NOT EXISTS\s+(\w+)", ct)
if m:
table_names.append(m.group(1))
assert "projection" not in [t.lower() for t in table_names]
def test_polkit_authorizes_one_binary(self):
"""Polkit authorizes exactly one binary: fenris-monitor. [8.5]"""
monitor_src = (FENRIS_PKG / "monitor.py").read_text()
assert "fenris-monitor" in monitor_src or "fenris_monitor" in monitor_src
collect_src = (FENRIS_PKG / "collect.py").read_text()
assert "polkit" not in collect_src.lower()
def test_no_hour_interpolation(self):
"""No absent hour is interpolated or fabricated. [5.3]"""
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert "interpolat" not in proj_src.lower()
assert "fabricat" not in proj_src.lower()
def test_fenris_sh_not_shipped(self):
"""fenris.sh is not shipped. [8.8]"""
repo_root = Path(__file__).parent.parent
assert not (repo_root / "fenris.sh").exists()
# ===================================================================
# CI-4: Required wording and six disclosures
# ===================================================================
class TestCI4WordingAndDisclosures:
"""Verify exact fixed phrases and disclosures in both views."""
def test_exactly_six_disclosures(self):
assert len(DISCLOSURES) == 6
def test_disclosure_1_endurance_not_failure(self):
assert "endurance projection" in DISCLOSURES[0].lower()
assert "hardware-failure" in DISCLOSURES[0].lower() or "failure date" in DISCLOSURES[0].lower()
def test_disclosure_2_vendor_specific(self):
assert "vendor-specific" in DISCLOSURES[1]
assert "255 is saturated" in DISCLOSURES[1]
def test_disclosure_3_warranty_not_failure(self):
assert "warranty" in DISCLOSURES[2].lower() or "endurance threshold" in DISCLOSURES[2].lower()
assert "failure threshold" in DISCLOSURES[2].lower()
def test_disclosure_4_duw_rounding(self):
assert "DUW" in DISCLOSURES[3]
assert "upward-rounded" in DISCLOSURES[3]
assert "NAND" in DISCLOSURES[3]
def test_disclosure_5_quality_depends(self):
assert "baseline provenance" in DISCLOSURES[4]
assert "future workload" in DISCLOSURES[4]
def test_disclosure_6_gaps_and_disabled(self):
assert "Gaps" in DISCLOSURES[5]
assert "deliberately disabled" in DISCLOSURES[5]
def test_disclosures_render_in_status(self):
output = format_disclosures()
assert output.startswith("Disclosures")
for i in range(1, 7):
assert "%d." % i in output
for d in DISCLOSURES:
assert d in output
def test_disclosures_render_in_tui(self):
tui_src = (FENRIS_PKG / "tui.py").read_text()
assert "format_disclosures" in tui_src
def test_zero_rate_phrase(self):
phrase = "no finite projection from this history"
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert phrase in proj_src
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
def test_unavailable_no_baseline_phrase(self):
phrase = "no applicable endurance baseline"
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert phrase in proj_src
def test_no_observations_phrase(self):
phrase = "no observations yet"
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
tui_src = (FENRIS_PKG / "tui.py").read_text()
assert phrase in tui_src.lower()
def test_config_error_phrase(self):
phrase = "configuration error:"
status_src = (FENRIS_PKG / "status.py").read_text()
assert phrase in status_src
def test_degraded_identity_phrase(self):
phrase = "controller identity unavailable"
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert phrase in proj_src
phrase2 = "replacement detection relies on write-counter continuity only"
assert phrase2 in proj_src
def test_scenario_range_only_spread(self):
proj_src = (FENRIS_PKG / "projection.py").read_text()
assert "confidence interval" not in proj_src.lower()
def test_no_percentage_in_confidence_rendering(self):
for name in ["projection.py", "tui.py", "status.py"]:
src = (FENRIS_PKG / name).read_text()
assert not re.search(r"\\d+%\\s*confidence", src, re.IGNORECASE), (
"Found XX%% confidence in %s" % name
)
def test_status_disclosures_accessible(self):
db = Path("/tmp/_ci4_test.db")
conn = init_store(db)
conn.close()
now = _clock()
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = render_status(store_path=db, clock_now=now,
query_services=True, query_journal=False,
show_disclosures=True)
assert "Disclosures" in result
assert "1." in result
assert "6." in result
db.unlink(missing_ok=True)
+52
View File
@@ -0,0 +1,52 @@
"""Visible geometry and evidence boundaries of the terminal volume plot."""
from fenris.activity_plot import VolumePoint, volume_plot
from fenris.themes import get_graph_colors
def plot(points, selected=-1, width=48, height=10):
return volume_plot(points, width, height, selected, get_graph_colors("chalktone"))
def test_trace_fits_viewport_and_uses_time_not_sample_index():
points = [VolumePoint(t, v, str(t)) for t, v in ((0, 0), (3, 50), (30, 100))]
text, columns, unit = plot(points)
assert unit == "B"
assert len(text.plain.splitlines()) == 10
assert all(len(line) <= 48 for line in text.plain.splitlines())
assert columns[1] - columns[0] < (columns[2] - columns[0]) / 5
assert any(0x2801 <= ord(c) <= 0x28ff for c in text.plain)
assert not any(c in text.plain for c in "█▒░")
assert "100.00" in text.plain and "0.00" in text.plain
def test_gap_breaks_trace_while_measured_zero_stays_on_axis():
text, columns, _ = plot([
VolumePoint(0, 0, "00:00"), VolumePoint(1, None, "00:03", "gap"),
VolumePoint(2, 100, "00:06"),
])
rows = text.plain.splitlines()
assert "?" in rows[-2]
# No invented intermediate dots on either side of the missing measurement.
for row in rows[:-2]:
assert all(c == " " for c in row[columns[0] + 1:columns[-1]])
assert 0x2801 <= ord(rows[-3][columns[0]]) <= 0x28ff
def test_partial_and_unallocated_totals_are_isolated_and_labelled():
text, columns, unit = plot([
VolumePoint(0, 1_000_000, "01", "partial"),
VolumePoint(1, 2_000_000, "02", "unallocated"),
VolumePoint(2, 3_000_000, "03"),
])
assert unit == "MB"
assert "~" in text.plain and "u" in text.plain
for row in text.plain.splitlines()[:-2]:
assert all(c == " " for c in row[columns[0] + 1:columns[1]])
def test_many_intervals_and_selection_fit_small_plot_without_losing_points():
points = [VolumePoint(i, i % 9 * 1_000_000, str(i)) for i in range(180)]
text, columns, _ = plot(points, selected=100, width=32, height=6)
assert len(columns) == len(points)
assert text.plain.splitlines()[-2][columns[100]] == "▼"
assert all(len(row) == 32 for row in text.plain.splitlines())
+94
View File
@@ -0,0 +1,94 @@
"""Activity-selection transitions independent from Textual widgets."""
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.activity_selection import ActivitySelection, IntervalIdentity
def _interval(n):
return IntervalIdentity(f"start-{n}", f"end-{n}")
def test_live_selection_starts_and_stays_on_newest_interval():
selection = ActivitySelection()
first = [_interval(1), _interval(2)]
selection.update_live(first)
assert selection.following_live
assert selection.selected_live_interval == first[-1]
second = first + [_interval(3)]
selection.update_live(second)
assert selection.selected_live_interval == second[-1]
def test_keyboard_and_mouse_inspection_pin_interval_identity():
intervals = [_interval(1), _interval(2), _interval(3)]
keyboard = ActivitySelection()
mouse = ActivitySelection()
keyboard.update_live(intervals)
keyboard.move_live(intervals, -1)
mouse.update_live(intervals)
mouse.inspect_live(intervals, intervals[1])
assert keyboard.selected_live_interval == mouse.selected_live_interval == intervals[1]
assert not keyboard.following_live
assert not mouse.following_live
refreshed = [_interval(2), _interval(3), _interval(4)]
keyboard.update_live(refreshed)
mouse.update_live(refreshed)
assert keyboard.selected_live_interval == mouse.selected_live_interval == intervals[1]
def test_expired_pin_explains_expiration_and_resumes_following():
selection = ActivitySelection()
intervals = [_interval(1), _interval(2), _interval(3)]
selection.update_live(intervals)
selection.inspect_live(intervals, intervals[0])
newest = _interval(4)
selection.update_live([_interval(2), _interval(3), newest])
assert selection.following_live
assert selection.selected_live_interval == newest
assert selection.live_interval_expired
def test_today_resets_expired_pin_and_read_write_toggle_keeps_identity():
selection = ActivitySelection()
intervals = [_interval(1), _interval(2)]
selection.update_live(intervals)
selection.inspect_live(intervals, intervals[0])
selection.update_live([intervals[1]])
newest = _interval(3)
selection.update_live([intervals[1], newest])
assert selection.live_interval_expired
selection.toggle_measure()
assert selection.measure == "read"
assert selection.selected_live_interval == newest
selection.select_history_date("2026-09-27")
selection.set_view("history")
selection.set_view("live")
assert selection.browse_date is None
assert selection.following_live
assert not selection.live_interval_expired
def test_live_identity_uses_both_interval_boundaries():
first = IntervalIdentity("2026-09-28T10:00:00Z", "2026-09-28T10:03:00Z")
changed_start = IntervalIdentity("2026-09-28T09:57:00Z", "2026-09-28T10:03:00Z")
selection = ActivitySelection()
selection.update_live([first])
selection.inspect_live([first], first)
selection.update_live([changed_start])
assert selection.following_live
assert selection.selected_live_interval == changed_start
assert selection.live_interval_expired
+308
View File
@@ -0,0 +1,308 @@
"""Release-note changelog extraction tests (DC-6, DC-7)."""
import importlib.util
import json
import subprocess
import sys
from pathlib import Path
import pytest
REPO_ROOT = Path(__file__).resolve().parent.parent
EXTRACTOR_PATH = REPO_ROOT / "scripts" / "extract_changelog.py"
RELEASE_REQUEST_PATH = REPO_ROOT / "scripts" / "release_request.py"
CHANGELOG_PATH = REPO_ROOT / "CHANGELOG.md"
def _extractor_module():
spec = importlib.util.spec_from_file_location("extract_changelog", EXTRACTOR_PATH)
assert spec and spec.loader
module = importlib.util.module_from_spec(spec)
sys.modules[spec.name] = module
spec.loader.exec_module(module)
return module
def test_extracts_the_requested_version_section_verbatim():
extractor = _extractor_module()
changelog = """# Changelog
## [Unreleased]
## [1.4.0] - 2026-09-10
### Added
- Show a release summary to consumers.
## [1.3.0] - 2026-09-01
### Fixed
- Preserve the observation history during upgrades.
"""
expected = """## [1.4.0] - 2026-09-10
### Added
- Show a release summary to consumers.
"""
assert extractor.extract_version_section(changelog, "1.4.0") == expected
def test_checked_in_changelog_keeps_unreleased_first_and_categories_limited():
lines = CHANGELOG_PATH.read_text(encoding="utf-8").splitlines()
unreleased = lines.index("## [Unreleased]")
version_headings = [
index for index, line in enumerate(lines)
if line.startswith("## [") and line != "## [Unreleased]"
]
first_version = version_headings[0] if version_headings else len(lines)
categories = [
line.removeprefix("### ")
for line in lines[unreleased + 1:first_version]
if line.startswith("### ")
]
assert unreleased < first_version
assert set(categories) <= {"Added", "Changed", "Fixed"}
def test_release_footer_verifies_the_clearsigned_checksum_asset():
footer = (REPO_ROOT / "packaging" / "release-footer.md").read_text(
encoding="utf-8"
)
assert "gpg --output SHA256SUMS --decrypt SHA256SUMS.asc" in footer
assert "sha256sum -c SHA256SUMS" in footer
def test_release_footer_configures_the_xbps_url_as_a_repository():
footer = (REPO_ROOT / "packaging" / "release-footer.md").read_text(
encoding="utf-8"
)
repository = (
"repository=https://git.bongbetic.com/xavierk/Fenris-xbps/"
"raw/branch/stable/x86_64"
)
assert repository in footer
assert "sudo xbps-install -M -S fenris" in footer
assert "sudo xbps-install -S https://git.bongbetic.com" not in footer
@pytest.mark.parametrize(
("changelog", "expected_error"),
[
("# Changelog\n\n## [Unreleased]\n", "missing"),
(
"# Changelog\n\n## [Unreleased]\n\n## [1.4.0] - 2026-09-10\n",
"empty",
),
(
"# Changelog\n\n## [Unreleased]\n\n## [1.4.0] - 2026-02-30\n\n- Add a note.\n",
"malformed release date",
),
],
)
def test_fails_closed_for_missing_empty_or_malformed_sections(
changelog, expected_error
):
extractor = _extractor_module()
with pytest.raises(extractor.ChangelogError, match=expected_error):
extractor.extract_version_section(changelog, "1.4.0")
def test_command_emits_a_workflow_error_and_nonzero_status(tmp_path):
changelog = tmp_path / "CHANGELOG.md"
changelog.write_text("# Changelog\n\n## [Unreleased]\n", encoding="utf-8")
result = subprocess.run(
[sys.executable, str(EXTRACTOR_PATH), str(changelog), "1.4.0"],
capture_output=True,
text=True,
check=False,
)
assert result.returncode != 0
assert result.stderr.startswith("::error::")
assert "missing" in result.stderr
def test_assembles_a_release_body_without_changing_the_section():
extractor = _extractor_module()
section = "## [1.4.0] - 2026-09-10\n\n### Added\n\n- Show a release summary.\n"
footer = "## Install\n\nUse the package channel.\n"
assert extractor.assemble_release_body(section, footer) == (
section + "\n" + footer
)
def test_command_can_write_the_complete_release_body(tmp_path):
changelog = tmp_path / "CHANGELOG.md"
changelog.write_text(
"# Changelog\n\n## [Unreleased]\n\n## [1.4.0] - 2026-09-10\n\n"
"### Added\n\n- Show a release summary.\n",
encoding="utf-8",
)
footer = tmp_path / "footer.md"
footer.write_text("## Install\n\nUse the package channel.\n", encoding="utf-8")
result = subprocess.run(
[
sys.executable,
str(EXTRACTOR_PATH),
str(changelog),
"1.4.0",
"--footer",
str(footer),
],
capture_output=True,
text=True,
check=False,
)
assert result.returncode == 0
assert result.stdout == (
"## [1.4.0] - 2026-09-10\n\n### Added\n\n- Show a release summary.\n\n"
"## Install\n\nUse the package channel.\n"
)
def test_release_request_command_reports_create_or_patch_decisions(tmp_path):
body = tmp_path / "release-body.md"
body.write_text("## [1.4.0] - 2026-09-10\n", encoding="utf-8")
create = subprocess.run(
[
sys.executable,
str(RELEASE_REQUEST_PATH),
"--version",
"1.4.0",
"--body-file",
str(body),
],
capture_output=True,
text=True,
check=False,
)
existing = tmp_path / "existing-release.json"
existing.write_text('{"id": 17, "assets": []}', encoding="utf-8")
patch = subprocess.run(
[
sys.executable,
str(RELEASE_REQUEST_PATH),
"--version",
"1.4.0",
"--body-file",
str(body),
"--existing-release",
str(existing),
],
capture_output=True,
text=True,
check=False,
)
assert create.returncode == patch.returncode == 0
assert json.loads(create.stdout) == {
"method": "POST",
"path": "/releases",
"payload": {
"tag_name": "v1.4.0",
"name": "v1.4.0",
"body": "## [1.4.0] - 2026-09-10\n",
},
}
assert json.loads(patch.stdout) == {
"method": "PATCH",
"path": "/releases/17",
"payload": {"body": "## [1.4.0] - 2026-09-10\n"},
}
def test_format_availability_section_with_both():
extractor = _extractor_module()
result = extractor.format_availability_section(
available=["Debian/Ubuntu (deb)", "Fedora/openSUSE (rpm)"],
withheld=["Void Linux (xbps) — pending host acceptance"],
)
assert "Available: Debian/Ubuntu (deb), Fedora/openSUSE (rpm)" in result
assert "Withheld: Void Linux (xbps) — pending host acceptance" in result
assert "## Package formats" in result
def test_format_availability_section_available_only():
extractor = _extractor_module()
result = extractor.format_availability_section(
available=["Debian/Ubuntu (deb)", "Fedora/openSUSE (rpm)", "Void Linux (xbps)"],
withheld=[],
)
assert "Available: Debian/Ubuntu (deb), Fedora/openSUSE (rpm), Void Linux (xbps)" in result
assert "Withheld" not in result
def test_format_availability_section_empty():
extractor = _extractor_module()
result = extractor.format_availability_section(available=[], withheld=[])
assert result == ""
def test_command_includes_format_availability(tmp_path):
changelog = tmp_path / "CHANGELOG.md"
changelog.write_text(
"# Changelog\n\n## [Unreleased]\n\n## [1.4.0] - 2026-09-10\n\n"
"### Added\n\n- Show a release summary.\n",
encoding="utf-8",
)
result = subprocess.run(
[
sys.executable,
str(EXTRACTOR_PATH),
str(changelog),
"1.4.0",
"--available",
"Debian/Ubuntu (deb)",
"--available",
"Fedora/openSUSE (rpm)",
"--withheld",
"Void Linux (xbps)",
],
capture_output=True,
text=True,
check=False,
)
assert result.returncode == 0
assert "## Package formats" in result.stdout
assert "Available: Debian/Ubuntu (deb), Fedora/openSUSE (rpm)" in result.stdout
assert "Withheld: Void Linux (xbps)" in result.stdout
def test_command_without_format_args_has_no_formats_section(tmp_path):
changelog = tmp_path / "CHANGELOG.md"
changelog.write_text(
"# Changelog\n\n## [Unreleased]\n\n## [1.4.0] - 2026-09-10\n\n"
"### Added\n\n- Show a release summary.\n",
encoding="utf-8",
)
result = subprocess.run(
[
sys.executable,
str(EXTRACTOR_PATH),
str(changelog),
"1.4.0",
],
capture_output=True,
text=True,
check=False,
)
assert result.returncode == 0
assert "## Package formats" not in result.stdout
+873
View File
@@ -0,0 +1,873 @@
"""Collector history tracer tests (issue #73).
Tests the end-to-end history pipeline: sample acquisition → interval
derivation → hour observation → day aggregate, with concurrent-read
safety, cross-hour handling, and display states.
Seams:
- write side: run_collection() → observation store
- read side: get_status(), compute_projection() → observation store
"""
import os
import sqlite3
from datetime import datetime, timedelta, timezone
from pathlib import Path
from typing import Any, Dict
import pytest
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.collector import run_collection, normalize_identity
from fenris.store import init_store, get_store_path, SCHEMA_VERSION
from fenris.monitoring_periods import ensure_period_open, close_period, get_open_period
from fenris.day_aggregate import derive_day, derive_all_days
# ---------------------------------------------------------------------------
# Fixtures
# ---------------------------------------------------------------------------
@pytest.fixture
def smartctl_fixture() -> Dict[str, Any]:
"""Minimal smartctl -a -j output with required fields."""
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0,
"temperature": 35,
"available_spare": 100,
"available_spare_threshold": 10,
"percentage_used": 5,
"data_units_written": 12345678,
"data_units_read": 9876543,
"power_on_hours": 8765,
"power_cycles": 1234,
"unsafe_shutdowns": 5,
"media_errors": 0,
"num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
@pytest.fixture
def sysfs_fixture_tree(tmp_path: Path) -> Path:
"""Create a minimal sysfs fixture tree with controller identity."""
ctrl_dir = tmp_path / "sys" / "class" / "nvme" / "nvme0"
ctrl_dir.mkdir(parents=True)
(ctrl_dir / "subsysnqn").write_text("nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc\n")
(ctrl_dir / "model").write_text("Samsung SSD 970 EVO Plus 1TB\n")
(ctrl_dir / "serial").write_text("S4EWNX0N123456\n")
(ctrl_dir / "firmware_rev").write_text("2B2QEXM7\n")
transport_dir = ctrl_dir / "transport"
transport_dir.mkdir()
(transport_dir / "address").write_text("0000:03:00.0")
(transport_dir / "trstring").write_text("pcie")
return tmp_path
@pytest.fixture
def config_fixture(tmp_path: Path) -> Dict[str, Any]:
"""Configuration fixture naming the device."""
return {
"device": "/dev/nvme0",
"store_path": str(tmp_path / "observations.db"),
}
class FakeClock:
"""Injected clock returning controlled time."""
def __init__(self, initial: datetime):
self.now = initial
def utcnow(self):
return self.now
def advance(self, **kwargs):
self.now = self.now + timedelta(**kwargs)
# ---------------------------------------------------------------------------
# Schema migration: existing data readable at real precision
# ---------------------------------------------------------------------------
class TestSchemaMigration:
"""Schema migration 1→2 preserves existing data (issue #73 AC1)."""
def test_migration_bumps_version(self, tmp_path):
"""Migration from v1 to v2 succeeds."""
from fenris.store import migrate_to_latest, SCHEMA_VERSION
# Create a v1 store directly (simulating pre-migration state)
db = tmp_path / "test.db"
conn = sqlite3.connect(str(db))
conn.execute("PRAGMA journal_mode=WAL")
# Create v1 schema manually
conn.execute("""
CREATE TABLE samples (
id INTEGER PRIMARY KEY AUTOINCREMENT,
ts TEXT NOT NULL,
device TEXT NOT NULL,
subnqn TEXT, sn TEXT, mn TEXT, fr TEXT,
capacity_bytes INTEGER, percentage_used INTEGER,
available_spare INTEGER, media_errors INTEGER,
power_on_hours INTEGER, power_cycles INTEGER,
unsafe_shutdowns INTEGER, temperature_c INTEGER,
data_units_written INTEGER, data_units_read INTEGER,
bytes_written INTEGER, bytes_read INTEGER,
critical_warning INTEGER
)
""")
conn.execute("""
CREATE TABLE hour_observations (
id INTEGER PRIMARY KEY AUTOINCREMENT,
hour TEXT NOT NULL UNIQUE,
active_seconds INTEGER DEFAULT 0, idle_seconds INTEGER DEFAULT 0,
powered_off_seconds INTEGER DEFAULT 0, unknown_seconds INTEGER DEFAULT 0,
bytes_written_delta INTEGER DEFAULT 0, bytes_read_delta INTEGER DEFAULT 0,
temperature_min INTEGER, temperature_avg REAL, temperature_max INTEGER,
sample_count INTEGER DEFAULT 0, coverage REAL DEFAULT 0.0
)
""")
conn.execute("""
CREATE TABLE day_aggregates (
id INTEGER PRIMARY KEY AUTOINCREMENT,
day TEXT NOT NULL UNIQUE,
active_seconds INTEGER DEFAULT 0, idle_seconds INTEGER DEFAULT 0,
powered_off_seconds INTEGER DEFAULT 0, unknown_seconds INTEGER DEFAULT 0,
bytes_written_delta INTEGER DEFAULT 0, bytes_read_delta INTEGER DEFAULT 0,
sample_count INTEGER DEFAULT 0, coverage REAL DEFAULT 0.0
)
""")
conn.execute("CREATE TABLE monitoring_periods (id INTEGER PRIMARY KEY AUTOINCREMENT, started_at TEXT NOT NULL, ended_at TEXT, end_cause TEXT)")
conn.execute("CREATE TABLE controller_segments (id INTEGER PRIMARY KEY AUTOINCREMENT, opened_at TEXT NOT NULL, identity_key TEXT, identity_degraded BOOLEAN DEFAULT 0, subnqn TEXT, sn TEXT, mn TEXT, fr TEXT, vid TEXT, ssvid TEXT, transport TEXT)")
conn.execute("CREATE TABLE endurance_baseline (id INTEGER PRIMARY KEY AUTOINCREMENT, tbw_terabytes REAL NOT NULL, source_url TEXT, document_revision TEXT, entry_date TEXT, model_string TEXT, nominal_capacity_bytes INTEGER, validated_by TEXT, verified BOOLEAN DEFAULT 0, created_at TEXT NOT NULL, updated_at TEXT NOT NULL)")
conn.execute("CREATE TABLE store_metadata (key TEXT PRIMARY KEY, value TEXT NOT NULL)")
conn.execute("PRAGMA user_version=1")
conn.execute("INSERT INTO samples (ts, device, mn, sn, fr, capacity_bytes, percentage_used, available_spare, media_errors, power_on_hours, power_cycles, unsafe_shutdowns, temperature_c, data_units_written, data_units_read, bytes_written, bytes_read, critical_warning) VALUES ('2026-09-01T12:00:00+00:00', '/dev/nvme0', 'Test', 'SN', 'FR', 1000000000000, 5, 100, 0, 1000, 100, 0, 35, 1000000, 500000, 512000000000, 256000000000, 0)")
conn.commit()
conn.close()
# Migrate
steps = migrate_to_latest(db)
assert steps == 5 # v1→v2→v3→v4→v5→v6
# Verify data preserved
conn = sqlite3.connect(str(db))
row = conn.execute("SELECT ts, mn FROM samples").fetchone()
version = conn.execute("PRAGMA user_version").fetchone()[0]
conn.close()
assert row[0] == "2026-09-01T12:00:00+00:00"
assert row[1] == "Test"
assert version == SCHEMA_VERSION
def test_newer_schema_refused(self, tmp_path):
"""Store with user_version > SCHEMA_VERSION is refused."""
db = tmp_path / "test.db"
conn = sqlite3.connect(str(db))
conn.execute("PRAGMA user_version=%d" % (SCHEMA_VERSION + 1))
conn.commit()
conn.close()
with pytest.raises(ValueError, match="newer Fenris"):
init_store(db)
def test_existing_data_preserved_after_migration(self, config_fixture,
smartctl_fixture,
sysfs_fixture_tree):
"""Existing sample data is not lost or modified by migration."""
clock = FakeClock(datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc))
# Write first sample
run_collection(smartctl_fixture, sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock)
conn = sqlite3.connect(config_fixture["store_path"])
row = conn.execute("SELECT ts, device, bytes_written FROM samples").fetchone()
conn.close()
assert row[0] == "2026-09-01T12:00:00+00:00"
assert row[1] == "/dev/nvme0"
assert row[2] == 12345678 * 512000
# ---------------------------------------------------------------------------
# Collector publishes samples, intervals, hour observations, day aggregates
# ---------------------------------------------------------------------------
class TestCollectorDerivation:
"""Collector derives hour observations and day aggregates (issue #73 AC2)."""
def _make_sample(self, duw_units: int, ts: str) -> Dict[str, Any]:
"""Build a smartctl fixture with specific DUW."""
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0,
"temperature": 35,
"available_spare": 100,
"available_spare_threshold": 10,
"percentage_used": 5,
"data_units_written": duw_units,
"data_units_read": 9876543,
"power_on_hours": 8765,
"power_cycles": 1234,
"unsafe_shutdowns": 5,
"media_errors": 0,
"num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
def test_same_hour_20mb_derives_hour_obs(self, config_fixture, sysfs_fixture_tree):
"""Two samples in same hour with 20 MB delta → hour_obs gets 20 MB."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
# DUW units: 12345678 * 512000 = ~6.3 TB; 20 MB = 20*1024*1024 / 512000 ≈ 40 units
duw1 = 12345678
duw2 = duw1 + 40 # ~20 MB more
clock1 = FakeClock(t1)
s1 = self._make_sample(duw1, t1.isoformat())
r1 = run_collection(s1, sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
assert r1["ok"]
clock2 = FakeClock(t2)
s2 = self._make_sample(duw2, t2.isoformat())
r2 = run_collection(s2, sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
assert r2["ok"]
# Check hour_observation was derived
conn = sqlite3.connect(config_fixture["store_path"])
hour = conn.execute(
"SELECT bytes_written_delta, sample_count FROM hour_observations WHERE hour LIKE '2026-09-01T12%'"
).fetchone()
conn.close()
assert hour is not None, "Hour observation should exist for 12:00"
assert hour[1] >= 2 # at least 2 samples contributed
# bytes_written_delta should be the 20 MB delta (40 * 512000 = 20480000)
assert hour[0] == 40 * 512000
def test_zero_delta_derives_hour_obs(self, config_fixture, sysfs_fixture_tree):
"""Two samples in same hour with no DUW change → 0 B written."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
duw = 12345678 # same for both
clock1 = FakeClock(t1)
s1 = self._make_sample(duw, t1.isoformat())
run_collection(s1, sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
clock2 = FakeClock(t2)
s2 = self._make_sample(duw, t2.isoformat())
run_collection(s2, sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
conn = sqlite3.connect(config_fixture["store_path"])
hour = conn.execute(
"SELECT bytes_written_delta FROM hour_observations WHERE hour LIKE '2026-09-01T12%'"
).fetchone()
conn.close()
assert hour is not None
assert hour[0] == 0
def test_cross_hour_100mb_unattributed(self, config_fixture, sysfs_fixture_tree):
"""Samples in different hours → delta is unattributed to any hour."""
t1 = datetime(2026, 9, 1, 11, 55, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
duw1 = 12345678
duw2 = duw1 + 200 # ~100 MB
clock1 = FakeClock(t1)
s1 = self._make_sample(duw1, t1.isoformat())
run_collection(s1, sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
clock2 = FakeClock(t2)
s2 = self._make_sample(duw2, t2.isoformat())
run_collection(s2, sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
conn = sqlite3.connect(config_fixture["store_path"])
# Hour observations should NOT contain the cross-hour delta
hour11 = conn.execute(
"SELECT bytes_written_delta FROM hour_observations WHERE hour LIKE '2026-09-01T11%'"
).fetchone()
hour12 = conn.execute(
"SELECT bytes_written_delta FROM hour_observations WHERE hour LIKE '2026-09-01T12%'"
).fetchone()
conn.close()
# Neither hour should have the full 100 MB delta attributed
# (they may have 0 or partial, but not 200*512000)
full_delta = 200 * 512000
if hour11 is not None:
assert hour11[0] != full_delta, "Hour 11 should not have full cross-hour delta"
if hour12 is not None:
assert hour12[0] != full_delta, "Hour 12 should not have full cross-hour delta"
def test_readonly_sees_consistent_snapshot(self, config_fixture, sysfs_fixture_tree):
"""A read-only reader sees valid pre- or post-publication snapshot."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
clock1 = FakeClock(t1)
run_collection(self._make_sample(12345678, t1.isoformat()),
sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
# Open read-only
ro_conn = sqlite3.connect(
"file:%s?mode=ro" % config_fixture["store_path"], uri=True
)
count_before = ro_conn.execute("SELECT COUNT(*) FROM samples").fetchone()[0]
ro_conn.close()
assert count_before == 1
# Write second sample
clock2 = FakeClock(t2)
run_collection(self._make_sample(12345718, t2.isoformat()),
sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
# Read-only reader sees 2 samples now
ro_conn2 = sqlite3.connect(
"file:%s?mode=ro" % config_fixture["store_path"], uri=True
)
count_after = ro_conn2.execute("SELECT COUNT(*) FROM samples").fetchone()[0]
ro_conn2.close()
assert count_after == 2
class TestCollectionAtomicity:
"""Collection exposes sample and derived evidence as one publication."""
@pytest.mark.asyncio
async def test_collection_is_visible_through_cli_and_tui_readers(
self, config_fixture, smartctl_fixture, sysfs_fixture_tree, monkeypatch
):
"""Ordinary readers observe matching published sample and activity evidence."""
from fenris.status import get_status, read_status
from fenris.tui import FenrisTuiApp
monkeypatch.setenv("TZ", "UTC")
service = {
"boot_enabled": True,
"timer_active": True,
"last_collect_ok": True,
}
monkeypatch.setattr("fenris.status.query_service_state", lambda: service)
now = datetime.now(timezone.utc).replace(
minute=35, second=0, microsecond=0,
)
first = {
**smartctl_fixture,
"nvme_smart_health_information_log": {
**smartctl_fixture["nvme_smart_health_information_log"],
"data_units_written": 12345678,
"data_units_read": 9876543,
},
}
second = {
**smartctl_fixture,
"nvme_smart_health_information_log": {
**smartctl_fixture["nvme_smart_health_information_log"],
"data_units_written": 12345698,
"data_units_read": 9876553,
},
}
sysfs_path = sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0"
assert run_collection(
first, sysfs_path, config_fixture, FakeClock(now - timedelta(minutes=5))
)["ok"] is True
assert run_collection(
second, sysfs_path, config_fixture, FakeClock(now)
)["ok"] is True
store_path = Path(config_fixture["store_path"])
with read_status(store_path, now, query_services=False) as (reader, composition):
assert reader is not None
assert composition.sample_count == 2
assert composition.day_count == 1
utc_day = reader.execute(
"SELECT bytes_written_delta, bytes_read_delta FROM day_aggregates"
).fetchone()
local_day = reader.execute(
"SELECT bytes_written, bytes_read FROM local_days"
).fetchone()
assert tuple(utc_day) == (10_240_000, 5_120_000)
assert tuple(local_day) == (10_240_000, 5_120_000)
cli_output = get_status(
store_path, now, query_services=True, query_journal=False
)
assert "Monitoring" in cli_output
app = FenrisTuiApp(store_path=store_path)
async with app.run_test(size=(100, 30)) as pilot:
await pilot.pause()
live_readout = str(app.query_one("#live-readout").render())
assert "W 0.010 GB" in live_readout
assert "R 0.005 GB" in live_readout
repeated_result = run_collection(
second, sysfs_path, config_fixture, FakeClock(now + timedelta(minutes=5))
)
assert repeated_result["ok"] is True
with read_status(
store_path, now + timedelta(minutes=5), query_services=False
) as (reader, composition):
assert reader is not None
assert composition.sample_count == 3
utc_day = reader.execute(
"SELECT bytes_written_delta, bytes_read_delta FROM day_aggregates"
).fetchone()
local_day = reader.execute(
"SELECT bytes_written, bytes_read FROM local_days"
).fetchone()
assert tuple(utc_day) == (10_240_000, 5_120_000)
assert tuple(local_day) == (10_240_000, 5_120_000)
def test_failed_local_day_publication_keeps_previous_publication(
self, config_fixture, smartctl_fixture, sysfs_fixture_tree, monkeypatch
):
"""A failed final derivation step leaves all prior reader state intact."""
from fenris.status import read_status
monkeypatch.setenv("TZ", "UTC")
now = datetime.now(timezone.utc).replace(second=0, microsecond=0)
first_clock = FakeClock(now)
first = {
**smartctl_fixture,
"nvme_smart_health_information_log": {
**smartctl_fixture["nvme_smart_health_information_log"],
"data_units_written": 12345678,
"data_units_read": 9876543,
},
}
first_result = run_collection(
first,
sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture,
first_clock,
)
assert first_result["ok"] is True, first_result
writer = sqlite3.connect(config_fixture["store_path"])
writer.execute(
"CREATE TRIGGER fail_local_day_publication "
"BEFORE INSERT ON local_days "
"BEGIN SELECT RAISE(ABORT, 'injected local-day publication failure'); END"
)
writer.commit()
writer.close()
next_time = now + timedelta(minutes=5)
second = {
**smartctl_fixture,
"nvme_smart_health_information_log": {
**smartctl_fixture["nvme_smart_health_information_log"],
"data_units_written": 12345698,
"data_units_read": 9876548,
},
}
failed_result = run_collection(
second,
sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture,
FakeClock(next_time),
)
assert failed_result["ok"] is False
assert "injected local-day publication failure" in failed_result["error"]
with read_status(
Path(config_fixture["store_path"]), next_time, query_services=False
) as (reader, composition):
assert reader is not None
assert composition.sample_count == 1
public_counts = reader.execute(
"SELECT (SELECT COUNT(*) FROM samples), "
"(SELECT COUNT(*) FROM controller_segments), "
"(SELECT COUNT(*) FROM monitoring_periods), "
"(SELECT COUNT(*) FROM hour_observations), "
"(SELECT COUNT(*) FROM day_aggregates), "
"(SELECT COUNT(*) FROM local_days)"
).fetchone()
assert tuple(public_counts) == (1, 1, 1, 0, 0, 0)
# ---------------------------------------------------------------------------
# Display states: awaiting first sample, awaiting another sample
# ---------------------------------------------------------------------------
class TestDisplayStates:
"""Display states for zero/one/two+ samples (issue #73 AC3)."""
def test_zero_samples_awaiting_first(self, tmp_path):
"""Zero samples → 'awaiting first sample' state."""
from fenris.status import get_status
db = tmp_path / "test.db"
init_store(db)
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
from unittest.mock import patch
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = get_status(store_path=db, clock_now=now,
query_services=False, query_journal=False)
assert "no observations yet" in result.lower() or "awaiting" in result.lower()
def test_one_sample_awaiting_another(self, tmp_path):
"""One sample → 'awaiting another sample' state."""
from fenris.status import get_status
db = tmp_path / "test.db"
conn = init_store(db)
conn.execute(
"INSERT INTO samples (ts, device, mn, sn, fr, capacity_bytes, "
"percentage_used, available_spare, media_errors, power_on_hours, "
"power_cycles, unsafe_shutdowns, temperature_c, "
"data_units_written, data_units_read, bytes_written, bytes_read, "
"critical_warning) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
("2026-09-01T12:00:00+00:00", "/dev/nvme0", "Test", "SN", "FR",
1000000000000, 5, 100, 0, 1000, 100, 0, 35,
1000000, 500000, 512000000000, 256000000000, 0),
)
conn.commit()
conn.close()
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
from unittest.mock import patch
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = get_status(store_path=db, clock_now=now,
query_services=False, query_journal=False)
# Should mention awaiting or insufficient data
lower = result.lower()
assert "awaiting" in lower or "another sample" in lower or "no projection" in lower
def test_one_sample_awaiting_another_in_tui(self, tmp_path):
"""One sample → TUI shows awaiting state."""
from fenris.tui import FenrisTuiApp
db = tmp_path / "test.db"
conn = init_store(db)
conn.execute(
"INSERT INTO samples (ts, device, mn, sn, fr, capacity_bytes, "
"percentage_used, available_spare, media_errors, power_on_hours, "
"power_cycles, unsafe_shutdowns, temperature_c, "
"data_units_written, data_units_read, bytes_written, bytes_read, "
"critical_warning) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
("2026-09-01T12:00:00+00:00", "/dev/nvme0", "Test", "SN", "FR",
1000000000000, 5, 100, 0, 1000, 100, 0, 35,
1000000, 500000, 512000000000, 256000000000, 0),
)
conn.commit()
conn.close()
now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
# Test the projection handles single sample
from fenris.projection import compute_projection, ConfidenceState
conn = sqlite3.connect(db)
proj = compute_projection(conn, now)
conn.close()
# With only one sample, projection should be unavailable
assert proj.confidence_state == ConfidenceState.UNSUPPORTED
def test_two_same_hour_samples_show_measured_usage(self, config_fixture, sysfs_fixture_tree):
"""Two compatible same-hour samples show measured usage."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
clock1 = FakeClock(t1)
run_collection(self._make_sample_helper(12345678), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
clock2 = FakeClock(t2)
run_collection(self._make_sample_helper(12345718), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
from fenris.status import get_status
from unittest.mock import patch
now = datetime(2026, 9, 1, 12, 10, 0, tzinfo=timezone.utc)
with patch("fenris.status.query_service_state", return_value={
"boot_enabled": False, "timer_active": False,
"last_collect_ok": None, "last_collect_age_s": None,
"last_collect_reason": None,
}):
result = get_status(store_path=Path(config_fixture["store_path"]),
clock_now=now, query_services=False, query_journal=False)
# Should not say "no observations" or "awaiting"
lower = result.lower()
assert "no observations yet" not in lower
def _make_sample_helper(self, duw_units: int) -> Dict[str, Any]:
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0, "temperature": 35,
"available_spare": 100, "available_spare_threshold": 10,
"percentage_used": 5, "data_units_written": duw_units,
"data_units_read": 9876543, "power_on_hours": 8765,
"power_cycles": 1234, "unsafe_shutdowns": 5,
"media_errors": 0, "num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
# ---------------------------------------------------------------------------
# Pause crossing and counter reset
# ---------------------------------------------------------------------------
class TestPauseCrossing:
"""Delta across monitoring period gap (issue #73 AC4)."""
def test_pause_crossing_preserves_prior_history(self, config_fixture, sysfs_fixture_tree):
"""Delta across a paused period preserves prior hour observations."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t_pause = datetime(2026, 9, 1, 13, 0, 0, tzinfo=timezone.utc)
t_resume = datetime(2026, 9, 1, 14, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 14, 5, 0, tzinfo=timezone.utc)
duw1 = 12345678
duw2 = duw1 + 100
# First sample (opens period)
clock1 = FakeClock(t1)
run_collection(self._make_sample_for_pause(duw1), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
# Pause
conn = sqlite3.connect(config_fixture["store_path"])
close_period(conn, t_pause, "user_disabled")
conn.close()
# Resume with new sample
clock_resume = FakeClock(t_resume)
run_collection(self._make_sample_for_pause(duw1), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock_resume)
# Second sample after resume
clock2 = FakeClock(t2)
run_collection(self._make_sample_for_pause(duw2), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
# Verify prior hour observations are intact
conn = sqlite3.connect(config_fixture["store_path"])
hour_12 = conn.execute(
"SELECT bytes_written_delta FROM hour_observations WHERE hour LIKE '2026-09-01T12%'"
).fetchone()
conn.close()
# Hour 12 should still have data from the first collection
assert hour_12 is not None, "Prior hour observation should be preserved"
def _make_sample_for_pause(self, duw_units: int) -> Dict[str, Any]:
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0, "temperature": 35,
"available_spare": 100, "available_spare_threshold": 10,
"percentage_used": 5, "data_units_written": duw_units,
"data_units_read": 9876543, "power_on_hours": 8765,
"power_cycles": 1234, "unsafe_shutdowns": 5,
"media_errors": 0, "num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
class TestCounterReset:
"""Counter reset / replacement opens new segment (issue #73 AC6)."""
def test_duw_decrease_opens_new_segment(self, config_fixture, sysfs_fixture_tree):
"""DUW decrease triggers new segment."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
duw1 = 12345678
duw2 = duw1 - 100 # decrease = reset
clock1 = FakeClock(t1)
run_collection(self._make_sample_for_reset(duw1), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
clock2 = FakeClock(t2)
run_collection(self._make_sample_for_reset(duw2), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
conn = sqlite3.connect(config_fixture["store_path"])
segments = conn.execute("SELECT COUNT(*) FROM controller_segments").fetchone()[0]
samples = conn.execute("SELECT COUNT(*) FROM samples").fetchone()[0]
conn.close()
# Should have 2 segments (new one opened for DUW decrease)
assert segments == 2
# Should have 2 samples
assert samples == 2
def _make_sample_for_reset(self, duw_units: int) -> Dict[str, Any]:
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0, "temperature": 35,
"available_spare": 100, "available_spare_threshold": 10,
"percentage_used": 5, "data_units_written": duw_units,
"data_units_read": 9876543, "power_on_hours": 8765,
"power_cycles": 1234, "unsafe_shutdowns": 5,
"media_errors": 0, "num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
# ---------------------------------------------------------------------------
# Derivation failure preserves prior history
# ---------------------------------------------------------------------------
class TestDerivationFailure:
"""Injected derivation failure preserves prior history (issue #73 AC6)."""
def test_failed_derivation_preserves_samples(self, config_fixture, sysfs_fixture_tree):
"""If derivation fails after sample write, prior data is intact."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
clock1 = FakeClock(t1)
run_collection(self._make_sample_for_failure(12345678),
sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
# Verify first sample exists
conn = sqlite3.connect(config_fixture["store_path"])
count = conn.execute("SELECT COUNT(*) FROM samples").fetchone()[0]
conn.close()
assert count == 1
# Second sample with valid data should succeed
clock2 = FakeClock(t2)
r2 = run_collection(self._make_sample_for_failure(12345718),
sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
assert r2["ok"]
# Both samples should exist
conn = sqlite3.connect(config_fixture["store_path"])
count = conn.execute("SELECT COUNT(*) FROM samples").fetchone()[0]
conn.close()
assert count == 2
def _make_sample_for_failure(self, duw_units: int) -> Dict[str, Any]:
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0, "temperature": 35,
"available_spare": 100, "available_spare_threshold": 10,
"percentage_used": 5, "data_units_written": duw_units,
"data_units_read": 9876543, "power_on_hours": 8765,
"power_cycles": 1234, "unsafe_shutdowns": 5,
"media_errors": 0, "num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
# ---------------------------------------------------------------------------
# Monitoring period is opened by collector
# ---------------------------------------------------------------------------
class TestMonitoringPeriod:
"""Collector ensures monitoring period is open (issue #73 AC2)."""
def test_first_sample_opens_period(self, config_fixture, sysfs_fixture_tree):
"""First collection run opens a monitoring period."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
clock1 = FakeClock(t1)
run_collection(self._make_sample_simple(), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
conn = sqlite3.connect(config_fixture["store_path"])
period = get_open_period(conn)
conn.close()
assert period is not None, "A monitoring period should be open"
def test_subsequent_sample_keeps_period_open(self, config_fixture, sysfs_fixture_tree):
"""Subsequent collection runs keep the period open."""
t1 = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 1, 12, 5, 0, tzinfo=timezone.utc)
clock1 = FakeClock(t1)
run_collection(self._make_sample_simple(), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock1)
clock2 = FakeClock(t2)
run_collection(self._make_sample_simple(), sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config_fixture, clock2)
conn = sqlite3.connect(config_fixture["store_path"])
period = get_open_period(conn)
periods_count = conn.execute("SELECT COUNT(*) FROM monitoring_periods").fetchone()[0]
conn.close()
assert period is not None
assert periods_count == 1 # Still only one period
def _make_sample_simple(self) -> Dict[str, Any]:
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0, "temperature": 35,
"available_spare": 100, "available_spare_threshold": 10,
"percentage_used": 5, "data_units_written": 12345678,
"data_units_read": 9876543, "power_on_hours": 8765,
"power_cycles": 1234, "unsafe_shutdowns": 5,
"media_errors": 0, "num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
+363
View File
@@ -0,0 +1,363 @@
"""Collector tracer bullet test.
Tests the thinnest complete write path through the system:
- Input: smartctl-JSON fixture, sysfs fixture tree, config fixture, injected clock
- Output: resulting store contents, run outcome
Seam: write side of the observation store database file.
"""
import json
import os
import sqlite3
import tempfile
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Dict, Generator
import pytest
# Add src to path for imports
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.collector import run_collection, AcquisitionError, InvariantViolationError
from fenris.store import init_store, get_store_path
# Fixtures
@pytest.fixture
def smartctl_fixture() -> Dict[str, Any]:
"""Minimal smartctl -a -j output with required fields."""
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155", "build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0,
"temperature": 35,
"available_spare": 100,
"available_spare_threshold": 10,
"percentage_used": 5,
"data_units_written": 12345678,
"data_units_read": 9876543,
"power_on_hours": 8765,
"power_cycles": 1234,
"unsafe_shutdowns": 5,
"media_errors": 0,
"num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
@pytest.fixture
def sysfs_fixture_tree(tmp_path: Path) -> Path:
"""Create a minimal sysfs fixture tree with controller identity."""
ctrl_dir = tmp_path / "sys" / "class" / "nvme" / "nvme0"
ctrl_dir.mkdir(parents=True)
# Controller identity files
(ctrl_dir / "subsysnqn").write_text("nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc\n")
(ctrl_dir / "model").write_text("Samsung SSD 970 EVO Plus 1TB\n")
(ctrl_dir / "serial").write_text("S4EWNX0N123456\n")
(ctrl_dir / "firmware_rev").write_text("2B2QEXM7\n")
# Transport info (optional, but we'll include it)
transport_dir = ctrl_dir / "transport"
transport_dir.mkdir()
(transport_dir / "address").write_text("0000:03:00.0")
(transport_dir / "trstring").write_text("pcie")
return tmp_path
@pytest.fixture
def config_fixture(tmp_path: Path) -> Dict[str, Any]:
"""Configuration fixture naming the device."""
return {
"device": "/dev/nvme0",
"store_path": str(tmp_path / "observations.db"),
}
@pytest.fixture
def clock_fixture():
"""Injected clock returning fixed time."""
class FakeClock:
def __init__(self):
self.now = datetime(2026, 9, 1, 12, 0, 0, tzinfo=timezone.utc)
def utcnow(self):
return self.now
return FakeClock()
# Test: Collector writes one well-formed sample
def test_collector_writes_one_sample(
smartctl_fixture: Dict[str, Any],
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given a smartctl fixture and sysfs fixture tree,
when the collector runs,
then one well-formed sample is written to the observation store."""
result = run_collection(
smartctl_data=smartctl_fixture,
sysfs_path=sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config=config_fixture,
clock=clock_fixture,
)
# Verify run succeeded
assert result["ok"] is True, f"Collection failed: {result.get('error')}"
# Verify store contents
conn = sqlite3.connect(config_fixture["store_path"])
cursor = conn.execute("SELECT COUNT(*) FROM samples")
count = cursor.fetchone()[0]
assert count == 1
cursor = conn.execute("SELECT * FROM samples")
row = cursor.fetchone()
assert row is not None
# Verify row contents match fixtures
# Row structure: id, ts, device, subnqn, sn, mn, fr, capacity_bytes,
# percentage_used, available_spare, media_errors, power_on_hours, power_cycles,
# unsafe_shutdowns, temperature_c, data_units_written, data_units_read,
# bytes_written, bytes_read, critical_warning
assert row[2] == "/dev/nvme0" # device
assert row[3] == "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc" # subnqn
assert row[4] == "S4EWNX0N123456" # sn
assert row[5] == "Samsung SSD 970 EVO Plus 1TB" # mn
assert row[6] == "2B2QEXM7" # fr
assert row[7] == 1024000000000 # capacity_bytes
assert row[8] == 5 # percentage_used
assert row[15] == 12345678 # data_units_written
assert row[17] == 12345678 * 512000 # bytes_written
conn.close()
# Test: Identity normalization applied exactly once at write time
def test_identity_normalization(
smartctl_fixture: Dict[str, Any],
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given sysfs identity with trailing spaces/newlines,
when the collector writes,
then identity is normalized exactly once at write time."""
# Create identity with trailing whitespace
identity = {
"subnqn": "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc \n",
"mn": "Samsung SSD 970 EVO Plus 1TB\n",
"sn": "S4EWNX0N123456\n",
"fr": "2B2QEXM7",
"transport": "pcie",
}
from fenris.collector import normalize_identity
# Normalize once
key1 = normalize_identity(identity)
# Normalize again - should be identical
key2 = normalize_identity(identity)
assert key1 == key2
assert key1 == "nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc"
assert "\n" not in key1
assert key1 == key1.rstrip() # No trailing whitespace
# Test: Any acquisition failure fails the whole run
def test_acquisition_failure_fails_run(
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given a smartctl fixture with missing fields,
when the collector runs,
then the whole run fails and writes nothing."""
# Missing required field
bad_smartctl = {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3]},
# Missing nvme_smart_health_information_log
}
result = run_collection(
smartctl_data=bad_smartctl,
sysfs_path=sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config=config_fixture,
clock=clock_fixture,
)
# Verify run failed
assert result["ok"] is False
assert "Missing required field" in result["error"]
# Verify nothing was written
if os.path.exists(config_fixture["store_path"]):
conn = sqlite3.connect(config_fixture["store_path"])
cursor = conn.execute("SELECT COUNT(*) FROM samples")
count = cursor.fetchone()[0]
assert count == 0
conn.close()
else:
# Store wasn't even created - also valid
pass
# Test: Store initializes with six entities
def test_store_initialization(config_fixture: Dict[str, Any]):
"""Given no existing store,
when the collector runs,
then the store is initialized with six entities."""
store_path = Path(config_fixture["store_path"])
# Store shouldn't exist yet
assert not store_path.exists()
# Initialize store
conn = init_store(store_path)
# Verify the observation, derived-history, and metadata tables exist.
cursor = conn.execute("SELECT name FROM sqlite_master WHERE type='table'")
tables = {row[0] for row in cursor.fetchall()}
expected_tables = {
"samples",
"hour_observations",
"day_aggregates",
"monitoring_periods",
"controller_segments",
"endurance_baseline",
}
# sqlite_sequence is a system table created by AUTOINCREMENT
expected_tables.add("sqlite_sequence")
expected_tables.add("store_metadata")
expected_tables.add("local_days")
expected_tables.add("pending_publications")
expected_tables.add("local_day_unallocated_evidence")
expected_tables.add("local_day_segment_totals")
assert expected_tables == tables
conn.close()
# Test: Schema versioning with PRAGMA user_version
def test_schema_versioning(config_fixture: Dict[str, Any]):
"""Given a store with unknown newer version,
when the collector runs,
then it refuses to proceed."""
store_path = Path(config_fixture["store_path"])
# Create a store with newer version
conn = sqlite3.connect(str(store_path))
conn.execute("PRAGMA journal_mode=WAL")
conn.execute("PRAGMA user_version=999") # Unknown newer version
conn.close()
# Try to initialize - should fail
with pytest.raises(ValueError, match="newer Fenris"):
init_store(store_path)
def test_schema_version_current(config_fixture: Dict[str, Any]):
"""Given a store with current version,
when the collector runs,
then it proceeds without migration."""
from fenris.store import SCHEMA_VERSION
store_path = Path(config_fixture["store_path"])
# Initialize store
conn1 = init_store(store_path)
conn1.close()
# Open again - should succeed
conn2 = init_store(store_path)
# Verify version is current
cursor = conn2.execute("PRAGMA user_version")
version = cursor.fetchone()[0]
assert version == SCHEMA_VERSION
conn2.close()
def test_schema_version_older(config_fixture: Dict[str, Any]):
"""Given a store with older version,
when the collector runs,
then it applies migrations and proceeds."""
from fenris.store import SCHEMA_VERSION
store_path = Path(config_fixture["store_path"])
# Create a store with older version
conn = sqlite3.connect(str(store_path))
conn.execute("PRAGMA journal_mode=WAL")
conn.execute("PRAGMA user_version=0") # Older version
conn.close()
# Initialize store - should apply migrations
conn = init_store(store_path)
# Verify version is current
cursor = conn.execute("PRAGMA user_version")
version = cursor.fetchone()[0]
assert version == SCHEMA_VERSION
conn.close()
# Test: Invariant-violating run writes nothing
def test_invariant_violation_writes_nothing(
smartctl_fixture: Dict[str, Any],
sysfs_fixture_tree: Path,
config_fixture: Dict[str, Any],
clock_fixture,
):
"""Given a smartctl fixture that would violate store invariants,
when the collector runs,
then it writes nothing and fails visibly."""
# Create a fixture that would cause invariant violation
# (negative bytes_written - we'll mock this)
bad_smartctl = smartctl_fixture.copy()
bad_smartctl["nvme_smart_health_information_log"] = {
**smartctl_fixture["nvme_smart_health_information_log"],
"data_units_written": -1, # This will cause negative bytes_written
}
result = run_collection(
smartctl_data=bad_smartctl,
sysfs_path=sysfs_fixture_tree / "sys" / "class" / "nvme" / "nvme0",
config=config_fixture,
clock=clock_fixture,
)
# Verify run failed
assert result["ok"] is False
assert "InvariantViolation" in result["error_type"]
# Verify nothing was written
if os.path.exists(config_fixture["store_path"]):
conn = sqlite3.connect(config_fixture["store_path"])
cursor = conn.execute("SELECT COUNT(*) FROM samples")
count = cursor.fetchone()[0]
assert count == 0
conn.close()
+332
View File
@@ -0,0 +1,332 @@
"""Complete observation day gate tests (issue #94).
Verifies that the endurance projection is withheld until at least one
complete local calendar day has been observed within a monitoring period.
Seams:
- compute_projection() → gate check via local_days table
- ProjectionResult.contributing_facts → "waiting for a full local observation day"
Acceptance criteria:
- Gate-1: No complete local day → UNSUPPORTED with waiting fact
- Gate-2: One complete local day → Limited confidence (if other conditions met)
- Gate-3: Partial days don't satisfy the gate
- Gate-4: CLI and TUI share the same gate via compute_projection()
"""
import sqlite3
from datetime import datetime, timedelta, timezone
from pathlib import Path
import pytest
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.store import init_store
from fenris.monitoring_periods import ensure_period_open
from fenris.projection import (
compute_projection, ConfidenceState, BaselineTier,
WARMING_COVERAGE_FLOOR,
)
@pytest.fixture
def store(tmp_path):
conn = init_store(tmp_path / "test.db")
yield conn
conn.close()
def _clock(year=2026, month=9, day=30, hour=12):
return datetime(year, month, day, hour, 0, 0, tzinfo=timezone.utc)
def _insert_baseline(conn, tbw_tb=1.0, verified=True):
conn.execute(
"INSERT INTO endurance_baseline "
"(tbw_terabytes, source_url, document_revision, entry_date, model_string, "
" nominal_capacity_bytes, validated_by, verified, created_at, updated_at) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(tbw_tb, "https://example.com/spec", "v1.0", "2026-01-01",
"Samsung SSD 970 EVO Plus 1TB", 1024000000000,
"machine_match" if verified else None, verified,
"2026-01-01T00:00:00+00:00", "2026-01-01T00:00:00+00:00"),
)
conn.commit()
def _insert_segment(conn, opened_at="2026-09-01T00:00:00+00:00"):
conn.execute(
"INSERT INTO controller_segments "
"(opened_at, identity_key, identity_degraded, subnqn, sn, mn, fr, vid, ssvid, transport) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(opened_at, "nqn.test", False, "nqn.test", "SN123",
"Samsung SSD 970 EVO Plus 1TB", "FW1", "0x144d", "0x144d", "pcie"),
)
conn.commit()
def _insert_day(conn, day, bw=1024*1024*100, coverage=0.95, samples=24):
conn.execute(
"INSERT INTO day_aggregates (day, active_seconds, idle_seconds, powered_off_seconds, "
"unknown_seconds, bytes_written_delta, bytes_read_delta, sample_count, coverage) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)",
(day, 3600, 0, 0, 0, bw, 0, samples, coverage),
)
conn.commit()
def _insert_sample(conn, ts, pu=5):
conn.execute(
"INSERT INTO samples (ts, device, data_units_written, data_units_read, "
"percentage_used, bytes_written, bytes_read, power_on_hours) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?)",
(ts, "/dev/nvme0n1", 1000000, 500000, pu, 512000000000, 256000000000, 8765),
)
conn.commit()
def _insert_local_day(conn, local_date, tz_name="UTC", tz_offset="+00:00",
utc_start=None, utc_end=None, bw=1024*1024*100,
br=0, coverage=0.95, samples=24, complete=True):
"""Insert a local_days row for testing the gate."""
if utc_start is None:
utc_start = local_date + "T00:00:00+00:00"
if utc_end is None:
# Next day
dt = datetime.strptime(local_date, "%Y-%m-%d") + timedelta(days=1)
utc_end = dt.strftime("%Y-%m-%dT00:00:00+00:00")
conn.execute(
"INSERT INTO local_days "
"(local_date, tz_name, tz_offset, utc_start, utc_end, "
" bytes_written, bytes_read, coverage, sample_count, complete, activity_intervals) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, 1)",
(local_date, tz_name, tz_offset, utc_start, utc_end,
bw, br, coverage, samples, complete),
)
conn.commit()
def _open_period(conn, start="2026-09-01T00:00:00+00:00"):
ensure_period_open(conn, datetime.fromisoformat(start))
# ---------------------------------------------------------------------------
# Gate-1: No complete local day → UNSUPPORTED with waiting fact
# ---------------------------------------------------------------------------
class TestGateNoCompleteDay:
"""Projection is unavailable before any complete local observation day."""
def test_no_local_days_unsupported(self, store):
"""With no local_days entries, projection is UNSUPPORTED."""
_insert_baseline(store)
_insert_segment(store)
_open_period(store)
# 14 days of UTC data — enough for normal projection, but no local_days
for i in range(14):
d = (datetime(2026, 9, 15) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("full local observation day" in f for f in result.contributing_facts)
assert result.headline_remaining_seconds is None
def test_only_partial_local_days_unsupported(self, store):
"""Partial (incomplete) local days don't satisfy the gate."""
_insert_baseline(store)
_insert_segment(store)
_open_period(store)
for i in range(14):
d = (datetime(2026, 9, 15) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
# Insert only incomplete local days
for i in range(5):
d = (datetime(2026, 9, 25) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_local_day(store, d, complete=False, coverage=0.3)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("full local observation day" in f for f in result.contributing_facts)
def test_gate_before_warming_check(self, store):
"""Gate fires even when warming would also block — gate has precedence."""
_insert_baseline(store)
_insert_segment(store)
_open_period(store)
# Only 3 days of data (below warming threshold)
for i in range(3):
d = (datetime(2026, 9, 27) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
# Insert one complete local day — but still below warming
_insert_local_day(store, "2026-09-29", complete=True)
result = compute_projection(store, _clock())
# Gate is satisfied (one complete day), but warming blocks Supported
# The key assertion: gate message should NOT appear when gate IS met
assert not any("full local observation day" in f for f in result.contributing_facts)
def test_complete_legacy_day_without_trusted_activity_does_not_open_gate(
self, store,
):
"""A complete flag cannot make unavailable local activity qualify."""
_insert_baseline(store)
_insert_segment(store)
_open_period(store)
for i in range(3):
day = (datetime(2026, 9, 27) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, day, bw=1024 * 1024 * 100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
_insert_local_day(store, "2026-09-29", complete=True)
store.execute(
"UPDATE local_days SET activity_precision = 'measured', activity_intervals = 0 "
"WHERE local_date = '2026-09-29'"
)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("full local observation day" in fact
for fact in result.contributing_facts)
def test_day_split_by_deliberate_pause_does_not_open_gate(self, store):
"""A complete-looking summary cannot span separate monitoring periods."""
_insert_baseline(store)
_insert_segment(store)
for i in range(20):
day = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, day, bw=1024 * 1024 * 100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
store.executemany(
"INSERT INTO monitoring_periods (started_at, ended_at, end_cause) "
"VALUES (?, ?, ?)",
[
("2026-09-29T00:00:00+00:00", "2026-09-29T12:00:00+00:00", "user_disabled"),
("2026-09-29T13:00:00+00:00", "2026-09-30T00:00:00+00:00", "user_disabled"),
("2026-09-30T00:00:00+00:00", None, None),
],
)
store.commit()
_insert_local_day(
store,
"2026-09-29",
utc_start="2026-09-29T00:00:00+00:00",
utc_end="2026-09-30T00:00:00+00:00",
complete=True,
)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("full local observation day" in fact
for fact in result.contributing_facts)
def test_partial_but_trusted_local_activity_opens_gate(self, store):
"""Known local intervals can coexist with an incomplete day total."""
_insert_baseline(store)
_insert_segment(store)
_open_period(store)
for i in range(3):
day = (datetime(2026, 9, 27) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, day, bw=1024 * 1024 * 100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
_insert_local_day(store, "2026-09-29", complete=True)
store.execute(
"UPDATE local_days SET activity_incomplete = 1 "
"WHERE local_date = '2026-09-29'"
)
result = compute_projection(store, _clock())
assert not any("full local observation day" in fact
for fact in result.contributing_facts)
# ---------------------------------------------------------------------------
# Gate-2: One complete local day → Limited confidence
# ---------------------------------------------------------------------------
class TestGateOneCompleteDay:
"""After one complete local day, projection can proceed with Limited confidence."""
def test_one_complete_day_allows_projection(self, store):
"""With one complete local day and valid baseline/rate, projection is Limited."""
_insert_baseline(store, tbw_tb=1.0, verified=True)
_insert_segment(store)
_open_period(store)
# 14 days of UTC data
for i in range(14):
d = (datetime(2026, 9, 15) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
# One complete local day
_insert_local_day(store, "2026-09-29", complete=True)
result = compute_projection(store, _clock())
# Gate satisfied — no "waiting" fact
assert not any("full local observation day" in f for f in result.contributing_facts)
# With only 14 days and other Limited factors, should be Limited or Supported
assert result.confidence_state in (ConfidenceState.LIMITED, ConfidenceState.SUPPORTED)
# Headline should exist (rate > 0, baseline exists)
assert result.headline_remaining_seconds is not None
def test_gate_fact_absent_when_satisfied(self, store):
"""The 'waiting for full day' fact does not appear when gate is met."""
_insert_baseline(store)
_insert_segment(store)
_open_period(store)
for i in range(20):
d = (datetime(2026, 9, 10) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
_insert_local_day(store, "2026-09-29", complete=True)
result = compute_projection(store, _clock())
assert not any("full local observation day" in f for f in result.contributing_facts)
assert result.confidence_state != ConfidenceState.UNSUPPORTED
# ---------------------------------------------------------------------------
# Gate-3: Partial first day doesn't satisfy the gate
# ---------------------------------------------------------------------------
class TestGatePartialFirstDay:
"""Starting monitoring at noon means the partial first day doesn't count."""
def test_partial_first_day_not_enough(self, store):
"""A single incomplete local day (started at noon) doesn't open the gate."""
_insert_baseline(store)
_insert_segment(store, opened_at="2026-09-29T12:00:00+00:00")
_open_period(store, start="2026-09-29T12:00:00+00:00")
# Only Sep 29 (partial) and Sep 30 (today, partial)
_insert_day(store, "2026-09-29", bw=1024*1024*100, coverage=0.5, samples=12)
_insert_day(store, "2026-09-30", bw=1024*1024*100, coverage=0.5, samples=12)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
# Only partial local days
_insert_local_day(store, "2026-09-29", complete=False, coverage=0.5)
_insert_local_day(store, "2026-09-30", complete=False, coverage=0.5)
result = compute_projection(store, _clock())
assert result.confidence_state == ConfidenceState.UNSUPPORTED
assert any("full local observation day" in f for f in result.contributing_facts)
# ---------------------------------------------------------------------------
# Gate-4: Multiple complete days also satisfy the gate
# ---------------------------------------------------------------------------
class TestGateMultipleCompleteDays:
"""Multiple complete local days satisfy the gate."""
def test_multiple_complete_days_satisfy_gate(self, store):
"""Several complete local days open the gate."""
_insert_baseline(store)
_insert_segment(store)
_open_period(store)
for i in range(14):
d = (datetime(2026, 9, 15) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_day(store, d, bw=1024*1024*100)
_insert_sample(store, "2026-09-30T10:00:00+00:00", pu=5)
# 7 complete local days
for i in range(7):
d = (datetime(2026, 9, 23) + timedelta(days=i)).strftime("%Y-%m-%d")
_insert_local_day(store, d, complete=True)
result = compute_projection(store, _clock())
assert not any("full local observation day" in f for f in result.contributing_facts)
assert result.headline_remaining_seconds is not None
+77
View File
@@ -0,0 +1,77 @@
"""The CLI and TUI cross the same terminal-attached action interface."""
import argparse
import runpy
import shlex
import subprocess
import sys
from pathlib import Path
from unittest.mock import patch
import pytest
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.control import MONITOR_HELPER, MonitorError, run_monitor
from fenris.tui import FenrisTuiApp
@pytest.mark.parametrize("uid", [0, 1000])
@pytest.mark.parametrize("args", [("enable", "--now"), ("disable", "--now"), ("collect",)])
def test_fixed_helper_and_terminal_attachment(uid, args):
with patch("fenris.control.os.geteuid", return_value=uid), \
patch("fenris.control.subprocess.run", return_value=subprocess.CompletedProcess([], 0)) as run:
run_monitor(*args)
expected = [MONITOR_HELPER, *args]
if uid:
expected.insert(0, "pkexec")
# Inherited stdin/out/err keep authentication on the user's terminal.
# No frontend deadline can cut short a valid 90-second Collection run.
run.assert_called_once_with(expected)
@pytest.mark.parametrize("code", [1, 126, 127, -15])
def test_failure_is_not_retried_and_root_equivalent_preserves_arguments(code):
args = ("baseline", "set", '{"model": "Drive $(whoami)", "tbw": 100}')
with patch("fenris.control.os.geteuid", return_value=1000), \
patch("fenris.control.subprocess.run", return_value=subprocess.CompletedProcess([], code)) as run:
with pytest.raises(MonitorError) as failure:
run_monitor(*args)
run.assert_called_once()
assert failure.value.exit_code == (code if code > 0 else 128 - code)
root_command = str(failure.value).split("run in your terminal: ")[1]
assert shlex.split(root_command) == ["sudo", MONITOR_HELPER, *args]
@pytest.mark.parametrize("uid,missing", [(0, MONITOR_HELPER), (1000, "pkexec")])
def test_missing_command_is_identified(uid, missing):
with patch("fenris.control.os.geteuid", return_value=uid), \
patch("fenris.control.subprocess.run", side_effect=FileNotFoundError(2, "Missing", missing)):
with pytest.raises(MonitorError, match="Command not found") as failure:
run_monitor("collect")
assert missing in str(failure.value)
assert failure.value.exit_code == 127
assert ("sudo" in str(failure.value)) == bool(uid)
def test_interrupt_reports_uncertain_outcome():
with patch("fenris.control.subprocess.run", side_effect=KeyboardInterrupt):
with pytest.raises(MonitorError, match="Check fenris status") as failure:
run_monitor("collect")
assert failure.value.exit_code == 130
def test_cli_and_tui_show_the_same_failure(tmp_path, capsys):
# The launcher may prepend an installed runtime while loading; isolate it.
with patch.object(sys, "path", sys.path.copy()):
cli = runpy.run_path(str(Path(__file__).parent.parent / "scripts" / "fenris"))
with patch("fenris.control.os.geteuid", return_value=1000), \
patch("fenris.control.subprocess.run", return_value=subprocess.CompletedProcess([], 126)):
with pytest.raises(SystemExit) as exit_info:
cli["cmd_monitor_resume"](argparse.Namespace())
assert exit_info.value.code == 126
cli_message = capsys.readouterr().err.strip()
app = FenrisTuiApp(store_path=tmp_path / "missing.db")
with patch.object(app, "suspend"), patch.object(app, "_refresh") as refresh, \
patch.object(app, "notify") as notify:
app.action_resume()
notify.assert_called_once_with(cli_message, severity="error")
refresh.assert_called_once()
+145
View File
@@ -0,0 +1,145 @@
"""Cross-day interval unattributed-byte preservation (issue #88).
Verifies that a cross-hour interval spanning midnight is stored ONCE as
shared boundary evidence, not duplicated into both days.
Seam: derive._add_unattributed_bytes() → day_aggregates.unattributed_bytes_*
"""
import sqlite3
from datetime import datetime, timedelta, timezone
from pathlib import Path
import pytest
import sys
sys.path.insert(0, str(Path(__file__).parent.parent / "src"))
from fenris.collector import run_collection
from fenris.store import init_store
from fenris.monitoring_periods import ensure_period_open
def _make_smartctl(duw: int, dur: int):
return {
"json_format_version": [1, 0],
"smartctl": {"version": [7, 3], "svn_revision": "5155",
"build_info": "(local build)"},
"nvme_smart_health_information_log": {
"critical_warning": 0, "temperature": 35,
"available_spare": 100, "available_spare_threshold": 10,
"percentage_used": 5, "data_units_written": duw,
"data_units_read": dur, "power_on_hours": 8765,
"power_cycles": 1234, "unsafe_shutdowns": 5,
"media_errors": 0, "num_err_log_entries": 0,
},
"user_capacity": {"bytes": 1024000000000, "units": "bytes"},
"model_name": "Samsung SSD 970 EVO Plus 1TB",
"serial_number": "S4EWNX0N123456",
"firmware_version": "2B2QEXM7",
}
@pytest.fixture
def sysfs_tree(tmp_path: Path) -> Path:
ctrl_dir = tmp_path / "sys" / "class" / "nvme" / "nvme0"
ctrl_dir.mkdir(parents=True)
(ctrl_dir / "subsysnqn").write_text(
"nqn.2014-08.org.nvmexpress:uuid:12345678-1234-1234-1234-123456789abc\n"
)
(ctrl_dir / "model").write_text("Samsung SSD 970 EVO Plus 1TB\n")
(ctrl_dir / "serial").write_text("S4EWNX0N123456\n")
(ctrl_dir / "firmware_rev").write_text("2B2QEXM7\n")
transport_dir = ctrl_dir / "transport"
transport_dir.mkdir()
(transport_dir / "address").write_text("0000:03:00.0")
(transport_dir / "trstring").write_text("pcie")
return tmp_path
class _Clock:
def __init__(self, initial):
self.now = initial
def utcnow(self):
return self.now
class TestCrossDayUnattributedNoDuplication:
"""Unattributed bytes from a midnight-spanning interval must be stored
once, not duplicated into both days (issue #88)."""
def test_midnight_spanning_interval_not_duplicated(
self, tmp_path, sysfs_tree,
):
"""Two samples spanning midnight: 23:55 UTC day1 → 00:05 UTC day2.
The unattributed bytes should appear once, not in both days.
"""
store = str(tmp_path / "obs.db")
cfg = {"device": "/dev/nvme0", "store_path": store}
sysfs_nvme = sysfs_tree / "sys" / "class" / "nvme" / "nvme0"
t1 = datetime(2026, 9, 1, 23, 55, 0, tzinfo=timezone.utc)
t2 = datetime(2026, 9, 2, 0, 5, 0, tzinfo=timezone.utc)
# +100 DUW, +60 DUR across midnight
r1 = run_collection(_make_smartctl(10000000, 8000000), sysfs_nvme, cfg, _Clock(t1))
assert r1["ok"]
r2 = run_collection(_make_smartctl(10000100, 8000060), sysfs_nvme, cfg, _Clock(t2))
assert r2["ok"]
conn = sqlite3.connect(store)
# Get unattributed bytes for each day
unattr_w_day1 = conn.execute(
"SELECT unattributed_bytes_written FROM day_aggregates WHERE day = ?",
("2026-09-01",)
).fetchone()
unattr_w_day2 = conn.execute(
"SELECT unattributed_bytes_written FROM day_aggregates WHERE day = ?",
("2026-09-02",)
).fetchone()
unattr_r_day1 = conn.execute(
"SELECT unattributed_bytes_read FROM day_aggregates WHERE day = ?",
("2026-09-01",)
).fetchone()
unattr_r_day2 = conn.execute(
"SELECT unattributed_bytes_read FROM day_aggregates WHERE day = ?",
("2026-09-02",)
).fetchone()
expected_bw = 100 * 512000 # 51200000
expected_br = 60 * 512000 # 30720000
# BUG: Current code adds the SAME bytes to BOTH days.
# After fix: only ONE day should have the unattributed bytes.
# The spec says: preserve once as shared boundary evidence.
# At least one day must have the unattributed bytes
total_unattr_w = (unattr_w_day1[0] if unattr_w_day1 else 0) + (unattr_w_day2[0] if unattr_w_day2 else 0)
total_unattr_r = (unattr_r_day1[0] if unattr_r_day1 else 0) + (unattr_r_day2[0] if unattr_r_day2 else 0)
# The total unattributed bytes across both days must equal
# the actual delta (not double)
assert total_unattr_w == expected_bw, (
f"Unattributed writes across both days should be {expected_bw}, "
f"got {total_unattr_w} (duplication detected)"
)
assert total_unattr_r == expected_br, (
f"Unattributed reads across both days should be {expected_br}, "
f"got {total_unattr_r} (duplication detected)"
)
# Neither day should have MORE than the actual delta
for day_label, uw, ur in [
("day1", unattr_w_day1, unattr_r_day1),
("day2", unattr_w_day2, unattr_r_day2),
]:
if uw is not None:
assert uw[0] <= expected_bw, (
f"{day_label} unattributed writes {uw[0]} exceeds delta {expected_bw}"
)
if ur is not None:
assert ur[0] <= expected_br, (
f"{day_label} unattributed reads {ur[0]} exceeds delta {expected_br}"
)
conn.close()
+177
View File
@@ -0,0 +1,177 @@
"""User-facing navigation, zoom, and plotted volume in the redesigned dashboard."""
from datetime import datetime, timedelta, timezone
from xml.etree import ElementTree
import pytest
from test_tui import (
_insert_baseline,
_insert_day,
_insert_local_day,
_insert_segment,
_open_period,
)
from fenris.store import init_store
from fenris.tui import FenrisTuiApp
NOW = datetime(2026, 9, 19, 12, tzinfo=timezone.utc)
@pytest.fixture
def dashboard(tmp_path, monkeypatch):
monkeypatch.setenv("XDG_CONFIG_HOME", str(tmp_path / "prefs"))
class Clock(datetime):
@classmethod
def now(cls, tz=None):
return NOW
monkeypatch.setattr("fenris.tui.datetime", Clock)
monkeypatch.setattr("fenris.status.query_service_state", lambda: {
"boot_enabled": True, "timer_active": True, "last_collect_ok": True,
"last_collect_age_s": 0, "last_collect_reason": None,
})
conn = init_store(tmp_path / "test.db")
_insert_segment(conn)
_insert_baseline(conn)
_open_period(conn)
for offset in range(18):
date = (NOW - timedelta(days=offset)).date().isoformat()
_insert_day(conn, date, bw=(offset + 1) * 1_000_000_000)
_insert_local_day(conn, date, bw=(offset + 1) * 1_000_000_000, br=2_000_000_000)
for index in range(61):
conn.execute(
"INSERT INTO samples (ts, device, bytes_written, bytes_read, segment_id) VALUES (?, ?, ?, ?, ?)",
((NOW - timedelta(minutes=(60 - index) * 3)).isoformat(),
"/dev/test", index * 1_000_000, index * 2_000_000, 1),
)
conn.execute(
"INSERT INTO hour_observations (hour, bytes_written_delta, bytes_read_delta, coverage, sample_count) "
"VALUES (?, ?, ?, ?, ?)", ("2026-09-18T12:00:00+00:00", 4_000_000, 8_000_000, 1, 20),
)
conn.commit()
conn.close()
return FenrisTuiApp(store_path=tmp_path / "test.db", refresh_interval_s=999)
def visible(app):
return " ".join("".join(ElementTree.fromstring(app.export_screenshot()).itertext()).split())
@pytest.mark.asyncio
@pytest.mark.parametrize("size", [(140, 44), (80, 24)])
async def test_live_plot_inspection_zoom_refresh_and_resize(dashboard, size):
app = dashboard
async with app.run_test(size=size) as pilot:
app.on_refresh_tick()
await pilot.pause()
assert app.theme == "fenris-chalktone"
graph = app.query_one("#live-activity")
assert app.focused is graph
await pilot.press("left", "w")
selected = str(app.query_one("#live-readout").render())
assert "11:54" in selected and "W 0.001 GB" in selected and "R 0.002 GB" in selected
assert "Reads" in str(app.query_one("#live-legend").render())
await pilot.press("z")
app.on_refresh_tick()
await pilot.pause()
assert str(app.query_one("#live-readout").render()) == selected
assert app.query_one("#activity-panel").region.width == size[0]
for fact in ("Freshness:", "Last collect:", "Boot:", "Timer:", "q Quit TUI"):
assert fact in visible(app)
await pilot.resize_terminal(100, 30)
await pilot.press("escape")
assert str(app.query_one("#live-readout").render()) == selected
assert app._zoomed_panel is None
assert app.query_one("#action-rail").region.bottom <= 30
@pytest.mark.asyncio
async def test_tabs_date_entry_and_hourly_inspection_keep_context(dashboard):
app = dashboard
async with app.run_test(size=(100, 36)) as pilot:
await pilot.click("#view-history")
await pilot.press("left", "enter")
assert app.query_one("#activity-tabs").active == "view-day"
assert "2026-09-18" in str(app.query_one("#local-day").render())
# Hour 23 is selected initially; move to the known hour 12.
await pilot.press(*(["left"] * 11))
readout = str(app.query_one("#bar-readout").render())
assert "12:00 UTC" in readout and "W 0.004 GB" in readout
await pilot.press("z", "w")
app.on_refresh_tick()
await pilot.pause()
assert str(app.query_one("#bar-readout").render()) == readout
assert "Reads" in str(app.query_one("#bar-legend").render())
await pilot.press("g")
await pilot.press(*list("2026-09-17"))
await pilot.press("escape")
assert str(app.query_one("#bar-readout").render()) == readout
assert app._zoomed_panel == "activity-panel"
await pilot.press("escape", "t")
assert app.query_one("#activity-tabs").active == "view-live"
@pytest.mark.asyncio
async def test_mouse_inspection_matches_time_axis(dashboard):
app = dashboard
async with app.run_test(size=(100, 36)) as pilot:
await pilot.pause()
graph = app.query_one("#live-activity")
x = graph._point_columns[0]
await pilot.click("#live-render", offset=(x, 1))
text = str(app.query_one("#live-readout").render())
assert "09:00 → 09:03 UTC" in text
@pytest.mark.asyncio
async def test_theme_control_persists_choice_and_quit_never_pauses(dashboard, monkeypatch):
calls = []
monkeypatch.setattr(dashboard, "_run_helper", lambda *args: calls.append(args))
async with dashboard.run_test(size=(80, 24)) as pilot:
await pilot.press("s")
assert dashboard.theme == "fenris-amber"
await pilot.press("q")
assert calls == []
from fenris.preferences import load_preferences
assert load_preferences()["theme"] == "amber"
@pytest.mark.asyncio
@pytest.mark.parametrize("size", [(100, 36), (70, 20)])
async def test_unallocated_volume_survives_day_with_missing_coverage(dashboard, size):
async with dashboard.run_test(size=size) as pilot:
await pilot.click("#view-history")
graph = dashboard.query_one("#usage-history")
graph.set_data([{
"day": "2026-09-18", "total_bytes": 2_000_000_000,
"unallocated_bytes": 2_000_000_000, "is_gap": True,
"is_partial": True,
}])
readout = str(dashboard.query_one("#bar-readout").render())
assert "W 2.000 GB unallocated" in readout
assert "R unavailable" in readout and "gap" in readout
if size[0] >= 80:
plot = str(dashboard.query_one("#bar-render").render())
assert any(0x2801 <= ord(c) <= 0x28ff for c in plot)
graph.measure = "read"
graph._refresh()
plot = str(dashboard.query_one("#bar-render").render())
assert not any(0x2801 <= ord(c) <= 0x28ff for c in plot)
@pytest.mark.asyncio
@pytest.mark.parametrize("state", ["gap", "future"])
async def test_small_terminal_hour_readout_does_not_invent_zero(dashboard, state):
async with dashboard.run_test(size=(70, 20)) as pilot:
await pilot.click("#view-day")
graph = dashboard.query_one("#usage-history")
graph.set_hour_data([{
"hour": "2026-09-18T12:00:00+00:00", "local_label": "12",
"is_gap": state == "gap", "is_future": state == "future",
}])
graph._hourly_selected = 0
graph._refresh_hourly()
readout = str(dashboard.query_one("#bar-readout").render())
assert "W unavailable · R unavailable" in readout
assert state in readout and "12:00 UTC" in readout

Some files were not shown because too many files have changed in this diff Show More