Map: CUDA-native seamless dictation with Handy parity on Linux #1

Open
opened 2026-08-12 05:36:16 +00:00 by xavierk · 0 comments
Owner

Destination

The route to a decision-ready spec for making Vocalinux a seamless, CUDA-native, Handy-parity dictation app on Linux. Reaching the end means every open decision below is resolved and recorded in tickets — ready for builders to implement, not implemented here. Concretely:

  • CUDA native: auto-detect a compatible NVIDIA CUDA card (RTX 2050 4GB reference), auto-select a VRAM-fitting whisper.cpp model with user notify+confirm, use CUDA backend with CPU fallback.
  • Terminal/paste: transcription auto-pasted into any focused field (terminal, browser, editor) on release, via the existing text injector — no terminal-specific handling.
  • Ubuntu 26.04: ship a .deb with correct dependencies so no init/compat crash surfaces; full packaging refresh (.deb, AppImage, installer) with clean dep manifests.
  • Handy seamlessness: hold-to-talk = record, release = paste (core loop); model picker + recording overlay; Wayland parity (ydotool/systemd, per-DE instructions).
  • Handy model parity: same model array as Handy — Whisper bins (small, medium-q4_1, turbo, large-v3-q5_0) plus Parakeet V2/V3/Unified.

Notes

  • Domain: Linux desktop dictation (GTK3/PyGObject, whisper.cpp, Vosk, Silero VAD, text injection via X11/Wayland).
  • Skills to consult: /grilling, /domain-modeling, /research (AFK), /prototype (HITL), /wayfinder (map/tickets).
  • Standing preferences: decision-map (plan, don't do); build on existing text injector; user consent before model auto-swap; any-focused-field paste (not terminal-only).
  • Tracker: Gitea via tea CLI. Map = this issue. Tickets are children with Part of #<map>.
  • Reference hardware: RTX 2050 4GB laptop GPU — the CUDA test card.

Decisions so far

Not yet specified

Fog — in-scope, too coarse to ticket yet:

  • How the CUDA auto-detect should distinguish "card present but no CUDA runtime" vs "no card" (nvidia-smi absent vs driver present).
  • Whether Parakeet needs a bundled runtime or can be a Python dependency; how to handle its ~478MB download.
  • How the recording overlay should behave on X11 vs Wayland (focus-steal risk — Handy disables overlay by default on Linux).
  • Whether auto-model-selection should consider both VRAM and CPU-only fallback (no GPU → which default?).
  • Exact Wayland tooling to recommend (ydotool systemd setup) and whether to bundle it.
  • How much of Handy's per-DE shortcut documentation to adopt.

Out of scope

  • Porting Handy's Rust/Tauri engine wholesale — Vocalinux stays GTK/Python, uses existing injector. (Decided in grilling.)
  • Terminal-specific paste handling — any focused field is the target.
  • Re-diagnosing a 26.04 crash that hasn't been observed yet — instead, pre-empt via .deb dependency correctness.
## Destination The route to a decision-ready spec for making Vocalinux a seamless, CUDA-native, Handy-parity dictation app on Linux. Reaching the end means every open decision below is resolved and recorded in tickets — ready for builders to implement, not implemented here. Concretely: - **CUDA native**: auto-detect a compatible NVIDIA CUDA card (RTX 2050 4GB reference), auto-select a VRAM-fitting whisper.cpp model with user notify+confirm, use CUDA backend with CPU fallback. - **Terminal/paste**: transcription auto-pasted into any focused field (terminal, browser, editor) on release, via the existing text injector — no terminal-specific handling. - **Ubuntu 26.04**: ship a .deb with correct dependencies so no init/compat crash surfaces; full packaging refresh (.deb, AppImage, installer) with clean dep manifests. - **Handy seamlessness**: hold-to-talk = record, release = paste (core loop); model picker + recording overlay; Wayland parity (ydotool/systemd, per-DE instructions). - **Handy model parity**: same model array as Handy — Whisper bins (small, medium-q4_1, turbo, large-v3-q5_0) plus Parakeet V2/V3/Unified. ## Notes - **Domain**: Linux desktop dictation (GTK3/PyGObject, whisper.cpp, Vosk, Silero VAD, text injection via X11/Wayland). - **Skills to consult**: `/grilling`, `/domain-modeling`, `/research` (AFK), `/prototype` (HITL), `/wayfinder` (map/tickets). - **Standing preferences**: decision-map (plan, don't do); build on existing text injector; user consent before model auto-swap; any-focused-field paste (not terminal-only). - Tracker: Gitea via `tea` CLI. Map = this issue. Tickets are children with `Part of #<map>`. - **Reference hardware**: RTX 2050 4GB laptop GPU — the CUDA test card. ## Decisions so far - [Decide Handy model-array parity (Whisper bins + Parakeet V2/V3/Unified)](https://git.bongbetic.com/xavierk/vocalinux-cuda/issues/5) — 3/4 Whisper bins load in pywhispercpp today (medium-q4_1 is legacy → use medium-q5_0); Parakeet needs new engine: onnx-asr Phase 1, transcribe-cpp Phase 2; 4GB fit = large-v3-turbo-q5_0 recommended, exclude fp16 turbo/full large. - [Decide CUDA auto-detection and VRAM-aware model selection](https://git.bongbetic.com/xavierk/vocalinux-cuda/issues/2) — nvidia-smi primary probe; RTX 2050 = CC 8.6 supported; auto-select ≥4GB → large-v3-turbo-q5_0; CPU force must use use_gpu=False (GGML_CUDA=0 env is non-functional); 0.75× headroom + OOM retry + notify+confirm. ## Not yet specified Fog — in-scope, too coarse to ticket yet: - How the CUDA auto-detect should distinguish "card present but no CUDA runtime" vs "no card" (nvidia-smi absent vs driver present). - Whether Parakeet needs a bundled runtime or can be a Python dependency; how to handle its ~478MB download. - How the recording overlay should behave on X11 vs Wayland (focus-steal risk — Handy disables overlay by default on Linux). - Whether auto-model-selection should consider both VRAM and CPU-only fallback (no GPU → which default?). - Exact Wayland tooling to recommend (ydotool systemd setup) and whether to bundle it. - How much of Handy's per-DE shortcut documentation to adopt. ## Out of scope - Porting Handy's Rust/Tauri engine wholesale — Vocalinux stays GTK/Python, uses existing injector. (Decided in grilling.) - Terminal-specific paste handling — any focused field is the target. - Re-diagnosing a 26.04 crash that hasn't been observed yet — instead, pre-empt via .deb dependency correctness.
xavierk added the wayfinder:map label 2026-08-12 05:36:16 +00:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: xavierk/vocalinux-cuda#1