Decide CUDA auto-detection and VRAM-aware model selection #2

Closed
opened 2026-08-12 05:36:39 +00:00 by xavierk · 1 comment
Owner

Blocked by: #5

Part of #1

Question

How should Vocalinux auto-detect a CUDA-compatible NVIDIA card (RTX 2050 4GB reference) and select a VRAM-fitting whisper.cpp model automatically, with user notify+confirm, using the CUDA backend and falling back to CPU when the card/runtime is absent?

Context

  • Repo already has backend detection (Vulkan → CUDA → CPU) in whispercpp_model_info.py and recognition_manager.py, plus installer NVIDIA detection in install.sh.
  • Reference card: RTX 2050 4GB laptop GPU. Whisper.cpp VRAM requirements per model (tiny→large-v3-turbo) need to be mapped to 4GB.
  • Decision: auto-swap model when current doesn't fit VRAM, with notify+confirm; distinguish "no card" vs "card but no CUDA runtime".

Resolution target

A decision on the detection flow (what signals, what thresholds, what fallback chain) and the VRAM→model mapping table for 4GB-class cards, ready for a builder to implement. Research ticket — resolved by a /research subagent.

Blocked by: #5 Part of #1 ## Question How should Vocalinux auto-detect a CUDA-compatible NVIDIA card (RTX 2050 4GB reference) and select a VRAM-fitting whisper.cpp model automatically, with user notify+confirm, using the CUDA backend and falling back to CPU when the card/runtime is absent? ## Context - Repo already has backend detection (Vulkan → CUDA → CPU) in `whispercpp_model_info.py` and `recognition_manager.py`, plus installer NVIDIA detection in `install.sh`. - Reference card: RTX 2050 4GB laptop GPU. Whisper.cpp VRAM requirements per model (tiny→large-v3-turbo) need to be mapped to 4GB. - Decision: auto-swap model when current doesn't fit VRAM, with notify+confirm; distinguish "no card" vs "card but no CUDA runtime". ## Resolution target A decision on the detection flow (what signals, what thresholds, what fallback chain) and the VRAM→model mapping table for 4GB-class cards, ready for a builder to implement. Research ticket — resolved by a /research subagent.
xavierk added the wayfinder:research label 2026-08-12 05:36:39 +00:00
Author
Owner

Resolved (research). Report: docs/research/cuda-detection-vram-model-selection.md on branch research/cuda-detection.

Answer: Use nvidia-smi --query-gpu=name,memory.total,compute_cap --format=csv,noheader,nounits as the primary probe (works on driver-without-runtime; torch.cuda is a false-negative there). pywhispercpp CUDA build detection via libggml-cuda*.so* is correct; authoritative check is Model.system_info(). RTX 2050 = compute capability 8.6, fully supported. Auto-selection: ≥8GB → turbo-q8_0; ≥4GB → large-v3-turbo-q5_0; ≥2GB → small; else base. CPU fallback default base/tiny. Fix: current CPU-force uses non-functional GGML_CUDA=0 env — must use context_params={"use_gpu": False}. Add 0.75× headroom guard + OOM retry ladder + notify+confirm.

**Resolved (research).** Report: `docs/research/cuda-detection-vram-model-selection.md` on branch `research/cuda-detection`. **Answer:** Use `nvidia-smi --query-gpu=name,memory.total,compute_cap --format=csv,noheader,nounits` as the primary probe (works on driver-without-runtime; `torch.cuda` is a false-negative there). pywhispercpp CUDA build detection via `libggml-cuda*.so*` is correct; authoritative check is `Model.system_info()`. RTX 2050 = compute capability 8.6, fully supported. Auto-selection: ≥8GB → turbo-q8_0; **≥4GB → large-v3-turbo-q5_0**; ≥2GB → small; else base. CPU fallback default base/tiny. Fix: current CPU-force uses non-functional `GGML_CUDA=0` env — must use `context_params={"use_gpu": False}`. Add 0.75× headroom guard + OOM retry ladder + notify+confirm.
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: xavierk/vocalinux-cuda#2