How should Vocalinux auto-detect a CUDA-compatible NVIDIA card (RTX 2050 4GB reference) and select a VRAM-fitting whisper.cpp model automatically, with user notify+confirm, using the CUDA backend and falling back to CPU when the card/runtime is absent?
Context
Repo already has backend detection (Vulkan → CUDA → CPU) in whispercpp_model_info.py and recognition_manager.py, plus installer NVIDIA detection in install.sh.
Reference card: RTX 2050 4GB laptop GPU. Whisper.cpp VRAM requirements per model (tiny→large-v3-turbo) need to be mapped to 4GB.
Decision: auto-swap model when current doesn't fit VRAM, with notify+confirm; distinguish "no card" vs "card but no CUDA runtime".
Resolution target
A decision on the detection flow (what signals, what thresholds, what fallback chain) and the VRAM→model mapping table for 4GB-class cards, ready for a builder to implement. Research ticket — resolved by a /research subagent.
Blocked by: #5
Part of #1
## Question
How should Vocalinux auto-detect a CUDA-compatible NVIDIA card (RTX 2050 4GB reference) and select a VRAM-fitting whisper.cpp model automatically, with user notify+confirm, using the CUDA backend and falling back to CPU when the card/runtime is absent?
## Context
- Repo already has backend detection (Vulkan → CUDA → CPU) in `whispercpp_model_info.py` and `recognition_manager.py`, plus installer NVIDIA detection in `install.sh`.
- Reference card: RTX 2050 4GB laptop GPU. Whisper.cpp VRAM requirements per model (tiny→large-v3-turbo) need to be mapped to 4GB.
- Decision: auto-swap model when current doesn't fit VRAM, with notify+confirm; distinguish "no card" vs "card but no CUDA runtime".
## Resolution target
A decision on the detection flow (what signals, what thresholds, what fallback chain) and the VRAM→model mapping table for 4GB-class cards, ready for a builder to implement. Research ticket — resolved by a /research subagent.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Blocked by: #5
Part of #1
Question
How should Vocalinux auto-detect a CUDA-compatible NVIDIA card (RTX 2050 4GB reference) and select a VRAM-fitting whisper.cpp model automatically, with user notify+confirm, using the CUDA backend and falling back to CPU when the card/runtime is absent?
Context
whispercpp_model_info.pyandrecognition_manager.py, plus installer NVIDIA detection ininstall.sh.Resolution target
A decision on the detection flow (what signals, what thresholds, what fallback chain) and the VRAM→model mapping table for 4GB-class cards, ready for a builder to implement. Research ticket — resolved by a /research subagent.
Resolved (research). Report:
docs/research/cuda-detection-vram-model-selection.mdon branchresearch/cuda-detection.Answer: Use
nvidia-smi --query-gpu=name,memory.total,compute_cap --format=csv,noheader,nounitsas the primary probe (works on driver-without-runtime;torch.cudais a false-negative there). pywhispercpp CUDA build detection vialibggml-cuda*.so*is correct; authoritative check isModel.system_info(). RTX 2050 = compute capability 8.6, fully supported. Auto-selection: ≥8GB → turbo-q8_0; ≥4GB → large-v3-turbo-q5_0; ≥2GB → small; else base. CPU fallback default base/tiny. Fix: current CPU-force uses non-functionalGGML_CUDA=0env — must usecontext_params={"use_gpu": False}. Add 0.75× headroom guard + OOM retry ladder + notify+confirm.