Decide Handy model-array parity (Whisper bins + Parakeet V2/V3/Unified) #5

Closed
opened 2026-08-12 05:36:43 +00:00 by xavierk · 1 comment
Owner

Part of #1

Question

What is the feasibility and approach for supporting the same model array as Handy — Whisper bins (small, medium-q4_1, turbo, large-v3-q5_0) plus Parakeet V2/V3/Unified — inside Vocalinux's whisper.cpp/Python architecture?

Context

  • Handy models: ggml-small.bin, whisper-medium-q4_1.bin, ggml-large-v3-turbo.bin, ggml-large-v3-q5_0.bin; Parakeet V2 (int8), V3 (int8), Unified EN 0.6B (GGUF Q8_0).
  • Vocalinux currently supports whisper.cpp models (tiny→large-v3-turbo, quantized) but no Parakeet.
  • Parakeet is a different inference engine (transcribe-rs/CPU-optimized; GGUF); adding it means a new backend — need feasibility (Python bindings? bundled runtime? ~478MB download handling?).

Resolution target

A decision on model-array parity: which models to add, how Parakeet integrates (backend, deps, download flow), and the VRAM/CPU trade-offs — ready for a builder. Research ticket — resolved by a /research subagent.

Part of #1 ## Question What is the feasibility and approach for supporting the same model array as Handy — Whisper bins (small, medium-q4_1, turbo, large-v3-q5_0) plus Parakeet V2/V3/Unified — inside Vocalinux's whisper.cpp/Python architecture? ## Context - Handy models: ggml-small.bin, whisper-medium-q4_1.bin, ggml-large-v3-turbo.bin, ggml-large-v3-q5_0.bin; Parakeet V2 (int8), V3 (int8), Unified EN 0.6B (GGUF Q8_0). - Vocalinux currently supports whisper.cpp models (tiny→large-v3-turbo, quantized) but no Parakeet. - Parakeet is a different inference engine (transcribe-rs/CPU-optimized; GGUF); adding it means a new backend — need feasibility (Python bindings? bundled runtime? ~478MB download handling?). ## Resolution target A decision on model-array parity: which models to add, how Parakeet integrates (backend, deps, download flow), and the VRAM/CPU trade-offs — ready for a builder. Research ticket — resolved by a /research subagent.
xavierk added the wayfinder:research label 2026-08-12 05:36:43 +00:00
Author
Owner

Resolved (research). Report: docs/research/handy-model-array-feasibility.md on branch research/handy-model-parity.

Answer: Handy's model array is adoptable. 3 of 4 Whisper .bin files are byte-identical mirrors of existing ggml files (small, large-v3-turbo, large-v3-q5_0) and load in pywhispercpp today. whisper-medium-q4_1.bin is a legacy Q4_1 ftype the current loader may reject — substitute medium-q5_0. All 3 Parakeet artifacts need a new engine: recommend onnx-asr (pure pip, numpy+onnxruntime, supports nemo-parakeet-tdt-0.6b-v2/v3 int8) as Phase 1, transcribe-cpp PyPI bindings when native wheels ship as Phase 2. 4GB VRAM fit set: small, medium-q5_0, large-v3-q5_0, large-v3-turbo-q5_0 (recommended); exclude fp16 turbo and full large-v3. Prefer permanent HF mirrors over blob.handy.computer (unversioned, no checksums).

**Resolved (research).** Report: `docs/research/handy-model-array-feasibility.md` on branch `research/handy-model-parity`. **Answer:** Handy's model array is adoptable. 3 of 4 Whisper .bin files are byte-identical mirrors of existing ggml files (small, large-v3-turbo, large-v3-q5_0) and load in pywhispercpp today. `whisper-medium-q4_1.bin` is a legacy Q4_1 ftype the current loader may reject — substitute `medium-q5_0`. All 3 Parakeet artifacts need a new engine: recommend `onnx-asr` (pure pip, numpy+onnxruntime, supports nemo-parakeet-tdt-0.6b-v2/v3 int8) as Phase 1, `transcribe-cpp` PyPI bindings when native wheels ship as Phase 2. 4GB VRAM fit set: small, medium-q5_0, large-v3-q5_0, **large-v3-turbo-q5_0 (recommended)**; exclude fp16 turbo and full large-v3. Prefer permanent HF mirrors over blob.handy.computer (unversioned, no checksums).
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: xavierk/vocalinux-cuda#5