Memory management
A local model lives in memory while the service runs: VRAM on a GPU, RAM on CPU. Cloud backends hold none.
What each model needs
Section titled “What each model needs”Roughly, at default settings:
| Model | Backend | Memory |
|---|---|---|
| Parakeet v3 · Orukeet | Parakeet | 1 GB |
| Parakeet v3 | Parakeet.cpp | 0.9 GB |
Whisper base |
whisper.cpp · faster-whisper | ~150 MB |
Whisper small |
whisper.cpp · faster-whisper | ~490 MB |
Whisper large-v3-turbo |
whisper.cpp · faster-whisper | 1.6 GB |
Whisper large-v3 |
whisper.cpp · faster-whisper | ~3 GB |
| Qwen3-ASR 0.6B | Qwen3-ASR | ~1.0 GB + context |
| Qwen3-ASR 1.7B | Qwen3-ASR | ~2.4 GB + context |
| Cohere Transcribe | Cohere Transcribe | 4 GB VRAM · 8 GB RAM on CPU |
| — | REST API · Realtime WebSocket | none |
Whisper figures are model sizes; expect some overhead on top. Qwen3-ASR’s context adds ~0.9 GB at the default 8192 tokens.
Spending less
Section titled “Spending less”- Smaller model. Whisper
baseorsmall, Qwen3-ASR0.6b-q8_0. - Quantize. Parakeet runs
int8by default (onnx_asr_quantization); keep it. faster-whisper’sautocompute type isint8;float16andfloat32cost more. - Keep Cohere on bfloat16. Its GPU default.
float32doubles the footprint. - Trim Qwen3-ASR’s context.
qwen3_asr_ctx_sizesizes the KV cache, ~112 KiB per token on the 1.7B.nullmeans llama.cpp’s 32000, ~3.5 GB. - Move to CPU. Frees VRAM, spends RAM.
- Go remote. A cloud or self-hosted backend keeps memory off this machine.
- Share the model.
hyprwhspr transcribeborrows an idle service’s model rather than loading a second.
Unload and reload
Section titled “Unload and reload”Need the GPU back for a game or a local LLM? Unload the model. The service stays up, shortcuts stay bound:
hyprwhspr model unload # Free the model's VRAM or RAMhyprwhspr model reload # Load it againWhile unloaded, recording is refused with a notification, and the Waybar tray shows a sleep icon. Works on every local backend; Qwen3-ASR stops its sidecar. Cloud backends have nothing to unload.
Bind both in ~/.config/hypr/hyprland.conf:
# Free the GPU before starting a local LLMbindd = SUPER ALT, U, Unload speech model, exec, hyprwhspr model unload
# Reclaim dictation when donebindd = SUPER ALT, L, Reload speech model, exec, hyprwhspr model reload