Choosing a backend
Quick pick by hardware:
- NVIDIA GPU → Cohere Transcribe
- AMD / Intel GPU → Parakeet.cpp · whisper.cpp (Vulkan) for 99 languages
- CPU only → Parakeet or faster-whisper
- ARM64 → Parakeet
- Chinese, Japanese or Korean → Qwen3-ASR
- No local setup → REST API
| Model | Engine | Runs on | Arch | Speed | Accuracy | Memory | Languages |
|---|---|---|---|---|---|---|---|
| Parakeet v3 | ONNX | CPU · NVIDIA | x64 · ARM64 | ●●● | ●●○ | 1 GB | 25 European |
| Orukeet † | ONNX | CPU · NVIDIA | x64 · ARM64 | ●●● | ●●○ | 1 GB | 25 European |
| Parakeet v3 † | Parakeet.cpp | CPU · Vulkan | x64 · ARM64 | ●●● | ●●○ | 0.9 GB | 25 European |
| Whisper turbo | whisper.cpp · faster-whisper | CPU · NVIDIA · Vulkan | x64 · ARM64 | ●●○ | ●●○ | 1.6 GB | 99 |
| Cohere Transcribe | PyTorch | CPU · NVIDIA | x64 · ARM64 | ●●● | ●●● | 4 GB | 14 |
| Qwen3-ASR 1.7B † | llama.cpp | CPU · Vulkan | x64 | ●●○ | ●●○ · ●●● CJK | 2.4 GB | 30 |
Speed on each model’s best hardware. Accuracy from the Open ASR Leaderboard. Memory is RAM or VRAM, roughly. ARM64 runs on CPU. Whisper ships smaller models: faster-whisper, whisper.cpp. † Experimental.
Cloud or your own server: REST API and Realtime WebSocket. Speed and accuracy follow the provider.
hyprwhspr setup auto picks Whisper for your hardware: faster-whisper on NVIDIA or CPU, whisper.cpp on AMD/Intel.
Model commands
Section titled “Model commands”hyprwhspr model commands route automatically to the configured local backend.
For cloud backends (rest-api and realtime-ws), model operations are not
applicable and exit nonzero.
hyprwhspr model status # Check if model is downloaded/cachedhyprwhspr model list # Show model info for active backendhyprwhspr model download [model] # Download or re-download modelhyprwhspr model unload # Free its memory; the service keeps runninghyprwhspr model reload # Reload the model after an unloadModels are downloaded automatically during hyprwhspr setup; use model download to re-download if needed. To free memory, see Memory management.