MODEL · ALIBABA · 30B TOTAL / 3B ACTIVE
Qwen3-Omni-30B-A3B-Instruct
The only locally-runnable open-weight model that does real-time streaming speech-out natively. 119 input languages, 10 speech-output languages (two voices: Chelsie, Ethan). Alibaba released hosted Qwen3.8-Omni-Flash and Qwen3.8-LiveTranslate on September 18–19, 2026, but neither has downloadable weights, so they do not replace this local model.
License: Custom Qwen license (HuggingFace-tagged "other" — verify terms before commercial use; not Apache 2.0) · Context: Multimodal tokens dominate the budget · Released: September 22, 2025
The decision in five lines
- The call
- Buy — for voice
- Best for
- voice
- Runs on
- 16 hardware picks fit (cheapest: Minisforum UM890 Pro · $463 barebone (no RAM / SSD); $991 with 32 GB / 1 TB — Minisforum US store, Sep 15 2026)
- Watch out
- Ollama-only workflows — Ollama has no native voice pipeline, so you can't use this model's speech-out through it at all.
- Evidence
- Estimated
- 30B total
- PARAMETERS
- MOE
- TYPE
- Multimodal
- CONTEXT
- ~16 GB (AWQ 4-bit, full stack)
- VRAM AT Q4
Where we recommend this
Every tier slot in the planner where this model is a top or alternate pick. Pulled live from planner.js — when the planner refreshes, this table stays current.
The call
The only locally-runnable open-weight model that does real-time streaming speech-out natively. 119 input languages, 10 speech-output languages (two voices: Chelsie, Ethan). Alibaba released hosted Qwen3.8-Omni-Flash and Qwen3.8-LiveTranslate on September 18–19, 2026, but neither has downloadable weights, so they do not replace this local model.
When not to use: Ollama-only workflows — Ollama has no native voice pipeline, so you can't use this model's speech-out through it at all.
Runner notes
Serve via vLLM or the QwenLM/Qwen3-Omni reference server behind Open-WebUI. Route text-only traffic separately if you don't need audio. AWQ-4bit (`cpatonn/Qwen3-Omni-30B-A3B-Instruct-AWQ-4bit`) fits ~16 GB VRAM. Qwen3.8-Omni-Flash (1M context, text/image/audio/video input) and Qwen3.8-LiveTranslate (60 languages) are Model Studio services only as of September 22, 2026.
Hardware that fits
Every hardware pick whose memory fits this model at the quant we recommend. Sorted cheapest-first — the top row is your best-value fit. Click through for the full buyer’s guide.
- Minisforum UM890 ProGood · 1.4× 32 GB DDR5 (shared) · $463 barebone (no RAM / SSD); $991 with 32 GB / 1 TB — Minisforum US store, Sep 15 2026
- NVIDIA RTX 3090 (used, single)Good · 1.4× 24 GB · $950–$1,200
- AMD Radeon RX 7900 XTXGood · 1.4× 24 GB · $1,300–$1,500 new (no refurb in stock Sep 15 2026)
- MacBook Air M5 24 GBRequires tweak · 1.2× 24 GB unified · $1,499–$1,899
- Mac Mini M4 Pro 24 GBRequires tweak · 1.2× 24 GB unified · $1,699 (M5 Pro successor, 24 GB / 512 GB, pre-order; M4 Pro discontinued Aug 25 2026)
- Dual RTX 3090 (used)Perfect · 2.7× 48 GB · $1,800–$2,500 all-in
- NVIDIA RTX 4090Good · 1.4× 24 GB · $2,200–$2,800
- M5 Pro MacBook Pro 48 GBPerfect · 1.8× 48 GB unified · $2,999–$3,599
- Framework Desktop (Ryzen AI Max+ 395)Perfect · 4.9× 128 GB unified · $3,449 (128 GB config)
- NVIDIA RTX A6000 (48 GB, used)Perfect · 2.7× 48 GB ECC · $3,500–$4,500
- Mac Studio M4 Max 64 GBPerfect · 2.4× 64 GB unified · $3,799 (M5 Max successor, 64 GB / 1 TB, pre-order; M4 Max discontinued Aug 25 2026)
- NVIDIA DGX SparkPerfect · 4.9× 128 GB unified · $4,699
- M5 Max MacBook Pro 64 GBPerfect · 2.4× 64 GB unified · ~$5,199 (est.; June 25 2026 increase)
- Mac Studio M3 Ultra 96 GBPerfect · 3.7× 96 GB unified · $5,499 (M5 Ultra successor, 96 GB / 1 TB, pre-order; M3 Ultra discontinued Aug 25 2026)
- NVIDIA RTX 5090Perfect · 1.8× 32 GB · $6,610–$7,000 (new, in stock)
- Dual RTX 5090Perfect · 3.6× 64 GB (2×32) · $13,900–$15,000 all-in (two cards alone are ~$13,220–$13,800 at the Sep 22 floor)
Next step
Find-by-model — see what hardware runs this→