USE CASE · IMAGE
Local image generation.
Qwen-Image-2.1 (September 20 2026) is the newest major open-weight image release: a unified 7B generator/editor with native 2K and transparent RGBA, but a non-commercial research licence. HiDream-O1-Image remains the commercially clean MIT quality pick; choose by licence as carefully as by image quality.
Verdict — Strong open-weight options at every tier above 8GB
HiDream-O1-Image (8B MIT) at full precision sits cleanly at the top here too. Z-Image-Turbo (Apache 2.0 community daily driver, sub-second on H800). FLUX.2 dev for absolute quality if non-commercial works.
What's the answer at each tier
HiDream-I1 Full FP16 (17B MIT) + FLUX.2 dev FP16 (32B non-commercial) + Qwen-Image-2.1 (7B visual model, research-only) for unified gen/edit and native transparency. Qwen-Image-2512 remains the Apache-2.0 commercial text-rendering option.
- FLUX.2 dev FP16 (~64 GB, no quant compromise) — Full-precision 32B at native FP16. Quality ceiling locally — no Q4/FP8 artifacts. Frontier-only because FP16 weights need 64+ GB resident.
- HiDream-I1 Full (17B, FP16, MIT) — April 2025 release; community testing shows it outperforms SDXL, DALL-E 3, and FLUX.1 on key benchmarks. MIT-licensed — clean for commercial redistribution. Frontier hardware unlocks the full FP16 weights.
- Wan 2.2 (A14B variants) — Community-standard local video generation in 2026. A14B variants need 16-24 GB minimum; frontier 96+ GB has headroom for 1080p / longer clips / batched runs. Pairs with ComfyUI workflow.
HiDream-O1-Image (8B MIT) at full precision sits cleanly at the top here too. Z-Image-Turbo (Apache 2.0 community daily driver, sub-second on H800). FLUX.2 dev for absolute quality if non-commercial works.
- Z-Image-Turbo (Apache 2.0) — Community daily driver for realism; 6B, 8-step inference, Apache 2.0 — commercial OK.
- FLUX.2 dev (non-commercial) — Quality ceiling at 32B; FLUX Non-Commercial license — commercial deployment needs paid BFL license.
- Qwen-Image-2.1 (7B, Qwen Research License) — Fresh unified gen+edit pick: strong multilingual text, native transparency and up to ten references. Research/evaluation only unless Qwen grants a commercial licence; use Qwen-Image-2512 for Apache-2.0 commercial work.
HiDream-O1-Image FP8 remains the MIT commercial-quality default at 16-24 GB. Qwen-Image-2.1 is the fresh research-only gen+edit pick, with no official VRAM table and component offload expected below 24 GB. Qwen-Image-2512 remains the Apache-2.0 commercial text-rendering path.
- HiDream-O1-Image (8B, MIT) — May 8, 2026 release. Pixel-space (no VAE, no disjoint text encoder) — debuted top-10 on Artificial Analysis T2I Arena. MIT-licensed 8B; one model handles T2I + edit + subject-driven personalization at up to 2,048².
- Z-Image-Turbo (Apache 2.0) — The dependable pick at this tier: 6B, 8-step inference, Apache 2.0, and community-proven. Promoted here in August 2026 after Microsoft withdrew Lens — a reminder that a permissive licence is worth little if the weights stop being downloadable.
- Qwen-Image-2.1 (7B, research-only) — Unified generation and editing with CPU-offload support, native 2K and transparent RGBA. No official VRAM table: expect component offload below 24 GB. Non-commercial Qwen Research License.
HiDream-O1-Image-Dev-2604 (distilled, 28-step) is the newer freshest pick. FLUX.2 klein 4B (Apache 2.0 commercial) remains the right Apache-clean default. HiDream-I1 Fast for higher quality at the same size class.
- FLUX.2 klein 4B (Apache 2.0) — BFL's first fully Apache-2.0 model; 4B distilled for fast inference on mid-tier GPUs; commercial OK.
- SANA-1.6B (non-commercial) — NVIDIA 4K-capable 1.6B; extremely fast on mid-tier GPUs; weights are NVIDIA NSCL v2 (non-commercial).
- Z-Image-Turbo (FP8, fits 8GB) — FP8 quant runs comfortably on an 8GB card with Apache 2.0 commercial license.
SANA-0.6B (NVIDIA NSCL v2 non-commercial) is fast on 6-8GB VRAM; Z-Image-Turbo at int4 fits the same window. SD 3.5 Medium is the SAI-licensed fallback.
- SANA-0.6B (non-commercial) — 0.6B params; <1s per 1024² on a 16GB laptop GPU; weights are NVIDIA NSCL v2 (non-commercial).
- Z-Image-Turbo (int4, 6GB) — Community int4 quant fits 6GB VRAM; Apache 2.0 — commercial-OK at this tier is rare.
- SD 3.5 Medium — 2.6B; ~10GB VRAM; reliable; SAI Community License (free under $1M revenue).
How to actually run it
ComfyUI is the standard runner — every major checkpoint has community node workflows. Forge for SD-family lineage. Diffusers Python directly for batch generation pipelines. Mac users: native MLX inference is well-supported for FLUX + SD families; HiDream-O1 Mac performance not yet community-benchmarked.
Watchouts
- License landscape is fractured. FLUX.2 dev is non-commercial; FLUX.2 klein 4B is Apache 2.0 but 9B is non-commercial; HiDream-I1 and HiDream-O1 are MIT; Qwen-Image-2.1 uses the non-commercial Qwen Research License while Qwen-Image-2512 is Apache 2.0; SANA is NVIDIA NSCL v2 non-commercial; SD 3.5 uses the SAI Community License. Read the exact checkpoint licence before commercial deployment.
- HiDream-O1 (May 2026) ships with a new pixel-space architecture (no VAE, no separate text encoders) — existing ComfyUI custom-node graphs for FLUX/SDXL won't port cleanly. Plan for 1-2 weeks of node ecosystem catch-up.
- Qwen-Image-2.1 shipped September 20, 2026 as the runnable unified successor to the never-released 2.0. Its LICENSE permits non-commercial research/evaluation only; commercial users should stay on Qwen-Image-2512 or obtain a separate Qwen licence.
- Top 5 of AA T2I Arena is all CLOSED models (GPT Image 2 Elo 1338, Nano Banana 2/Pro, MAI-Image-2). Top open-weight peaks at Elo 1184. For absolute quality first-try, cloud still leads.
When cloud still wins
You need GPT Image 2 / Nano Banana / Midjourney-class consistency on the first try, or you don't want to manage a ComfyUI workflow. Local image gen is genuinely good now but the iteration loop is faster on cloud for one-off creative work. For volume + privacy + specific style fine-tuning, local wins.
Hardware that fits this use case
Related guides
Next step
Try the planner with Image generation preselected→The planner pulls all six dimensions together — your hardware, your VRAM/RAM, your GPU family, your context, and your priorities — and returns specific picks with fit badges.
Notes flagged for next refresh
Wan 2.2 (A14B + variants) is text-to-VIDEO, not image — flagged for a separate /use-cases/video/ page when built. SANA-WM (NVIDIA, May 16 2026) is also video — 2.6B world-model, single-RTX-5090 60-sec 720p. HiDream-O1-Image-Dev-2604 is the freshest distilled image checkpoint (~3 days old at time of writing).