Viggle Turbo v0.3 — 6-step Qwen-Image-2.1
Text-to-image and image editing in 6 steps: about 5× faster than the 40-step Qwen-Image-2.1 and very competitive with it in quality — see the Comparison tab. Leave the references empty for text-to-image, or add up to 6 to edit, compose or transfer style.
Model card, weights and ComfyUI workflows: Viggle/Qwen-Image-2.1-viggle-turbo
A distilled Qwen-Image-2.1 that runs with no classifier-free guidance. On most prompts it is hard to tell apart from the base model; small, dense text and complicated edits (multi-reference composition, face swaps, identity-preserving edits) can still fall short of it. The Comparison tab has 32 examples of the official Qwen Space side by side, turbo vs base, same seed.
v0.3 (2026-09-29): at 6 steps, less grain than v0.2.1 and a little softer on fine texture. We think 6 steps is close to its capacity: every further gain we found cost something elsewhere. 9 steps runs 7 turbo steps and lets the base model finish the last two: finer detail and small text right more often, at about 1.4–1.5× the time of 6 steps. v0.2.1 can still be picked under Advanced settings. Weights: v0.3 LoRA r256 (and v0.2.1) from Viggle/Qwen-Image-2.1-viggle-turbo.
The menu switches to the editing sizes as soon as a reference is attached. Custom sizes above 2048² area (1536² when editing) are scaled down, keeping the ratio.
| Prompt | References (optional, up to 6) | Result |
|---|
viggle-turbo (6 steps) vs Qwen-Image-2.1 (40 steps) on 32 of the 37 examples of the official Qwen/Qwen-Image-2.1 Space. Both models get the same prompt, input images and seed, with prompt enhancement off, one sample each and no seed picking. At 5.0× less time the turbo is very competitive with the base model; the clearest gaps are small, dense text (9 steps often narrows it) and complicated edits, where the base is still ahead. Pick an example on the left, then drag the slider. The 5 dense-text examples open on the 9-step turbo, the others on 6 steps (both are in every menu), and the menus also offer Qwen's own API reference where the Space ships one.
Median time, editing: 3.0 s vs 14.2 s · text-to-image (~4 MP): 4.8 s vs 26.1 s · all 32 examples: 5.0× faster (one NVIDIA B200, one pipeline call each)
Transparent anime bride · 透明婚纱角色立绘
1664 × 2496 · left: viggle-turbo · 6 steps · 4.8 s · right: Qwen-Image-2.1 base · 40 steps · 26.0 s · drag the handle to compare
- Turbo: Qwen-Image-2.1 + the viggle-turbo v0.3 LoRA (rank 256), 6 steps on sigmas
[1, 0.9375, 0.875, 0.75, 0.5, 0.25], no classifier-free guidance, exactly as the Generate tab runs it. - Turbo, 9 steps: the Generate tab's 9-step setting: 7 turbo steps, then the LoRA is switched off and the base model takes the last two. Finer detail than 6 steps and small text right more often; about 1.4–1.5× the time of 6 steps.
- Base: Qwen-Image-2.1, 40 steps, default scheduler,
true_cfg_scale=1.0(the pipeline default). - Both: diffusers
QwenImage21Pipelinein bf16, seed 42. Input images are encoded at 1024² pixels; the output keeps the aspect ratio of Qwen's reference at 2048² pixels for text-to-image and 1536² for editing, in multiples of 32. - Time: one pipeline call on one NVIDIA B200 (text encoding, denoising and VAE decoding), after warm-up.
- Left out (5 of 37): three Chinese infographics and a subtitled storyboard whose short prompts leave the on-image text to Qwen's prompt enhancement (without it both models fill them with made-up characters), and a storyboard that asks for a named third-party character.
- Qwen API reference: the output bundled with the example in Qwen's Space, made with Qwen's API, which may differ from the open weights. It is omitted for the 6 text-to-image examples Qwen generated with prompt enhancement on.
Prompts, input images and reference outputs come from the Qwen/Qwen-Image-2.1 Space (Qwen Research License).
Model: Viggle/Qwen-Image-2.1-viggle-turbo · Built with Qwen — distilled from Qwen/Qwen-Image-2.1, which is released under the Qwen RESEARCH LICENSE AGREEMENT (non-commercial: research or evaluation purposes only). This demo inherits that restriction.