Viggle Turbo v0.2.1 — 6-step Qwen-Image-2.1
Text-to-image and image editing in 6 steps: about 5× faster than the 40-step Qwen-Image-2.1 and very competitive with it in quality — see the Comparison tab. Leave the references empty for text-to-image, or add up to 6 to edit, compose or transfer style.
Model card, weights and ComfyUI workflows: Viggle/Qwen-Image-2.1-viggle-turbo
A DMD-distilled student of Qwen-Image-2.1, sampled with no classifier-free guidance. On most prompts it is hard to tell apart from the base model; small, dense text (8 steps narrows the gap) and complicated edits (multi-reference composition, face swaps, identity-preserving edits) can still fall short of it. The Comparison tab has 32 examples of the official Qwen Space side by side, turbo vs base, same seed.
v0.2.1 (2026-09-24): the step-700 checkpoint of the v0.2 run on the 6-step schedule — intra-prompt diversity 0.98× the base model, 0% composition drift; earlier versions and the numbers are on the model card. Weights: v0.2.1 LoRA r256 from Viggle/Qwen-Image-2.1-viggle-turbo.
The menu switches to the editing sizes as soon as a reference is attached. Custom sizes above 2048² area (1536² when editing) are scaled down, keeping the ratio.
| Prompt | References (optional, up to 6) | Result |
|---|
viggle-turbo (6 steps) vs Qwen-Image-2.1 (40 steps) on 32 of the 37 examples of the official Qwen/Qwen-Image-2.1 Space. Both models get the same prompt, input images and seed, with prompt enhancement off, one sample each and no seed picking. At 5.0× less time the turbo is very competitive with the base model; the clearest gaps are small, dense text (8 steps narrows it) and complicated edits, where the base is still ahead. Pick an example on the left, then drag the slider. The 5 dense-text examples open on the 8-step turbo (6 steps is in the menu), and the menus also offer Qwen's own API reference where the Space ships one.
Median time, editing: 3.0 s vs 14.2 s · text-to-image (~4 MP): 4.7 s vs 26.1 s · all 32 examples: 5.0× faster (one NVIDIA B200, one pipeline call each)
Transparent anime bride · 透明婚纱角色立绘
1664 × 2496 · left: viggle-turbo · 6 steps · 4.7 s · right: Qwen-Image-2.1 base · 40 steps · 26.0 s · drag the handle to compare
- Turbo: Qwen-Image-2.1 + the viggle-turbo v0.2.1 LoRA (rank 256), 6 steps on sigmas
[1, 0.9375, 0.875, 0.75, 0.5, 0.25], no classifier-free guidance, exactly as the Generate tab runs it. - Turbo, 8 steps (the 5 dense-text examples): sigmas
[1, 0.9375, 0.875, 0.75, 0.625, 0.5, 0.25, 0.125]. At 6 steps small Latin text can print twice, like a double exposure (decided in the 0.75 → 0.5 step), and small Chinese strokes can get colour blotches (the last 0.25 → 0 step); each added sigma splits one of the two. In our OCR tests 8 steps raise word recall from 0.71 to 0.83 on the academic infographic's caption (16 seeds) and from 0.76 to 0.90 on the architecture board's Chinese labels (4 seeds); the base model scores 0.95 on both. - Base: Qwen-Image-2.1, 40 steps, default scheduler,
true_cfg_scale=1.0(the pipeline default). - Both: diffusers
QwenImage21Pipelinein bf16, seed 42. Input images are encoded at 1024² pixels; the output keeps the aspect ratio of Qwen's reference at 2048² pixels for text-to-image and 1536² for editing, in multiples of 32. - Time: one pipeline call on one NVIDIA B200 (text encoding, denoising and VAE decoding), after warm-up.
- Left out (5 of 37): three Chinese infographics and a subtitled storyboard whose short prompts leave the on-image text to Qwen's prompt enhancement (without it both models fill them with made-up characters), and a storyboard that asks for a named third-party character.
- Qwen API reference: the output bundled with the example in Qwen's Space, made with Qwen's API, which may differ from the open weights. It is omitted for the 6 text-to-image examples Qwen generated with prompt enhancement on.
Prompts, input images and reference outputs come from the Qwen/Qwen-Image-2.1 Space (Qwen Research License).
Model: Viggle/Qwen-Image-2.1-viggle-turbo · Built with Qwen — distilled from Qwen/Qwen-Image-2.1, which is released under the Qwen RESEARCH LICENSE AGREEMENT (non-commercial: research or evaluation purposes only). This demo inherits that restriction.