# Qwen-Image 2.1 Qwen-Image 2.1 encodes the prompt and any condition images together with a Qwen3-VL model, then denoises the target image with a single-stream block-causal transformer. See [`QwenImage21Transformer2DModel`](../models/qwenimage21_transformer2d) for block-causal attention, the attention processors, and `causal_condition`. The defaults are the values Qwen recommends: 40 steps and no guidance. Pass a `negative_prompt` together with `true_cfg_scale > 1` to turn classifier-free guidance on, which doubles the work per step. ```python import torch from diffusers import QwenImage21Pipeline pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16).to("cuda") # Text-to-image image = pipe("A capybara wearing a wizard hat, oil painting").images[0] image.save("t2i.png") # Image-conditioned editing edited = pipe("Move it to a snowy mountain top", image=image).images[0] edited.save("edit.png") ``` ## Multiple condition images Pass a list to `image` and every entry becomes its own block in the joint sequence: the Qwen3-VL encoder sees them as vision context and the VAE contributes their latent tokens. Block-causal attention keeps each block internally bidirectional while letting later blocks and the target image attend to the earlier ones, so the order you pass them in is the order the model reads them. ```python edited = pipe("Put the flowers from the first image into the second scene", image=[flowers, scene]).images[0] ``` ## Faster attention with flex_attention The default `QwenImage21AttnProcessor` runs the block-causal prefill as one attention call per prefix segment. It needs no compilation and works on any PyTorch build. `QwenImage21FlexAttnProcessor` expresses the same mask as a single `flex_attention` call, which is faster once the model is **_compiled_**. > [!TIP] > Compile the model when you switch to the flex processor. An uncompiled `flex_attention` materializes the full > attention score matrix in fp32, which is much slower and runs out of memory at high resolution. ```python from diffusers.models.transformers.transformer_qwenimage21 import QwenImage21FlexAttnProcessor pipe.transformer.set_attn_processor(QwenImage21FlexAttnProcessor()) pipe.transformer.compile() ``` ## Loading single-file checkpoints ```python import torch from diffusers import QwenImage21Pipeline, QwenImage21Transformer2DModel transformer = QwenImage21Transformer2DModel.from_single_file( "https://huggingface.co/Comfy-Org/Qwen-Image-2.1/blob/main/diffusion_models/qwen_image_2.1_bf16.safetensors", dtype=torch.bfloat16, ) pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", transformer=transformer, dtype=torch.bfloat16).to( "cuda" ) ``` ## Sampling sigmas Model authors can configure a default sampling grid with `sample_sigmas` in the pipeline config. When you load a released checkpoint, its default grid and scheduler settings are restored automatically. To experiment with a different grid at runtime, pass `sigmas` to the pipeline call: ```python # Use the checkpoint's default sampling grid. image = pipe(prompt).images[0] # Override the default grid for this call. image = pipe(prompt, sigmas=[1.0, 0.8, 0.5, 0.2]).images[0] ``` The custom grid above illustrates the API; generation quality depends on the checkpoint and grid. Sigma lists exclude the terminal sigma, which the scheduler appends. Explicit `sigmas` override the configured `sample_sigmas`, and either list determines the number of steps instead of `num_inference_steps`. If neither is provided, the pipeline uses `num_inference_steps` to generate the schedule. The scheduler applies its configured processing to either grid. ## QwenImage21Pipeline [[autodoc]] QwenImage21Pipeline - all - __call__ ## QwenImagePipelineOutput [[autodoc]] pipelines.qwenimage.pipeline_output.QwenImagePipelineOutput