Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/source/en/_toctree.yml
Original file line number Diff line number Diff line change
Expand Up @@ -849,4 +849,4 @@
- local: api/video_processor
title: Video Processor
title: Internal classes
title: API
title: API
6 changes: 6 additions & 0 deletions docs/source/en/api/cache.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,3 +52,9 @@ Cache methods speedup diffusion transformers by storing and reusing intermediate
[[autodoc]] SeaCacheConfig

[[autodoc]] apply_sea_cache

## TextKVCacheConfig

[[autodoc]] TextKVCacheConfig

[[autodoc]] apply_text_kv_cache
2 changes: 1 addition & 1 deletion docs/source/en/api/pipelines/hunyuandit.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,7 +36,7 @@ HunyuanDiT has the following components:

## Optimization

You can optimize the pipeline's runtime and memory consumption with torch.compile and feed-forward chunking. To learn about other optimization methods, check out the [Speed up inference](../../optimization/fp16) and [Reduce memory usage](../../optimization/memory) guides.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think precision and compilation as a title is less exciting as a reader in the sense that it doesn't tell me that it's related to speeding up of inference.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think it's a little strange to broadly name it "Accelerate inference" but then it doesn't include things like caching or the attention backends, which may cause users to overlook those last two options. so i'd rather keep it as "Precision and compilation"

You can optimize the pipeline's runtime and memory consumption with torch.compile and feed-forward chunking. To learn about other optimization methods, see [Optimize and scale](../../stable_diffusion#optimization-techniques).

### Inference

Expand Down
3 changes: 1 addition & 2 deletions docs/source/en/api/pipelines/marigold.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,8 +81,7 @@ The following is a summary of the recommended checkpoints, all of which produce
> Make sure to check out the Schedulers [guide](../../using-diffusers/schedulers) to learn how to explore the tradeoff
> between scheduler speed and quality, and see the [reuse components across pipelines](../../using-diffusers/loading#reusing-models-in-multiple-pipelines) section to learn how to
> efficiently load the same components into multiple pipelines.
> Also, to know more about reducing the memory usage of this pipeline, refer to the ["Reduce memory usage"] section
> [here](./stable_diffusion/svd#reduce-memory-usage).
> Also, to know more about reducing the memory usage of this pipeline, refer to the [Reduce memory usage](../../optimization/memory) guide.

> [!WARNING]
> Marigold pipelines were designed and tested with the scheduler embedded in the model checkpoint.
Expand Down
2 changes: 1 addition & 1 deletion docs/source/en/conceptual/philosophy.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ In a nutshell, Diffusers is built to be a natural extension of PyTorch. Therefor

## Usability over Performance

- While Diffusers has many built-in performance-enhancing features (see [Memory and Speed](https://huggingface.co/docs/diffusers/optimization/fp16)), models are always loaded with the highest precision and lowest optimization. Therefore, by default diffusion pipelines are always instantiated on CPU with float32 precision if not otherwise defined by the user. This ensures usability across different platforms and accelerators and means that no complex installations are required to run the library.
- While Diffusers has many built-in performance-enhancing features (see [Optimize and scale](https://huggingface.co/docs/diffusers/stable_diffusion#optimization-techniques)), models are always loaded with the highest precision and lowest optimization. Therefore, by default diffusion pipelines are always instantiated on CPU with float32 precision if not otherwise defined by the user. This ensures usability across different platforms and accelerators and means that no complex installations are required to run the library.
- Diffusers aims to be a **light-weight** package and therefore has very few required dependencies, but many soft dependencies that can improve performance (such as `accelerate`, `safetensors`, `onnx`, etc...). We strive to keep the library as lightweight as possible so that it can be added without much concern as a dependency on other packages.
- Diffusers prefers simple, self-explainable code over condensed, magic code. This means that short-hand code syntaxes such as lambda functions, and advanced PyTorch operators are often not desired.

Expand Down
132 changes: 60 additions & 72 deletions docs/source/en/optimization/attention_backends.md

Large diffs are not rendered by default.

176 changes: 98 additions & 78 deletions docs/source/en/optimization/cache.md

Large diffs are not rendered by default.

147 changes: 31 additions & 116 deletions docs/source/en/optimization/fp16.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion docs/source/en/optimization/speed-memory-optims.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@ an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express o
specific language governing permissions and limitations under the License.
-->

# Compiling and offloading quantized models
# Quantize, compile, and offload

Optimizing models often involves trade-offs between [inference speed](./fp16) and [memory-usage](./memory). For instance, while [caching](./cache) can boost inference speed, it also increases memory consumption since it needs to store the outputs of intermediate attention layers. A more balanced optimization strategy combines quantizing a model, [torch.compile](./fp16#torchcompile) and various [offloading methods](./memory#offloading).

Expand Down
109 changes: 33 additions & 76 deletions docs/source/en/stable_diffusion.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,123 +10,80 @@ an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express o
specific language governing permissions and limitations under the License.
-->

[[open-in-colab]]
# Overview

# Basic performance
Diffusion inference is computationally expensive, and you often run a [`DiffusionPipeline`] more than once before you like the result. This page provides an overview of the main Diffusers optimization techniques, what they do, and when to use them.

Diffusion is a random process that is computationally demanding. You may need to run the [`DiffusionPipeline`] several times before getting a desired output. That's why it's important to carefully balance generation speed and memory usage in order to iterate faster,
## Starter path

This guide recommends some basic performance tips for using the [`DiffusionPipeline`]. Refer to the Inference Optimization section docs such as [Accelerate inference](./optimization/fp16) or [Reduce memory usage](./optimization/memory) for more detailed performance guides.
When the model fits on one GPU, start with this baseline load. Set `dtype` and place the pipeline on an accelerator. Reach for model CPU offload only when memory is tight. You could also speed up inference with fewer steps or a faster scheduler.

## Memory usage

Reducing the amount of memory used indirectly speeds up generation and can help a model fit on device.

The [`~DiffusionPipeline.enable_model_cpu_offload`] method moves a model to the CPU when it is not in use to save GPU memory.
When you omit `dtype`, Diffusers loads components in `float32`. Pass `dtype=torch.bfloat16` (or `torch.float16` if bfloat16 is unsupported), then place the pipeline on an accelerator with `pipeline.to("cuda")`.

```py
import torch
from diffusers import DiffusionPipeline

pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
device_map="cuda" # or "mps", "xpu", "cpu"
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
)
pipeline.enable_model_cpu_offload()
pipeline.to("cuda") # or "mps", "xpu"

prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
pipeline(prompt).images[0]
print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB")
```

## Inference speed
If the pipeline does not fit, or memory is tight, call [`~DiffusionPipeline.enable_model_cpu_offload`] instead of keeping everything on the GPU. It places the active model on the GPU and keeps the other components on the CPU.

Denoising is the most computationally demanding process during diffusion. Methods that optimizes this process accelerates inference speed. Try the following methods for a speed up.
Skip it when the model fits. Offloading is slower when you do not need it.

- Add `device_map="cuda"` to place the pipeline on a GPU. Placing a model on an accelerator, like a GPU, increases speed because it performs computations in parallel.
- Set `dtype=torch.bfloat16` to execute the pipeline in half-precision. Reducing the data type precision increases speed because it takes less time to perform computations in a lower precision.
For more offloading options, see [Reduce memory usage](./optimization/memory#offloading).

```py
import torch
import time
from diffusers import DiffusionPipeline, DPMSolverMultistepScheduler

pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
device_map="cuda
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
)
pipeline.enable_model_cpu_offload()
```

- Use a faster scheduler, such as [`DPMSolverMultistepScheduler`], which only requires ~20-25 steps.
- Set `num_inference_steps` to a lower value. Reducing the number of inference steps reduces the overall number of computations. However, this can result in lower generation quality.
Lower latency with fewer `num_inference_steps` or a faster scheduler such as [`DPMSolverMultistepScheduler`]. That usually speeds up generation but can reduce image quality versus a slower, higher-quality scheduler. See [Optimization techniques](#optimization-techniques) below for more speed techniques.

```py
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(pipeline.scheduler.config)
import time
from diffusers import DPMSolverMultistepScheduler

prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
pipeline.scheduler = DPMSolverMultistepScheduler.from_config(pipeline.scheduler.config)

start_time = time.perf_counter()
image = pipeline(prompt).images[0]
image = pipeline(prompt, num_inference_steps=25).images[0]
end_time = time.perf_counter()

print(f"Image generation took {end_time - start_time:.3f} seconds")
```

## Generation quality

Many modern diffusion models deliver high-quality images out-of-the-box. However, you can still improve generation quality by trying the following.

- Try a more detailed and descriptive prompt. Include details such as the image medium, subject, style, and aesthetic. A negative prompt may also help by guiding a model away from undesirable features by using words like low quality or blurry.

```py
import torch
from diffusers import DiffusionPipeline

pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
device_map="cuda" # or "mps", "xpu", "cpu"
)

prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
negative_prompt = "low quality, blurry, ugly, poor details"
pipeline(prompt, negative_prompt=negative_prompt).images[0]
```
## Optimization techniques

For more details about creating better prompts, take a look at the [Prompt techniques](./using-diffusers/weighted_prompts) doc.
When the starter path is not enough, use these techniques. If inference is too slow, start with `torch.compile`, caching, or attention backends. If you are out of memory, start with offloading or quantization.

- Try a different scheduler, like [`HeunDiscreteScheduler`] or [`LMSDiscreteScheduler`], that gives up generation speed for quality.
Faster inference:

```py
import torch
from diffusers import DiffusionPipeline, HeunDiscreteScheduler
- [torch.compile](./optimization/fp16#torchcompile) — Compile the UNet, transformer, or VAE into optimized kernels.
- [Regional compilation](./optimization/fp16#regional-compilation) — Compile repeated blocks to cut `torch.compile` cold-start latency and reuse compiled artifacts.
- [Caching](./optimization/cache) — Reuse intermediates across denoising steps when you want more speed and can spend memory.
- [Attention backends](./optimization/attention_backends) — Swap Diffusers attention implementations through a unified API when attention is the bottleneck.
- [Kernels](./optimization/fp16#kernels) — Load optimized Hub compute kernels (attention and custom CUDA ops such as RMSNorm or RoPE) when you need hardware-specific speedups beyond stock PyTorch.

pipeline = DiffusionPipeline.from_pretrained(
"stabilityai/stable-diffusion-xl-base-1.0",
dtype=torch.bfloat16,
device_map="cuda" # or "mps", "xpu", "cpu"
)
pipeline.scheduler = HeunDiscreteScheduler.from_config(pipeline.scheduler.config)
Less memory:

prompt = """
cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California
highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain
"""
negative_prompt = "low quality, blurry, ugly, poor details"
pipeline(prompt, negative_prompt=negative_prompt).images[0]
```
- [Offloading](./optimization/memory#offloading) — Move inactive models or layers to the CPU with CPU, model, or group offloading.
- [Quantization](./quantization/overview) — Load smaller weights to cut memory (some backends also speed up inference). [GGUF](./quantization/gguf) is a common starting point.
- [VAE slicing](./optimization/memory#vae-slicing) and [VAE tiling](./optimization/memory#vae-tiling) — Decode large batches or high-resolution images in pieces to lower peak memory.

## Next steps
Both:

Diffusers offers more advanced and powerful optimizations such as [group-offloading](./optimization/memory#group-offloading) and [regional compilation](./optimization/fp16#regional-compilation). To learn more about how to maximize performance, take a look at the Inference Optimization section.
- [Quantize, compile, and offload](./optimization/speed-memory-optims) — Combine quantization, `torch.compile`, and offloading when one technique is not enough.
Original file line number Diff line number Diff line change
Expand Up @@ -304,4 +304,4 @@ pipeline = AutoPipelineForText2Image.from_pretrained("stable-diffusion-v1-5/stab
pipeline.unet = torch.compile(pipeline.unet, mode="reduce-overhead", fullgraph=True)
```

For more tips on how to optimize your code to save memory and speed up inference, read the [Accelerate inference](../optimization/fp16) and [Reduce memory usage](../optimization/memory) guides.
For more tips on how to optimize your code to save memory and speed up inference, see [Optimize and scale](../stable_diffusion#optimization-techniques).
2 changes: 1 addition & 1 deletion docs/source/en/using-diffusers/img2img.md
Original file line number Diff line number Diff line change
Expand Up @@ -593,4 +593,4 @@ With [`torch.compile`](../optimization/fp16#torchcompile), you can boost your in
pipeline.unet = torch.compile(pipeline.unet, mode="reduce-overhead", fullgraph=True)
```

To learn more, take a look at the [Reduce memory usage](../optimization/memory) and [Accelerate inference](../optimization/fp16) guides.
To learn more, take a look at [Optimize and scale](../stable_diffusion#optimization-techniques).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

../stable_diffusion#optimization-techniques

Do we still have a reason keep these under "stable_diffusion". I don't see any reason to.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stable_diffusion.md is actually the overview, so it may be nice to point users toward an overview of optimization techniques for a particular task and then they can choose what they want

2 changes: 1 addition & 1 deletion docs/source/en/using-diffusers/inpaint.md
Original file line number Diff line number Diff line change
Expand Up @@ -797,4 +797,4 @@ To speed-up your inference code even more, use [`torch_compile`](../optimization
pipeline.unet = torch.compile(pipeline.unet, mode="reduce-overhead", fullgraph=True)
```

Learn more in the [Reduce memory usage](../optimization/memory) and [Accelerate inference](../optimization/fp16) guides.
Learn more in [Optimize and scale](../stable_diffusion#optimization-techniques).
Loading