diff --git a/docs/source/en/_toctree.yml b/docs/source/en/_toctree.yml index 4b4e4247195d..d12740b44174 100644 --- a/docs/source/en/_toctree.yml +++ b/docs/source/en/_toctree.yml @@ -849,4 +849,4 @@ - local: api/video_processor title: Video Processor title: Internal classes - title: API + title: API \ No newline at end of file diff --git a/docs/source/en/api/cache.md b/docs/source/en/api/cache.md index 5d0d16585013..67160cbafa4b 100644 --- a/docs/source/en/api/cache.md +++ b/docs/source/en/api/cache.md @@ -52,3 +52,9 @@ Cache methods speedup diffusion transformers by storing and reusing intermediate [[autodoc]] SeaCacheConfig [[autodoc]] apply_sea_cache + +## TextKVCacheConfig + +[[autodoc]] TextKVCacheConfig + +[[autodoc]] apply_text_kv_cache diff --git a/docs/source/en/api/pipelines/hunyuandit.md b/docs/source/en/api/pipelines/hunyuandit.md index 8d056bc61f7c..957a25737cef 100644 --- a/docs/source/en/api/pipelines/hunyuandit.md +++ b/docs/source/en/api/pipelines/hunyuandit.md @@ -36,7 +36,7 @@ HunyuanDiT has the following components: ## Optimization -You can optimize the pipeline's runtime and memory consumption with torch.compile and feed-forward chunking. To learn about other optimization methods, check out the [Speed up inference](../../optimization/fp16) and [Reduce memory usage](../../optimization/memory) guides. +You can optimize the pipeline's runtime and memory consumption with torch.compile and feed-forward chunking. To learn about other optimization methods, see [Optimize and scale](../../stable_diffusion#optimization-techniques). ### Inference diff --git a/docs/source/en/api/pipelines/marigold.md b/docs/source/en/api/pipelines/marigold.md index d5a1c6bd2675..c6c3125b2d2d 100644 --- a/docs/source/en/api/pipelines/marigold.md +++ b/docs/source/en/api/pipelines/marigold.md @@ -81,8 +81,7 @@ The following is a summary of the recommended checkpoints, all of which produce > Make sure to check out the Schedulers [guide](../../using-diffusers/schedulers) to learn how to explore the tradeoff > between scheduler speed and quality, and see the [reuse components across pipelines](../../using-diffusers/loading#reusing-models-in-multiple-pipelines) section to learn how to > efficiently load the same components into multiple pipelines. -> Also, to know more about reducing the memory usage of this pipeline, refer to the ["Reduce memory usage"] section -> [here](./stable_diffusion/svd#reduce-memory-usage). +> Also, to know more about reducing the memory usage of this pipeline, refer to the [Reduce memory usage](../../optimization/memory) guide. > [!WARNING] > Marigold pipelines were designed and tested with the scheduler embedded in the model checkpoint. diff --git a/docs/source/en/conceptual/philosophy.md b/docs/source/en/conceptual/philosophy.md index 3d7f6c691c92..ad9b7f96df9b 100644 --- a/docs/source/en/conceptual/philosophy.md +++ b/docs/source/en/conceptual/philosophy.md @@ -21,7 +21,7 @@ In a nutshell, Diffusers is built to be a natural extension of PyTorch. Therefor ## Usability over Performance -- While Diffusers has many built-in performance-enhancing features (see [Memory and Speed](https://huggingface.co/docs/diffusers/optimization/fp16)), models are always loaded with the highest precision and lowest optimization. Therefore, by default diffusion pipelines are always instantiated on CPU with float32 precision if not otherwise defined by the user. This ensures usability across different platforms and accelerators and means that no complex installations are required to run the library. +- While Diffusers has many built-in performance-enhancing features (see [Optimize and scale](https://huggingface.co/docs/diffusers/stable_diffusion#optimization-techniques)), models are always loaded with the highest precision and lowest optimization. Therefore, by default diffusion pipelines are always instantiated on CPU with float32 precision if not otherwise defined by the user. This ensures usability across different platforms and accelerators and means that no complex installations are required to run the library. - Diffusers aims to be a **light-weight** package and therefore has very few required dependencies, but many soft dependencies that can improve performance (such as `accelerate`, `safetensors`, `onnx`, etc...). We strive to keep the library as lightweight as possible so that it can be added without much concern as a dependency on other packages. - Diffusers prefers simple, self-explainable code over condensed, magic code. This means that short-hand code syntaxes such as lambda functions, and advanced PyTorch operators are often not desired. diff --git a/docs/source/en/optimization/attention_backends.md b/docs/source/en/optimization/attention_backends.md index a33fbe89815f..d7eac88888d9 100644 --- a/docs/source/en/optimization/attention_backends.md +++ b/docs/source/en/optimization/attention_backends.md @@ -14,35 +14,35 @@ specific language governing permissions and limitations under the License. --> > [!NOTE] > The attention dispatcher is an experimental feature. Please open an issue if you have any feedback or encounter any problems. -Diffusers provides several optimized attention algorithms that are more memory and computationally efficient through it's *attention dispatcher*. The dispatcher acts as a router for managing and switching between different attention implementations and provides a unified interface for interacting with them. +Most Diffusers transformer models route attention through an *attention dispatcher* so you can switch optimized backends behind one API. The dispatcher manages registered implementations and exposes a unified call path for them. Some models, such as many autoencoders, don't use the dispatcher because of their internals, and switching the backend has no effect on their attention layers. -Refer to the table below for an overview of the available attention families and to the [Available backends](#available-backends) section for a more complete list. +Refer to the table below for an overview of the available attention families and to the [Available backends](#available-backends) section for a more complete list. The fastest backend depends on the model, GPU, dtype, and input shape. | attention family | main feature | |---|---| | FlashAttention | minimizes memory reads/writes through tiling and recomputation | | AI Tensor Engine for ROCm | FlashAttention implementation optimized for AMD ROCm accelerators | | SageAttention | quantizes attention to int8 | +| FlexAttention | PyTorch FlexAttention | | PyTorch native | built-in PyTorch implementation using [scaled_dot_product_attention](./fp16#scaled-dot-product-attention) | | xFormers | memory-efficient attention with support for various attention kernels | -This guide will show you how to set and use the different attention backends. +Hub backends (the `*_hub` names) only need the [Kernels](https://github.com/huggingface/kernels) library and download the kernel on first use. Other backends need their own package, such as `flash-attn` or `sageattention`. The [Available backends](#available-backends) table lists the requirements Diffusers checks when you enable a backend. -## set_attention_backend +## Set a backend on the model -The [`~ModelMixin.set_attention_backend`] method iterates through all the modules in the model and sets the appropriate attention backend to use. The attention backend setting persists until [`~ModelMixin.reset_attention_backend`] is called. +The [`~ModelMixin.set_attention_backend`] method walks the model’s attention layers and applies the chosen backend on each one. It also sets the dispatcher’s process-wide active backend to the same value. -The example below demonstrates how to enable the `_flash_3_hub` implementation for FlashAttention-3 from the [`kernels`](https://github.com/huggingface/kernels) library, which allows you to instantly use optimized compute kernels from the Hub without requiring any setup. +[`~ModelMixin.reset_attention_backend`] clears the backend on attention layers only. It does not clear the process-wide active backend. For a temporary switch that restores the previous active backend on exit, use the [attention_backend](#try-a-backend-temporarily) context manager. -> [!NOTE] -> FlashAttention-3 requires Ampere GPUs at a minimum. +The example below enables `_flash_3_hub` (FlashAttention-3 from the Hub) with `device_map="cuda"`. ```py import torch from diffusers import QwenImagePipeline pipeline = QwenImagePipeline.from_pretrained( - "Qwen/Qwen-Image", dtype=torch.bfloat16, device_map="cuda" # or "mps", "xpu", "cpu" + "Qwen/Qwen-Image", dtype=torch.bfloat16, device_map="cuda" ) pipeline.transformer.set_attention_backend("_flash_3_hub") @@ -53,22 +53,19 @@ highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous pipeline(prompt).images[0] ``` -To restore the default attention backend, call [`~ModelMixin.reset_attention_backend`]. - -```py -pipeline.transformer.reset_attention_backend() -``` +> [!NOTE] +> The non-Hub FlashAttention-3 backends (`_flash_3`, `_flash_varlen_3`) require building FlashAttention-3 from source and will be deprecated soon. Use `_flash_3_hub` or `_flash_3_varlen_hub` instead. -## attention_backend context manager +## Try a backend temporarily -The [attention_backend](https://github.com/huggingface/diffusers/blob/5e181eddfe7e44c1444a2511b0d8e21d177850a0/src/diffusers/models/attention_dispatch.py#L225) context manager temporarily sets an attention backend for a model within the context. Outside the context, the default attention (PyTorch's native scaled dot product attention) is used. This is useful if you want to use different backends for different parts of a pipeline or if you want to test the different backends. +The [`attention_backend`] context manager sets the process-wide active backend for the duration of the block and restores the previous backend when the block exits. Use it to try a backend for one call without leaving a permanent backend applied from [`~ModelMixin.set_attention_backend`]. ```py import torch -from diffusers import QwenImagePipeline +from diffusers import QwenImagePipeline, attention_backend pipeline = QwenImagePipeline.from_pretrained( - "Qwen/Qwen-Image", dtype=torch.bfloat16, device_map="cuda" # or "mps", "xpu", "cpu" + "Qwen/Qwen-Image", dtype=torch.bfloat16, device_map="cuda" ) prompt = """ cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California @@ -80,42 +77,46 @@ with attention_backend("_flash_3_hub"): ``` > [!TIP] -> Most attention backends support `torch.compile` without graph breaks and can be used to further speed up inference. +> Most attention backends work with `torch.compile`. Whether that speeds up your pipeline depends on the model and backend. See [Precision and compilation](./fp16#torchcompile). ## Trusting remote kernels -Hub backends and other kernel-backed features (such as [GGUF](../quantization/gguf) and [Nunchaku Lite](../quantization/nunchaku)) download compute kernels from the Hub with [`kernels`](https://github.com/huggingface/kernels) and execute their code locally. +Hub backends need the [Kernels](https://github.com/huggingface/kernels) library first, and then Diffusers fetches the Hub kernel on first use. -By default, `kernels` only loads a kernel when its publisher is a trusted kernel publisher on the Hub. Kernels published under the [`kernels-community`](https://huggingface.co/kernels-community) organization are trusted, so Diffusers loads them without any additional configuration. The `_flash_3_hub`, `flash_hub`, `sage_hub`, and the other Hub attention backends all resolve to `kernels-community` repositories. +Hub attention backends download compute kernels with the [Kernels](https://github.com/huggingface/kernels) library and run them locally. Most Hub attention names (`_flash_3_hub`, `flash_hub`, and the other `*_hub` backends) resolve to the [kernels-community](https://huggingface.co/kernels-community) organization. That organization is a trusted publisher in Kernels, so Diffusers loads those attention kernels without setting `DIFFUSERS_TRUST_REMOTE_KERNELS`. -Kernels from any other publisher are not vetted. Loading one downloads and runs code that Diffusers cannot vouch for, so Diffusers keeps it disabled unless you explicitly opt in with the `DIFFUSERS_TRUST_REMOTE_KERNELS` environment variable. When set, Diffusers forwards `trust_remote_code=True` to `kernels` so it loads kernels from untrusted publishers too. +> [!NOTE] +> The SageAttention Hub backends (`sage_hub` and `sage_blackwell_hub`) load from the [SageAttention](https://huggingface.co/SageAttention) organization instead. Set `DIFFUSERS_TRUST_REMOTE_KERNELS=true` to use them. + +Other kernel-backed features such as [GGUF](../quantization/gguf) and [Nunchaku Lite](../quantization/nunchaku) can pull kernels from publishers outside kernels-community. Those paths stay blocked unless you opt in with `DIFFUSERS_TRUST_REMOTE_KERNELS`. When set, Diffusers forwards `trust_remote_code=True` to Kernels so untrusted publishers can load too. ```bash export DIFFUSERS_TRUST_REMOTE_KERNELS=true ``` -Only enable this after inspecting the kernel repository, since it grants the downloaded code the ability to run on your machine. Without it, loading a kernel from an untrusted publisher raises an error. Diffusers performs this check itself, so it also applies to `kernels<0.14.0`, which predates the `trust_remote_code` argument. Setting `DIFFUSERS_DISABLE_REMOTE_CODE=true` disables remote code globally and takes precedence over `DIFFUSERS_TRUST_REMOTE_KERNELS`. +Only enable this after inspecting the kernel repository. Without it, loading a kernel from an untrusted publisher raises an error. Diffusers performs this check itself, so it also applies to `kernels<0.14.0`, which predates the `trust_remote_code` argument. Setting `DIFFUSERS_DISABLE_REMOTE_CODE=true` disables remote code globally and takes precedence over `DIFFUSERS_TRUST_REMOTE_KERNELS`. ## Checks -The attention dispatcher includes debugging checks that catch common errors before they cause problems. +The attention dispatcher can run debugging checks before each dispatched attention call. Which checks run depends on the constraints registered for the active backend. 1. Device checks verify that query, key, and value tensors live on the same device. -2. Data type checks confirm tensors have matching dtypes and use either bfloat16 or float16. +2. Data type checks, where registered, confirm matching dtypes and often require `bfloat16` or `float16`. 3. Shape checks validate tensor dimensions and prevent mixing attention masks with causal flags. -Enable these checks by setting the `DIFFUSERS_ATTN_CHECKS` environment variable. Checks add overhead to every attention operation, so they're disabled by default. +Enable checks with the `DIFFUSERS_ATTN_CHECKS` environment variable. Checks add overhead, so they are disabled by default. ```bash export DIFFUSERS_ATTN_CHECKS=yes ``` -The checks are run now before every attention operation. +With checks on, Diffusers runs those constraints before every dispatched attention call. The low-level example below calls [`dispatch_attention_fn`] directly. Pipeline inference does not need that import. It only needs the backend set via [`~ModelMixin.set_attention_backend`] or [`attention_backend`]. ```py import torch +from diffusers.models.attention_dispatch import attention_backend, dispatch_attention_fn -query = torch.randn(1, 10, 8, 64, dtype=torch.bfloat16, device="cuda") # or "mps", "xpu", "cpu" +query = torch.randn(1, 10, 8, 64, dtype=torch.bfloat16, device="cuda") key = torch.randn(1, 10, 8, 64, dtype=torch.bfloat16, device="cuda") value = torch.randn(1, 10, 8, 64, dtype=torch.bfloat16, device="cuda") @@ -127,49 +128,36 @@ except Exception as e: print(f"✗ Flash Attention failed: {e}") ``` -You can also configure the registry directly. - -```py -from diffusers.models.attention_dispatch import _AttentionBackendRegistry - -_AttentionBackendRegistry._checks_enabled = True -``` - ## Available backends -Refer to the table below for a complete list of available attention backends and their variants. - -
-Expand - -| Backend Name | Family | Description | -|--------------|--------|-------------| -| `native` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | Default backend using PyTorch's scaled_dot_product_attention | -| `flex` | [FlexAttention](https://docs.pytorch.org/docs/stable/nn.attention.flex_attention.html#module-torch.nn.attention.flex_attention) | PyTorch FlexAttention implementation | -| `_native_cudnn` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | CuDNN-optimized attention | -| `_native_efficient` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | Memory-efficient attention | -| `_native_flash` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | PyTorch's FlashAttention | -| `_native_math` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | Math-based attention (fallback) | -| `_native_npu` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | NPU-optimized attention | -| `_native_xla` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | XLA-optimized attention | -| `flash` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-2 | -| `flash_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-2 from kernels | -| `flash_varlen` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention | -| `flash_varlen_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention from kernels | -| `aiter_fa2_hub` | [AI Tensor Engine for ROCm](https://github.com/ROCm/aiter) | FlashAttention-2 for AMD ROCm from kernels | -| `flash_4_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-4 | -| `_flash_3` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-3 | -| `_flash_varlen_3` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention-3 | -| `_flash_3_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-3 from kernels | -| `_flash_3_varlen_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention-3 from kernels | -| `sage` | [SageAttention](https://github.com/thu-ml/SageAttention) | Quantized attention (INT8 QK) | -| `sage_hub` | [SageAttention](https://github.com/thu-ml/SageAttention) | Quantized attention (INT8 QK) from kernels | -| `sage_blackwell_hub` | [SageAttention](https://github.com/thu-ml/SageAttention) | SageAttention3 FP4 attention for SM120 Blackwell GPUs from kernels | -| `sage_varlen` | [SageAttention](https://github.com/thu-ml/SageAttention) | Variable length SageAttention | -| `_sage_qk_int8_pv_fp8_cuda` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP8 PV (CUDA) | -| `_sage_qk_int8_pv_fp8_cuda_sm90` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP8 PV (SM90) | -| `_sage_qk_int8_pv_fp16_cuda` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP16 PV (CUDA) | -| `_sage_qk_int8_pv_fp16_triton` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP16 PV (Triton) | -| `xformers` | [xFormers](https://github.com/facebookresearch/xformers) | Memory-efficient attention | - -
+Refer to the table below for a complete list of available attention backends and their variants. Diffusers checks package availability and version pins when you enable a backend. + +| Backend Name | Family | Description | Prerequisite | +|--------------|--------|-------------|--------------| +| `native` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | Default backend using PyTorch's scaled_dot_product_attention | None | +| `flex` | [FlexAttention](https://docs.pytorch.org/docs/stable/nn.attention.flex_attention.html#module-torch.nn.attention.flex_attention) | PyTorch FlexAttention | `torch>=2.5.0` | +| `_native_cudnn` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | CuDNN-optimized attention | CUDA + CuDNN | +| `_native_efficient` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | Memory-efficient attention | None beyond PyTorch | +| `_native_flash` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | PyTorch's FlashAttention | CUDA | +| `_native_math` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | Math-based attention (fallback) | None | +| `_native_npu` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | NPU-optimized attention | `torch_npu` | +| `_native_xla` | [PyTorch native](https://docs.pytorch.org/docs/stable/generated/torch.nn.attention.SDPBackend.html#torch.nn.attention.SDPBackend) | XLA-optimized attention | `torch_xla>=2.2` | +| `flash` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-2 | `flash-attn>=2.6.3` | +| `flash_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-2 from Hub kernels | `kernels>=0.12` | +| `flash_varlen` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention | `flash-attn>=2.6.3` | +| `flash_varlen_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention from Hub kernels | `kernels>=0.12` | +| `aiter_fa2_hub` | [AI Tensor Engine for ROCm](https://github.com/ROCm/aiter) | FlashAttention-2 for AMD ROCm from Hub kernels (`bfloat16`) | `kernels>=0.12`, ROCm | +| `flash_4_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-4 from Hub kernels | `kernels>=0.12.3` | +| `_flash_3` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-3 (local; deprecated soon) | Build FA3 from source | +| `_flash_varlen_3` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention-3 (local; deprecated soon) | Build FA3 from source | +| `_flash_3_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | FlashAttention-3 from Hub kernels | `kernels>=0.12` | +| `_flash_3_varlen_hub` | [FlashAttention](https://github.com/Dao-AILab/flash-attention) | Variable length FlashAttention-3 from Hub kernels | `kernels>=0.12` | +| `sage` | [SageAttention](https://github.com/thu-ml/SageAttention) | Quantized attention (INT8 QK) | `sageattention>=2.1.1` | +| `sage_hub` | [SageAttention](https://github.com/thu-ml/SageAttention) | Quantized attention (INT8 QK) from Hub kernels | `kernels>=0.12`, `DIFFUSERS_TRUST_REMOTE_KERNELS=true` | +| `sage_blackwell_hub` | [SageAttention](https://github.com/thu-ml/SageAttention) | SageAttention3 FP4 attention for SM120 Blackwell GPUs from Hub kernels | `kernels>=0.12`, `DIFFUSERS_TRUST_REMOTE_KERNELS=true` | +| `sage_varlen` | [SageAttention](https://github.com/thu-ml/SageAttention) | Variable length SageAttention | `sageattention>=2.1.1` | +| `_sage_qk_int8_pv_fp8_cuda` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP8 PV (CUDA) | `sageattention>=2.1.1` | +| `_sage_qk_int8_pv_fp8_cuda_sm90` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP8 PV (SM90) | `sageattention>=2.1.1`; SM90 | +| `_sage_qk_int8_pv_fp16_cuda` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP16 PV (CUDA) | `sageattention>=2.1.1` | +| `_sage_qk_int8_pv_fp16_triton` | [SageAttention](https://github.com/thu-ml/SageAttention) | INT8 QK + FP16 PV (Triton) | `sageattention>=2.1.1` | +| `xformers` | [xFormers](https://github.com/facebookresearch/xformers) | Memory-efficient attention | `xformers>=0.0.29` | diff --git a/docs/source/en/optimization/cache.md b/docs/source/en/optimization/cache.md index 3a771336e4b7..4d07e8111627 100644 --- a/docs/source/en/optimization/cache.md +++ b/docs/source/en/optimization/cache.md @@ -11,50 +11,57 @@ specific language governing permissions and limitations under the License. --> # Caching -Caching accelerates inference by storing and reusing intermediate outputs of different layers, such as attention and feedforward layers, instead of performing the entire computation at each inference step. It significantly improves generation speed at the expense of more memory and doesn't require additional training. +Caching reuses intermediate layer outputs across denoising steps to speed up inference. It uses more memory and doesn't need training. Enable a method with a config on a transformer that supports caching. -This guide shows you how to use the caching methods supported in Diffusers. +## Choose a cache method -## Pyramid Attention Broadcast +Pick a method depending on how much config you will set, and the fit you need. Each method balances speed, memory, and how closely outputs match the uncached run differently. Compare outputs with and without the cache on your own prompts before relying on a method. + +| Method | Use when | Notes | +|--------|----------|-------| +| Text KV Cache | NucleusMoE image, reuse text K/V projections across steps | Currently hooks NucleusMoE image blocks | +| SeaCache | Cosmos 3 video generation | Settings often don't transfer across models | +| FirstBlockCache | Want one main speed/quality knob on a registered transformer | One `threshold` controls how often steps are skipped | +| MagCache | Have magnitude ratios for your checkpoint and scheduler, or will calibrate first | Ratios are checkpoint and scheduler specific | +| TaylorSeer | Want to predict later activations from earlier steps | Higher `max_order` uses more memory | +| PAB | Video, willing to tune attention reuse (block and timestep skip ranges per attention kind) | Skip ranges need tuning per model | +| FasterCache | Like PAB, plus optional CFG-branch skipping | Experimental | -[Pyramid Attention Broadcast (PAB)](https://huggingface.co/papers/2408.12588) is based on the observation that attention outputs aren't that different between successive timesteps of the generation process. The attention differences are smallest in the cross attention layers and are generally cached over a longer timestep range. This is followed by temporal attention and spatial attention layers. +## Pyramid Attention Broadcast -> [!TIP] -> Not all video models have three types of attention (cross, temporal, and spatial)! +[Pyramid Attention Broadcast (PAB)](https://huggingface.co/papers/2408.12588) approximates attention across denoising steps by reusing attention outputs for some blocks and timesteps instead of recomputing every step. Config separates attention kinds (spatial, temporal, cross) when the model has them. Not every video model exposes all three, and set only the ranges that match the blocks you have. -PAB can be combined with other techniques like sequence parallelism and classifier-free guidance parallelism (data parallelism) for near real-time video generation. +Each kind uses a `*_attention_block_skip_range` (how often to recompute vs reuse within the window) and a `*_attention_timestep_skip_range` (which denoising timesteps may skip). You must pass `current_timestep_callback` so the hook can read the pipeline’s current timestep. Wider or more aggressive skips usually mean more speed and more quality risk. -Set up and pass a [`PyramidAttentionBroadcastConfig`] to a pipeline's transformer to enable it. The `spatial_attention_block_skip_range` controls how often to skip attention calculations in the spatial attention blocks and the `spatial_attention_timestep_skip_range` is the range of timesteps to skip. Take care to choose an appropriate range because a smaller interval can lead to slower inference speeds and a larger interval can result in lower generation quality. +Pass a [`PyramidAttentionBroadcastConfig`] to enable it. ```python import torch from diffusers import CogVideoXPipeline, PyramidAttentionBroadcastConfig -pipeline = CogVideoXPipeline.from_pretrained("THUDM/CogVideoX-5b", dtype=torch.bfloat16) -pipeline.to("cuda") # or "mps", "xpu", "cpu" +pipe = CogVideoXPipeline.from_pretrained("THUDM/CogVideoX-5b", dtype=torch.bfloat16) +pipe.to("cuda") # or "mps", "xpu", "cpu" config = PyramidAttentionBroadcastConfig( spatial_attention_block_skip_range=2, spatial_attention_timestep_skip_range=(100, 800), current_timestep_callback=lambda: pipe.current_timestep, ) -pipeline.transformer.enable_cache(config) +pipe.transformer.enable_cache(config) ``` ## FasterCache -[FasterCache](https://huggingface.co/papers/2410.19355) caches and reuses attention features similar to [PAB](#pyramid-attention-broadcast) since output differences are small for each successive timestep. +[FasterCache](https://huggingface.co/papers/2410.19355) caches and reuses attention features similar to [PAB](#pyramid-attention-broadcast). It can also skip the unconditional branch under classifier-free guidance and estimate it from the conditional branch when successive latents are redundant enough. -This method may also choose to skip the unconditional branch prediction, when using classifier-free guidance for sampling (common in most base models), and estimate it from the conditional branch prediction if there is significant redundancy in the predicted latent outputs between successive timesteps. - -Set up and pass a [`FasterCacheConfig`] to a pipeline's transformer to enable it. +Pass a [`FasterCacheConfig`] to enable it. Like PAB, set `*_attention_block_skip_range` and `*_attention_timestep_skip_range` for the attention kinds you have, plus the CFG-branch skip options when you want them. ```python import torch from diffusers import CogVideoXPipeline, FasterCacheConfig -pipe line= CogVideoXPipeline.from_pretrained("THUDM/CogVideoX-5b", dtype=torch.bfloat16) -pipeline.to("cuda") # or "mps", "xpu", "cpu" +pipe = CogVideoXPipeline.from_pretrained("THUDM/CogVideoX-5b", dtype=torch.bfloat16) +pipe.to("cuda") # or "mps", "xpu", "cpu" config = FasterCacheConfig( spatial_attention_block_skip_range=2, @@ -62,45 +69,22 @@ config = FasterCacheConfig( current_timestep_callback=lambda: pipe.current_timestep, attention_weight_callback=lambda _: 0.3, unconditional_batch_skip_range=5, - unconditional_batch_timestep_skip_range=(-1, 781), + unconditional_batch_timestep_skip_range=(-1, 641), tensor_format="BFCHW", ) -pipeline.transformer.enable_cache(config) +pipe.transformer.enable_cache(config) ``` ## SeaCache -[SeaCache](https://huggingface.co/papers/2602.18993) compares Spectral-Evolution-Aware (SEA) indicators between -successive denoising steps. When the accumulated indicator change remains below a threshold, it skips the expensive -transformer block stack and predicts its output from cached residuals. Build the indicator from the visual latents that -form the generated output. Include clean conditioning frames when they are part of that output trajectory, as in -image-to-video and video-to-video generation. Exclude separate visual hints that condition the generation but are not -part of the output. Text conditioning is excluded because it is not a visual latent. - -Cosmos 3 Transfer packs control hints as separate visual sequences, so its adapter excludes them from the indicator. -Control-CFG branches compare the same output trajectory while retaining their own cached residuals. Control hints still -condition the transformer. - -The implementation provides built-in adapters for the following models: +[SeaCache](https://huggingface.co/papers/2602.18993) compares Spectral Evolution Aware (SEA) indicators between successive denoising steps. When the accumulated change stays under a threshold, it skips the transformer block stack and predicts the output from cached residuals. The method is approximate and designed for video generation. -- **Cosmos 3** is the primary optimized and benchmarked integration. It caches the complete decoder stack through a - post-normalization boundary. -- **Wan T2V** uses the generic repeated-block path in eager mode. This integration demonstrates how another - single-stream video transformer can provide raw vision latents to SeaCache; it is not a claim that the same cache - parameters are optimal for Wan or that other Wan variants are supported. +Built-in adapters for SeaCache include: -Other video transformers can integrate with the generic path when they use `CacheMixin`, expose a recognized repeated -block list, and register the block input/output layout in `TransformerBlockRegistry`. The pipeline must enter a -`cache_context` for every transformer call, attach `step_index`, `sigma`, and `num_inference_steps`, and use separate -context names for independent trajectories such as conditional and unconditional guidance. Pass a `raw_vision_callback` -that returns the visual latents forming the generated output when no built-in adapter is available. Validate output -quality and tune the cache parameters for each model and scheduler; support and benchmark results do not transfer -automatically from Cosmos 3. +- Cosmos 3 is the primary optimized and benchmarked integration. +- Wan T2V uses the generic repeated-block path as a demo for how to provide the raw vision latents to SeaCache. The same cache parameters may not transfer to Wan or other Wan variants. -### Cosmos 3 - -SeaCache is disabled by default. Enable it on the transformer; the Cosmos 3 denoising loop attaches the active -scheduler step, sigma, and step count to each `cache_context` call, so no extra wiring is needed: +Enable SeaCache on the transformer. The Cosmos 3 denoising loop attaches scheduler step, sigma, and step count to each `cache_context`, so no extra parameters are needed. ```python from diffusers import Cosmos3OmniPipeline, SeaCacheConfig @@ -109,35 +93,33 @@ pipe = Cosmos3OmniPipeline.from_pretrained("nvidia/Cosmos3-Nano") pipe.transformer.enable_cache(SeaCacheConfig(threshold=0.2, max_consecutive_cached=2)) ``` -This model-level API works with [`Cosmos3OmniPipeline`], [`Cosmos3OmniModularPipeline`], and -[`Cosmos3DistilledModularPipeline`]. SeaCache is an approximate optimization and may change generated outputs. Call -`pipe.transformer.disable_cache()` when you need every denoising step to execute the full transformer. +SeaCache may change outputs. Call `pipe.transformer.disable_cache()` when you need every step to run the full transformer. The same enable call works with [`Cosmos3OmniPipeline`], [`Cosmos3OmniModularPipeline`], and [`Cosmos3DistilledModularPipeline`]. + +To use SeaCache with another video transformer, follow [Add caching to a new model](#add-caching-to-a-new-model). SeaCache also needs `step_index`, `sigma`, and `num_inference_steps` in every `cache_context`, and a `raw_vision_callback` when no built-in adapter exists. Tune parameters per model and scheduler. ## FirstBlockCache -[FirstBlock Cache](https://huggingface.co/docs/diffusers/main/en/api/cache#diffusers.FirstBlockCacheConfig) checks how much the early layers of the denoiser changes from one timestep to the next. If the change is small, the model skips the expensive later layers and reuses the previous output. +[`FirstBlockCacheConfig`] checks how much the early layers of the denoiser change from one timestep to the next. If the change is small, the model skips the expensive later layers and reuses the previous output. -```py +Enable it through `enable_cache` so `disable_cache` and `is_cache_enabled` stay in sync. The default `threshold` is `0.05`. A higher value such as `0.2` skips more often for extra speed, but generation quality may drop. + +```python import torch -from diffusers import DiffusionPipeline -from diffusers.hooks import apply_first_block_cache, FirstBlockCacheConfig +from diffusers import DiffusionPipeline, FirstBlockCacheConfig -pipeline = DiffusionPipeline.from_pretrained( +pipe = DiffusionPipeline.from_pretrained( "Qwen/Qwen-Image", dtype=torch.bfloat16 ) -apply_first_block_cache(pipeline.transformer, FirstBlockCacheConfig(threshold=0.2)) +pipe.transformer.enable_cache(FirstBlockCacheConfig(threshold=0.2)) ``` -## TaylorSeer Cache - -[TaylorSeer Cache](https://huggingface.co/papers/2403.06923) accelerates diffusion inference by using Taylor series expansions to approximate and cache intermediate activations across denoising steps. The method predicts future outputs based on past computations, reusing them at specified intervals to reduce redundant calculations. -This caching mechanism delivers strong results with minimal additional memory overhead. For detailed performance analysis, see [our findings here](https://github.com/huggingface/diffusers/pull/12648#issuecomment-3610615080). +## TaylorSeer Cache -To enable TaylorSeer Cache, create a [`TaylorSeerCacheConfig`] and pass it to your pipeline's transformer: +[TaylorSeer Cache](https://huggingface.co/papers/2503.06923) accelerates diffusion inference with Taylor series expansions across denoising steps. It predicts later-step activations from earlier ones and reuses those predictions for several steps so the transformer does less full work. - `cache_interval`: Number of steps to reuse cached outputs before performing a full forward pass - `disable_cache_before_step`: Initial steps that use full computations to gather data for approximations -- `max_order`: Approximation accuracy (in theory, higher values improve quality but increase memory usage but we recommend it should be set to `1`) +- `max_order`: Higher Taylor orders can be more accurate but use more memory. Keep this at `1` unless you have a reason to change it. ```python import torch @@ -159,18 +141,19 @@ pipe.transformer.enable_cache(config) ## MagCache -[MagCache](https://github.com/Zehong-Ma/MagCache) accelerates inference by skipping transformer blocks based on the magnitude of the residual update. It observes that the magnitude of updates (Output - Input) decays predictably over the diffusion process. By accumulating an "error budget" based on pre-computed magnitude ratios, it dynamically decides when to skip computation and reuse the previous residual. +[MagCache](https://github.com/Zehong-Ma/MagCache) skips transformer blocks from the residual update magnitude. Update magnitudes decay predictably over denoising, and MagCache tracks an error budget from precomputed magnitude ratios (`mag_ratios`) to decide when reuse is safe. Those ratios are checkpoint and scheduler-specific. Ratios from a high step count can be interpolated down to fewer steps. -MagCache relies on **Magnitude Ratios** (`mag_ratios`), which describe this decay curve. These ratios are specific to the model checkpoint and scheduler. +MagCache follows two steps: -To use MagCache, you typically follow a two-step process: **Calibration** and **Inference**. +1. Calibration: Run inference once with `calibrate=True`. The hook measures residual magnitudes and prints the calculated ratios. +2. Inference: Disable the calibration cache, then pass those ratios to `MagCacheConfig` for acceleration. -1. **Calibration**: Run inference once with `calibrate=True`. The hook will measure the residual magnitudes and print the calculated ratios to the console. -2. **Inference**: Pass these ratios to `MagCacheConfig` to enable acceleration. +Classifier-free guidance may affect calibration. Pipelines that use true CFG with sequential contexts, such as Flux when `true_cfg_scale > 1`, enter `cache_context("cond")` and `cache_context("uncond")` separately. Calibration may print one array per context, but you should use the conditional array in most cases. Pipelines that batch CFG by concatenating conditional and unconditional inputs (for example, CogVideoX) produce a single joint array you can use directly. ```python import torch from diffusers import FluxPipeline, MagCacheConfig +from diffusers.hooks.mag_cache import FLUX_MAG_RATIOS pipe = FluxPipeline.from_pretrained( "black-forest-labs/FLUX.1-schnell", @@ -184,27 +167,64 @@ pipe.transformer.enable_cache(calib_config) # Run a prompt to trigger calibration pipe("A cat playing chess", num_inference_steps=4) -# Logs will print something like: "MagCache Calibration Results: [1.0, 1.37, 0.97, 0.87]" +# Prints: [MagCache] Calibration Complete. Copy these values to MagCacheConfig(mag_ratios=...): # 2. Inference Step -# Apply the specific ratios obtained from calibration for optimized speed. -# Note: For Flux models, you can also import defaults: -# from diffusers.hooks.mag_cache import FLUX_MAG_RATIOS +# Disable calibration hooks before enabling MagCache for inference. +pipe.transformer.disable_cache() + +# Apply ratios from calibration, or use the Flux defaults: +# mag_ratios=FLUX_MAG_RATIOS mag_config = MagCacheConfig( mag_ratios=[1.0, 1.37, 0.97, 0.87], num_inference_steps=4 ) -pipe.transformer.enable_cache(mag_config) +pipe.transformer.enable_cache(mag_config) image = pipe("A cat playing chess", num_inference_steps=4).images[0] ``` -> [!NOTE] -> `mag_ratios` represent the model's intrinsic magnitude decay curve. Ratios calibrated for a high number of steps (e.g., 50) can be reused for lower step counts (e.g., 20). The implementation uses interpolation to map the curve to the current number of inference steps. +## Text KV Cache + +[`TextKVCacheConfig`] computes the text key and value projections once and reuses them at every denoising step instead of recomputing them. [`apply_text_kv_cache`] only currently hooks `NucleusMoEImageTransformerBlock`. Enable it with `enable_cache`. + +```python +import torch +from diffusers import NucleusMoEImagePipeline, TextKVCacheConfig + +pipe = NucleusMoEImagePipeline.from_pretrained( + "NucleusAI/NucleusMoE-Image", dtype=torch.bfloat16 +) +pipe.to("cuda") # or "mps", "xpu", "cpu" + +pipe.transformer.enable_cache(TextKVCacheConfig()) + +image = pipe("A cat holding a sign that says hello world", num_inference_steps=50).images[0] +``` + +## Add caching to a new model + +Cache methods attach to a model through hooks, and the hooks depend on the model class, its transformer blocks, and the pipeline's denoising loop. Set up all three before calling `enable_cache` on a new model. + +1. Inherit from [`CacheMixin`] in the transformer class. It adds `enable_cache`, `disable_cache`, and `cache_context`. +2. Register the output layout of the transformer block in `_register_transformer_blocks_metadata` in `hooks/_helpers.py`. Methods that skip or reuse whole blocks, such as FirstBlockCache, MagCache, and SeaCache, use it to find the hidden states in the block's outputs. + + ```py + TransformerBlockRegistry.register( + model_class=MyTransformerBlock, + metadata=TransformerBlockMetadata( + return_hidden_states_index=0, + return_encoder_hidden_states_index=None, + ), + ) + ``` + +3. Wrap each transformer call in the pipeline's denoising loop with `cache_context`. Use separate names, such as `"cond"` and `"uncond"`, when classifier-free guidance runs the conditional and unconditional branches as separate calls, so each branch keeps its own cache state. -> [!TIP] -> For pipelines that run Classifier-Free Guidance sequentially (like Kandinsky 5.0), the calibration log might print two arrays: one for the Conditional pass and one for the Unconditional pass. In most cases, you should use the first array (Conditional). + ```py + with self.transformer.cache_context("cond"): + noise_pred = self.transformer(hidden_states=latents, timestep=timestep, ...)[0] + ``` -> [!TIP] -> For pipelines that run Classifier-Free Guidance in a **batched** manner (like SDXL or Flux), the `hidden_states` processed by the model contain both conditional and unconditional branches concatenated together. The calibration process automatically accounts for this, producing a single array of ratios that represents the joint behavior. You can use this resulting array directly without modification. +PAB and FasterCache also need `current_timestep_callback` so the hooks can read the current timestep. Check the output quality on the new model, because cache settings tuned for one model often don't transfer to another. diff --git a/docs/source/en/optimization/fp16.md b/docs/source/en/optimization/fp16.md index d3a73aac8595..6204b5d556c4 100644 --- a/docs/source/en/optimization/fp16.md +++ b/docs/source/en/optimization/fp16.md @@ -10,22 +10,18 @@ an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express o specific language governing permissions and limitations under the License. --> -# Accelerate inference +# Precision and compilation -Diffusion models are slow at inference because generation is an iterative process where noise is gradually refined into an image or video over a certain number of "steps". To speedup this process, you can try experimenting with different [schedulers](../api/schedulers/overview), reduce the precision of the model weights for faster computations, use more memory-efficient attention mechanisms, and more. - -Combine and use these techniques together to make inference faster than using any single technique on its own. - -This guide will go over how to accelerate inference. +Lower precision and compilation are two ways to speed up Diffusers inference. Load weights in `bfloat16` or `float16`, then compile the denoiser (the UNet or the transformer) with `torch.compile` or regional compilation. ## Model data type -The precision and data type of the model weights affect inference speed because a higher precision requires more memory to load and more time to perform the computations. PyTorch loads model weights in float32 or full precision by default, so changing the data type is a simple way to quickly get faster inference. +The precision and data type of the model weights affect inference speed because a higher precision requires more memory to load and more time to perform the computations. Diffusers loads model weights in `float32` when you omit `dtype`, so changing the data type is a simple way to quickly get faster inference. -bfloat16 is similar to float16 but it is more robust to numerical errors. Hardware support for bfloat16 varies, but most modern GPUs are capable of supporting bfloat16. +`bfloat16` is similar to `float16` but it is more robust to numerical errors. Hardware support for `bfloat16` varies, but most modern GPUs are capable of supporting `bfloat16`. ```py import torch @@ -42,7 +38,7 @@ pipeline(prompt, num_inference_steps=30).images[0] -float16 is similar to bfloat16 but may be more prone to numerical errors. +`float16` is similar to `bfloat16` but may be more prone to numerical errors. ```py import torch @@ -59,9 +55,9 @@ pipeline(prompt, num_inference_steps=30).images[0] -[TensorFloat-32 (tf32)](https://blogs.nvidia.com/blog/2020/05/14/tensorfloat-32-precision-format/) mode is supported on NVIDIA Ampere GPUs and it computes the convolution and matrix multiplication operations in tf32. Storage and other operations are kept in float32. This enables significantly faster computations when combined with bfloat16 or float16. +[TensorFloat-32](https://blogs.nvidia.com/blog/2020/05/14/tensorfloat-32-precision-format/) (`tf32`) mode is supported on NVIDIA Ampere and newer GPUs and it computes the convolution and matrix multiplication operations in `tf32`. Storage and other operations are kept in `float32`. It speeds up operations that run in `float32`, so it helps most when the pipeline or parts of it stay in `float32`. -PyTorch only enables tf32 mode for convolutions by default and you'll need to explicitly enable it for matrix multiplications. +PyTorch only enables `tf32` mode for convolutions by default and you'll need to explicitly enable it for matrix multiplications. ```py import torch @@ -70,7 +66,7 @@ from diffusers import StableDiffusionXLPipeline torch.backends.cuda.matmul.allow_tf32 = True pipeline = StableDiffusionXLPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.bfloat16 + "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float32 ).to("cuda") # or "mps", "xpu", "cpu" prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" @@ -84,31 +80,11 @@ Refer to the [mixed precision training](https://huggingface.co/docs/transformers ## Scaled dot product attention -> [!TIP] -> Memory-efficient attention optimizes for inference speed *and* [memory usage](./memory#memory-efficient-attention)! - -[Scaled dot product attention (SDPA)](https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html) implements several attention backends, [FlashAttention](https://github.com/Dao-AILab/flash-attention), [xFormers](https://github.com/facebookresearch/xformers), and a native C++ implementation. It automatically selects the most optimal backend for your hardware. - -SDPA is enabled by default if you're using PyTorch >= 2.0 and no additional changes are required to your code. You could try experimenting with other attention backends though if you'd like to choose your own. The example below uses the [torch.nn.attention.sdpa_kernel](https://pytorch.org/docs/stable/generated/torch.nn.attention.sdpa_kernel.html) context manager to enable efficient attention. - -```py -from torch.nn.attention import SDPBackend, sdpa_kernel -import torch -from diffusers import StableDiffusionXLPipeline - -pipeline = StableDiffusionXLPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.bfloat16 -).to("cuda") # or "mps", "xpu", "cpu" - -prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" - -with sdpa_kernel(SDPBackend.EFFICIENT_ATTENTION): - image = pipeline(prompt, num_inference_steps=30).images[0] -``` +[Scaled dot product attention (SDPA)](https://pytorch.org/docs/stable/generated/torch.nn.functional.scaled_dot_product_attention.html) is the default attention on PyTorch 2.0 and later, through [`AttnProcessor2_0`] or the attention dispatcher's `native` backend. PyTorch picks an SDPA kernel for your hardware. For FlashAttention, SageAttention, xFormers, Hub kernels, and other backends, see [Attention backends](./attention_backends). ## torch.compile -[torch.compile](https://pytorch.org/tutorials/intermediate/torch_compile_tutorial.html) accelerates inference by compiling PyTorch code and operations into optimized kernels. Diffusers typically compiles the more compute-intensive models like the UNet, transformer, or VAE. +[torch.compile](https://pytorch.org/tutorials/intermediate/torch_compile_tutorial.html) accelerates inference by compiling PyTorch code and operations into optimized kernels. You typically compile the denoiser that dominates runtime, either `pipeline.unet` or `pipeline.transformer`, and sometimes the VAE as well. Enable the following compiler settings for maximum speed (refer to the [full list](https://github.com/pytorch/pytorch/blob/main/torch/_inductor/config.py) for more options). @@ -122,22 +98,19 @@ torch._inductor.config.epilogue_fusion = False torch._inductor.config.coordinate_descent_check_all_directions = True ``` -Load and compile the UNet and VAE. There are several different modes you can choose from, but `"max-autotune"` optimizes for the fastest speed by compiling to a CUDA graph. CUDA graphs effectively reduces the overhead by launching multiple GPU operations through a single CPU operation. - -> [!TIP] -> With PyTorch 2.3.1, you can control the caching behavior of torch.compile. This is particularly beneficial for compilation modes like `"max-autotune"` which performs a grid-search over several compilation flags to find the optimal configuration. Learn more in the [Compile Time Caching in torch.compile](https://pytorch.org/tutorials/recipes/torch_compile_caching_tutorial.html) tutorial. +Load and compile the denoiser and VAE. Use `pipeline.unet` on UNet pipelines such as Stable Diffusion XL, or `pipeline.transformer` on Flux and DiT-style pipelines. There are several different modes you can choose from, but `"max-autotune"` searches for the fastest kernels and uses CUDA graphs. CUDA graphs reduce the overhead by launching multiple GPU operations through a single CPU operation. -Changing the memory layout to [channels_last](./memory#torchchannelslast) also optimizes memory and inference speed. +Changing the memory layout to [channels_last](./memory#torchchannelslast) can also speed up convolution-heavy models like UNets and VAEs. Benchmark it first, because some models run slower. ```py pipeline = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16 ).to("cuda") # or "mps", "xpu", "cpu" pipeline.unet.to(memory_format=torch.channels_last) -pipeline.vae.to(memory_format=torch.channels_last) pipeline.unet = torch.compile( pipeline.unet, mode="max-autotune", fullgraph=True ) +pipeline.vae.to(memory_format=torch.channels_last) pipeline.vae.decode = torch.compile( pipeline.vae.decode, mode="max-autotune", @@ -148,13 +121,10 @@ prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" pipeline(prompt, num_inference_steps=30).images[0] ``` -Compilation is slow the first time, but once compiled, it is significantly faster. Try to only use the compiled pipeline on the same type of inference operations. Calling the compiled pipeline on a different image size retriggers compilation which is slow and inefficient. +Compilation is slow the first time, especially with `"max-autotune"`, because the compiler searches for kernels and keeps using the same compiled pipeline object afterward (see [Compile Time Caching](https://docs.pytorch.org/tutorials/recipes/torch_compile_caching_tutorial.html) and [Caching Configuration](https://docs.pytorch.org/tutorials/recipes/torch_compile_caching_configuration_tutorial.html) for how to configure the caching behavior). Calling the compiled pipeline on a different image size triggers recompilation. ### Dynamic shape compilation -> [!TIP] -> Make sure to always use the nightly version of PyTorch for better support. - `torch.compile` keeps track of input shapes and conditions, and if these are different, it recompiles the model. For example, if a model is compiled on a 1024x1024 resolution image and used on an image with a different resolution, it triggers recompilation. To avoid recompilation, add `dynamic=True` to try and generate a more dynamic kernel to avoid recompilation when conditions change. @@ -175,12 +145,11 @@ Feel free to open an issue if dynamic compilation doesn't work as expected for a ### Regional compilation [Regional compilation](https://docs.pytorch.org/tutorials/recipes/regional_compilation.html) trims cold-start latency by only compiling the *small and frequently-repeated block(s)* of a model - typically a transformer layer - and enables reusing compiled artifacts for every subsequent occurrence. -For many diffusion architectures, this delivers the same runtime speedups as full-graph compilation and reduces compile time by 8–10x. +For many diffusion architectures, this delivers comparable runtime speedups to full-graph compilation and can cut compile time by up to 8–10x. -Use the [`~ModelMixin.compile_repeated_blocks`] method, a helper that wraps `torch.compile`, on any component such as the transformer model as shown below. +Use the [`~ModelMixin.compile_repeated_blocks`] method, a helper that wraps `torch.compile`, on the denoiser. Call it on `pipeline.unet` or `pipeline.transformer`, depending on the pipeline. ```py -# pip install -U diffusers import torch from diffusers import StableDiffusionXLPipeline @@ -188,25 +157,22 @@ pipeline = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16, ).to("cuda") # or "mps", "xpu", "cpu" - -# compile only the repeated transformer layers inside the UNet pipeline.unet.compile_repeated_blocks(fullgraph=True) ``` -To enable regional compilation for a new model, add a `_repeated_blocks` attribute to a model class containing the class names (as strings) of the blocks you want to compile. +To enable regional compilation for a new model, set `_repeated_blocks` to the class names of the repeated blocks (strings). For SDXL's UNet, use `BasicTransformerBlock`. Transformer denoisers list their own repeated block class names the same way. ```py -class MyUNet(ModelMixin): - _repeated_blocks = ("Transformer2DModel",) # ← compiled by default +_repeated_blocks = ["BasicTransformerBlock"] ``` -> [!TIP] -> For more regional compilation examples, see the reference [PR](https://github.com/huggingface/diffusers/pull/11705). - There is also a [compile_regions](https://github.com/huggingface/accelerate/blob/273799c85d849a1954a4f2e65767216eb37fa089/src/accelerate/utils/other.py#L78) method in [Accelerate](https://huggingface.co/docs/accelerate/index) that automatically selects candidate blocks in a model to compile. The remaining graph is compiled separately. This is useful for quick experiments because there aren't as many options for you to set which blocks to compile or adjust compilation flags. +```bash +pip install -U accelerate +``` + ```py -# pip install -U accelerate import torch from diffusers import StableDiffusionXLPipeline from accelerate.utils import compile_regions @@ -221,26 +187,12 @@ pipeline.unet = compile_regions(pipeline.unet, mode="reduce-overhead", fullgraph ### Graph breaks -It is important to specify `fullgraph=True` in torch.compile to ensure there are no graph breaks in the underlying model. This allows you to take advantage of torch.compile without any performance degradation. For the UNet and VAE, this changes how you access the return variables. - -```diff -- latents = unet( -- latents, timestep=timestep, encoder_hidden_states=prompt_embeds --).sample - -+ latents = unet( -+ latents, timestep=timestep, encoder_hidden_states=prompt_embeds, return_dict=False -+)[0] -``` +Set `fullgraph=True` so torch.compile raises an error on a graph break instead of silently splitting the graph, which reduces the speedup. ### GPU sync -The `step()` function is [called](https://github.com/huggingface/diffusers/blob/1d686bac8146037e97f3fd8c56e4063230f71751/src/diffusers/pipelines/stable_diffusion_xl/pipeline_stable_diffusion_xl.py#L1228) on the scheduler each time after the denoiser makes a prediction, and the `sigmas` variable is [indexed](https://github.com/huggingface/diffusers/blob/1d686bac8146037e97f3fd8c56e4063230f71751/src/diffusers/schedulers/scheduling_euler_discrete.py#L476). When placed on the GPU, it introduces latency because of the communication sync between the CPU and GPU. It becomes more evident when the denoiser has already been compiled. - -In general, the `sigmas` should [stay on the CPU](https://github.com/huggingface/diffusers/blob/35a969d297cba69110d175ee79c59312b9f49e1e/src/diffusers/schedulers/scheduling_euler_discrete.py#L240) to avoid the communication sync and latency. +After each denoiser prediction, the pipeline calls the scheduler's `step` method, which indexes `sigmas`. Keeping `sigmas` on the GPU can force a CPU↔GPU sync on every step. That cost is easy to miss until the denoiser is compiled. Prefer leaving `sigmas` on the CPU (Diffusers schedulers such as Euler already do this). -> [!TIP] -> Refer to the [torch.compile and Diffusers: A Hands-On Guide to Peak Performance](https://pytorch.org/blog/torch-compile-and-diffusers-a-hands-on-guide-to-peak-performance/) blog post for maximizing performance with `torch.compile` for diffusion models. ### Benchmarks @@ -250,12 +202,9 @@ The [diffusers-torchao](https://github.com/sayakpaul/diffusers-torchao#benchmark ## Kernels -[Kernels](https://huggingface.co/docs/kernels/index) is a library for building, distributing, and loading optimized compute kernels on the [Hub](https://huggingface.co/kernels-community). It supports [attention](./attention_backends#setattentionbackend) kernels and custom CUDA kernels for operations like RMSNorm, GEGLU, RoPE, and AdaLN. +[Kernels](https://huggingface.co/docs/kernels/index) is a library for building, distributing, and loading optimized compute kernels on the [Hub](https://huggingface.co/kernels-community). It supports [attention](./attention_backends#set-a-backend-on-the-model) kernels and custom CUDA kernels for operations like RMSNorm, GEGLU, RoPE, and AdaLN. -The [Diffusers Pipeline Integration](https://github.com/huggingface/kernels/blob/main/skills/cuda-kernels/references/diffusers-integration.md) guide shows how to integrate a kernel with the [add cuda-kernels](https://github.com/huggingface/kernels/blob/main/skills/cuda-kernels/SKILL.md) skill. This skill enables an agent, like Claude or Codex, to write custom kernels targeted towards a specific model and your hardware. - -> [!TIP] -> Install the [add cuda-kernels](https://github.com/huggingface/kernels/blob/main/skills/cuda-kernels/SKILL.md) skill to teach an agent how to write a kernel. The [Custom kernels for all from Codex and Claude](https://huggingface.co/blog/custom-cuda-kernels-agent-skills) blog post covers this in more detail. +The [Diffusers Pipeline Integration](https://github.com/huggingface/kernels/blob/main/skills/cuda-kernels/references/diffusers-integration.md) guide shows how to integrate a kernel with the [add cuda-kernels](https://github.com/huggingface/kernels/blob/main/skills/cuda-kernels/SKILL.md) skill. This skill enables an agent, like Claude or Codex, to write custom kernels targeted towards a specific model and your hardware. The [Custom kernels for all from Codex and Claude](https://huggingface.co/blog/custom-cuda-kernels-agent-skills) post has more detail. For example, a custom RMSNorm kernel (generated by the `add cuda-kernels` skill) with [torch.compile](#torchcompile) speeds up LTX-Video generation 1.43x on an H100. @@ -268,46 +217,12 @@ For example, a custom RMSNorm kernel (generated by the `add cuda-kernels` skill) ## Dynamic quantization -[Dynamic quantization](https://pytorch.org/tutorials/recipes/recipes/dynamic_quantization.html) improves inference speed by reducing precision to enable faster math operations. This particular type of quantization determines how to scale the activations based on the data at runtime rather than using a fixed scaling factor. As a result, the scaling factor is more accurately aligned with the data. - -The example below applies [dynamic int8 quantization](https://pytorch.org/tutorials/recipes/recipes/dynamic_quantization.html) to the UNet and VAE with the [torchao](../quantization/torchao) library. - -> [!TIP] -> Refer to our [torchao](../quantization/torchao) docs to learn more about how to use the Diffusers torchao integration. - -Configure the compiler tags for maximum speed. - -```py -import torch -from torchao import apply_dynamic_quant -from diffusers import StableDiffusionXLPipeline - -torch._inductor.config.conv_1x1_as_mm = True -torch._inductor.config.coordinate_descent_tuning = True -torch._inductor.config.epilogue_fusion = False -torch._inductor.config.coordinate_descent_check_all_directions = True -torch._inductor.config.force_fuse_int_mm_with_mul = True -torch._inductor.config.use_mixed_mm = True -``` - -Filter out some linear layers in the UNet and VAE which don't benefit from dynamic quantization with the [dynamic_quant_filter_fn](https://github.com/huggingface/diffusion-fast/blob/0f169640b1db106fe6a479f78c1ed3bfaeba3386/utils/pipeline_utils.py#L16). - -```py -pipeline = StableDiffusionXLPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.bfloat16 -).to("cuda") # or "mps", "xpu", "cpu" - -apply_dynamic_quant(pipeline.unet, dynamic_quant_filter_fn) -apply_dynamic_quant(pipeline.vae, dynamic_quant_filter_fn) - -prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" -pipeline(prompt, num_inference_steps=30).images[0] -``` +Dynamic int8 activation quantization is available in [torchao](../quantization/torchao). Use [`TorchAoConfig`] with a config such as [`Int8DynamicActivationInt8WeightConfig`](https://docs.pytorch.org/ao/stable/api_reference/generated/torchao.quantization.Int8DynamicActivationInt8WeightConfig.html), or torchao's `quantize_` API. Do not use the older `apply_dynamic_quant` helper. ## Fused projection matrices > [!WARNING] -> The [fuse_qkv_projections](https://github.com/huggingface/diffusers/blob/58431f102cf39c3c8a569f32d71b2ea8caa461e1/src/diffusers/pipelines/pipeline_utils.py#L2034) method is experimental and support is limited to mostly Stable Diffusion pipelines. Take a look at this [PR](https://github.com/huggingface/diffusers/pull/6179) to learn more about how to enable it for other pipelines +> [`~DiffusionPipeline.fuse_qkv_projections`] is experimental. An input is projected into three subspaces, represented by the projection matrices Q, K, and V, in an attention block. These projections are typically calculated separately, but you can horizontally combine these into a single matrix and perform the projection in a single step. It increases the size of the matrix multiplications of the input projections and also improves the impact of quantization. @@ -315,10 +230,10 @@ An input is projected into three subspaces, represented by the projection matric pipeline.fuse_qkv_projections() ``` -## Resources +## Next steps - Read the [Presenting Flux Fast: Making Flux go brrr on H100s](https://pytorch.org/blog/presenting-flux-fast-making-flux-go-brrr-on-h100s/) blog post to learn more about how you can combine all of these optimizations with [TorchInductor](https://docs.pytorch.org/docs/stable/torch.compiler.html) and [AOTInductor](https://docs.pytorch.org/docs/stable/torch.compiler_aot_inductor.html) for a ~2.5x speedup using recipes from [flux-fast](https://github.com/huggingface/flux-fast). These recipes support AMD hardware and [Flux.1 Kontext Dev](https://huggingface.co/black-forest-labs/FLUX.1-Kontext-dev). -- Read the [torch.compile and Diffusers: A Hands-On Guide to Peak Performance](https://pytorch.org/blog/torch-compile-and-diffusers-a-hands-on-guide-to-peak-performance/) blog post -to maximize performance when using `torch.compile`. + +- Refer to the [torch.compile and Diffusers: A Hands-On Guide to Peak Performance](https://pytorch.org/blog/torch-compile-and-diffusers-a-hands-on-guide-to-peak-performance/) blog post for maximizing performance with `torch.compile` for diffusion models. diff --git a/docs/source/en/optimization/speed-memory-optims.md b/docs/source/en/optimization/speed-memory-optims.md index 21dd5b6ea1db..f22b28527c32 100644 --- a/docs/source/en/optimization/speed-memory-optims.md +++ b/docs/source/en/optimization/speed-memory-optims.md @@ -10,7 +10,7 @@ an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express o specific language governing permissions and limitations under the License. --> -# Compiling and offloading quantized models +# Quantize, compile, and offload Optimizing models often involves trade-offs between [inference speed](./fp16) and [memory-usage](./memory). For instance, while [caching](./cache) can boost inference speed, it also increases memory consumption since it needs to store the outputs of intermediate attention layers. A more balanced optimization strategy combines quantizing a model, [torch.compile](./fp16#torchcompile) and various [offloading methods](./memory#offloading). diff --git a/docs/source/en/stable_diffusion.md b/docs/source/en/stable_diffusion.md index fd368f5b1654..e04a40684d55 100644 --- a/docs/source/en/stable_diffusion.md +++ b/docs/source/en/stable_diffusion.md @@ -10,123 +10,80 @@ an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express o specific language governing permissions and limitations under the License. --> -[[open-in-colab]] +# Overview -# Basic performance +Diffusion inference is computationally expensive, and you often run a [`DiffusionPipeline`] more than once before you like the result. This page provides an overview of the main Diffusers optimization techniques, what they do, and when to use them. -Diffusion is a random process that is computationally demanding. You may need to run the [`DiffusionPipeline`] several times before getting a desired output. That's why it's important to carefully balance generation speed and memory usage in order to iterate faster, +## Starter path -This guide recommends some basic performance tips for using the [`DiffusionPipeline`]. Refer to the Inference Optimization section docs such as [Accelerate inference](./optimization/fp16) or [Reduce memory usage](./optimization/memory) for more detailed performance guides. +When the model fits on one GPU, start with this baseline load. Set `dtype` and place the pipeline on an accelerator. Reach for model CPU offload only when memory is tight. You could also speed up inference with fewer steps or a faster scheduler. -## Memory usage - -Reducing the amount of memory used indirectly speeds up generation and can help a model fit on device. - -The [`~DiffusionPipeline.enable_model_cpu_offload`] method moves a model to the CPU when it is not in use to save GPU memory. +When you omit `dtype`, Diffusers loads components in `float32`. Pass `dtype=torch.bfloat16` (or `torch.float16` if bfloat16 is unsupported), then place the pipeline on an accelerator with `pipeline.to("cuda")`. ```py import torch from diffusers import DiffusionPipeline pipeline = DiffusionPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", - dtype=torch.bfloat16, - device_map="cuda" # or "mps", "xpu", "cpu" + "stabilityai/stable-diffusion-xl-base-1.0", + dtype=torch.bfloat16, ) -pipeline.enable_model_cpu_offload() +pipeline.to("cuda") # or "mps", "xpu" prompt = """ cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain """ pipeline(prompt).images[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") ``` -## Inference speed +If the pipeline does not fit, or memory is tight, call [`~DiffusionPipeline.enable_model_cpu_offload`] instead of keeping everything on the GPU. It places the active model on the GPU and keeps the other components on the CPU. -Denoising is the most computationally demanding process during diffusion. Methods that optimizes this process accelerates inference speed. Try the following methods for a speed up. +Skip it when the model fits. Offloading is slower when you do not need it. -- Add `device_map="cuda"` to place the pipeline on a GPU. Placing a model on an accelerator, like a GPU, increases speed because it performs computations in parallel. -- Set `dtype=torch.bfloat16` to execute the pipeline in half-precision. Reducing the data type precision increases speed because it takes less time to perform computations in a lower precision. +For more offloading options, see [Reduce memory usage](./optimization/memory#offloading). ```py -import torch -import time -from diffusers import DiffusionPipeline, DPMSolverMultistepScheduler - pipeline = DiffusionPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", - dtype=torch.bfloat16, - device_map="cuda + "stabilityai/stable-diffusion-xl-base-1.0", + dtype=torch.bfloat16, ) +pipeline.enable_model_cpu_offload() ``` -- Use a faster scheduler, such as [`DPMSolverMultistepScheduler`], which only requires ~20-25 steps. -- Set `num_inference_steps` to a lower value. Reducing the number of inference steps reduces the overall number of computations. However, this can result in lower generation quality. +Lower latency with fewer `num_inference_steps` or a faster scheduler such as [`DPMSolverMultistepScheduler`]. That usually speeds up generation but can reduce image quality versus a slower, higher-quality scheduler. See [Optimization techniques](#optimization-techniques) below for more speed techniques. ```py -pipeline.scheduler = DPMSolverMultistepScheduler.from_config(pipeline.scheduler.config) +import time +from diffusers import DPMSolverMultistepScheduler -prompt = """ -cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California -highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain -""" +pipeline.scheduler = DPMSolverMultistepScheduler.from_config(pipeline.scheduler.config) start_time = time.perf_counter() -image = pipeline(prompt).images[0] +image = pipeline(prompt, num_inference_steps=25).images[0] end_time = time.perf_counter() print(f"Image generation took {end_time - start_time:.3f} seconds") ``` -## Generation quality - -Many modern diffusion models deliver high-quality images out-of-the-box. However, you can still improve generation quality by trying the following. - -- Try a more detailed and descriptive prompt. Include details such as the image medium, subject, style, and aesthetic. A negative prompt may also help by guiding a model away from undesirable features by using words like low quality or blurry. - - ```py - import torch - from diffusers import DiffusionPipeline - - pipeline = DiffusionPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", - dtype=torch.bfloat16, - device_map="cuda" # or "mps", "xpu", "cpu" - ) - - prompt = """ - cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California - highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain - """ - negative_prompt = "low quality, blurry, ugly, poor details" - pipeline(prompt, negative_prompt=negative_prompt).images[0] - ``` +## Optimization techniques - For more details about creating better prompts, take a look at the [Prompt techniques](./using-diffusers/weighted_prompts) doc. +When the starter path is not enough, use these techniques. If inference is too slow, start with `torch.compile`, caching, or attention backends. If you are out of memory, start with offloading or quantization. -- Try a different scheduler, like [`HeunDiscreteScheduler`] or [`LMSDiscreteScheduler`], that gives up generation speed for quality. +Faster inference: - ```py - import torch - from diffusers import DiffusionPipeline, HeunDiscreteScheduler +- [torch.compile](./optimization/fp16#torchcompile) — Compile the UNet, transformer, or VAE into optimized kernels. +- [Regional compilation](./optimization/fp16#regional-compilation) — Compile repeated blocks to cut `torch.compile` cold-start latency and reuse compiled artifacts. +- [Caching](./optimization/cache) — Reuse intermediates across denoising steps when you want more speed and can spend memory. +- [Attention backends](./optimization/attention_backends) — Swap Diffusers attention implementations through a unified API when attention is the bottleneck. +- [Kernels](./optimization/fp16#kernels) — Load optimized Hub compute kernels (attention and custom CUDA ops such as RMSNorm or RoPE) when you need hardware-specific speedups beyond stock PyTorch. - pipeline = DiffusionPipeline.from_pretrained( - "stabilityai/stable-diffusion-xl-base-1.0", - dtype=torch.bfloat16, - device_map="cuda" # or "mps", "xpu", "cpu" - ) - pipeline.scheduler = HeunDiscreteScheduler.from_config(pipeline.scheduler.config) +Less memory: - prompt = """ - cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California - highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain - """ - negative_prompt = "low quality, blurry, ugly, poor details" - pipeline(prompt, negative_prompt=negative_prompt).images[0] - ``` +- [Offloading](./optimization/memory#offloading) — Move inactive models or layers to the CPU with CPU, model, or group offloading. +- [Quantization](./quantization/overview) — Load smaller weights to cut memory (some backends also speed up inference). [GGUF](./quantization/gguf) is a common starting point. +- [VAE slicing](./optimization/memory#vae-slicing) and [VAE tiling](./optimization/memory#vae-tiling) — Decode large batches or high-resolution images in pieces to lower peak memory. -## Next steps +Both: -Diffusers offers more advanced and powerful optimizations such as [group-offloading](./optimization/memory#group-offloading) and [regional compilation](./optimization/fp16#regional-compilation). To learn more about how to maximize performance, take a look at the Inference Optimization section. \ No newline at end of file +- [Quantize, compile, and offload](./optimization/speed-memory-optims) — Combine quantization, `torch.compile`, and offloading when one technique is not enough. diff --git a/docs/source/en/using-diffusers/conditional_image_generation.md b/docs/source/en/using-diffusers/conditional_image_generation.md index fe1602567db9..aeee424b7106 100644 --- a/docs/source/en/using-diffusers/conditional_image_generation.md +++ b/docs/source/en/using-diffusers/conditional_image_generation.md @@ -304,4 +304,4 @@ pipeline = AutoPipelineForText2Image.from_pretrained("stable-diffusion-v1-5/stab pipeline.unet = torch.compile(pipeline.unet, mode="reduce-overhead", fullgraph=True) ``` -For more tips on how to optimize your code to save memory and speed up inference, read the [Accelerate inference](../optimization/fp16) and [Reduce memory usage](../optimization/memory) guides. +For more tips on how to optimize your code to save memory and speed up inference, see [Optimize and scale](../stable_diffusion#optimization-techniques). diff --git a/docs/source/en/using-diffusers/img2img.md b/docs/source/en/using-diffusers/img2img.md index e689f6d9fce6..461e94ad7815 100644 --- a/docs/source/en/using-diffusers/img2img.md +++ b/docs/source/en/using-diffusers/img2img.md @@ -593,4 +593,4 @@ With [`torch.compile`](../optimization/fp16#torchcompile), you can boost your in pipeline.unet = torch.compile(pipeline.unet, mode="reduce-overhead", fullgraph=True) ``` -To learn more, take a look at the [Reduce memory usage](../optimization/memory) and [Accelerate inference](../optimization/fp16) guides. +To learn more, take a look at [Optimize and scale](../stable_diffusion#optimization-techniques). diff --git a/docs/source/en/using-diffusers/inpaint.md b/docs/source/en/using-diffusers/inpaint.md index 6d7d174e2bd1..ee1c4161be2c 100644 --- a/docs/source/en/using-diffusers/inpaint.md +++ b/docs/source/en/using-diffusers/inpaint.md @@ -797,4 +797,4 @@ To speed-up your inference code even more, use [`torch_compile`](../optimization pipeline.unet = torch.compile(pipeline.unet, mode="reduce-overhead", fullgraph=True) ``` -Learn more in the [Reduce memory usage](../optimization/memory) and [Accelerate inference](../optimization/fp16) guides. +Learn more in [Optimize and scale](../stable_diffusion#optimization-techniques).