From 7a201122e12e31e3a5fa56565a8ac13d5456c1ce Mon Sep 17 00:00:00 2001 From: stevhliu Date: Fri, 25 Sep 2026 15:38:36 -0700 Subject: [PATCH 1/3] memory --- docs/source/en/optimization/memory.md | 214 +++++++++++++++----------- 1 file changed, 120 insertions(+), 94 deletions(-) diff --git a/docs/source/en/optimization/memory.md b/docs/source/en/optimization/memory.md index a68628946015..37d38d44997a 100644 --- a/docs/source/en/optimization/memory.md +++ b/docs/source/en/optimization/memory.md @@ -12,16 +12,24 @@ specific language governing permissions and limitations under the License. # Reduce memory usage -Modern diffusion models like [Flux](../api/pipelines/flux) and [Wan](../api/pipelines/wan) have billions of parameters that take up a lot of memory on your hardware for inference. This is challenging because common GPUs often don't have sufficient memory. To overcome the memory limitations, you can use more than one GPU (if available), offload some of the pipeline components to the CPU, and more. +Modern diffusion models have billions of parameters, which often exceeds the memory available on a consumer GPU. Diffusers provides several techniques to reduce memory usage, such as distributing a model across multiple GPUs, offloading components to the CPU, and storing weights in lower precision. -This guide will show you how to reduce your memory usage. +Choose a technique based on your hardware and workload. + +| Situation | Technique | +|---|---| +| You have more than one GPU | [Multiple GPUs](#multiple-gpus) | +| You generate several images per prompt | [VAE slicing](#vae-slicing) | +| You generate high-resolution images | [VAE tiling](#vae-tiling) | +| The model doesn't fit on a single GPU | [Offloading](#offloading) | +| You want to store weights in lower precision | [Layerwise casting](#layerwise-casting) | > [!TIP] -> Keep in mind these techniques may need to be adjusted depending on the model. For example, a transformer-based diffusion model may not benefit equally from these memory optimizations as a UNet-based model. +> Results vary by model. For example, a transformer-based model may not benefit from some of these techniques as much as a UNet-based model. ## Multiple GPUs -If you have access to more than one GPU, there a few options for efficiently loading and distributing a large model across your hardware. These features are supported by the [Accelerate](https://huggingface.co/docs/accelerate/index) library, so make sure it is installed first. +If you have access to more than one GPU, there are a few options for efficiently loading and distributing a large model across your hardware. These features are supported by the [Accelerate](https://huggingface.co/docs/accelerate/index) library, so make sure it is installed. ```bash pip install -U accelerate @@ -29,9 +37,9 @@ pip install -U accelerate ### Sharded checkpoints -Loading large checkpoints in several shards in useful because the shards are loaded one at a time. This keeps memory usage low, only requiring enough memory for the model size and the largest shard size. We recommend sharding when the fp32 checkpoint is greater than 5GB. The default shard size is 5GB. +A sharded checkpoint splits a large checkpoint into several smaller files that are loaded one at a time. Peak memory only needs to fit the model and the largest shard, instead of the model and the entire checkpoint. Sharding is recommended when the fp32 checkpoint is larger than 5GB. The default shard size is 10GB. -Shard a checkpoint in [`~DiffusionPipeline.save_pretrained`] with the `max_shard_size` parameter. +Shard a checkpoint with the `max_shard_size` parameter in [`~ModelMixin.save_pretrained`]. ```py from diffusers import AutoModel @@ -42,7 +50,7 @@ unet = AutoModel.from_pretrained( unet.save_pretrained("sdxl-unet-sharded", max_shard_size="5GB") ``` -Now you can use the sharded checkpoint, instead of the regular checkpoint, to save memory. +Load the sharded checkpoint in place of the original to reduce peak memory while loading. ```py import torch @@ -60,10 +68,10 @@ pipeline = StableDiffusionXLPipeline.from_pretrained( ### Device placement -> [!WARNING] -> Device placement is an experimental feature and the API may change. Only the `balanced` strategy is supported at the moment. We plan to support additional mapping strategies in the future. +The `device_map` parameter controls how the model components in a pipeline or the layers in an individual model are distributed across devices. -The `device_map` parameter controls how the model components in a pipeline or the layers in an individual model are distributed across devices. +> [!WARNING] +> Device placement is an experimental feature and the API may change. At the pipeline level, `device_map` accepts `"balanced"`, or a single device such as `"cuda"` or `"cpu"`. @@ -72,7 +80,7 @@ The `balanced` device placement strategy evenly splits the pipeline across all a ```py import torch -from diffusers import AutoModel, StableDiffusionXLPipeline +from diffusers import StableDiffusionXLPipeline pipeline = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", @@ -85,13 +93,12 @@ You can inspect a pipeline's device map with `hf_device_map`. ```py print(pipeline.hf_device_map) -{'unet': 1, 'vae': 1, 'safety_checker': 0, 'text_encoder': 0} ``` -The `device_map` is useful for loading large models, such as the Flux diffusion transformer which has 12.5B parameters. Set it to `"auto"` to automatically distribute a model across the fastest device first before moving to slower devices. Refer to the [Sharded checkpoints](#sharded-checkpoints) section for more details. +Set `device_map="auto"` to distribute the layers of a large model across devices. The fastest device is filled first before moving to slower devices. ```py import torch @@ -114,9 +121,7 @@ print(transformer.hf_device_map) -When designing your own `device_map`, it should be a dictionary of a model's specific module name or layer and a device identifier (an integer for GPUs, `cpu` for CPUs, and `disk` for disk). - -Call `hf_device_map` on a model to see how model layers are distributed and then design your own. +A custom `device_map` is a dictionary that maps module names to devices, where a device is an integer for a GPU, `"cpu"`, or `"disk"`. Print a model's `hf_device_map` to see how its layers are distributed, and use it as a starting point to design your own. ```py print(transformer.hf_device_map) @@ -141,13 +146,13 @@ transformer = AutoModel.from_pretrained( ) ``` -Pass a dictionary mapping maximum memory usage to each device to enforce a limit. If a device is not in `max_memory`, it is ignored and pipeline components won't be distributed to it. +To cap how much memory each GPU uses, pass `max_memory`, a dictionary that maps each device to a limit. Components aren't placed on GPUs you leave out, and a component that doesn't fit under a GPU's limit is placed on the CPU instead. ```py import torch -from diffusers import AutoModel, StableDiffusionXLPipeline +from diffusers import StableDiffusionXLPipeline -max_memory = {0:"1GB", 1:"1GB"} +max_memory = {0: "16GB", 1: "16GB"} pipeline = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", dtype=torch.float16, @@ -156,12 +161,12 @@ pipeline = StableDiffusionXLPipeline.from_pretrained( ) ``` -Diffusers uses the maximum memory of all devices by default, but if they don't fit on the GPUs, then you'll need to use a single GPU and offload to the CPU with the methods below. +By default, Diffusers uses all available memory on each GPU. Components that don't fit on a GPU are placed on the CPU. If most of the pipeline ends up on the CPU, try a single GPU with one of the offloading methods below instead. -- [`~DiffusionPipeline.enable_model_cpu_offload`] only works on a single GPU but a very large model may not fit on it -- [`~DiffusionPipeline.enable_sequential_cpu_offload`] may work but it is extremely slow and also limited to a single GPU +- [`~DiffusionPipeline.enable_model_cpu_offload`] moves one whole model to the GPU at a time. It's faster, but each model must fit on a single GPU. +- [`~DiffusionPipeline.enable_sequential_cpu_offload`] moves one submodule to the GPU at a time. It uses the least memory, but it's very slow. -Use the [`~DiffusionPipeline.reset_device_map`] method to reset the `device_map`. This is necessary if you want to use methods like `.to()`, [`~DiffusionPipeline.enable_sequential_cpu_offload`], and [`~DiffusionPipeline.enable_model_cpu_offload`] on a pipeline that was device-mapped. +Before calling `.to()`, `enable_sequential_cpu_offload`, or `enable_model_cpu_offload` on a device-mapped pipeline, reset its device map with [`~DiffusionPipeline.reset_device_map`]. ```py pipeline.reset_device_map() @@ -169,15 +174,30 @@ pipeline.reset_device_map() ## VAE slicing -VAE slicing saves memory by splitting large batches of inputs into a single batch of data and separately processing them. This method works best when generating more than one image at a time. +VAE slicing splits a batch of latents into single latents and decodes them one at a time. The decoded images are concatenated back into a batch at the end. Peak memory stays close to the cost of decoding one image regardless of batch size, which makes slicing useful when generating several images at once. It has no effect on single-image batches. -For example, if you're generating 4 images at once, decoding would increase peak activation memory by 4x. VAE slicing reduces this by only decoding 1 image at a time instead of all 4 images at once. +```text +Without slicing: one decode for the whole batch -Call [`~AutoencoderKL.enable_slicing`] to enable sliced VAE. You can expect a small increase in performance when decoding multi-image batches and no performance impact for single-image batches. + [ z1 | z2 | z3 | z4 ] --> VAE decode --> [ img1 | img2 | img3 | img4 ] + (peak memory: 4 images) + +With slicing: decode one latent at a time, then concatenate + + [ z1 ] --> VAE decode --> [ img1 ] --+ + [ z2 ] --> VAE decode --> [ img2 ] --+ + [ z3 ] --> VAE decode --> [ img3 ] --+ + [ z4 ] --> VAE decode --> [ img4 ] --+ + (peak memory: 1 image) | + v + [ img1 | img2 | img3 | img4 ] +``` + +Call [`~AutoencoderKL.enable_slicing`] to enable VAE slicing. ```py import torch -from diffusers import AutoModel, StableDiffusionXLPipeline +from diffusers import StableDiffusionXLPipeline pipeline = StableDiffusionXLPipeline.from_pretrained( "stabilityai/stable-diffusion-xl-base-1.0", @@ -185,17 +205,35 @@ pipeline = StableDiffusionXLPipeline.from_pretrained( ).to("cuda") # or "mps", "xpu", "cpu" pipeline.vae.enable_slicing() pipeline(["An astronaut riding a horse on Mars"]*32).images[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") +print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") ``` > [!WARNING] -> The [`AutoencoderKLWan`] and [`AsymmetricAutoencoderKL`] classes don't support slicing. +> [`AsymmetricAutoencoderKL`] doesn't support slicing. ## VAE tiling -VAE tiling saves memory by dividing an image into smaller overlapping tiles instead of processing the entire image at once. This also reduces peak memory usage because the GPU is only processing a tile at a time. +VAE tiling splits a latent into overlapping tiles and decodes each tile separately. The overlapping edges are blended together to stitch the tiles into the final image. Peak memory depends on the tile size instead of the full image size, which makes tiling useful for generating high-resolution images. + +```text +1. Split the latent into tiles 2. Decode each tile 3. Blend the overlaps + that overlap by 25% separately and stitch + + +--------+----+--------+ + | tile 1 |####| tile 2 | tile 1 --> decode +-------------------+ + | |####| | tile 2 --> decode | | + +--------+----+--------+ --> tile 3 --> decode --> | full-resolution | + |########|####|########| tile 4 --> decode | image | + +--------+----+--------+ | | + | tile 3 |####| tile 4 | (peak memory: 1 tile) +-------------------+ + | |####| | + +--------+----+--------+ + #### = region shared by neighboring tiles +``` + +Tiles are decoded separately, so tone may vary slightly from tile to tile, but there shouldn't be any obvious seams. Tiling also applies when encoding an image into latents, and only activates when the input is larger than the VAE's `sample_size`. -Call [`~AutoencoderKL.enable_tiling`] to enable VAE tiling. The generated image may have some tone variation from tile-to-tile because they're decoded separately, but there shouldn't be any obvious seams between the tiles. Tiling is disabled for resolutions lower than a pre-specified (but configurable) limit. For example, this limit is 512x512 for the VAE in [`StableDiffusionPipeline`]. +Call [`~AutoencoderKL.enable_tiling`] to enable VAE tiling. ```py import torch @@ -210,26 +248,24 @@ pipeline.vae.enable_tiling() init_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/img2img-sdxl-init.png") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" pipeline(prompt, image=init_image, strength=0.5).images[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") +print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") ``` > [!WARNING] -> [`AutoencoderKLWan`] and [`AsymmetricAutoencoderKL`] don't support tiling. +> [`AsymmetricAutoencoderKL`] doesn't support tiling. ## Offloading -Offloading strategies move not currently active layers or models to the CPU to avoid increasing GPU memory. These strategies can be combined with quantization and torch.compile to balance inference speed and memory usage. - -Refer to the [Compile and offloading quantized models](./speed-memory-optims) guide for more details. +Offloading keeps inactive layers or models on the CPU and moves them to the GPU only when they're needed. You can combine offloading with quantization and torch.compile to balance inference speed and memory usage. -### CPU offloading +Refer to the [Compiling and offloading quantized models](./speed-memory-optims) guide for more details. -CPU offloading selectively moves weights from the GPU to the CPU. When a component is required, it is transferred to the GPU and when it isn't required, it is moved to the CPU. This method works on submodules rather than whole models. It saves memory by avoiding storing the entire model on the GPU. +### Sequential CPU offloading -CPU offloading dramatically reduces memory usage, but it is also **extremely slow** because submodules are passed back and forth multiple times between devices. It can often be impractical due to how slow it is. +Sequential CPU offloading keeps weights on the CPU and moves each submodule to the GPU only when it runs. The entire model is never on the GPU at once, so sequential offloading uses the least memory of the offloading methods. It's also the slowest because submodules are transferred between devices many times during inference, which often makes it impractical. > [!WARNING] -> Don't move the pipeline to CUDA before calling [`~DiffusionPipeline.enable_sequential_cpu_offload`], otherwise the amount of memory saved is only minimal (refer to this [issue](https://github.com/huggingface/diffusers/issues/1934) for more details). This is a stateful operation that installs hooks on the model. +> Don't move the pipeline to CUDA before calling `enable_sequential_cpu_offload`, otherwise the memory savings are minimal. Refer to [issue #1934](https://github.com/huggingface/diffusers/issues/1934) for more details. Sequential offloading is stateful and installs hooks on the model. Call [`~DiffusionPipeline.enable_sequential_cpu_offload`] to enable it on a pipeline. @@ -250,15 +286,15 @@ pipeline( num_inference_steps=4, max_sequence_length=256, ).images[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") +print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") ``` ### Model offloading -Model offloading moves entire models to the GPU instead of selectively moving *some* layers or model components. One of the main pipeline models, usually the text encoder, UNet, and VAE, is placed on the GPU while the other components are held on the CPU. Components like the UNet that run multiple times stays on the GPU until its completely finished and no longer needed. This eliminates the communication overhead of [CPU offloading](#cpu-offloading) and makes model offloading a faster alternative. The tradeoff is memory savings won't be as large. +Model offloading moves whole models to the GPU instead of individual submodules. Only one model, such as the text encoder, denoiser (UNet or transformer), or VAE, is on the GPU at a time while the other models stay on the CPU. A model that runs multiple times, like the denoiser, stays on the GPU until it finishes. Model offloading avoids the transfer overhead of [sequential CPU offloading](#sequential-cpu-offloading), which makes it faster, but the memory savings are smaller. > [!WARNING] -> Keep in mind that if models are reused outside the pipeline after hookes have been installed (see [Removing Hooks](https://huggingface.co/docs/accelerate/en/package_reference/big_modeling#accelerate.hooks.remove_hook_from_module) for more details), you need to run the entire pipeline and models in the expected order to properly offload them. This is a stateful operation that installs hooks on the model. +> Model offloading is stateful and installs hooks on each model. If you call a model outside the pipeline, run the models in the pipeline's order so they're offloaded correctly, or [remove the hooks](https://huggingface.co/docs/accelerate/en/package_reference/big_modeling#accelerate.hooks.remove_hook_from_module) first. Call [`~DiffusionPipeline.enable_model_cpu_offload`] to enable it on a pipeline. @@ -279,24 +315,24 @@ pipeline( num_inference_steps=4, max_sequence_length=256, ).images[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") +print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") ``` -[`~DiffusionPipeline.enable_model_cpu_offload`] also helps when you're using the [`~StableDiffusionXLPipeline.encode_prompt`] method on its own to generate the text encoders hidden state. +Model offloading also helps when you call [`~StableDiffusionXLPipeline.encode_prompt`] on its own, because only the text encoders are moved to the GPU. ### Group offloading -Group offloading moves groups of internal layers ([torch.nn.ModuleList](https://pytorch.org/docs/stable/generated/torch.nn.ModuleList.html) or [torch.nn.Sequential](https://pytorch.org/docs/stable/generated/torch.nn.Sequential.html)) to the CPU. It uses less memory than [model offloading](#model-offloading) and it is faster than [CPU offloading](#cpu-offloading) because it reduces communication overhead. +Group offloading moves groups of internal layers ([torch.nn.ModuleList](https://pytorch.org/docs/stable/generated/torch.nn.ModuleList.html) or [torch.nn.Sequential](https://pytorch.org/docs/stable/generated/torch.nn.Sequential.html)) to the CPU. It uses less memory than [model offloading](#model-offloading) and it is faster than [sequential CPU offloading](#sequential-cpu-offloading) because it reduces communication overhead. > [!WARNING] -> Group offloading may not work with all models if the forward implementation contains weight-dependent device casting of inputs because it may clash with group offloading's device casting mechanism. +> Group offloading may not work with models whose forward pass moves inputs to the weights' device, because that conflicts with how group offloading moves tensors. -Enable group offloading by configuring the `offload_type` parameter to `block_level` or `leaf_level`. +Set `offload_type` to `block_level` or `leaf_level` to choose how layers are grouped. -- `block_level` offloads groups of layers based on the `num_blocks_per_group` parameter. For example, if `num_blocks_per_group=2` on a model with 40 layers, 2 layers are onloaded and offloaded at a time (20 total onloads/offloads). This drastically reduces memory requirements. -- `leaf_level` offloads individual layers at the lowest level and is equivalent to [CPU offloading](#cpu-offloading). But it can be made faster if you use streams without giving up inference speed. +- `block_level` offloads groups of layers, and `num_blocks_per_group` sets the size of each group. For example, `num_blocks_per_group=2` on a model with 40 layers creates 20 groups and moves 2 layers to the GPU at a time. +- `leaf_level` offloads each individual layer, similar to [sequential CPU offloading](#sequential-cpu-offloading), but it can be much faster with [CUDA streams](#cuda-stream). -Group offloading is supported for entire pipelines or individual models. Applying group offloading to the entire pipeline is the easiest option while selectively applying it to individual models gives users more flexibility to use different offloading techniques for different models. +Group offloading is supported for entire pipelines or individual models. Apply it to the whole pipeline for the simplest setup, or to individual models to mix offloading techniques. @@ -306,7 +342,6 @@ Call [`~DiffusionPipeline.enable_group_offload`] on a pipeline. ```py import torch from diffusers import CogVideoXPipeline -from diffusers.hooks import apply_group_offloading from diffusers.utils import export_to_video onload_device = torch.device("cuda") @@ -320,16 +355,9 @@ pipeline.enable_group_offload( use_stream=True ) -prompt = ( - "A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. " - "The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other " - "pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, " - "casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. " - "The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical " - "atmosphere of this unique musical performance." -) +prompt = "A panda playing a tiny acoustic guitar in a bamboo forest" video = pipeline(prompt=prompt, guidance_scale=6, num_inference_steps=50).frames[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") +print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") export_to_video(video, "output.mp4", fps=8) ``` @@ -355,16 +383,9 @@ pipeline.vae.enable_group_offload(onload_device=onload_device, offload_type="lea # Use the apply_group_offloading method for other model components apply_group_offloading(pipeline.text_encoder, onload_device=onload_device, offload_type="block_level", num_blocks_per_group=2) -prompt = ( - "A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. " - "The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other " - "pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, " - "casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. " - "The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical " - "atmosphere of this unique musical performance." -) +prompt = "A panda playing a tiny acoustic guitar in a bamboo forest" video = pipeline(prompt=prompt, guidance_scale=6, num_inference_steps=50).frames[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") +print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") export_to_video(video, "output.mp4", fps=8) ``` @@ -373,24 +394,24 @@ export_to_video(video, "output.mp4", fps=8) #### CUDA stream -The `use_stream` parameter can be activated for CUDA devices that support asynchronous data transfer streams to reduce overall execution time compared to [CPU offloading](#cpu-offloading). It overlaps data transfer and computation by using layer prefetching. The next layer to be executed is loaded onto the GPU while the current layer is still being executed. It can increase CPU memory significantly so ensure you have 2x the amount of memory as the model size. +Set `use_stream=True` on CUDA devices to prefetch the next layer onto the GPU while the current layer is still running. Overlapping data transfer and computation makes group offloading much faster than [sequential CPU offloading](#sequential-cpu-offloading). Streams pin tensors in CPU memory, so make sure you have about twice the model size in system RAM. Set `record_stream=True` for more of a speedup at the cost of slightly increased memory usage. Refer to the [torch.Tensor.record_stream](https://pytorch.org/docs/stable/generated/torch.Tensor.record_stream.html) docs to learn more. > [!TIP] -> When `use_stream=True` on VAEs with tiling enabled, make sure to do a dummy forward pass (possible with dummy inputs as well) before inference to avoid device mismatch errors. This may not work on all implementations, so feel free to open an issue if you encounter any problems. +> If a VAE has tiling enabled and `use_stream=True`, run a forward pass with dummy inputs before inference to avoid device mismatch errors. Open an [issue](https://github.com/huggingface/diffusers/issues) if this doesn't work for your model. -If you're using `block_level` group offloading with `use_stream` enabled, the `num_blocks_per_group` parameter should be set to `1`, otherwise a warning will be raised. +Streams require `num_blocks_per_group=1` with `block_level` offloading. Other values log a warning and are reset to `1`. ```py pipeline.transformer.enable_group_offload(onload_device=onload_device, offload_device=offload_device, offload_type="leaf_level", use_stream=True, record_stream=True) ``` -The `low_cpu_mem_usage` parameter can be set to `True` to reduce CPU memory usage when using streams during group offloading. It is best for `leaf_level` offloading and when CPU memory is bottlenecked. Memory is saved by creating pinned tensors on the fly instead of pre-pinning them. However, this may increase overall execution time. +Set `low_cpu_mem_usage=True` to reduce CPU memory usage with streams. Tensors are pinned on the fly instead of all at once up front, which saves CPU memory but can increase inference time. It works best with `leaf_level` offloading when CPU memory is the bottleneck. #### Offloading to disk -Group offloading can consume significant system memory depending on the model size. On systems with limited memory, try group offloading onto the disk as a secondary memory. +Group offloading can use a lot of system memory depending on the model size. On systems with limited RAM, offload to disk instead. Set the `offload_to_disk_path` argument in either [`~ModelMixin.enable_group_offload`] or [`~hooks.apply_group_offloading`] to offload the model to the disk. @@ -400,19 +421,18 @@ pipeline.transformer.enable_group_offload(onload_device=onload_device, offload_d apply_group_offloading(pipeline.text_encoder, onload_device=onload_device, offload_type="block_level", num_blocks_per_group=2, offload_to_disk_path="path/to/disk") ``` -Refer to these [two](https://github.com/huggingface/diffusers/pull/11682#issue-3129365363) [tables](https://github.com/huggingface/diffusers/pull/11682#issuecomment-2955715126) to compare the speed and memory trade-offs. +Compare the speed and memory trade-offs in the [disk offloading benchmark](https://github.com/huggingface/diffusers/pull/11682#issue-3129365363) and the [follow-up benchmark](https://github.com/huggingface/diffusers/pull/11682#issuecomment-2955715126). ## Layerwise casting -> [!TIP] -> Combine layerwise casting with [group offloading](#group-offloading) for even more memory savings. +Layerwise casting stores weights in a smaller data format (for example, `torch.float8_e4m3fn` and `torch.float8_e5m2`) to use less memory and upcasts those weights to a higher precision like `torch.float16` or `torch.bfloat16` for computation. By default, positional embeddings, patch embeddings, normalization layers, and the input and output projections are skipped because storing them in fp8 can degrade generation quality. -Layerwise casting stores weights in a smaller data format (for example, `torch.float8_e4m3fn` and `torch.float8_e5m2`) to use less memory and upcasts those weights to a higher precision like `torch.float16` or `torch.bfloat16` for computation. Certain layers (normalization and modulation related weights) are skipped because storing them in fp8 can degrade generation quality. +Combine layerwise casting with [group offloading](#group-offloading) for even more memory savings. > [!WARNING] -> Layerwise casting may not work with all models if the forward implementation contains internal typecasting of weights. The current implementation of layerwise casting assumes the forward pass is independent of the weight precision and the input datatypes are always specified in `compute_dtype` (see [here](https://github.com/huggingface/transformers/blob/7f5077e53682ca855afc826162b204ebf809f1f9/src/transformers/models/t5/modeling_t5.py#L294-L299) for an incompatible implementation). +> Layerwise casting may not work with all models if the forward implementation contains internal typecasting of weights. The current implementation of layerwise casting assumes the forward pass is independent of the weight precision and the input datatypes are always specified in `compute_dtype` (see the [T5 implementation in Transformers](https://github.com/huggingface/transformers/blob/7f5077e53682ca855afc826162b204ebf809f1f9/src/transformers/models/t5/modeling_t5.py#L294-L299) for an incompatible example). > -> Layerwise casting may also fail on custom modeling implementations with [PEFT](https://huggingface.co/docs/peft/index) layers. There are some checks available but they are not extensively tested or guaranteed to work in all cases. +> Layerwise casting may also fail on custom modeling implementations with [PEFT](https://huggingface.co/docs/peft/index) layers. Diffusers includes some checks for this, but they aren't extensively tested and may not catch every case. Call [`~ModelMixin.enable_layerwise_casting`] to set the storage and computation datatypes. @@ -432,20 +452,13 @@ pipeline = CogVideoXPipeline.from_pretrained("THUDM/CogVideoX-5b", transformer=transformer, dtype=torch.bfloat16 ).to("cuda") # or "mps", "xpu", "cpu" -prompt = ( - "A panda, dressed in a small, red jacket and a tiny hat, sits on a wooden stool in a serene bamboo forest. " - "The panda's fluffy paws strum a miniature acoustic guitar, producing soft, melodic tunes. Nearby, a few other " - "pandas gather, watching curiously and some clapping in rhythm. Sunlight filters through the tall bamboo, " - "casting a gentle glow on the scene. The panda's face is expressive, showing concentration and joy as it plays. " - "The background includes a small, flowing stream and vibrant green foliage, enhancing the peaceful and magical " - "atmosphere of this unique musical performance." -) +prompt = "A panda playing a tiny acoustic guitar in a bamboo forest" video = pipeline(prompt=prompt, guidance_scale=6, num_inference_steps=50).frames[0] -print(f"Max memory reserved: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") +print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} GB") export_to_video(video, "output.mp4", fps=8) ``` -The [`~hooks.apply_layerwise_casting`] method can also be used if you need more control and flexibility. It can be partially applied to model layers by calling it on specific internal modules. Use the `skip_modules_pattern` or `skip_modules_classes` parameters to specify modules to avoid, such as the normalization and modulation layers. +For more control, use [`~hooks.apply_layerwise_casting`]. Call it on specific internal modules to apply layerwise casting to only part of a model, and use `skip_modules_pattern` or `skip_modules_classes` to exclude modules such as normalization layers. ```python import torch @@ -463,18 +476,25 @@ apply_layerwise_casting( transformer, storage_dtype=torch.float8_e4m3fn, compute_dtype=torch.bfloat16, - skip_modules_classes=["norm"], + skip_modules_pattern=["norm"], non_blocking=True, ) ``` ## torch.channels_last -[torch.channels_last](https://pytorch.org/tutorials/intermediate/memory_format_tutorial.html) flips how tensors are stored from `(batch size, channels, height, width)` to `(batch size, height, width, channels)`. This aligns the tensors with how the hardware sequentially accesses the tensors stored in memory and avoids skipping around in memory to access the pixel values. +[torch.channels_last](https://pytorch.org/tutorials/intermediate/memory_format_tutorial.html) changes how tensors are stored in memory from `(batch size, channels, height, width)` to `(batch size, height, width, channels)`. Storing each pixel's channels next to each other matches how many GPU kernels read memory, which mainly speeds up inference rather than reducing memory. -Not all operators currently support the channels-last format and may result in worst performance, but it is still worth trying. +channels_last only affects 4D tensors, so it benefits convolution-based models like UNets and VAEs. Not all operators support the channels-last format, and some models may run slower with it, so benchmark it on your model first. ```py +import torch +from diffusers import StableDiffusionPipeline + +pipeline = StableDiffusionPipeline.from_pretrained( + "stable-diffusion-v1-5/stable-diffusion-v1-5", dtype=torch.float16 +).to("cuda") + print(pipeline.unet.conv_out.state_dict()["weight"].stride()) # (2880, 9, 3, 1) pipeline.unet.to(memory_format=torch.channels_last) # in-place operation print( @@ -485,3 +505,9 @@ print( ## Memory-efficient attention Diffusers supports multiple memory-efficient attention backends (FlashAttention, xFormers, SageAttention, and more) through [`~ModelMixin.set_attention_backend`]. Refer to the [Attention backends](./attention_backends) guide to learn how to switch between them. + +## Next steps + +- Combine offloading with quantization and torch.compile in the [Compiling and offloading quantized models](./speed-memory-optims) guide. +- Reduce memory further with [quantization](../quantization/overview). +- Switch attention implementations in the [Attention backends](./attention_backends) guide. From f07c5bdac8ef6b0a0f0539e561674543d444b7a3 Mon Sep 17 00:00:00 2001 From: stevhliu Date: Fri, 25 Sep 2026 17:06:12 -0700 Subject: [PATCH 2/3] all the optims --- .../en/optimization/speed-memory-optims.md | 57 +++++++++---------- 1 file changed, 27 insertions(+), 30 deletions(-) diff --git a/docs/source/en/optimization/speed-memory-optims.md b/docs/source/en/optimization/speed-memory-optims.md index 21dd5b6ea1db..9bfaedcd9878 100644 --- a/docs/source/en/optimization/speed-memory-optims.md +++ b/docs/source/en/optimization/speed-memory-optims.md @@ -12,28 +12,27 @@ specific language governing permissions and limitations under the License. # Compiling and offloading quantized models -Optimizing models often involves trade-offs between [inference speed](./fp16) and [memory-usage](./memory). For instance, while [caching](./cache) can boost inference speed, it also increases memory consumption since it needs to store the outputs of intermediate attention layers. A more balanced optimization strategy combines quantizing a model, [torch.compile](./fp16#torchcompile) and various [offloading methods](./memory#offloading). +Quantization, [torch.compile](./fp16#torchcompile), and [offloading](./memory#offloading) can be combined to balance [inference speed](./fp16) and [memory usage](./memory). Quantization reduces the memory needed to store weights, torch.compile speeds up inference, and offloading keeps inactive layers or models on the CPU until they're needed. Other techniques trade one for the other. For example, [caching](./cache) speeds up inference but increases memory usage because it stores intermediate outputs. > [!TIP] -> Check the [torch.compile](./fp16#torchcompile) guide to learn more about compilation and how they can be applied here. For example, regional compilation can significantly reduce compilation time without giving up any speedups. +> Refer to the [torch.compile](./fp16#torchcompile) guide to learn more about compilation. For example, [regional compilation](./fp16#regional-compilation) significantly reduces compilation time without giving up the speedup. -For image generation, combining quantization and [model offloading](./memory#model-offloading) can often give the best trade-off between quality, speed, and memory. Group offloading is not as effective for image generation because it is usually not possible to *fully* overlap data transfer if the compute kernel finishes faster. This results in some communication overhead between the CPU and GPU. +The offloading method to combine with quantization depends on the workload. -For video generation, combining quantization and [group-offloading](./memory#group-offloading) tends to be better because video models are more compute-bound. +- For image generation, use [model offloading](./memory#model-offloading). Image models do less compute per layer, so with group offloading, the current layer often finishes before the next layer has transferred and the GPU waits on the CPU. +- For video generation, use [group offloading](./memory#group-offloading). Video models are more compute-bound, so data transfer overlaps with computation. -The table below provides a comparison of optimization strategy combinations and their impact on latency and memory-usage for Flux. +The table below shows the latency and memory usage of each combination on Flux. -| combination | latency (s) | memory-usage (GB) | +| Combination | Latency (s) | Memory usage (GB) | |---|---|---| | quantization | 32.602 | 14.9453 | | quantization, torch.compile | 25.847 | 14.9448 | | quantization, torch.compile, model CPU offloading | 32.312 | 12.2369 | -These results are benchmarked on Flux with a RTX 4090. The transformer and text_encoder components are quantized. Refer to the benchmarking script if you're interested in evaluating your own model. +Benchmarked on Flux with an RTX 4090, with the `transformer` and `text_encoder_2` (T5) components quantized. Use the benchmarking script to evaluate your own model. -This guide will show you how to compile and offload a quantized model with [bitsandbytes](../quantization/bitsandbytes#torchcompile). Make sure you are using [PyTorch nightly](https://pytorch.org/get-started/locally/) and the latest version of bitsandbytes. - -While we use bitsandbytes in this example, other quantization backends such as [TorchAO](../quantization/torchao.md) also support these features. +The examples below use [bitsandbytes](../quantization/bitsandbytes#torchcompile), but other quantization backends, such as [TorchAO](../quantization/torchao), also support compilation and offloading. Install the latest version of bitsandbytes. [PyTorch nightly](https://pytorch.org/get-started/locally/) is also recommended. ```bash pip install -U bitsandbytes @@ -41,9 +40,9 @@ pip install -U bitsandbytes ## Quantization and torch.compile -Start by [quantizing](../quantization/overview) a model to reduce the memory required for storage and [compiling](./fp16#torchcompile) it to accelerate inference. +[Quantize](../quantization/overview) a model to reduce the memory needed to store its weights, then [compile](./fp16#torchcompile) it to speed up inference. -Configure the [Dynamo](https://docs.pytorch.org/docs/stable/torch.compiler_dynamo_overview.html) `capture_dynamic_output_shape_ops = True` to handle dynamic outputs when compiling bitsandbytes models. +Set `torch._dynamo.config.capture_dynamic_output_shape_ops = True` so [Dynamo](https://docs.pytorch.org/docs/stable/torch.compiler_dynamo_overview.html) can compile bitsandbytes ops whose output shapes depend on the input. ```py import torch @@ -67,23 +66,21 @@ pipeline = DiffusionPipeline.from_pretrained( # compile pipeline.transformer.to(memory_format=torch.channels_last) pipeline.transformer.compile(mode="max-autotune", fullgraph=True) -pipeline(""" - cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California - highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain -""" +pipeline( + "cinematic film still of a cat sipping a margarita in a pool in Palm Springs, California, highly detailed, high budget hollywood movie, cinemascope, moody, epic, gorgeous, film grain" ).images[0] ``` ## Quantization, torch.compile, and offloading -In addition to quantization and torch.compile, try offloading if you need to reduce memory-usage further. Offloading moves various layers or model components from the CPU to the GPU as needed for computations. +Add offloading to quantization and torch.compile to reduce memory usage further. Offloading keeps layers or model components on the CPU and moves them to the GPU only when they're needed. -Configure the [Dynamo](https://docs.pytorch.org/docs/stable/torch.compiler_dynamo_overview.html) `cache_size_limit` during offloading to avoid excessive recompilation and set `capture_dynamic_output_shape_ops = True` to handle dynamic outputs when compiling bitsandbytes models. +Raise the [Dynamo](https://docs.pytorch.org/docs/stable/torch.compiler_dynamo_overview.html) `cache_size_limit` to avoid excessive recompilation with offloading, and set `capture_dynamic_output_shape_ops = True` to compile bitsandbytes ops whose output shapes depend on the input. -[Model CPU offloading](./memory#model-offloading) moves an individual pipeline component, like the transformer model, to the GPU when it is needed for computation. Otherwise, it is offloaded to the CPU. +[Model offloading](./memory#model-offloading) moves a whole pipeline component, like the transformer, to the GPU only when it's needed for computation. Otherwise, the component stays on the CPU. ```py import torch @@ -103,7 +100,7 @@ pipeline = DiffusionPipeline.from_pretrained( "black-forest-labs/FLUX.1-dev", quantization_config=pipeline_quant_config, dtype=torch.bfloat16, -).to("cuda") # or "mps", "xpu", "cpu" +) # model CPU offloading pipeline.enable_model_cpu_offload() @@ -118,18 +115,15 @@ pipeline( -[Group offloading](./memory#group-offloading) moves the internal layers of an individual pipeline component, like the transformer model, to the GPU for computation and offloads it when it's not required. At the same time, it uses the [CUDA stream](./memory#cuda-stream) feature to prefetch the next layer for execution. - -By overlapping computation and data transfer, it is faster than model CPU offloading while also saving memory. +[Group offloading](./memory#group-offloading) moves the internal layers of a component, like the transformer, to the GPU only when they run. With `use_stream=True`, it uses [CUDA streams](./memory#cuda-stream) to prefetch the next layer while the current one runs. For compute-bound video models, this overlap makes group offloading faster than model offloading while also using less memory. ```py # pip install ftfy import torch -from diffusers import AutoModel, DiffusionPipeline +from diffusers import DiffusionPipeline from diffusers.hooks import apply_group_offloading from diffusers.utils import export_to_video from diffusers.quantizers import PipelineQuantizationConfig -from transformers import UMT5EncoderModel torch._dynamo.config.cache_size_limit = 1000 torch._dynamo.config.capture_dynamic_output_shape_ops = True @@ -141,14 +135,11 @@ pipeline_quant_config = PipelineQuantizationConfig( components_to_quantize=["transformer", "text_encoder"], ) -text_encoder = UMT5EncoderModel.from_pretrained( - "Wan-AI/Wan2.1-T2V-14B-Diffusers", subfolder="text_encoder", dtype=torch.bfloat16 -) pipeline = DiffusionPipeline.from_pretrained( "Wan-AI/Wan2.1-T2V-14B-Diffusers", quantization_config=pipeline_quant_config, dtype=torch.bfloat16, -).to("cuda") # or "mps", "xpu", "cpu" +) # group offloading onload_device = torch.device("cuda") @@ -202,4 +193,10 @@ export_to_video(output, "output.mp4", fps=16) ``` - \ No newline at end of file + + +## Next steps + +- Learn more about each offloading method in the [Reduce memory usage](./memory) guide. +- Speed up inference further with the [torch.compile](./fp16#torchcompile) guide. +- Compare quantization backends in the [quantization overview](../quantization/overview). From fce14c6a0f811785fa518f52716988dd0793c672 Mon Sep 17 00:00:00 2001 From: stevhliu Date: Tue, 29 Sep 2026 09:33:11 -0700 Subject: [PATCH 3/3] soften --- docs/source/en/optimization/memory.md | 14 +++++++------- docs/source/en/optimization/speed-memory-optims.md | 8 ++++---- 2 files changed, 11 insertions(+), 11 deletions(-) diff --git a/docs/source/en/optimization/memory.md b/docs/source/en/optimization/memory.md index 37d38d44997a..23b721f411e0 100644 --- a/docs/source/en/optimization/memory.md +++ b/docs/source/en/optimization/memory.md @@ -164,7 +164,7 @@ pipeline = StableDiffusionXLPipeline.from_pretrained( By default, Diffusers uses all available memory on each GPU. Components that don't fit on a GPU are placed on the CPU. If most of the pipeline ends up on the CPU, try a single GPU with one of the offloading methods below instead. - [`~DiffusionPipeline.enable_model_cpu_offload`] moves one whole model to the GPU at a time. It's faster, but each model must fit on a single GPU. -- [`~DiffusionPipeline.enable_sequential_cpu_offload`] moves one submodule to the GPU at a time. It uses the least memory, but it's very slow. +- [`~DiffusionPipeline.enable_sequential_cpu_offload`] moves one submodule to the GPU at a time. It uses the least GPU memory, but it's very slow. Before calling `.to()`, `enable_sequential_cpu_offload`, or `enable_model_cpu_offload` on a device-mapped pipeline, reset its device map with [`~DiffusionPipeline.reset_device_map`]. @@ -174,7 +174,7 @@ pipeline.reset_device_map() ## VAE slicing -VAE slicing splits a batch of latents into single latents and decodes them one at a time. The decoded images are concatenated back into a batch at the end. Peak memory stays close to the cost of decoding one image regardless of batch size, which makes slicing useful when generating several images at once. It has no effect on single-image batches. +VAE slicing splits a batch of latents into single latents and decodes them one at a time. The decoded images are concatenated back into a batch at the end. Peak decoding memory stays close to the cost of decoding one image, which makes slicing useful when generating several images at once. It has no effect on single-image batches. ```text Without slicing: one decode for the whole batch @@ -213,7 +213,7 @@ print(f"Max memory allocated: {torch.cuda.max_memory_allocated() / 1024**3:.2f} ## VAE tiling -VAE tiling splits a latent into overlapping tiles and decodes each tile separately. The overlapping edges are blended together to stitch the tiles into the final image. Peak memory depends on the tile size instead of the full image size, which makes tiling useful for generating high-resolution images. +VAE tiling splits a latent into overlapping tiles and decodes each tile separately. The overlapping edges are blended together to stitch the tiles into the final image. Peak decoding memory depends mostly on the tile size instead of the full image size, which makes tiling useful for generating high-resolution images. ```text 1. Split the latent into tiles 2. Decode each tile 3. Blend the overlaps @@ -262,7 +262,7 @@ Refer to the [Compiling and offloading quantized models](./speed-memory-optims) ### Sequential CPU offloading -Sequential CPU offloading keeps weights on the CPU and moves each submodule to the GPU only when it runs. The entire model is never on the GPU at once, so sequential offloading uses the least memory of the offloading methods. It's also the slowest because submodules are transferred between devices many times during inference, which often makes it impractical. +Sequential CPU offloading keeps weights on the CPU and moves each submodule to the GPU only when it runs. The entire model is never on the GPU at once, so sequential offloading uses the least GPU memory of the offloading methods. It's also the slowest because submodules are transferred between devices many times during inference, which often makes it impractical. > [!WARNING] > Don't move the pipeline to CUDA before calling `enable_sequential_cpu_offload`, otherwise the memory savings are minimal. Refer to [issue #1934](https://github.com/huggingface/diffusers/issues/1934) for more details. Sequential offloading is stateful and installs hooks on the model. @@ -322,7 +322,7 @@ Model offloading also helps when you call [`~StableDiffusionXLPipeline.encode_pr ### Group offloading -Group offloading moves groups of internal layers ([torch.nn.ModuleList](https://pytorch.org/docs/stable/generated/torch.nn.ModuleList.html) or [torch.nn.Sequential](https://pytorch.org/docs/stable/generated/torch.nn.Sequential.html)) to the CPU. It uses less memory than [model offloading](#model-offloading) and it is faster than [sequential CPU offloading](#sequential-cpu-offloading) because it reduces communication overhead. +Group offloading moves groups of internal layers ([torch.nn.ModuleList](https://pytorch.org/docs/stable/generated/torch.nn.ModuleList.html) or [torch.nn.Sequential](https://pytorch.org/docs/stable/generated/torch.nn.Sequential.html)) to the CPU. It usually uses less memory than [model offloading](#model-offloading) and runs faster than [sequential CPU offloading](#sequential-cpu-offloading) because it reduces communication overhead. > [!WARNING] > Group offloading may not work with models whose forward pass moves inputs to the weights' device, because that conflicts with how group offloading moves tensors. @@ -394,7 +394,7 @@ export_to_video(video, "output.mp4", fps=8) #### CUDA stream -Set `use_stream=True` on CUDA devices to prefetch the next layer onto the GPU while the current layer is still running. Overlapping data transfer and computation makes group offloading much faster than [sequential CPU offloading](#sequential-cpu-offloading). Streams pin tensors in CPU memory, so make sure you have about twice the model size in system RAM. +Set `use_stream=True` on CUDA devices to prefetch the next layer onto the GPU while the current layer is still running. Overlapping data transfer and computation can make group offloading much faster than [sequential CPU offloading](#sequential-cpu-offloading). Streams create a pinned copy of each weight in CPU memory, so system RAM usage can reach about twice the model size. Set `record_stream=True` for more of a speedup at the cost of slightly increased memory usage. Refer to the [torch.Tensor.record_stream](https://pytorch.org/docs/stable/generated/torch.Tensor.record_stream.html) docs to learn more. @@ -485,7 +485,7 @@ apply_layerwise_casting( [torch.channels_last](https://pytorch.org/tutorials/intermediate/memory_format_tutorial.html) changes how tensors are stored in memory from `(batch size, channels, height, width)` to `(batch size, height, width, channels)`. Storing each pixel's channels next to each other matches how many GPU kernels read memory, which mainly speeds up inference rather than reducing memory. -channels_last only affects 4D tensors, so it benefits convolution-based models like UNets and VAEs. Not all operators support the channels-last format, and some models may run slower with it, so benchmark it on your model first. +channels_last only affects 4D tensors, so it can benefit convolution-based models like UNets and VAEs. Not all operators support the channels-last format, and some models may run slower with it, so benchmark it on your model first. ```py import torch diff --git a/docs/source/en/optimization/speed-memory-optims.md b/docs/source/en/optimization/speed-memory-optims.md index 9bfaedcd9878..a2c2431ebc52 100644 --- a/docs/source/en/optimization/speed-memory-optims.md +++ b/docs/source/en/optimization/speed-memory-optims.md @@ -15,12 +15,12 @@ specific language governing permissions and limitations under the License. Quantization, [torch.compile](./fp16#torchcompile), and [offloading](./memory#offloading) can be combined to balance [inference speed](./fp16) and [memory usage](./memory). Quantization reduces the memory needed to store weights, torch.compile speeds up inference, and offloading keeps inactive layers or models on the CPU until they're needed. Other techniques trade one for the other. For example, [caching](./cache) speeds up inference but increases memory usage because it stores intermediate outputs. > [!TIP] -> Refer to the [torch.compile](./fp16#torchcompile) guide to learn more about compilation. For example, [regional compilation](./fp16#regional-compilation) significantly reduces compilation time without giving up the speedup. +> Refer to the [torch.compile](./fp16#torchcompile) guide to learn more about compilation. For example, [regional compilation](./fp16#regional-compilation) can significantly reduce compilation time while keeping a comparable speedup. The offloading method to combine with quantization depends on the workload. -- For image generation, use [model offloading](./memory#model-offloading). Image models do less compute per layer, so with group offloading, the current layer often finishes before the next layer has transferred and the GPU waits on the CPU. -- For video generation, use [group offloading](./memory#group-offloading). Video models are more compute-bound, so data transfer overlaps with computation. +- For image generation, [model offloading](./memory#model-offloading) usually works best. Image models do less compute per layer, so with group offloading, the current layer often finishes before the next layer has transferred and the GPU waits on the CPU. +- For video generation, [group offloading](./memory#group-offloading) usually works better. Video models are more compute-bound, so data transfer can overlap with computation. The table below shows the latency and memory usage of each combination on Flux. @@ -115,7 +115,7 @@ pipeline( -[Group offloading](./memory#group-offloading) moves the internal layers of a component, like the transformer, to the GPU only when they run. With `use_stream=True`, it uses [CUDA streams](./memory#cuda-stream) to prefetch the next layer while the current one runs. For compute-bound video models, this overlap makes group offloading faster than model offloading while also using less memory. +[Group offloading](./memory#group-offloading) moves the internal layers of a component, like the transformer, to the GPU only when they run. With `use_stream=True`, it uses [CUDA streams](./memory#cuda-stream) to prefetch the next layer while the current one runs. For compute-bound video models, this overlap can make group offloading faster than model offloading while also using less memory. ```py # pip install ftfy