Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
1 change: 1 addition & 0 deletions .github/workflows/pr_modular_tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,7 @@ jobs:
python utils/check_dummies.py
python utils/check_support_list.py
python utils/check_forward_call_docstrings.py
python utils/check_return_annotations.py
make deps_table_check_updated
- name: Check if failure
if: ${{ failure() }}
Expand Down
1 change: 1 addition & 0 deletions .github/workflows/pr_tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -78,6 +78,7 @@ jobs:
python utils/check_dummies.py
python utils/check_support_list.py
python utils/check_forward_call_docstrings.py
python utils/check_return_annotations.py
make deps_table_check_updated
- name: Check if failure
if: ${{ failure() }}
Expand Down
1 change: 1 addition & 0 deletions .github/workflows/pr_tests_gpu.yml
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,7 @@ jobs:
python utils/check_dummies.py
python utils/check_support_list.py
python utils/check_forward_call_docstrings.py
python utils/check_return_annotations.py
make deps_table_check_updated
- name: Check if failure
if: ${{ failure() }}
Expand Down
5 changes: 5 additions & 0 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ repo-consistency:
python utils/check_repo.py
python utils/check_inits.py
python utils/check_forward_call_docstrings.py
python utils/check_return_annotations.py

# this target runs checks on all files

Expand Down Expand Up @@ -80,6 +81,10 @@ modular-autodoctrings:
check-forward-call-docstrings:
python utils/check_forward_call_docstrings.py

# Verify forward() / __call__() have return type annotations
check-return-annotations:
python utils/check_return_annotations.py

# Run tests for the library

test:
Expand Down
15 changes: 11 additions & 4 deletions docs/source/en/optimization/cache.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,8 +72,14 @@ pipeline.transformer.enable_cache(config)

[SeaCache](https://huggingface.co/papers/2602.18993) compares Spectral-Evolution-Aware (SEA) indicators between
successive denoising steps. When the accumulated indicator change remains below a threshold, it skips the expensive
transformer block stack and predicts its output from cached residuals. The indicator is computed from the raw vision
latents, including clean conditioning frames for image-to-video generation.
transformer block stack and predicts its output from cached residuals. Build the indicator from the visual latents that
form the generated output. Include clean conditioning frames when they are part of that output trajectory, as in
image-to-video and video-to-video generation. Exclude separate visual hints that condition the generation but are not
part of the output. Text conditioning is excluded because it is not a visual latent.

Cosmos 3 Transfer packs control hints as separate visual sequences, so its adapter excludes them from the indicator.
Control-CFG branches compare the same output trajectory while retaining their own cached residuals. Control hints still
condition the transformer.

The implementation provides built-in adapters for the following models:

Expand All @@ -87,8 +93,9 @@ Other video transformers can integrate with the generic path when they use `Cach
block list, and register the block input/output layout in `TransformerBlockRegistry`. The pipeline must enter a
`cache_context` for every transformer call, attach `step_index`, `sigma`, and `num_inference_steps`, and use separate
context names for independent trajectories such as conditional and unconditional guidance. Pass a `raw_vision_callback`
that returns the noisy vision latents when no built-in adapter is available. Validate output quality and tune the cache
parameters for each model and scheduler; support and benchmark results do not transfer automatically from Cosmos 3.
that returns the visual latents forming the generated output when no built-in adapter is available. Validate output
quality and tune the cache parameters for each model and scheduler; support and benchmark results do not transfer
automatically from Cosmos 3.

### Cosmos 3

Expand Down
4 changes: 2 additions & 2 deletions examples/community/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -3526,7 +3526,7 @@ from controlnet_aux.midas import MidasDetector
from PIL import Image

from diffusers import AutoencoderKL, ControlNetModel, MultiAdapter, T2IAdapter
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.utils import load_image
from examples.community.pipeline_stable_diffusion_xl_controlnet_adapter import (
StableDiffusionXLControlNetAdapterPipeline,
Expand Down Expand Up @@ -3591,7 +3591,7 @@ from controlnet_aux.midas import MidasDetector
from PIL import Image

from diffusers import AutoencoderKL, ControlNetModel, MultiAdapter, T2IAdapter
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.utils import load_image
from examples.community.pipeline_stable_diffusion_xl_controlnet_adapter_inpaint import (
StableDiffusionXLControlNetAdapterInpaintPipeline,
Expand Down
2 changes: 1 addition & 1 deletion examples/community/fresco_v2v.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,9 +29,9 @@
from diffusers.loaders import StableDiffusionLoraLoaderMixin, TextualInversionLoaderMixin
from diffusers.models import AutoencoderKL, ControlNetModel, ImageProjection, UNet2DConditionModel
from diffusers.models.attention_processor import AttnProcessor2_0
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.lora import adjust_lora_scale_text_encoder
from diffusers.models.unets.unet_2d_condition import UNet2DConditionOutput
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.controlnet.pipeline_controlnet_img2img import StableDiffusionControlNetImg2ImgPipeline
from diffusers.pipelines.stable_diffusion import StableDiffusionPipelineOutput
from diffusers.pipelines.stable_diffusion.safety_checker import StableDiffusionSafetyChecker
Expand Down
2 changes: 1 addition & 1 deletion examples/community/pipeline_animatediff_controlnet.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,10 +24,10 @@
from diffusers.image_processor import PipelineImageInput, VaeImageProcessor
from diffusers.loaders import IPAdapterMixin, StableDiffusionLoraLoaderMixin, TextualInversionLoaderMixin
from diffusers.models import AutoencoderKL, ControlNetModel, ImageProjection, UNet2DConditionModel, UNetMotionModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.lora import adjust_lora_scale_text_encoder
from diffusers.models.unets.unet_motion_model import MotionAdapter
from diffusers.pipelines.animatediff.pipeline_output import AnimateDiffPipelineOutput
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.pipeline_utils import DiffusionPipeline, StableDiffusionMixin
from diffusers.schedulers import (
DDIMScheduler,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -25,8 +25,8 @@
from diffusers.image_processor import PipelineImageInput, VaeImageProcessor
from diffusers.loaders import FromSingleFileMixin, StableDiffusionXLLoraLoaderMixin, TextualInversionLoaderMixin
from diffusers.models import AutoencoderKL, ControlNetModel, MultiAdapter, T2IAdapter, UNet2DConditionModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.lora import adjust_lora_scale_text_encoder
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.pipeline_utils import DiffusionPipeline, StableDiffusionMixin
from diffusers.pipelines.stable_diffusion_xl.pipeline_output import StableDiffusionXLPipelineOutput
from diffusers.schedulers import KarrasDiffusionSchedulers
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,8 @@
T2IAdapter,
UNet2DConditionModel,
)
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.lora import adjust_lora_scale_text_encoder
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.pipeline_utils import StableDiffusionMixin
from diffusers.pipelines.stable_diffusion_xl.pipeline_output import StableDiffusionXLPipelineOutput
from diffusers.schedulers import KarrasDiffusionSchedulers
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@
from diffusers import StableDiffusionXLControlNetImg2ImgPipeline
from diffusers.image_processor import PipelineImageInput
from diffusers.models import ControlNetModel
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.stable_diffusion_xl import StableDiffusionXLPipelineOutput
from diffusers.utils import (
deprecate,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@
from diffusers import StableDiffusionXLControlNetPipeline
from diffusers.image_processor import PipelineImageInput
from diffusers.models import ControlNetModel
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.stable_diffusion_xl import StableDiffusionXLPipelineOutput
from diffusers.utils import (
deprecate,
Expand Down
2 changes: 1 addition & 1 deletion examples/community/rerender_a_video.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@
from diffusers.image_processor import VaeImageProcessor
from diffusers.models import AutoencoderKL, ControlNetModel, UNet2DConditionModel
from diffusers.models.attention_processor import Attention, AttnProcessor
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.controlnet.pipeline_controlnet_img2img import StableDiffusionControlNetImg2ImgPipeline
from diffusers.pipelines.stable_diffusion.safety_checker import StableDiffusionSafetyChecker
from diffusers.schedulers import KarrasDiffusionSchedulers
Expand Down
2 changes: 1 addition & 1 deletion examples/community/stable_diffusion_controlnet_img2img.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@
from transformers import CLIPImageProcessor, CLIPTextModel, CLIPTokenizer

from diffusers import AutoencoderKL, ControlNetModel, UNet2DConditionModel, logging
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.pipeline_utils import DiffusionPipeline, StableDiffusionMixin
from diffusers.pipelines.stable_diffusion import StableDiffusionPipelineOutput, StableDiffusionSafetyChecker
from diffusers.schedulers import KarrasDiffusionSchedulers
Expand Down
2 changes: 1 addition & 1 deletion examples/community/stable_diffusion_controlnet_inpaint.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
from transformers import CLIPImageProcessor, CLIPTextModel, CLIPTokenizer

from diffusers import AutoencoderKL, ControlNetModel, UNet2DConditionModel, logging
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.pipeline_utils import DiffusionPipeline, StableDiffusionMixin
from diffusers.pipelines.stable_diffusion import StableDiffusionPipelineOutput, StableDiffusionSafetyChecker
from diffusers.schedulers import KarrasDiffusionSchedulers
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,8 @@
from diffusers import StableDiffusionControlNetPipeline
from diffusers.models import ControlNetModel
from diffusers.models.attention import BasicTransformerBlock
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.unets.unet_2d_blocks import CrossAttnDownBlock2D, CrossAttnUpBlock2D, DownBlock2D, UpBlock2D
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.stable_diffusion import StableDiffusionPipelineOutput
from diffusers.utils import logging
from diffusers.utils.torch_utils import is_compiled_module, randn_tensor
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@
from diffusers.image_processor import PipelineImageInput
from diffusers.models import ControlNetModel
from diffusers.models.attention import BasicTransformerBlock
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.unets.unet_2d_blocks import CrossAttnDownBlock2D, CrossAttnUpBlock2D, DownBlock2D, UpBlock2D
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.stable_diffusion_xl.pipeline_output import StableDiffusionXLPipelineOutput
from diffusers.utils import PIL_INTERPOLATION, deprecate, logging, replace_example_docstring
from diffusers.utils.torch_utils import is_compiled_module, is_torch_version, randn_tensor
Expand Down
2 changes: 1 addition & 1 deletion examples/research_projects/anytext/anytext.py
Original file line number Diff line number Diff line change
Expand Up @@ -53,9 +53,9 @@
TextualInversionLoaderMixin,
)
from diffusers.models import AutoencoderKL, ControlNetModel, ImageProjection, UNet2DConditionModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.lora import adjust_lora_scale_text_encoder
from diffusers.models.modeling_utils import ModelMixin
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.pipeline_utils import DiffusionPipeline, StableDiffusionMixin
from diffusers.pipelines.stable_diffusion.pipeline_output import StableDiffusionPipelineOutput
from diffusers.pipelines.stable_diffusion.safety_checker import StableDiffusionSafetyChecker
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -30,8 +30,8 @@
from diffusers.image_processor import PipelineImageInput, VaeImageProcessor
from diffusers.loaders import FromSingleFileMixin, StableDiffusionLoraLoaderMixin, TextualInversionLoaderMixin
from diffusers.models import AutoencoderKL, ControlNetModel, UNet2DConditionModel
from diffusers.models.controlnets.multicontrolnet import MultiControlNetModel
from diffusers.models.lora import adjust_lora_scale_text_encoder
from diffusers.pipelines.controlnet.multicontrolnet import MultiControlNetModel
from diffusers.pipelines.pipeline_utils import DiffusionPipeline
from diffusers.pipelines.stable_diffusion.pipeline_output import StableDiffusionPipelineOutput
from diffusers.pipelines.stable_diffusion.safety_checker import StableDiffusionSafetyChecker
Expand Down
4 changes: 0 additions & 4 deletions src/diffusers/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -747,9 +747,7 @@
"LTXPipeline",
"LucyEditPipeline",
"Lumina2Pipeline",
"Lumina2Text2ImgPipeline",
"LuminaPipeline",
"LuminaText2ImgPipeline",
"MarigoldDepthPipeline",
"MarigoldIntrinsicsPipeline",
"MarigoldNormalsPipeline",
Expand Down Expand Up @@ -1604,9 +1602,7 @@
LTXPipeline,
LucyEditPipeline,
Lumina2Pipeline,
Lumina2Text2ImgPipeline,
LuminaPipeline,
LuminaText2ImgPipeline,
MarigoldDepthPipeline,
MarigoldIntrinsicsPipeline,
MarigoldNormalsPipeline,
Expand Down
14 changes: 8 additions & 6 deletions src/diffusers/hooks/sea_cache.py
Original file line number Diff line number Diff line change
Expand Up @@ -64,8 +64,9 @@ class SeaCacheConfig:
power_exp (`float`, defaults to `3.0`):
Exponent of the SEA clean-signal power prior. SeaCache uses `3.0` for video features.
raw_vision_callback (`Callable`, *optional*):
Advanced model adapter returning raw vision latents with shape `(C, T, H, W)`. When omitted, a built-in
adapter is used if one is available.
Advanced model adapter returning the visual latents forming the generated output, each with shape `(C, T,
H, W)`. Include clean conditioning frames within the output trajectory, but exclude separate visual hints
that are not part of the output. When omitted, a built-in adapter is used if one is available.

Example:
```python
Expand Down Expand Up @@ -326,7 +327,6 @@ def _prepare_cosmos3_raw_vision_metadata(
return None

raw_vision = []
has_noisy_vision = False
for latent, noisy_frame_indexes in zip(vision_tokens, vision_noisy_frame_indexes):
if not isinstance(latent, torch.Tensor) or not isinstance(noisy_frame_indexes, torch.Tensor):
return None
Expand All @@ -340,10 +340,12 @@ def _prepare_cosmos3_raw_vision_metadata(
noisy_frame_indexes = noisy_frame_indexes.flatten().to(device=latent.device, dtype=torch.long)
if torch.any(noisy_frame_indexes < 0) or torch.any(noisy_frame_indexes >= latent.shape[1]):
return None
has_noisy_vision = has_noisy_vision or noisy_frame_indexes.numel() > 0
raw_vision.append(latent)
# A sequence with noisy frames belongs to the generated output. Keep that sequence whole so clean conditioning
# frames remain in the indicator, but exclude separate clean hints that are not part of the output.
if noisy_frame_indexes.numel() > 0:
raw_vision.append(latent)

return raw_vision if raw_vision and has_noisy_vision else None
return raw_vision or None


def _prepare_wan_t2v_raw_vision_metadata(
Expand Down
49 changes: 1 addition & 48 deletions src/diffusers/loaders/__init__.py
Original file line number Diff line number Diff line change
@@ -1,56 +1,9 @@
from typing import TYPE_CHECKING

from ..utils import DIFFUSERS_SLOW_IMPORT, _LazyModule, deprecate
from ..utils import DIFFUSERS_SLOW_IMPORT, _LazyModule
from ..utils.import_utils import is_peft_available, is_torch_available, is_transformers_available


def text_encoder_lora_state_dict(text_encoder):
deprecate(
"text_encoder_load_state_dict in `models`",
"0.27.0",
"`text_encoder_lora_state_dict` is deprecated and will be removed in 0.27.0. Make sure to retrieve the weights using `get_peft_model`. See https://huggingface.co/docs/peft/v0.6.2/en/quicktour#peftmodel for more information.",
)
state_dict = {}

for name, module in text_encoder_attn_modules(text_encoder):
for k, v in module.q_proj.lora_linear_layer.state_dict().items():
state_dict[f"{name}.q_proj.lora_linear_layer.{k}"] = v

for k, v in module.k_proj.lora_linear_layer.state_dict().items():
state_dict[f"{name}.k_proj.lora_linear_layer.{k}"] = v

for k, v in module.v_proj.lora_linear_layer.state_dict().items():
state_dict[f"{name}.v_proj.lora_linear_layer.{k}"] = v

for k, v in module.out_proj.lora_linear_layer.state_dict().items():
state_dict[f"{name}.out_proj.lora_linear_layer.{k}"] = v

return state_dict


if is_transformers_available():

def text_encoder_attn_modules(text_encoder):
deprecate(
"text_encoder_attn_modules in `models`",
"0.27.0",
"`text_encoder_lora_state_dict` is deprecated and will be removed in 0.27.0. Make sure to retrieve the weights using `get_peft_model`. See https://huggingface.co/docs/peft/v0.6.2/en/quicktour#peftmodel for more information.",
)
from transformers import CLIPTextModel, CLIPTextModelWithProjection

attn_modules = []

if isinstance(text_encoder, (CLIPTextModel, CLIPTextModelWithProjection)):
for i, layer in enumerate(text_encoder.text_model.encoder.layers):
name = f"text_model.encoder.layers.{i}.self_attn"
mod = layer.self_attn
attn_modules.append((name, mod))
else:
raise ValueError(f"do not know how to get attention modules for: {text_encoder.__class__.__name__}")

return attn_modules


_import_structure = {}

if is_torch_available():
Expand Down
22 changes: 22 additions & 0 deletions src/diffusers/loaders/lora_conversion_utils.py
Original file line number Diff line number Diff line change
Expand Up @@ -1582,6 +1582,28 @@ def _convert_fal_kontext_lora_to_diffusers(original_state_dict):
f"{original_block_prefix}final_layer.linear.{lora_key}.bias"
)

# Some fal-kontext LoRAs carry the global embedder keys (time_in, vector_in, txt_in, img_in, guidance_in)
# without the `base_model.model.` prefix the block keys use.
for lora_key in ["lora_A", "lora_B"]:
for src, dst in [
(f"time_in.in_layer.{lora_key}.weight", f"time_text_embed.timestep_embedder.linear_1.{lora_key}.weight"),
(f"time_in.out_layer.{lora_key}.weight", f"time_text_embed.timestep_embedder.linear_2.{lora_key}.weight"),
(f"vector_in.in_layer.{lora_key}.weight", f"time_text_embed.text_embedder.linear_1.{lora_key}.weight"),
(f"vector_in.out_layer.{lora_key}.weight", f"time_text_embed.text_embedder.linear_2.{lora_key}.weight"),
(f"txt_in.{lora_key}.weight", f"context_embedder.{lora_key}.weight"),
(f"img_in.{lora_key}.weight", f"x_embedder.{lora_key}.weight"),
(
f"guidance_in.in_layer.{lora_key}.weight",
f"time_text_embed.guidance_embedder.linear_1.{lora_key}.weight",
),
(
f"guidance_in.out_layer.{lora_key}.weight",
f"time_text_embed.guidance_embedder.linear_2.{lora_key}.weight",
),
]:
if src in original_state_dict:
converted_state_dict[dst] = original_state_dict.pop(src)

if len(original_state_dict) > 0:
raise ValueError(f"`original_state_dict` should be empty at this point but has {original_state_dict.keys()=}.")

Expand Down
Loading
Loading