Skip to content

Stop passing fps to the transformer in Cosmos2VideoToWorldPipeline - #14940

Open
UniversePeak wants to merge 1 commit into
huggingface:mainfrom
UniversePeak:fix-cosmos2-v2w-fps-rope-modulation
Open

UniversePeak wants to merge 1 commit into
huggingface:mainfrom
UniversePeak:fix-cosmos2-v2w-fps-rope-modulation

Conversation

@UniversePeak

Copy link
Copy Markdown

What does this PR do?

Fixes #14768 (the second option suggested there).

The released Cosmos-Predict2 checkpoints ship with rope_enable_fps_modulation=False (config_video2world.py, same in the 14B and text-to-image nets), so natively they use raw integer temporal RoPE positions and ignore FPS entirely. Cosmos2VideoToWorldPipeline was the only Predict2-family pipeline still passing fps into the transformer, whose CosmosRotaryPosEmbed scales temporal positions by base_fps / fps = 24 / 16 = 1.5x on every default run, and the playback fps argument changed the generated content.

This PR stops passing fps to the transformer and documents fps as a playback-only argument. I went with this instead of porting enable_fps_modulation into CosmosTransformer3DModel because:

  • No published nvidia/Cosmos-Predict2-* config can contain a key that does not exist in the class today, so register_to_config would hand every released checkpoint the class default; the flag would carry no per-checkpoint information.
  • CosmosTransformer3DModel is shared with Cosmos-Predict1, which always applies the modulation (its pipelines pass fps). A default of False would silently break Predict1 and True keeps this bug, so the pipeline is where the family difference belongs.

Note that this changes the output of every default Cosmos2VideoToWorldPipeline run: on this branch fps=16 reproduces what fps=24 produced before, i.e. the un-modulated temporal RoPE the checkpoints were trained with.

Tests

  • Added test_fps_does_not_change_generation: the pipeline output for fps 16, 24 and 30 must be identical. On main it fails with max abs diff 0.0824; with this change it passes.
  • pytest tests/pipelines/cosmos/test_cosmos2_video2world.py tests/models/transformers/test_models_transformer_cosmos.py tests/pipelines/cosmos/test_cosmos.py tests/pipelines/cosmos/test_cosmos_video2world.py — 98 passed, 78 skipped on CPU.
  • ruff check and ruff format --check clean on the touched files.

  • This PR fixes a bug or fixes the issue mentioned in the title
  • I did read the contributing guideline
  • I did update the relevant documentation
  • I did write new tests for the change

The released Cosmos-Predict2 checkpoints ship with
rope_enable_fps_modulation=False, so natively they use raw integer
temporal RoPE positions and ignore FPS entirely. The diffusers port of
Cosmos2VideoToWorldPipeline always passed fps (default 16) to
CosmosTransformer3DModel, whose CosmosRotaryPosEmbed scales temporal
positions by base_fps / fps = 24 / 16 = 1.5x for every default run,
diverging from the reference implementation and making the playback fps
change the generated content.

Cosmos2TextToImagePipeline and the Cosmos 2.5 pipelines already pass no
fps, and the transformer class is shared with Cosmos-Predict1, whose
pipelines still pass it, so the pipeline is the right place to encode
this difference. fps stays a playback-only argument of the pipeline.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

1 participant