Skip to content

Add SPMD types in TorchTitan devlog - #40

Merged
pianpwk merged 3 commits into
mainfrom
gh/pianpwk/1/head
Sep 2, 2026
Merged

pianpwk merged 3 commits into
mainfrom
gh/pianpwk/1/head

Conversation

@pianpwk

@pianpwk pianpwk commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Stack from ghstack (oldest at bottom):

[ghstack-poisoned]
[ghstack-poisoned]
@meta-cla meta-cla Bot added the cla signed label Aug 28, 2026

@aditvenk aditvenk left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is great!

| --- | --- | ---: | ---: | --- |
| Llama 3, FSDP=4, TP=2 | Eager | 47,393 | 67,777 | **+43.0%** |
| Llama 3, FSDP=4, TP=2 | Eager + CUDA graphs | 235,810 | 255,280 | **+8.3%** |
| Llama 3, FSDP=4, TP=2 | Compile (`inductor`) | 104,738 | 120,154 | **+14.7%** |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@pianpwk -- how do we see this big perf improvement in this case? I'd think inductor compilation removes any Dtensor overhead.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I updated, we still see more overhead that inductor doesn't remove because compile is regional; in this case chunked loss wrapper isn't compiled, only inner loss fn

pianpwk added a commit that referenced this pull request Aug 31, 2026
ghstack-source-id: a657fde
Pull-Request: #40
[ghstack-poisoned]
pianpwk added a commit that referenced this pull request Sep 2, 2026
ghstack-source-id: c1ef634
Pull-Request: #40
@pianpwk
pianpwk changed the base branch from gh/pianpwk/1/base to main September 2, 2026 17:34
@pianpwk
pianpwk merged commit 3ab45f3 into main Sep 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants