Skip to content

Build the nightly docker with remote BuildKit on OSDC, and move it to CUDA 13.0 - #2712

Open
huydhn wants to merge 2 commits into
mainfrom
osdc/nightly-docker-buildkit
Open

huydhn wants to merge 2 commits into
mainfrom
osdc/nightly-docker-buildkit

Conversation

@huydhn

@huydhn huydhn commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Migrates build-nightly-docker.yml off EC2. docker build/tag/push all need a daemon, which OSDC runners do not have; the build moves to the remote BuildKit pool via test-infra's docker-build-remote-buildkit action.

  • The dev<date> tag came from docker run-ing the built image and importing torch. Instead the Dockerfile writes the version into the image and a scratch version-export stage exposes just that file, so buildx can pull it out with --output type=local. A trailing FROM base AS image keeps the default build target the real image, so a plain docker build still works.
  • That means two buildx calls, cached against a registry :buildcache — the pool can hand out a different builder each time and the second call must not rebuild install.py.
  • linux.4xlarge → mt-l-x86iavx512-16-128. The runner now only ships the context and waits; the build runs on buildkitd.

CUDA 13.0

Separate, pre-existing bug — this nightly has been failing on main since 2026-06-05, and would fail the same way on EC2. The cu128 index stopped being published in April 2026; all that is left is one torch (dev20260408) and torchvisions pinning dev20260407, which was pruned, so nothing resolves. cu129 is dead the same way, cu126 is stale since early September. cu130 is live and resolves to a consistent dev20260928 set.

Bumping CUDA_VERSION is the whole change — the index URL is derived from it and nvidia/cuda:13.0.2-devel-ubuntu22.04 exists. Happy to split this into its own PR if you would rather land it ahead of the migration.

Other bugs found

  • path: under pull_request was a typo for paths:. Unknown keys are ignored, so the filter never applied and this build ran on every PR.
  • The workflow-level HUGGING_FACE_HUB_TOKEN was dead — docker build cannot see host env without --build-arg or --secret.

Testing

https://github.com/pytorch/benchmark/actions/runs/36674543964/job/109756551782?pr=2712

An OSDC runner has no host docker daemon, so `docker build`, `docker tag` and
`docker push` all fail there. OSDC instead runs a per-architecture buildkitd in
every cluster; test-infra's docker-build-remote-buildkit action registers it as
a remote buildx builder and retries the connect-phase failures a cold pool
produces.

Recovering the dev<date> tag is the part that does not port directly. It came
from `docker run`-ing the freshly built image and importing torch, which needs
the daemon. Instead the Dockerfile now writes the version into the image and a
scratch `version-export` stage exposes just that file, so buildx can export it
with --output type=local. A trailing `FROM base AS image` keeps the default
build target the real image rather than the scratch stage, so a plain
`docker build` still behaves.

That does mean two buildx calls. They are cached against a registry
:buildcache, since the pool can hand out a different builder each time and the
second call must not rebuild install.py from scratch.

Also:
- `path:` under pull_request was a typo for `paths:`. GitHub ignores unknown
  keys, so the filter never applied and this build ran on every PR.
- Dropped the workflow-level HUGGING_FACE_HUB_TOKEN. `docker build` cannot see
  host environment without --build-arg or --secret, so it was never reaching
  the build.
- linux.4xlarge -> mt-l-x86iavx512-16-128 per pytorch/pytorch .github/arc.yaml,
  though the runner now only ships the context and waits; the build itself
  happens on buildkitd.

Authored with Claude Code.
The nightly build has been failing since 2026-06-05, independently of where it
runs: uv cannot resolve the install at all.

    error: No solution found when resolving dependencies
      cause: Because there is no version of torch==2.12.0.dev20260407 and all
      versions of torchvision depend on torch==2.12.0.dev20260407 ...

The cu128 nightly index stopped being published in April 2026. What is left in
it is one torch (dev20260408) and torchvisions that pin dev20260407, which was
pruned, so there is no consistent set to resolve to. cu129 is dead the same way
and cu126 has not been updated since early September.

cu130 is the live one:

    torch==2.15.0.dev20260928+cu130
    torchvision==0.30.0.dev20260928+cu130
    torchaudio==2.11.0.dev20260928+cu130
    torchao==0.19.0.dev20260928+cu130

resolved with `uv pip compile --python-platform x86_64-unknown-linux-gnu
--python-version 3.12` against each index; cu128 reproduces the CI error and
cu130 gives the set above. Bumping CUDA_VERSION is the whole change: the index
URL is derived from it, and nvidia/cuda:13.0.2-devel-ubuntu22.04 exists, so the
base image tag needs no edit.

Authored with Claude Code.
@huydhn
huydhn deployed to docker-s3-upload September 30, 2026 05:41 — with GitHub Actions Active
@huydhn huydhn changed the title Build the nightly docker with remote BuildKit on OSDC Build the nightly docker with remote BuildKit on OSDC, and move it to CUDA 13.0 Sep 30, 2026
@meta-codesync

meta-codesync Bot commented Sep 30, 2026

Copy link
Copy Markdown

@huydhn has imported this pull request. If you are a Meta employee, you can view this in D122505504.

@huydhn
huydhn marked this pull request as ready for review September 30, 2026 06:05

This branch had an error being deployed

1 failed deployment
docker-s3-upload — fa89f5d2 Deployed Sep 30, 2026 by huydhn via Test cpu #1690
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant