Conversation
An OSDC runner has no host docker daemon, so `docker build`, `docker tag` and `docker push` all fail there. OSDC instead runs a per-architecture buildkitd in every cluster; test-infra's docker-build-remote-buildkit action registers it as a remote buildx builder and retries the connect-phase failures a cold pool produces. Recovering the dev<date> tag is the part that does not port directly. It came from `docker run`-ing the freshly built image and importing torch, which needs the daemon. Instead the Dockerfile now writes the version into the image and a scratch `version-export` stage exposes just that file, so buildx can export it with --output type=local. A trailing `FROM base AS image` keeps the default build target the real image rather than the scratch stage, so a plain `docker build` still behaves. That does mean two buildx calls. They are cached against a registry :buildcache, since the pool can hand out a different builder each time and the second call must not rebuild install.py from scratch. Also: - `path:` under pull_request was a typo for `paths:`. GitHub ignores unknown keys, so the filter never applied and this build ran on every PR. - Dropped the workflow-level HUGGING_FACE_HUB_TOKEN. `docker build` cannot see host environment without --build-arg or --secret, so it was never reaching the build. - linux.4xlarge -> mt-l-x86iavx512-16-128 per pytorch/pytorch .github/arc.yaml, though the runner now only ships the context and waits; the build itself happens on buildkitd. Authored with Claude Code.
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:33 — with
GitHub Actions
Error
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:33 — with
GitHub Actions
Error
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:33 — with
GitHub Actions
Failure
The nightly build has been failing since 2026-06-05, independently of where it
runs: uv cannot resolve the install at all.
error: No solution found when resolving dependencies
cause: Because there is no version of torch==2.12.0.dev20260407 and all
versions of torchvision depend on torch==2.12.0.dev20260407 ...
The cu128 nightly index stopped being published in April 2026. What is left in
it is one torch (dev20260408) and torchvisions that pin dev20260407, which was
pruned, so there is no consistent set to resolve to. cu129 is dead the same way
and cu126 has not been updated since early September.
cu130 is the live one:
torch==2.15.0.dev20260928+cu130
torchvision==0.30.0.dev20260928+cu130
torchaudio==2.11.0.dev20260928+cu130
torchao==0.19.0.dev20260928+cu130
resolved with `uv pip compile --python-platform x86_64-unknown-linux-gnu
--python-version 3.12` against each index; cu128 reproduces the CI error and
cu130 gives the set above. Bumping CUDA_VERSION is the whole change: the index
URL is derived from it, and nvidia/cuda:13.0.2-devel-ubuntu22.04 exists, so the
base image tag needs no edit.
Authored with Claude Code.
huydhn
requested a deployment
to
docker-s3-upload
September 30, 2026 05:41 — with
GitHub Actions
Queued
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:41 — with
GitHub Actions
Failure
|
@huydhn has imported this pull request. If you are a Meta employee, you can view this in D122505504. |
huydhn
marked this pull request as ready for review
September 30, 2026 06:05
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Migrates
build-nightly-docker.ymloff EC2.docker build/tag/pushall need a daemon, which OSDC runners do not have; the build moves to the remote BuildKit pool via test-infra'sdocker-build-remote-buildkitaction.dev<date>tag came fromdocker run-ing the built image and importing torch. Instead the Dockerfile writes the version into the image and ascratchversion-exportstage exposes just that file, so buildx can pull it out with--output type=local. A trailingFROM base AS imagekeeps the default build target the real image, so a plaindocker buildstill works.:buildcache— the pool can hand out a different builder each time and the second call must not rebuildinstall.py.linux.4xlarge→mt-l-x86iavx512-16-128. The runner now only ships the context and waits; the build runs on buildkitd.CUDA 13.0
Separate, pre-existing bug — this nightly has been failing on main since 2026-06-05, and would fail the same way on EC2. The cu128 index stopped being published in April 2026; all that is left is one torch (
dev20260408) and torchvisions pinningdev20260407, which was pruned, so nothing resolves. cu129 is dead the same way, cu126 is stale since early September. cu130 is live and resolves to a consistentdev20260928set.Bumping
CUDA_VERSIONis the whole change — the index URL is derived from it andnvidia/cuda:13.0.2-devel-ubuntu22.04exists. Happy to split this into its own PR if you would rather land it ahead of the migration.Other bugs found
path:underpull_requestwas a typo forpaths:. Unknown keys are ignored, so the filter never applied and this build ran on every PR.HUGGING_FACE_HUB_TOKENwas dead —docker buildcannot see host env without--build-argor--secret.Testing
https://github.com/pytorch/benchmark/actions/runs/36674543964/job/109756551782?pr=2712