Conversation
Drops the docker wrapper. On EC2 the runner is a bare host, so the job had to pull ghcr.io/pytorch/torchbench:latest, `docker run` it with the workspace bind-mounted at /tmp/workspace, and `docker exec` the two .ci scripts into it. An OSDC runner has no host docker daemon and every job already runs inside a container, so the same thing is expressed by naming that image in `container:` and running the scripts directly. The image sets PATH to its own /workspace/benchmark/.venv, so `python3` still resolves to the venv. Runners follow pytorch/pytorch's .github/arc.yaml mapping: linux.aws.a100 -> mt-l-x86iavx512-11-125-a100 (1x A100, 11 vCPU / 125 GiB), and linux.24xlarge (c5.24xlarge, 96/192) -> mt-l-x86iavx512-94-192. `--gpus all` comes from the matrix so only the cuda leg gets it. The HF step is the standard read-only-mount dance, same shape as test-infra's linux_job_v3: models come from /mnt/hf_cache, everything HF writes anyway goes to RUNNER_TEMP. Two docker flags are dropped because a pod cannot set them per job: --shm-size=32g (OSDC gives every pod a 2Gi /dev/shm, which is what pytorch's own docker CI runs with) and --cap-add=SYS_PTRACE --security-opt seccomp=unconfined. `environment: docker-s3-upload` stays -- HUGGING_FACE_HUB_TOKEN is scoped to that environment, not the repo, which is also why this cannot use linux_job_v3: a job calling a reusable workflow cannot declare an environment. Authored with Claude Code.
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:34 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:34 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 05:34 — with
GitHub Actions
Failure
huydhn
added a commit
to pytorch/test-infra
that referenced
this pull request
Sep 30, 2026
cache_dir rejected a bare name outright, so seeding anything torchbench needs
was impossible:
ValueError: cannot parse repo id 'albert-base-v2'
gpt2, t5-base, t5-large, albert-base-v2, distilbert-base-uncased and
xlm-roberta-base are all real repo ids in that form, and the hub caches them as
models--<name> with no org segment. That is not the same directory as the
namespaced alias -- models--gpt2 and models--openai-community--gpt2 are
distinct -- so a cache holding one still misses a job asking for the other,
which is exactly how pytorch/benchmark#2711 failed:
OSError: [Errno 30] Read-only file system:
'/mnt/hf_cache/hub/models--albert-base-v2'
with models--albert--albert-base-v2 sitting right there in the same bucket.
--from-hub had to learn the same shape: it unpacked the directory into exactly
three parts, which a bare id does not have.
A bare *dataset* id is still rejected. There is no namespace to copy from, so
that one really is a typo.
The test asserting bare names are rejected was asserting my own bug, so it is
replaced rather than kept.
Authored with Claude Code.
Same shape as pytorch/TensorRT#4745: HF_HOME, HF_HUB_CACHE and HF_DATASETS_CACHE all point at RUNNER_TEMP, so nothing touches the read-only /mnt/hf_cache mount. Reading from the mount is a poor fit here. install.py pulls the whole hf_* model set, and most of those are the legacy un-namespaced ids -- gpt2, t5-base, t5-large, albert-base-v2, distilbert-base-uncased, xlm-roberta-base. Those cache as models--<name>, a different directory from the namespaced alias models--<org>--<name>, so a bucket holding one still misses a job asking for the other. Keeping twelve repos in that exact form in step across four regions is a standing obligation for one workflow, and the failure mode is the bad one: green in whichever region the PR happened to land in, broken on main from the region it did not. Authored with Claude Code.
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 06:20 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 06:20 — with
GitHub Actions
Failure
huydhn
had a problem deploying
to
docker-s3-upload
September 30, 2026 06:20 — with
GitHub Actions
Failure
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Migrates
pr-test.ymloff EC2.docker run/docker execwrapper. OSDC runners have no host docker daemon and every job already runs in a container, soghcr.io/pytorch/torchbench:latestmoves intocontainer:and the two.ci/torchbenchscripts run directly. The image puts its own venv onPATH, sopython3still resolves there.linux.aws.a100→mt-l-x86iavx512-11-125-a100,linux.24xlarge→mt-l-x86iavx512-94-192, per pytorch/pytorch.github/arc.yaml.--gpus allcomes from the matrix so only the cuda leg gets it.RUNNER_TEMPrather than reading the shared/mnt/hf_cachemount, same as Build the Linux wheels on OSDC runners TensorRT#4745.Not using
linux_job_v3becauseHUGGING_FACE_HUB_TOKENis scoped to thedocker-s3-uploadenvironment, and a job that calls a reusable workflow cannot declare an environment.Two flags are dropped, neither settable per job on a pod:
--shm-size=32g— OSDC gives every pod a 2Gi/dev/shm, which is what pytorch's own docker CI runs with.--cap-add=SYS_PTRACE --security-opt seccomp=unconfined.Why not the shared HF cache
install.pypulls the wholehf_*model set, and most of those are the legacy un-namespaced ids —gpt2,t5-base,t5-large,albert-base-v2,distilbert-base-uncased,xlm-roberta-base. Those cache asmodels--<name>, a different directory from the namespaced aliasmodels--<org>--<name>, so a bucket holding one still misses a job asking for the other. Keeping twelve repos in that exact form in step across four regions is a standing obligation for one workflow, and the failure mode is the bad one: green in whichever region the PR landed in, broken on main from the region it did not.Testing
pull_requestruns both legs here.