Skip to content

Run the PR test on OSDC runners - #2711

Draft
huydhn wants to merge 2 commits into
mainfrom
osdc/pr-test-arc
Draft

huydhn wants to merge 2 commits into
mainfrom
osdc/pr-test-arc

Conversation

@huydhn

@huydhn huydhn commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Migrates pr-test.yml off EC2.

  • Drops the docker run / docker exec wrapper. OSDC runners have no host docker daemon and every job already runs in a container, so ghcr.io/pytorch/torchbench:latest moves into container: and the two .ci/torchbench scripts run directly. The image puts its own venv on PATH, so python3 still resolves there.
  • linux.aws.a100 → mt-l-x86iavx512-11-125-a100, linux.24xlarge → mt-l-x86iavx512-94-192, per pytorch/pytorch .github/arc.yaml. --gpus all comes from the matrix so only the cuda leg gets it.
  • HF models download fresh into RUNNER_TEMP rather than reading the shared /mnt/hf_cache mount, same as Build the Linux wheels on OSDC runners TensorRT#4745.

Not using linux_job_v3 because HUGGING_FACE_HUB_TOKEN is scoped to the docker-s3-upload environment, and a job that calls a reusable workflow cannot declare an environment.

Two flags are dropped, neither settable per job on a pod:

  • --shm-size=32g — OSDC gives every pod a 2Gi /dev/shm, which is what pytorch's own docker CI runs with.
  • --cap-add=SYS_PTRACE --security-opt seccomp=unconfined.

Why not the shared HF cache

install.py pulls the whole hf_* model set, and most of those are the legacy un-namespaced ids — gpt2, t5-base, t5-large, albert-base-v2, distilbert-base-uncased, xlm-roberta-base. Those cache as models--<name>, a different directory from the namespaced alias models--<org>--<name>, so a bucket holding one still misses a job asking for the other. Keeping twelve repos in that exact form in step across four regions is a standing obligation for one workflow, and the failure mode is the bad one: green in whichever region the PR landed in, broken on main from the region it did not.

Testing

pull_request runs both legs here.

Drops the docker wrapper. On EC2 the runner is a bare host, so the job had to
pull ghcr.io/pytorch/torchbench:latest, `docker run` it with the workspace
bind-mounted at /tmp/workspace, and `docker exec` the two .ci scripts into it.
An OSDC runner has no host docker daemon and every job already runs inside a
container, so the same thing is expressed by naming that image in `container:`
and running the scripts directly. The image sets PATH to its own
/workspace/benchmark/.venv, so `python3` still resolves to the venv.

Runners follow pytorch/pytorch's .github/arc.yaml mapping: linux.aws.a100 ->
mt-l-x86iavx512-11-125-a100 (1x A100, 11 vCPU / 125 GiB), and linux.24xlarge
(c5.24xlarge, 96/192) -> mt-l-x86iavx512-94-192. `--gpus all` comes from the
matrix so only the cuda leg gets it.

The HF step is the standard read-only-mount dance, same shape as
test-infra's linux_job_v3: models come from /mnt/hf_cache, everything HF
writes anyway goes to RUNNER_TEMP.

Two docker flags are dropped because a pod cannot set them per job:
--shm-size=32g (OSDC gives every pod a 2Gi /dev/shm, which is what pytorch's
own docker CI runs with) and --cap-add=SYS_PTRACE --security-opt
seccomp=unconfined.

`environment: docker-s3-upload` stays -- HUGGING_FACE_HUB_TOKEN is scoped to
that environment, not the repo, which is also why this cannot use
linux_job_v3: a job calling a reusable workflow cannot declare an environment.

Authored with Claude Code.
@meta-cla meta-cla Bot added the cla signed label Sep 30, 2026
huydhn added a commit to pytorch/test-infra that referenced this pull request Sep 30, 2026
cache_dir rejected a bare name outright, so seeding anything torchbench needs
was impossible:

    ValueError: cannot parse repo id 'albert-base-v2'

gpt2, t5-base, t5-large, albert-base-v2, distilbert-base-uncased and
xlm-roberta-base are all real repo ids in that form, and the hub caches them as
models--<name> with no org segment. That is not the same directory as the
namespaced alias -- models--gpt2 and models--openai-community--gpt2 are
distinct -- so a cache holding one still misses a job asking for the other,
which is exactly how pytorch/benchmark#2711 failed:

    OSError: [Errno 30] Read-only file system:
        '/mnt/hf_cache/hub/models--albert-base-v2'

with models--albert--albert-base-v2 sitting right there in the same bucket.

--from-hub had to learn the same shape: it unpacked the directory into exactly
three parts, which a bare id does not have.

A bare *dataset* id is still rejected. There is no namespace to copy from, so
that one really is a typo.

The test asserting bare names are rejected was asserting my own bug, so it is
replaced rather than kept.

Authored with Claude Code.
Same shape as pytorch/TensorRT#4745: HF_HOME, HF_HUB_CACHE and
HF_DATASETS_CACHE all point at RUNNER_TEMP, so nothing touches the read-only
/mnt/hf_cache mount.

Reading from the mount is a poor fit here. install.py pulls the whole hf_*
model set, and most of those are the legacy un-namespaced ids -- gpt2,
t5-base, t5-large, albert-base-v2, distilbert-base-uncased, xlm-roberta-base.
Those cache as models--<name>, a different directory from the namespaced alias
models--<org>--<name>, so a bucket holding one still misses a job asking for
the other. Keeping twelve repos in that exact form in step across four regions
is a standing obligation for one workflow, and the failure mode is the bad one:
green in whichever region the PR happened to land in, broken on main from the
region it did not.

Authored with Claude Code.

This branch had an error being deployed

1 failed deployment
docker-s3-upload — 92a0846d Deployed Sep 30, 2026 by huydhn via Test cpu #1691
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant