Conversation
haarchri
force-pushed
the
feature/baremetal-amd
branch
from
September 9, 2026 13:28
06e3317 to
03d64b9
Compare
haarchri
force-pushed
the
feature/baremetal-amd
branch
from
September 30, 2026 18:48
86b5fc5 to
0061a2a
Compare
…ence clusters, composing multi-vendor GPU serving stacks Signed-off-by: Christopher Haar <christopher.haar@upbound.io>
haarchri
force-pushed
the
feature/baremetal-amd
branch
from
September 30, 2026 19:10
43a8047 to
c19401e
Compare
haarchri
marked this pull request as ready for review
September 30, 2026 19:12
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description of your changes
Design: #432
Vultr' biggest GPUs (8x MI355X/MI325X) are bare metal plans with no managedKubernetes in front, and Modelplane could neither provision metal nor install an AMD GPU stack.
A new K3sCluster XR turns machines into a cluster: it installs a k3s server on the control plane machine and joins workers as agents over SSH via provider-k3s, publishing the kubeconfig through the same
status.secretscontract as the other cluster XRs. AVultrBaremetalClusterXR provisions theservers on top, one CPU-only management server plus the GPU pools, SSH key registration, and cloud-init opening the k3s ports (Vultrs Ubuntu imagesfirewall everything but SSH) and composes the K3sCluster from their IPs. pools are fixed size and only the Standard stack is supported :)
InferenceCluster gains a VultrBaremetal source with an activation policy covering the Vultr bare metal and k3s provider kinds.
The serving stack implements the multi-accelerator design (#432): the VultrBaremetal component list carries both the AMD GPU operator (DRA mode) and the NVIDIA operator + DRA driver, vendor-tagged and filtered by the accelerators the InferenceClasses name. Unsupported class/cloud pairings fail early with UnsupportedDevices conditions.
Validated end to end on Vultr with a
vbm-256c-3072gb-8-mi325x-gpuplan (8x MI325X): servers provisioned, k3s v1.34.11 installed over SSH, the worker joined with pool labels and vendor taint, the serving stack installed, the AMDDRA driver published ResourceSlices (8 devices, 262128Mi each), and a vLLM engine (rocm/vllm, CDNA build) served a chat completion through the gateway with the GPU bound via a DRA claim. That run is recorded as the Qwen2.5-0.5B recipe, the first AMD entry in the catalog - with the exact manifests from the run.
Deliberately not surfaced as a quickstart: bare metal GPU plans bill thousands of dollars per month, are region-gated, and are often out of stock (MI355X currently has no deployable locations at all), so the flow ships as
examples/vultr-baremetalplus the recipe rather than a getting-started path.The management server is a single k3s server and the data plane rides public IPs...
I have:
nix flake check(or./nix.sh flake check) and made sure it passes.git commit -s.Tested: