Skip to content

Design: multiple accelerator vendors in the serving stack - #432

Open
haarchri wants to merge 1 commit into
modelplaneai:mainfrom
haarchri:design/serving-stack-accelerator-vendors
Open

haarchri wants to merge 1 commit into
modelplaneai:mainfrom
haarchri:design/serving-stack-accelerator-vendors

Conversation

@haarchri

@haarchri haarchri commented Sep 8, 2026

Copy link
Copy Markdown
Collaborator

Description of your changes

The serving stack we propose in #422 is NVIDIA-only by construction: every clouds component list installs the NVIDIA operator or DRA driver, the GPU taint and toleration are nvidia.com/gpu, and nothing checks that the devices an InferenceClass declares match the stack its cluster gets.
AMD Instinct GPUs are reaching the clouds Modelplane targets, and TPU and Trainium sit behind the same gap: a cluster may carry devices from more than one vendor, and the stack must install each vendors components exactly where that vendor's devices are.

The design treats the vendor as data that already exists: a class devices[].driver names it (gpu.amd.com, gpu.nvidia.com), the cluster composition derives the vendor set onto the ServingStack, and the component join filters vendor-tagged components to the vendors present. Mismatches fail early in the composition functions, on the class, on the class/cluster pairing, and in the join. The aicr-generated cloud halves are deliberately untouched until a second vendor is installable on a managed cloud.

Design only, no implementation. Feedback wanted on the tag semantics ("part of one vendors device stack", untagged always installs), the per-cloud vendor table and its mirrors in the two functions that need it, and on deferring the generate.py vendor classification.

Fixes #

I have:

  • Read and followed Modelplane's contribution process.
  • Run nix flake check (or ./nix.sh flake check) and made sure it passes.
  • Added or updated tests covering any composition function changes.
  • Signed off every commit with git commit -s.

Signed-off-by: Christopher Haar <christopher.haar@upbound.io>

@dennis-upbound dennis-upbound left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The vendor tag needs to cover the metrics exporter, not just the driver.

DCGM ships as an NVIDIA-stack component (nvidia-dcgm, per the allowlist in #456), so tagging it NVIDIA is correct — and it means an AMD-only cluster installs no GPU exporter at all. The telemetry design in #363 names DCGM as the source for temperature, throttling, ECC and interconnect errors, which is the fault taxonomy behind a drain. Those go silent on AMD, and nothing says so.

AMD's GPU Operator ships a Device Metrics Exporter with Prometheus-format temperature, utilization, memory and power, so the fix looks like the driver case:

Chart(key="amd-device-metrics-exporter", accelerator_vendor="AMD", ...)

Worth a line in the design saying a vendor's stack includes its exporter. Keeping modelplane_gpu_* meaning the same thing whichever exporter produced it is #363's problem, not this one — I'll take that side.

dennis-upbound added a commit to dennis-upbound/modelplane that referenced this pull request Sep 22, 2026
The Upbound Inference environment runs this stack against real traffic
and records what it found. Five things from it.

The GPU exporter is vendor-dependent, which modelplaneai#432 makes true rather than
hypothetical: DCGM on NVIDIA, the Device Metrics Exporter on AMD, chosen
by the serving stack alongside the driver. This document had named DCGM
as the source, which would have left an AMD cluster with no GPU series
and no statement that it had. modelplane_gpu_* means the same thing
whichever exporter produced it, which is the point of renaming at all.

Three metrics that environment's dashboard has and this set did not.
Termination reason is the important one: a response cut off at the token
cap reports length and still returns 200 with a well-formed body, and
every layer above reads it as success. A timeout severing a stream looks
the same from the other side. That environment has now been truncated
twice, once by a route timeout and once by a load balancer, and neither
showed in a status code. Tool-call parse outcome fails the same quiet
way, returning the unparsed call as text. Tensor-pipe activity is the
third, and is closer to useful for inference than the graphics-engine
figure beside it.

Two traps, both observed rather than reasoned about. An engine
container's port is unnamed by default, so selecting it by name matches
nothing and reports nothing, which is how that environment's first
PodMonitor scraped no engines at all. Modelplane has to name it, and
matching by number instead would find the sidecar on a disaggregated
pod. And vLLM's model_name is the LLM while DCGM's modelName is the
card: one underscore and one capital apart, and a join across them pairs
every model with every GPU type.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Dennis Ramdass <dennis@upbound.io>

This branch was successfully deployed

1 active deployment
Preview – modelplane-docs — 3fe0dab3 Deployed Sep 8, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants