Conversation
Signed-off-by: Christopher Haar <christopher.haar@upbound.io>
dennis-upbound
left a comment
There was a problem hiding this comment.
The vendor tag needs to cover the metrics exporter, not just the driver.
DCGM ships as an NVIDIA-stack component (nvidia-dcgm, per the allowlist in #456), so tagging it NVIDIA is correct — and it means an AMD-only cluster installs no GPU exporter at all. The telemetry design in #363 names DCGM as the source for temperature, throttling, ECC and interconnect errors, which is the fault taxonomy behind a drain. Those go silent on AMD, and nothing says so.
AMD's GPU Operator ships a Device Metrics Exporter with Prometheus-format temperature, utilization, memory and power, so the fix looks like the driver case:
Chart(key="amd-device-metrics-exporter", accelerator_vendor="AMD", ...)Worth a line in the design saying a vendor's stack includes its exporter. Keeping modelplane_gpu_* meaning the same thing whichever exporter produced it is #363's problem, not this one — I'll take that side.
The Upbound Inference environment runs this stack against real traffic and records what it found. Five things from it. The GPU exporter is vendor-dependent, which modelplaneai#432 makes true rather than hypothetical: DCGM on NVIDIA, the Device Metrics Exporter on AMD, chosen by the serving stack alongside the driver. This document had named DCGM as the source, which would have left an AMD cluster with no GPU series and no statement that it had. modelplane_gpu_* means the same thing whichever exporter produced it, which is the point of renaming at all. Three metrics that environment's dashboard has and this set did not. Termination reason is the important one: a response cut off at the token cap reports length and still returns 200 with a well-formed body, and every layer above reads it as success. A timeout severing a stream looks the same from the other side. That environment has now been truncated twice, once by a route timeout and once by a load balancer, and neither showed in a status code. Tool-call parse outcome fails the same quiet way, returning the unparsed call as text. Tensor-pipe activity is the third, and is closer to useful for inference than the graphics-engine figure beside it. Two traps, both observed rather than reasoned about. An engine container's port is unnamed by default, so selecting it by name matches nothing and reports nothing, which is how that environment's first PodMonitor scraped no engines at all. Modelplane has to name it, and matching by number instead would find the sidecar on a disaggregated pod. And vLLM's model_name is the LLM while DCGM's modelName is the card: one underscore and one capital apart, and a join across them pairs every model with every GPU type. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Dennis Ramdass <dennis@upbound.io>
Description of your changes
The serving stack we propose in #422 is NVIDIA-only by construction: every clouds component list installs the NVIDIA operator or DRA driver, the GPU taint and toleration are
nvidia.com/gpu, and nothing checks that the devices anInferenceClassdeclares match the stack its cluster gets.AMD Instinct GPUs are reaching the clouds Modelplane targets, and TPU and Trainium sit behind the same gap: a cluster may carry devices from more than one vendor, and the stack must install each vendors components exactly where that vendor's devices are.
The design treats the vendor as data that already exists: a class
devices[].drivernames it (gpu.amd.com,gpu.nvidia.com), the cluster composition derives the vendor set onto theServingStack, and the component join filters vendor-tagged components to the vendors present. Mismatches fail early in the composition functions, on the class, on the class/cluster pairing, and in the join. The aicr-generated cloud halves are deliberately untouched until a second vendor is installable on a managed cloud.Design only, no implementation. Feedback wanted on the tag semantics ("part of one vendors device stack", untagged always installs), the per-cloud vendor table and its mirrors in the two functions that need it, and on deferring the generate.py vendor classification.
Fixes #
I have:
nix flake check(or./nix.sh flake check) and made sure it passes.git commit -s.