From 383c3394620eb4c3bebd88f99f4fd5a678ffdc2b Mon Sep 17 00:00:00 2001 From: spinloop-agent Date: Mon, 5 Oct 2026 20:20:15 +0100 Subject: [PATCH 1/3] docs(openspec): propose macOS GPU utilisation in metrics --- .../macos-gpu-utilisation/.openspec.yaml | 2 + .../changes/macos-gpu-utilisation/design.md | 98 ++++++++++++++++ .../changes/macos-gpu-utilisation/proposal.md | 59 ++++++++++ .../specs/engine-metrics/spec.md | 110 ++++++++++++++++++ .../specs/remote-metrics-bar-format/spec.md | 56 +++++++++ .../changes/macos-gpu-utilisation/tasks.md | 22 ++++ 6 files changed, 347 insertions(+) create mode 100644 openspec/changes/macos-gpu-utilisation/.openspec.yaml create mode 100644 openspec/changes/macos-gpu-utilisation/design.md create mode 100644 openspec/changes/macos-gpu-utilisation/proposal.md create mode 100644 openspec/changes/macos-gpu-utilisation/specs/engine-metrics/spec.md create mode 100644 openspec/changes/macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md create mode 100644 openspec/changes/macos-gpu-utilisation/tasks.md diff --git a/openspec/changes/macos-gpu-utilisation/.openspec.yaml b/openspec/changes/macos-gpu-utilisation/.openspec.yaml new file mode 100644 index 00000000..0ca5fbe8 --- /dev/null +++ b/openspec/changes/macos-gpu-utilisation/.openspec.yaml @@ -0,0 +1,2 @@ +schema: spec-driven +created: 2026-09-26 diff --git a/openspec/changes/macos-gpu-utilisation/design.md b/openspec/changes/macos-gpu-utilisation/design.md new file mode 100644 index 00000000..c00d2206 --- /dev/null +++ b/openspec/changes/macos-gpu-utilisation/design.md @@ -0,0 +1,98 @@ +# Design + +## Context + +`internal/metrics.Collector` is a host-command collector with an injectable +runner and platform: each stat is one command plus one pure parser +(`nvidia-smi`/`vmstat`/`free` on Linux, `top`/`sysctl`/`vm_stat` on macOS), +and every figure is optional — a missing source is omitted silently, a +failing one is reported (`engine-metrics` spec). The GPU slot is the only +one with no macOS branch: `gpus()` returns `nil, nil` off Linux. + +Verified on the development host (Apple M5 Max, macOS): +`ioreg -rd1 -c IOAccelerator -w 0` answers quickly with the AGX accelerator +node carrying + +``` +"PerformanceStatistics" = {... "Renderer Utilization %"=3, + "Device Utilization %"=91, "TiledSceneBytes"=..., ...} +"model" = "Apple M5 Max" +``` + +and no GPU memory total (unified memory) and no temperature. A host with no +accelerator service answers with empty output and exit 0. + +## Goals / Non-Goals + +**Goals:** +- One utilisation figure per accelerator, per sampler tick, at the same + trust level as the existing figures (fixture-testable, no privileges). +- Absent GPU memory/temperature stay absent through collection, history, + and every renderer. + +**Non-Goals:** +- GPU memory attribution on Apple Silicon (unified memory has no per-GPU + total; any denominator would be the RAM bar restated). +- Per-process GPU attribution, MACb/ANE figures, Intel-Mac-specific + validation, and the remote Lambda path (Linux-only, unchanged). + +## Decisions + +**Source: `ioreg -rd1 -c IOAccelerator -w 0` over the alternatives.** +`top` has no GPU line on current macOS; `powermetrics` needs root; a direct +IOKit call needs cgo, which would break the pure-Go leaf and cross-compile. +`ioreg` matches the package's existing shape (host command + text parser + +injectable runner) and is already the standard trick GPU monitors use. +`-w 0` turns line-wrapping off so the dictionary parses intact when stdout +is a pipe, not a tty. + +**Parse the text, don't switch to plist XML.** A line-oriented scan of each +service block (`+-o …` headers) for `"PerformanceStatistics"`, then +`"Device Utilization %"`, falls out of the same fixture style as +`ParseGPUStats`; `-a` XML would need a plist decoder for three keys. +Per block: utilisation from `Device Utilization %` (the device-wide figure — +`Renderer`/`Tiler` are per-engine halves, not summed); name from `"model"`, +falling back to the service name; index = order of appearance, so the rare +multi-accelerator Mac still yields one `GpuStat` each with the labels the +formats already derive from `Index`. The parser also accepts the AMD +drivers' `"Device Utilization (%)"` inside a `"Performance Statistics"` +(spaced) dictionary as a fallback for Intel-Mac AMD cards — cheap, but not +verifiable on this host. A block with neither key contributes no GPU. +Empty output parses to `nil, nil` — the silent-absent case; a non-zero exit +is a present-and-failing error and surfaces as `gpu: …`, exactly like +`nvidia-smi` today. + +**Absence in the renderers follows the shape's own rule.** `currentGPUMem` +returns nil (not `0`) when `MemoryTotal == 0`, so `barSeriesList` drops the +series when no retained sample carried a total either — the GPU then draws +its util line alone on every bar/gauge surface, dashboard tiles included. +`renderGPUTable` builds its `mem=`/`temp=` segments conditionally, and its +multi-GPU totals line guards the memory ratio against a summed-zero total. +Treating `Temperature == 0` as absent is safe: real GPU readings under load +never sit at zero. + +**Sampling cost stays where the sampler already absorbs it.** The daemon +collects on its background tick, never inline in a handler, so the +accelerator dump (tens of KB, sub-second, on a loaded host — same precedent +as macOS CPU's `top -l 1`, see the `systemSample` comment) is invisible to +API latency. No cadence change. + +## Risks / Trade-offs + +- Apple could rename the dictionary or keys in a future macOS → the parser + yields no GPU and the host degrades exactly as it does today: silently + absent, not an error. +- `Device Utilization %` is what the driver samples; it may lag a + sub-tick burst. Acceptable at a 15s cadence, where every existing figure + is equally a sample. +- Dropping the zeroed `GPU mem` series changes existing gauge output on any + hypothetical source that reports a zero total today; the new specs adopt + that case deliberately (a zero bar for an unknown total was the lie). + +## Migration Plan + +No migration: additive on the collection side, and the render change is +covered by the two delta specs. Stale `GpuStat` wire data from older daemons +(NVIDIA) carries totals and renders unchanged. Update the now-false comments +naming this gap (`collect.go` header and `gpus()`, the `metrics` package +doc) as part of the implementation. diff --git a/openspec/changes/macos-gpu-utilisation/proposal.md b/openspec/changes/macos-gpu-utilisation/proposal.md new file mode 100644 index 00000000..3d87a463 --- /dev/null +++ b/openspec/changes/macos-gpu-utilisation/proposal.md @@ -0,0 +1,59 @@ +# Proposal + +## Why + +GPU utilisation is collected only from `nvidia-smi`, so a macOS host — the +ordinary home of a local `spinloop serve`/daemon node, and where most harness +users run — reports engine stats and CPU/RAM with no GPU figure at all +(issue #217; the collector itself flags the gap in `internal/metrics`). The +metrics plumbing already treats every GPU figure as optional, so filling the +gap is one collector branch plus one parser. + +## What Changes + +- On macOS the collector gains a GPU source: the IORegistry accelerator node + (via `ioreg`), from which it reports GPU utilisation and the GPU's name + (e.g. "Apple M5 Max"). Absent an accelerator node (a VM, say), the GPU + stat stays absent and silent, as it is today. +- GPU memory and temperature are not reported on macOS — Apple Silicon has + unified memory with no per-GPU total, and temperature needs root — so the + renderers stop presenting an absent figure as a zero: no "GPU mem" series + where no GPU ever reports a total, and the table format's `mem=`/`temp=` + segments are omitted rather than drawn `0B/0B`/`0C`. +- History, bar/gauge/table formats, the dashboard, and the wire shape are + otherwise unchanged: existing series and formats pick the GPU up with no + format-specific work. + +## Capabilities + +### New Capabilities + +_None._ + +### Modified Capabilities + +- `engine-metrics`: "System stats collection" gains a macOS GPU source + (utilisation and name from the IORegistry accelerator; memory/temperature + omitted where the host reports none). "Graceful platform degradation" + replaces the scenario that a macOS host lacks GPU stats — macOS now has a + GPU source — with the host-without-an-accelerator case staying silent. +- `remote-metrics-bar-format`: the GPU memory series (bar and gauge) is drawn + only where the reading or a retained sample reports a GPU memory total; a + utilisation-only GPU draws its util series alone. Table GPU lines omit the + memory and temperature segments when the reading carries none. + +## Impact + +- `internal/metrics`: `collect.go` darwin branch in `gpus()`, a new IORegistry + parser in `parse.go`; fixtures in `metrics_test.go`. +- `cmd/spinloop/metrics_render.go`: `currentGPUMem`/`barSeriesList` series + selection and `renderGPUTable` segment omission; dashboard and serve view + inherit through the shared renderers. +- `docs/maintainer/internals.md` note on the sampling cost precedent if + warranted; no API, OpenAPI, or wire-shape change (`GpuStat` fields already + carry the zero-means-absent convention). + +Assumption (from the issue's scope, "Sample the GPU util"): utilisation and +name only on macOS; GPU memory and temperature stay absent. Intel/AMD macOS +spellings of the IORegistry keys are accepted by the parser as a cheap +fallback but are out of testable scope here. diff --git a/openspec/changes/macos-gpu-utilisation/specs/engine-metrics/spec.md b/openspec/changes/macos-gpu-utilisation/specs/engine-metrics/spec.md new file mode 100644 index 00000000..79a30ae5 --- /dev/null +++ b/openspec/changes/macos-gpu-utilisation/specs/engine-metrics/spec.md @@ -0,0 +1,110 @@ +# Spec Delta + +## MODIFIED Requirements + +### Requirement: System stats collection + +The system SHALL collect system statistics from the host: GPU utilization, +GPU memory used/total, CPU utilization, and RAM used/total. On hosts with +NVIDIA GPUs the GPU figures SHALL be sourced from `nvidia-smi`. On macOS +hosts the GPU SHALL be sourced from the host's I/O Kit accelerator service, +which yields GPU utilization and the GPU's name; where that service reports +no memory total or no temperature — the ordinary case on a host with +unified memory — those figures SHALL be absent rather than reported as +zero. The collected values SHALL use the same units as the existing remote +stats pipeline (bytes for memory, percentages for utilization) so existing +rendering applies unchanged. + +#### Scenario: NVIDIA host reports GPU stats + +- **WHEN** metrics are collected on a host where `nvidia-smi` is available +- **THEN** the result includes GPU utilization and GPU memory used/total in + bytes + +#### Scenario: macOS host reports GPU utilisation + +- **WHEN** metrics are collected on a macOS host whose accelerator service + reports device utilization +- **THEN** the result includes a GPU with that utilization percentage and + the accelerator's model name, and omits GPU memory and temperature rather + than reporting them as zero + +#### Scenario: CPU and RAM are always attempted + +- **WHEN** metrics are collected on any supported host +- **THEN** the result includes CPU utilization and RAM used/total when the + platform provides them + +### Requirement: Graceful platform degradation + +When a system stat's source is unavailable on the host (for example +`nvidia-smi` on a machine without NVIDIA tooling, or no accelerator service +at all on a virtualized guest), the collector SHALL omit that stat and +return the remainder, rather than failing the collection. The absence SHALL +be distinguishable from a zero value in the collected result. + +A source that is *present and failing* SHALL be distinguished from one that +is absent. Where the collector has an address to query and the query fails, +it SHALL report that failure among the collected errors, naming what it +tried, so a misdirected or broken collector is visible rather than +presenting as an engine that has simply served nothing. An absent source +SHALL remain silent: reporting the routine absence of a source as an error +would bury the failures worth seeing. + +#### Scenario: macOS host lacks GPU stats + +- **WHEN** metrics are collected on a macOS host that names no I/O Kit + accelerator service, such as a virtualized guest +- **THEN** the result includes engine stats and available CPU/RAM figures, + omits GPU stats, and reports no error + +#### Scenario: Missing command omits only its section + +- **WHEN** one system stat source is missing but others are present +- **THEN** only the missing stat is absent from the result + +#### Scenario: A failing scrape is reported, not hidden + +- **WHEN** the collector has an engine address to query and the query fails +- **THEN** the result omits the engine's counters and reports an error naming + the address it tried + +#### Scenario: An engine with no metrics endpoint stays silent + +- **WHEN** the engine exposes no metrics endpoint, so there is no address to + query +- **THEN** the result omits the engine's counters and reports no error + +## ADDED Requirements + +### Requirement: Absent GPU figures render as absent + +A GPU reading that carries no memory total or no temperature SHALL render +without those figures, rather than as zeros: a resource-series view SHALL +draw no memory series for a GPU whose current reading and retained history +both report no memory total, and a table GPU line SHALL omit its memory and +temperature segments when the reading carries none. The GPU's utilization +SHALL still render. This keeps the shape's own rule — an absent figure is +not a zero — true of the rendered output as it is of the collected one. + +#### Scenario: A utilisation-only GPU draws one series + +- **WHEN** a bar or gauge view renders a GPU whose readings never carry a + memory total +- **THEN** the output includes the GPU utilization series and no GPU memory + series + +#### Scenario: A GPU with memory still draws both series + +- **WHEN** a bar or gauge view renders a GPU whose current reading or any + retained sample carries a memory total +- **THEN** the output includes both the utilization and the memory series, + as before + +#### Scenario: Table omits the segments a reading lacks + +- **WHEN** the table format renders a GPU line whose reading carries no + memory total and no temperature +- **THEN** the line shows the GPU's name and utilization with no `mem=` and + no `temp=` segment, and a GPU reading that does carry them shows them as + before diff --git a/openspec/changes/macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md b/openspec/changes/macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md new file mode 100644 index 00000000..72e68b7e --- /dev/null +++ b/openspec/changes/macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md @@ -0,0 +1,56 @@ +# Spec Delta + +## MODIFIED Requirements + +### Requirement: Bar format output + +The system SHALL support a `--format=bar` option that renders each resource series as a sparkline drawn from the history the on-instance daemon retains: a left-aligned label, one glyph per sample, and the latest value as a right-aligned percentage. The glyphs SHALL be Unicode block elements of one grade per utilisation level, so a series reads as a line of bars across the window. The series drawn SHALL be the same set the gauge format draws: CPU, RAM, each GPU's utilisation, and each GPU's memory where that GPU's current reading or retained history reports a memory total — a GPU that never reports one draws no memory series, with the same per-GPU labelling throughout. + +#### Scenario: Bar format displays CPU utilization + +- **WHEN** the user runs `spinloop remote metrics --format=bar` with a running instance that has CPU data and a retained history +- **THEN** the output includes a row labelled "CPU" whose glyphs are the sampled CPU utilisation across the window and whose trailing figure is the latest sample's percentage + +#### Scenario: Bar format displays RAM utilization + +- **WHEN** the user runs `spinloop remote metrics --format=bar` with a running instance that has memory data +- **THEN** the output includes a row labelled "RAM" whose glyphs are the sampled used/total memory ratio across the window and whose trailing figure is the latest ratio + +#### Scenario: Bar format displays GPU utilization + +- **WHEN** the user runs `spinloop remote metrics --format=bar` with a running instance that has GPU data whose readings carry a memory total +- **THEN** the output includes rows labelled "GPU util" and "GPU mem" (or "GPU N util"/"GPU N mem" for multiple GPUs), each drawn from the retained history + +#### Scenario: Bar format omits a GPU memory series with no total + +- **WHEN** the user runs `spinloop remote metrics --format=bar` with GPU data whose current reading and retained history report no memory total +- **THEN** the output includes "GPU util" and no "GPU mem" row + +#### Scenario: Bar format header line + +- **WHEN** the user runs `spinloop remote metrics --format=bar` with a running instance +- **THEN** the first line shows the environment, state, instance type, and model ID separated by double spaces + +### Requirement: Gauge format + +The system SHALL support a `--format=gauge` option that renders each resource series as a horizontal progress gauge: a left-aligned label, a filled portion using block characters, an unfilled portion using light shade characters, and a right-aligned percentage value. The gauge draws the current reading only — it carries no history. The series drawn SHALL be CPU, RAM, each GPU's utilisation, and each GPU's memory where the current reading reports a memory total — a GPU that reports none draws no memory gauge — with the same labels the bar format uses. The gauge fill SHALL be colour-coded on the bar format's thresholds: green for values at or below 80%, yellow for values from 80% to 90%, and red for values above 90%, with the colour reset after the filled portion so the unfilled characters and percentage appear in the terminal's default colour. + +#### Scenario: Gauge format displays CPU utilization + +- **WHEN** the user runs `spinloop remote metrics --format=gauge` with a running instance that has CPU data +- **THEN** the output includes a gauge labelled "CPU" with filled and unfilled segments proportional to the current utilization + +#### Scenario: Gauge format displays GPU utilisation + +- **WHEN** the user runs `spinloop remote metrics --format=gauge` with a running instance that has GPU data whose readings carry a memory total +- **THEN** the output includes gauges labelled "GPU util" and "GPU mem" (or "GPU N util"/"GPU N mem" for multiple GPUs) + +#### Scenario: Gauge omits a GPU memory gauge with no total + +- **WHEN** the user runs `spinloop remote metrics --format=gauge` with a GPU whose current reading reports no memory total +- **THEN** the output includes "GPU util" and no "GPU mem" gauge + +#### Scenario: Gauge colours the fill + +- **WHEN** a gauge's current value is 95% +- **THEN** its filled segment appears in red, and its unfilled segment and percentage appear in the terminal's default colour diff --git a/openspec/changes/macos-gpu-utilisation/tasks.md b/openspec/changes/macos-gpu-utilisation/tasks.md new file mode 100644 index 00000000..b12e85a4 --- /dev/null +++ b/openspec/changes/macos-gpu-utilisation/tasks.md @@ -0,0 +1,22 @@ +# Tasks + +## 1. macOS GPU collection (internal/metrics) + +- [ ] 1.1 Add `ParseIOAcceleratorGPU(out string) []GpuStat` to `internal/metrics/parse.go`: scan `+-o` service blocks for `PerformanceStatistics`/`Performance Statistics` dictionaries, read `Device Utilization %` (accepting the AMD `Device Utilization (%)` spelling), name from `"model"` with the service name as fallback, index by order of appearance, memory and temperature left zero; verify with parser tests covering an Apple AGX fixture, an AMD-spelling fixture, a two-accelerator fixture, a block lacking the utilization key, and empty output yielding no GPUs. +- [ ] 1.2 Give `Collector.gpus()` a darwin branch running `ioreg -rd1 -c IOAccelerator -w 0` through the injectable runner: parsed GPUs on success, silent nil on empty output, reported `gpu:` error on a non-zero exit; verify with collector tests feeding an `ioreg` fixture through the runner map, a silent-absence test, and an exit-failure test, keeping the Linux and CPU/RAM collector tests green. +- [ ] 1.3 Update the comments that record the gap — the `internal/metrics` package doc, the `Collector` header, and the `gpus()` "issue #47" note — to describe the macOS accelerator source and its absent memory/temperature; verify with `go vet ./internal/metrics` and `go test ./internal/metrics/...` passing. + +## 2. Rendering absent GPU figures (cmd/spinloop) + +- [ ] 2.1 Make `currentGPUMem` in `cmd/spinloop/metrics_render.go` return nil when `MemoryTotal == 0`, so `barSeriesList` drops a GPU's memory series only when no current reading and no retained sample carries a total; verify with tests that a utilisation-only GPU yields `GPU util` with no `GPU mem` in bar, combined, and gauge surfaces, and that an NVIDIA-shaped reading still yields both. +- [ ] 2.2 Make `renderGPUTable` emit `mem=` and `temp=` segments only when the reading carries a total and a temperature respectively, and the multi-GPU totals line drop its `total mem:` figure when the summed total is zero; verify with table-format tests for a utilisation-only GPU, a full NVIDIA-shaped GPU, and a mixed pair keeping the totals line whole. +- [ ] 2.3 Run the dashboard, serve-view, fleet, and metrics-render test suites and fix any golden outputs that expected the zeroed `GPU mem` line or `0B/0B` segments, confirming the shared renderers needed no format-specific change in dashboard or fleet code. + +## 3. Documentation + +- [ ] 3.1 Grep `docs/` for claims that GPU figures are Linux/NVIDIA-only or that macOS nodes show no GPU, and update the affected pages (metrics/dashboard guides) to state that macOS reports utilisation only; verify each edited command or output sample still matches the code's actual output. + +## 4. Integration checks + +- [ ] 4.1 Run `gofmt -l .`, `go vet ./...`, and `go test ./... -cover`, verifying a clean format, a clean vet, and total coverage at or above 80%. +- [ ] 4.2 Build the binary (`go build -o spinloop ./cmd/spinloop`) and run `spinloop metrics --format=bar` and `--format=table` against a local macOS daemon with a running engine, verifying a live `GPU util` series labelled with the accelerator model and no `GPU mem` line or `mem=`/`temp=` segments. From e695a6c8fdb0f469557ab467575eba382055d84d Mon Sep 17 00:00:00 2001 From: spinloop-agent Date: Mon, 5 Oct 2026 20:24:25 +0100 Subject: [PATCH 2/3] feat(metrics): report GPU utilisation on macOS Read the I/O Kit accelerator through ioreg for utilisation and name. Memory and temperature are absent on macOS, so the renderers omit them rather than drawing zeros. --- cmd/spinloop/metrics_render.go | 29 +++--- cmd/spinloop/metrics_render_gpu_test.go | 78 +++++++++++++++ docs/commands/serve.md | 4 +- docs/http-api.md | 2 +- internal/metrics/collect.go | 20 +++- internal/metrics/metrics.go | 9 +- internal/metrics/metrics_test.go | 98 ++++++++++++++++++- internal/metrics/parse.go | 56 +++++++++++ .../changes/macos-gpu-utilisation/tasks.md | 18 ++-- 9 files changed, 284 insertions(+), 30 deletions(-) create mode 100644 cmd/spinloop/metrics_render_gpu_test.go diff --git a/cmd/spinloop/metrics_render.go b/cmd/spinloop/metrics_render.go index 702a0cb8..515a978d 100644 --- a/cmd/spinloop/metrics_render.go +++ b/cmd/spinloop/metrics_render.go @@ -349,15 +349,13 @@ func currentGPUUtil(gpus []metrics.GpuStat, idx int) *float64 { return nil } -// currentGPUMem is the GPU's memory ratio from the current reading, 0 where -// the GPU reports no total — the gauge's own rule for that case. +// currentGPUMem is the GPU's memory ratio from the current reading, nil where +// the reading names no such GPU or the GPU reports no memory total (macOS), so +// no memory series is drawn for it unless retained history carries one. func currentGPUMem(gpus []metrics.GpuStat, idx int) *float64 { for _, g := range gpus { - if g.Index == idx { - pct := 0.0 - if g.MemoryTotal > 0 { - pct = float64(g.MemoryUsed) / float64(g.MemoryTotal) * 100 - } + if g.Index == idx && g.MemoryTotal > 0 { + pct := float64(g.MemoryUsed) / float64(g.MemoryTotal) * 100 return &pct } } @@ -448,8 +446,14 @@ func renderGPUTable(w io.Writer, gpus []metrics.GpuStat) { } fmt.Fprintln(w) for _, g := range gpus { - fmt.Fprintf(w, " GPU %d: %s util=%d%% mem=%s/%s temp=%dC\n", - g.Index, g.Name, g.Utilization, formatBytes(g.MemoryUsed), formatBytes(g.MemoryTotal), g.Temperature) + fmt.Fprintf(w, " GPU %d: %s util=%d%%", g.Index, g.Name, g.Utilization) + if g.MemoryTotal > 0 { + fmt.Fprintf(w, " mem=%s/%s", formatBytes(g.MemoryUsed), formatBytes(g.MemoryTotal)) + } + if g.Temperature > 0 { + fmt.Fprintf(w, " temp=%dC", g.Temperature) + } + fmt.Fprintln(w) } if len(gpus) > 1 { var totalUtil, totalMemUsed, totalMemTotal int64 @@ -459,8 +463,11 @@ func renderGPUTable(w io.Writer, gpus []metrics.GpuStat) { totalMemTotal += g.MemoryTotal } avgUtil := int(totalUtil) / len(gpus) - fmt.Fprintf(w, " avg util: %d%% total mem: %s/%s\n", - avgUtil, formatBytes(totalMemUsed), formatBytes(totalMemTotal)) + fmt.Fprintf(w, " avg util: %d%%", avgUtil) + if totalMemTotal > 0 { + fmt.Fprintf(w, " total mem: %s/%s", formatBytes(totalMemUsed), formatBytes(totalMemTotal)) + } + fmt.Fprintln(w) } } diff --git a/cmd/spinloop/metrics_render_gpu_test.go b/cmd/spinloop/metrics_render_gpu_test.go new file mode 100644 index 00000000..12ef737e --- /dev/null +++ b/cmd/spinloop/metrics_render_gpu_test.go @@ -0,0 +1,78 @@ +package main + +import ( + "bytes" + "strings" + "testing" + + "github.com/spinloop-ai/spinloop/internal/metrics" +) + +// A GPU that reports utilisation only (macOS) draws its utilisation series and +// no memory series, on the bar and gauge surfaces alike. +func TestUtilisationOnlyGPUDrawsNoMemorySeries(t *testing.T) { + macGPU := []metrics.GpuStat{{Index: 0, Name: "Apple M5 Max", Utilization: 56}} + history := []metrics.HistorySample{ + {GPUs: []metrics.HistoryGPU{{Index: 0, Util: 40}}}, + {GPUs: []metrics.HistoryGPU{{Index: 0, Util: 56}}}, + } + var bar, gauge bytes.Buffer + renderStatBars(&bar, nil, nil, macGPU, history, barLineW) + renderStatGauges(&gauge, nil, nil, macGPU) + for name, out := range map[string]string{"bar": bar.String(), "gauge": gauge.String()} { + if !strings.Contains(out, "GPU util") || !strings.Contains(out, " 56%") { + t.Errorf("%s: missing GPU util: %q", name, out) + } + if strings.Contains(out, "GPU mem") { + t.Errorf("%s: drew a GPU mem series: %q", name, out) + } + } +} + +// A GPU that reports a memory total still draws both series. +func TestGPUWithMemoryTotalDrawsBothSeries(t *testing.T) { + gpus := []metrics.GpuStat{{Index: 0, Name: "NVIDIA L40S", Utilization: 12, + MemoryUsed: 8 << 30, MemoryTotal: 16 << 30, Temperature: 42}} + var bar, gauge bytes.Buffer + renderStatBars(&bar, nil, nil, gpus, nil, barLineW) + renderStatGauges(&gauge, nil, nil, gpus) + for name, out := range map[string]string{"bar": bar.String(), "gauge": gauge.String()} { + if !strings.Contains(out, "GPU util") || !strings.Contains(out, "GPU mem") || !strings.Contains(out, " 50%") { + t.Errorf("%s: %q", name, out) + } + } +} + +func TestRenderGPUTableOmitsAbsentFigures(t *testing.T) { + mac := metrics.GpuStat{Index: 0, Name: "Apple M5 Max", Utilization: 56} + nvidia := metrics.GpuStat{Index: 0, Name: "NVIDIA L40S", Utilization: 12, + MemoryUsed: 8 << 30, MemoryTotal: 16 << 30, Temperature: 42} + + var b bytes.Buffer + renderGPUTable(&b, []metrics.GpuStat{mac}) + if got, want := b.String(), "\n GPU 0: Apple M5 Max util=56%\n"; got != want { + t.Errorf("utilisation-only = %q, want %q", got, want) + } + + b.Reset() + renderGPUTable(&b, []metrics.GpuStat{nvidia}) + if got, want := b.String(), "\n GPU 0: NVIDIA L40S util=12% mem=8.0 GB/16.0 GB temp=42C\n"; got != want { + t.Errorf("full reading = %q, want %q", got, want) + } + + // Two utilisation-only GPUs: the totals line drops its memory figure. + other := mac + other.Index = 1 + b.Reset() + renderGPUTable(&b, []metrics.GpuStat{mac, other}) + if strings.Contains(b.String(), "mem") || !strings.Contains(b.String(), " avg util: 56%\n") { + t.Errorf("two utilisation-only GPUs = %q", b.String()) + } + + // A mixed pair keeps the totals line whole. + b.Reset() + renderGPUTable(&b, []metrics.GpuStat{nvidia, other}) + if !strings.Contains(b.String(), "avg util: 34% total mem: 8.0 GB/16.0 GB") { + t.Errorf("mixed pair = %q", b.String()) + } +} diff --git a/docs/commands/serve.md b/docs/commands/serve.md index 8b30fbfc..fdc73296 100644 --- a/docs/commands/serve.md +++ b/docs/commands/serve.md @@ -381,7 +381,9 @@ Because the client sets the key, it knows the key — which is what it gives the agent it launches. `/v1/status` reports only *that* a key is required, never what it is, and no endpoint returns it. A supervised engine gets its own `/metrics` endpoint switched on (llama.cpp `--metrics`), which is where the token counters come from; GPU -readings need `nvidia-smi` (no Apple GPU source yet). +readings come from `nvidia-smi` on Linux and from the I/O Kit accelerator +(`ioreg`) on macOS, where only utilisation and the GPU name are reported — no +GPU memory or temperature. Those counters are also read every 15 seconds in the background, so `/v1/status` and `/v1/metrics` can both report `lastActiveAt` and diff --git a/docs/http-api.md b/docs/http-api.md index 0f3cb31d..59e6aaa6 100644 --- a/docs/http-api.md +++ b/docs/http-api.md @@ -57,7 +57,7 @@ Stops the engine. ### GET `/v1/metrics` Returns the current metrics: - Token usage counters (from the engine's Prometheus `/metrics` endpoint) -- Host system metrics (GPU, CPU, RAM) +- Host system metrics (GPU, CPU, RAM). On macOS a GPU carries utilisation and name only; its `memoryUsed`, `memoryTotal` and `temperature` are `0`, meaning the host reports none - `history`, the daemon's retained system readings — one per sampler tick while an engine ran, each a 0–100% figure per series (`t` time, `c` CPU, `m` memory, `g` per-GPU utilisation and memory), covering at most the last 10 minutes - `lastActiveAt` and `idleSeconds`, the same pair `/v1/status` reports diff --git a/internal/metrics/collect.go b/internal/metrics/collect.go index b69f2496..26daa71d 100644 --- a/internal/metrics/collect.go +++ b/internal/metrics/collect.go @@ -14,8 +14,13 @@ var nvidiaSMIArgs = []string{ "--format=csv,noheader,nounits", } +// ioregAcceleratorArgs dumps the I/O Kit accelerator services. -w 0 turns line +// wrapping off so each property stays on one line when stdout is a pipe. +var ioregAcceleratorArgs = []string{"-rd1", "-c", "IOAccelerator", "-w", "0"} + // Collector gathers system stats with host commands: nvidia-smi/vmstat/free -// on Linux; sysctl, vm_stat and top on macOS (no GPU source there yet). Both +// on Linux; sysctl, vm_stat, top and ioreg on macOS, where the GPU is the I/O +// Kit accelerator service (utilisation and name, no memory or temperature). Both // the command runner and the platform are injectable so tests feed fixture // output for either platform without running anything. type Collector struct { @@ -71,8 +76,17 @@ func reportable(err error) bool { } func (c *Collector) gpus(ctx context.Context) ([]GpuStat, error) { - if c.goos() != "linux" { - // No GPU source off Linux yet (Apple GPU stats are issue #47). + switch c.goos() { + case "darwin": + // A host with no accelerator service prints nothing, which parses to + // no GPUs. + out, err := c.run(ctx, "ioreg", ioregAcceleratorArgs...) + if err != nil { + return nil, err + } + return ParseIOAcceleratorGPU(out), nil + case "linux": + default: return nil, nil } out, err := c.run(ctx, "nvidia-smi", nvidiaSMIArgs...) diff --git a/internal/metrics/metrics.go b/internal/metrics/metrics.go index 99af154f..4877f02a 100644 --- a/internal/metrics/metrics.go +++ b/internal/metrics/metrics.go @@ -4,8 +4,9 @@ // the parsers here are ports of that Lambda's, kept value-for-value compatible // so the remote path can later delegate to a daemon running this code and the // TypeScript collectors can be deleted. Every stat is optional — a host -// without a source for one (no nvidia-smi, say) simply omits it, which is how -// a macOS node reports engine stats without GPU figures. +// without a source for one (no nvidia-smi, say) simply omits it. A macOS node +// reports GPU utilisation and name from the I/O Kit accelerator service, and +// no GPU memory or temperature. package metrics // Stats is the collected state of one serving host: what is running, its @@ -75,7 +76,9 @@ type TokenStats struct { Requests int `json:"requests"` } -// GpuStat holds per-GPU metrics from nvidia-smi. +// GpuStat holds per-GPU metrics from nvidia-smi, or from the I/O Kit +// accelerator on macOS, where MemoryUsed, MemoryTotal and Temperature are zero +// because the host reports none. type GpuStat struct { Index int `json:"index"` Name string `json:"name"` diff --git a/internal/metrics/metrics_test.go b/internal/metrics/metrics_test.go index 172ca2b3..16bca61d 100644 --- a/internal/metrics/metrics_test.go +++ b/internal/metrics/metrics_test.go @@ -203,16 +203,82 @@ func TestCollectorLinux(t *testing.T) { } } +const ioregAGXFixture = `+-o AGXAcceleratorG16G + { + "IOClass" = "AGXAcceleratorG16G" + "PerformanceStatistics" = {"In use system memory (driver)"=0,"Tiler Utilization %"=32,"Renderer Utilization %"=32,"Device Utilization %"=56,"In use system memory"=991494144} + "model" = "Apple M5 Max" + "gpu-core-count" = 40 + } + +` + +const ioregAMDFixture = `+-o AMDRadeonX6000_AMDRadeonAccelerator + { + "model" = "AMD Radeon Pro 5500M" + "PerformanceStatistics" = {"Device Utilization (%)"=7,"vramFreeBytes"=123} + } +` + +const ioregNoUtilFixture = `+-o IntelAccelerator + { + "model" = "Intel UHD Graphics" + "PerformanceStatistics" = {"Alloc system memory"=1} + } +` + +func TestParseIOAcceleratorGPU(t *testing.T) { + t.Run("apple", func(t *testing.T) { + gpus := ParseIOAcceleratorGPU(ioregAGXFixture) + want := []GpuStat{{Index: 0, Name: "Apple M5 Max", Utilization: 56}} + if len(gpus) != 1 || gpus[0] != want[0] { + t.Errorf("gpus = %+v, want %+v", gpus, want) + } + }) + t.Run("amd spelling", func(t *testing.T) { + gpus := ParseIOAcceleratorGPU(ioregAMDFixture) + if len(gpus) != 1 || gpus[0].Name != "AMD Radeon Pro 5500M" || gpus[0].Utilization != 7 { + t.Errorf("gpus = %+v", gpus) + } + }) + t.Run("two accelerators", func(t *testing.T) { + gpus := ParseIOAcceleratorGPU(ioregAMDFixture + ioregAGXFixture) + if len(gpus) != 2 || gpus[0].Index != 0 || gpus[1].Index != 1 || + gpus[0].Utilization != 7 || gpus[1].Utilization != 56 { + t.Errorf("gpus = %+v", gpus) + } + }) + t.Run("block without utilisation is skipped", func(t *testing.T) { + gpus := ParseIOAcceleratorGPU(ioregNoUtilFixture + ioregAGXFixture) + if len(gpus) != 1 || gpus[0].Index != 0 || gpus[0].Name != "Apple M5 Max" { + t.Errorf("gpus = %+v", gpus) + } + }) + t.Run("service name stands in for a missing model", func(t *testing.T) { + out := "+-o SomeAccel \n \"PerformanceStatistics\" = {\"Device Utilization %\"=3}\n" + gpus := ParseIOAcceleratorGPU(out) + if len(gpus) != 1 || gpus[0].Name != "SomeAccel" { + t.Errorf("gpus = %+v", gpus) + } + }) + t.Run("empty output", func(t *testing.T) { + if gpus := ParseIOAcceleratorGPU(""); gpus != nil { + t.Errorf("gpus = %+v, want none", gpus) + } + }) +} + func TestCollectorDarwin(t *testing.T) { c := &Collector{GOOS: "darwin", Run: fixtureRunner(map[string]string{ "top": topFixture, "sysctl": "34359738368\n", "vm_stat": vmStatFixture, + "ioreg": ioregAGXFixture, }, nil)} var stats Stats c.System(context.Background(), &stats) - if stats.GPUs != nil { - t.Errorf("darwin reported GPUs: %+v", stats.GPUs) + if len(stats.GPUs) != 1 || stats.GPUs[0].Utilization != 56 || stats.GPUs[0].MemoryTotal != 0 { + t.Errorf("darwin GPUs = %+v", stats.GPUs) } if stats.CPU == nil || stats.Memory == nil { t.Errorf("stats = %+v", stats) @@ -222,6 +288,34 @@ func TestCollectorDarwin(t *testing.T) { } } +func TestCollectorDarwinNoAcceleratorIsSilent(t *testing.T) { + c := &Collector{GOOS: "darwin", Run: fixtureRunner(map[string]string{ + "top": topFixture, + "sysctl": "34359738368\n", + "vm_stat": vmStatFixture, + "ioreg": "", + }, nil)} + var stats Stats + c.System(context.Background(), &stats) + if stats.GPUs != nil || len(stats.Errors) != 0 { + t.Errorf("GPUs = %+v, errors = %v; want neither", stats.GPUs, stats.Errors) + } +} + +func TestCollectorDarwinIoregFailureIsReported(t *testing.T) { + c := &Collector{GOOS: "darwin", Run: func(_ context.Context, name string, _ ...string) (string, error) { + if name == "ioreg" { + return "", errors.New("exit status 1") + } + return "", exec.ErrNotFound + }} + var stats Stats + c.System(context.Background(), &stats) + if len(stats.Errors) != 1 || !strings.HasPrefix(stats.Errors[0], "gpu: ") { + t.Errorf("errors = %v, want one gpu error", stats.Errors) + } +} + func TestCollectorMissingCommandIsSilent(t *testing.T) { c := &Collector{GOOS: "linux", Run: fixtureRunner(map[string]string{ "vmstat": vmstatFixture, diff --git a/internal/metrics/parse.go b/internal/metrics/parse.go index c87c974f..7c36577b 100644 --- a/internal/metrics/parse.go +++ b/internal/metrics/parse.go @@ -44,6 +44,62 @@ func ParseGPUStats(out string) []GpuStat { return gpus } +var ( + // ioregServiceHeader matches the line that opens one service in + // `ioreg -r` output, e.g. "+-o AGXAcceleratorG16G ". + ioregServiceHeader = regexp.MustCompile(`^\s*\+-o\s+(\S+)`) + // ioregModel matches the `"model" = "Apple M5 Max"` property. + ioregModel = regexp.MustCompile(`^\s*"model"\s*=\s*"([^"]*)"`) + // ioregUtilisation matches the device-wide utilisation figure inside a + // performance statistics dictionary: Apple's "Device Utilization %" and + // the AMD drivers' "Device Utilization (%)". + ioregUtilisation = regexp.MustCompile(`"Device Utilization (?:%|\(%\))"\s*=\s*(\d+)`) +) + +// ParseIOAcceleratorGPU parses the output of +// +// ioreg -rd1 -c IOAccelerator -w 0 +// +// into one GpuStat per accelerator service that reports a device utilisation. +// The name is the service's "model" property, falling back to the service +// name; the index is the order of appearance. Memory and temperature are left +// zero: Apple Silicon has no GPU memory total and temperature needs root. A +// service with no utilisation figure contributes no GPU, and empty output +// yields none. +func ParseIOAcceleratorGPU(out string) []GpuStat { + var gpus []GpuStat + var service, model string + util := -1 + flush := func() { + if util >= 0 { + name := model + if name == "" { + name = service + } + gpus = append(gpus, GpuStat{Index: len(gpus), Name: name, Utilization: util}) + } + service, model, util = "", "", -1 + } + for _, line := range strings.Split(out, "\n") { + if m := ioregServiceHeader.FindStringSubmatch(line); m != nil { + flush() + service = m[1] + continue + } + if m := ioregModel.FindStringSubmatch(line); m != nil { + model = m[1] + } + if !strings.Contains(line, "erformance") { + continue + } + if m := ioregUtilisation.FindStringSubmatch(line); m != nil { + util = atoiOrZero(m[1]) + } + } + flush() + return gpus +} + // ParseVmstatCPU parses `vmstat 1 2` output, whose last line is the sampled // interval. The idle column (id) is field 15 of the standard 17-column layout; // utilization is 100 - idle. Returns nil when the output is not vmstat's. diff --git a/openspec/changes/macos-gpu-utilisation/tasks.md b/openspec/changes/macos-gpu-utilisation/tasks.md index b12e85a4..7a0e7bf6 100644 --- a/openspec/changes/macos-gpu-utilisation/tasks.md +++ b/openspec/changes/macos-gpu-utilisation/tasks.md @@ -2,21 +2,21 @@ ## 1. macOS GPU collection (internal/metrics) -- [ ] 1.1 Add `ParseIOAcceleratorGPU(out string) []GpuStat` to `internal/metrics/parse.go`: scan `+-o` service blocks for `PerformanceStatistics`/`Performance Statistics` dictionaries, read `Device Utilization %` (accepting the AMD `Device Utilization (%)` spelling), name from `"model"` with the service name as fallback, index by order of appearance, memory and temperature left zero; verify with parser tests covering an Apple AGX fixture, an AMD-spelling fixture, a two-accelerator fixture, a block lacking the utilization key, and empty output yielding no GPUs. -- [ ] 1.2 Give `Collector.gpus()` a darwin branch running `ioreg -rd1 -c IOAccelerator -w 0` through the injectable runner: parsed GPUs on success, silent nil on empty output, reported `gpu:` error on a non-zero exit; verify with collector tests feeding an `ioreg` fixture through the runner map, a silent-absence test, and an exit-failure test, keeping the Linux and CPU/RAM collector tests green. -- [ ] 1.3 Update the comments that record the gap — the `internal/metrics` package doc, the `Collector` header, and the `gpus()` "issue #47" note — to describe the macOS accelerator source and its absent memory/temperature; verify with `go vet ./internal/metrics` and `go test ./internal/metrics/...` passing. +- [x] 1.1 Add `ParseIOAcceleratorGPU(out string) []GpuStat` to `internal/metrics/parse.go`: scan `+-o` service blocks for `PerformanceStatistics`/`Performance Statistics` dictionaries, read `Device Utilization %` (accepting the AMD `Device Utilization (%)` spelling), name from `"model"` with the service name as fallback, index by order of appearance, memory and temperature left zero; verify with parser tests covering an Apple AGX fixture, an AMD-spelling fixture, a two-accelerator fixture, a block lacking the utilization key, and empty output yielding no GPUs. +- [x] 1.2 Give `Collector.gpus()` a darwin branch running `ioreg -rd1 -c IOAccelerator -w 0` through the injectable runner: parsed GPUs on success, silent nil on empty output, reported `gpu:` error on a non-zero exit; verify with collector tests feeding an `ioreg` fixture through the runner map, a silent-absence test, and an exit-failure test, keeping the Linux and CPU/RAM collector tests green. +- [x] 1.3 Update the comments that record the gap — the `internal/metrics` package doc, the `Collector` header, and the `gpus()` "issue #47" note — to describe the macOS accelerator source and its absent memory/temperature; verify with `go vet ./internal/metrics` and `go test ./internal/metrics/...` passing. ## 2. Rendering absent GPU figures (cmd/spinloop) -- [ ] 2.1 Make `currentGPUMem` in `cmd/spinloop/metrics_render.go` return nil when `MemoryTotal == 0`, so `barSeriesList` drops a GPU's memory series only when no current reading and no retained sample carries a total; verify with tests that a utilisation-only GPU yields `GPU util` with no `GPU mem` in bar, combined, and gauge surfaces, and that an NVIDIA-shaped reading still yields both. -- [ ] 2.2 Make `renderGPUTable` emit `mem=` and `temp=` segments only when the reading carries a total and a temperature respectively, and the multi-GPU totals line drop its `total mem:` figure when the summed total is zero; verify with table-format tests for a utilisation-only GPU, a full NVIDIA-shaped GPU, and a mixed pair keeping the totals line whole. -- [ ] 2.3 Run the dashboard, serve-view, fleet, and metrics-render test suites and fix any golden outputs that expected the zeroed `GPU mem` line or `0B/0B` segments, confirming the shared renderers needed no format-specific change in dashboard or fleet code. +- [x] 2.1 Make `currentGPUMem` in `cmd/spinloop/metrics_render.go` return nil when `MemoryTotal == 0`, so `barSeriesList` drops a GPU's memory series only when no current reading and no retained sample carries a total; verify with tests that a utilisation-only GPU yields `GPU util` with no `GPU mem` in bar, combined, and gauge surfaces, and that an NVIDIA-shaped reading still yields both. +- [x] 2.2 Make `renderGPUTable` emit `mem=` and `temp=` segments only when the reading carries a total and a temperature respectively, and the multi-GPU totals line drop its `total mem:` figure when the summed total is zero; verify with table-format tests for a utilisation-only GPU, a full NVIDIA-shaped GPU, and a mixed pair keeping the totals line whole. +- [x] 2.3 Run the dashboard, serve-view, fleet, and metrics-render test suites and fix any golden outputs that expected the zeroed `GPU mem` line or `0B/0B` segments, confirming the shared renderers needed no format-specific change in dashboard or fleet code. ## 3. Documentation -- [ ] 3.1 Grep `docs/` for claims that GPU figures are Linux/NVIDIA-only or that macOS nodes show no GPU, and update the affected pages (metrics/dashboard guides) to state that macOS reports utilisation only; verify each edited command or output sample still matches the code's actual output. +- [x] 3.1 Grep `docs/` for claims that GPU figures are Linux/NVIDIA-only or that macOS nodes show no GPU, and update the affected pages (metrics/dashboard guides) to state that macOS reports utilisation only; verify each edited command or output sample still matches the code's actual output. ## 4. Integration checks -- [ ] 4.1 Run `gofmt -l .`, `go vet ./...`, and `go test ./... -cover`, verifying a clean format, a clean vet, and total coverage at or above 80%. -- [ ] 4.2 Build the binary (`go build -o spinloop ./cmd/spinloop`) and run `spinloop metrics --format=bar` and `--format=table` against a local macOS daemon with a running engine, verifying a live `GPU util` series labelled with the accelerator model and no `GPU mem` line or `mem=`/`temp=` segments. +- [x] 4.1 Run `gofmt -l .`, `go vet ./...`, and `go test ./... -cover`, verifying a clean format, a clean vet, and total coverage at or above 80%. +- [x] 4.2 Build the binary (`go build -o spinloop ./cmd/spinloop`) and run `spinloop metrics --format=bar` and `--format=table` against a local macOS daemon with a running engine, verifying a live `GPU util` series labelled with the accelerator model and no `GPU mem` line or `mem=`/`temp=` segments. From 03befdb05a2d537f6aa3940e9fad491fe5d71f6d Mon Sep 17 00:00:00 2001 From: spinloop-agent Date: Mon, 5 Oct 2026 22:59:46 +0100 Subject: [PATCH 3/3] chore: archive macos-gpu-utilisation change --- .../.openspec.yaml | 0 .../design.md | 0 .../proposal.md | 0 .../specs/engine-metrics/spec.md | 0 .../specs/remote-metrics-bar-format/spec.md | 0 .../tasks.md | 0 openspec/specs/engine-metrics/spec.md | 80 +++++++++++++++---- .../specs/remote-metrics-bar-format/spec.md | 21 +++-- 8 files changed, 79 insertions(+), 22 deletions(-) rename openspec/changes/{macos-gpu-utilisation => archive/2026-10-05-macos-gpu-utilisation}/.openspec.yaml (100%) rename openspec/changes/{macos-gpu-utilisation => archive/2026-10-05-macos-gpu-utilisation}/design.md (100%) rename openspec/changes/{macos-gpu-utilisation => archive/2026-10-05-macos-gpu-utilisation}/proposal.md (100%) rename openspec/changes/{macos-gpu-utilisation => archive/2026-10-05-macos-gpu-utilisation}/specs/engine-metrics/spec.md (100%) rename openspec/changes/{macos-gpu-utilisation => archive/2026-10-05-macos-gpu-utilisation}/specs/remote-metrics-bar-format/spec.md (100%) rename openspec/changes/{macos-gpu-utilisation => archive/2026-10-05-macos-gpu-utilisation}/tasks.md (100%) diff --git a/openspec/changes/macos-gpu-utilisation/.openspec.yaml b/openspec/changes/archive/2026-10-05-macos-gpu-utilisation/.openspec.yaml similarity index 100% rename from openspec/changes/macos-gpu-utilisation/.openspec.yaml rename to openspec/changes/archive/2026-10-05-macos-gpu-utilisation/.openspec.yaml diff --git a/openspec/changes/macos-gpu-utilisation/design.md b/openspec/changes/archive/2026-10-05-macos-gpu-utilisation/design.md similarity index 100% rename from openspec/changes/macos-gpu-utilisation/design.md rename to openspec/changes/archive/2026-10-05-macos-gpu-utilisation/design.md diff --git a/openspec/changes/macos-gpu-utilisation/proposal.md b/openspec/changes/archive/2026-10-05-macos-gpu-utilisation/proposal.md similarity index 100% rename from openspec/changes/macos-gpu-utilisation/proposal.md rename to openspec/changes/archive/2026-10-05-macos-gpu-utilisation/proposal.md diff --git a/openspec/changes/macos-gpu-utilisation/specs/engine-metrics/spec.md b/openspec/changes/archive/2026-10-05-macos-gpu-utilisation/specs/engine-metrics/spec.md similarity index 100% rename from openspec/changes/macos-gpu-utilisation/specs/engine-metrics/spec.md rename to openspec/changes/archive/2026-10-05-macos-gpu-utilisation/specs/engine-metrics/spec.md diff --git a/openspec/changes/macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md b/openspec/changes/archive/2026-10-05-macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md similarity index 100% rename from openspec/changes/macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md rename to openspec/changes/archive/2026-10-05-macos-gpu-utilisation/specs/remote-metrics-bar-format/spec.md diff --git a/openspec/changes/macos-gpu-utilisation/tasks.md b/openspec/changes/archive/2026-10-05-macos-gpu-utilisation/tasks.md similarity index 100% rename from openspec/changes/macos-gpu-utilisation/tasks.md rename to openspec/changes/archive/2026-10-05-macos-gpu-utilisation/tasks.md diff --git a/openspec/specs/engine-metrics/spec.md b/openspec/specs/engine-metrics/spec.md index 349b7818..956cab76 100644 --- a/openspec/specs/engine-metrics/spec.md +++ b/openspec/specs/engine-metrics/spec.md @@ -13,7 +13,9 @@ design: a host with no source for one omits it rather than erroring, which is how a machine without `nvidia-smi` reports engine stats and no GPU figures. The shape is kept value-for-value compatible with what the existing `spinloop remote metrics` formatters render. + ## Requirements + ### Requirement: Engine stats collection The system SHALL collect token and request statistics from a running engine by @@ -67,10 +69,14 @@ back to a configured base URL, and failing that to the engine's default. The system SHALL collect system statistics from the host: GPU utilization, GPU memory used/total, CPU utilization, and RAM used/total. On hosts with -NVIDIA GPUs the GPU figures SHALL be sourced from `nvidia-smi`. The collected -values SHALL use the same units as the existing remote stats pipeline (bytes -for memory, percentages for utilization) so existing rendering applies -unchanged. +NVIDIA GPUs the GPU figures SHALL be sourced from `nvidia-smi`. On macOS +hosts the GPU SHALL be sourced from the host's I/O Kit accelerator service, +which yields GPU utilization and the GPU's name; where that service reports +no memory total or no temperature — the ordinary case on a host with +unified memory — those figures SHALL be absent rather than reported as +zero. The collected values SHALL use the same units as the existing remote +stats pipeline (bytes for memory, percentages for utilization) so existing +rendering applies unchanged. #### Scenario: NVIDIA host reports GPU stats @@ -78,6 +84,14 @@ unchanged. - **THEN** the result includes GPU utilization and GPU memory used/total in bytes +#### Scenario: macOS host reports GPU utilisation + +- **WHEN** metrics are collected on a macOS host whose accelerator service + reports device utilization +- **THEN** the result includes a GPU with that utilization percentage and + the accelerator's model name, and omits GPU memory and temperature rather + than reporting them as zero + #### Scenario: CPU and RAM are always attempted - **WHEN** metrics are collected on any supported host @@ -87,22 +101,23 @@ unchanged. ### Requirement: Graceful platform degradation When a system stat's source is unavailable on the host (for example -`nvidia-smi` on a machine without NVIDIA tooling, or Linux-only commands on -macOS), the collector SHALL omit that stat and return the remainder, rather -than failing the collection. The absence SHALL be distinguishable from a zero -value in the collected result. - -A source that is *present and failing* SHALL be distinguished from one that is -absent. Where the collector has an address to query and the query fails, it -SHALL report that failure among the collected errors, naming what it tried, so -a misdirected or broken collector is visible rather than presenting as an -engine that has simply served nothing. An absent source SHALL remain silent: -reporting the routine absence of a source as an error would bury the failures -worth seeing. +`nvidia-smi` on a machine without NVIDIA tooling, or no accelerator service +at all on a virtualized guest), the collector SHALL omit that stat and +return the remainder, rather than failing the collection. The absence SHALL +be distinguishable from a zero value in the collected result. + +A source that is *present and failing* SHALL be distinguished from one that +is absent. Where the collector has an address to query and the query fails, +it SHALL report that failure among the collected errors, naming what it +tried, so a misdirected or broken collector is visible rather than +presenting as an engine that has simply served nothing. An absent source +SHALL remain silent: reporting the routine absence of a source as an error +would bury the failures worth seeing. #### Scenario: macOS host lacks GPU stats -- **WHEN** metrics are collected on a macOS host +- **WHEN** metrics are collected on a macOS host that names no I/O Kit + accelerator service, such as a virtualized guest - **THEN** the result includes engine stats and available CPU/RAM figures, omits GPU stats, and reports no error @@ -166,3 +181,34 @@ it still answers when work last happened. - **THEN** the result carries the last-active time and idle duration even though the running-engine figures are absent +### Requirement: Absent GPU figures render as absent + +A GPU reading that carries no memory total or no temperature SHALL render +without those figures, rather than as zeros: a resource-series view SHALL +draw no memory series for a GPU whose current reading and retained history +both report no memory total, and a table GPU line SHALL omit its memory and +temperature segments when the reading carries none. The GPU's utilization +SHALL still render. This keeps the shape's own rule — an absent figure is +not a zero — true of the rendered output as it is of the collected one. + +#### Scenario: A utilisation-only GPU draws one series + +- **WHEN** a bar or gauge view renders a GPU whose readings never carry a + memory total +- **THEN** the output includes the GPU utilization series and no GPU memory + series + +#### Scenario: A GPU with memory still draws both series + +- **WHEN** a bar or gauge view renders a GPU whose current reading or any + retained sample carries a memory total +- **THEN** the output includes both the utilization and the memory series, + as before + +#### Scenario: Table omits the segments a reading lacks + +- **WHEN** the table format renders a GPU line whose reading carries no + memory total and no temperature +- **THEN** the line shows the GPU's name and utilization with no `mem=` and + no `temp=` segment, and a GPU reading that does carry them shows them as + before diff --git a/openspec/specs/remote-metrics-bar-format/spec.md b/openspec/specs/remote-metrics-bar-format/spec.md index 977ac73b..8358cd09 100644 --- a/openspec/specs/remote-metrics-bar-format/spec.md +++ b/openspec/specs/remote-metrics-bar-format/spec.md @@ -3,10 +3,12 @@ ## Purpose Define the bar graph output format for `spinloop remote metrics` with colour-coded resource utilization indicators. + ## Requirements + ### Requirement: Bar format output -The system SHALL support a `--format=bar` option that renders each resource series as a sparkline drawn from the history the on-instance daemon retains: a left-aligned label, one glyph per sample, and the latest value as a right-aligned percentage. The glyphs SHALL be Unicode block elements of one grade per utilisation level, so a series reads as a line of bars across the window. The series drawn SHALL be the same set the gauge format draws: CPU, RAM, and each GPU's utilisation and memory, with the same per-GPU labelling. +The system SHALL support a `--format=bar` option that renders each resource series as a sparkline drawn from the history the on-instance daemon retains: a left-aligned label, one glyph per sample, and the latest value as a right-aligned percentage. The glyphs SHALL be Unicode block elements of one grade per utilisation level, so a series reads as a line of bars across the window. The series drawn SHALL be the same set the gauge format draws: CPU, RAM, each GPU's utilisation, and each GPU's memory where that GPU's current reading or retained history reports a memory total — a GPU that never reports one draws no memory series, with the same per-GPU labelling throughout. #### Scenario: Bar format displays CPU utilization @@ -20,9 +22,14 @@ The system SHALL support a `--format=bar` option that renders each resource seri #### Scenario: Bar format displays GPU utilization -- **WHEN** the user runs `spinloop remote metrics --format=bar` with a running instance that has GPU data +- **WHEN** the user runs `spinloop remote metrics --format=bar` with a running instance that has GPU data whose readings carry a memory total - **THEN** the output includes rows labelled "GPU util" and "GPU mem" (or "GPU N util"/"GPU N mem" for multiple GPUs), each drawn from the retained history +#### Scenario: Bar format omits a GPU memory series with no total + +- **WHEN** the user runs `spinloop remote metrics --format=bar` with GPU data whose current reading and retained history report no memory total +- **THEN** the output includes "GPU util" and no "GPU mem" row + #### Scenario: Bar format header line - **WHEN** the user runs `spinloop remote metrics --format=bar` with a running instance @@ -100,7 +107,7 @@ than shown empty or zeroed. ### Requirement: Gauge format -The system SHALL support a `--format=gauge` option that renders each resource series as a horizontal progress gauge: a left-aligned label, a filled portion using block characters, an unfilled portion using light shade characters, and a right-aligned percentage value. The gauge draws the current reading only — it carries no history. The series drawn SHALL be CPU, RAM, and each GPU's utilisation and memory, with the same labels the bar format uses. The gauge fill SHALL be colour-coded on the bar format's thresholds: green for values at or below 80%, yellow for values from 80% to 90%, and red for values above 90%, with the colour reset after the filled portion so the unfilled characters and percentage appear in the terminal's default colour. +The system SHALL support a `--format=gauge` option that renders each resource series as a horizontal progress gauge: a left-aligned label, a filled portion using block characters, an unfilled portion using light shade characters, and a right-aligned percentage value. The gauge draws the current reading only — it carries no history. The series drawn SHALL be CPU, RAM, each GPU's utilisation, and each GPU's memory where the current reading reports a memory total — a GPU that reports none draws no memory gauge — with the same labels the bar format uses. The gauge fill SHALL be colour-coded on the bar format's thresholds: green for values at or below 80%, yellow for values from 80% to 90%, and red for values above 90%, with the colour reset after the filled portion so the unfilled characters and percentage appear in the terminal's default colour. #### Scenario: Gauge format displays CPU utilization @@ -109,9 +116,14 @@ The system SHALL support a `--format=gauge` option that renders each resource se #### Scenario: Gauge format displays GPU utilisation -- **WHEN** the user runs `spinloop remote metrics --format=gauge` with a running instance that has GPU data +- **WHEN** the user runs `spinloop remote metrics --format=gauge` with a running instance that has GPU data whose readings carry a memory total - **THEN** the output includes gauges labelled "GPU util" and "GPU mem" (or "GPU N util"/"GPU N mem" for multiple GPUs) +#### Scenario: Gauge omits a GPU memory gauge with no total + +- **WHEN** the user runs `spinloop remote metrics --format=gauge` with a GPU whose current reading reports no memory total +- **THEN** the output includes "GPU util" and no "GPU mem" gauge + #### Scenario: Gauge colours the fill - **WHEN** a gauge's current value is 95% @@ -166,4 +178,3 @@ information. - **WHEN** two adjacent series are both at their maximum across the window - **THEN** each row's tallest glyph leaves the top eighth of its cell unfilled, so a blank strip separates the rows rather than the rows reading as one solid block -