Give an Amplifier session real control of a desktop — Windows (via WSL2), macOS, or Linux (X11) — on this machine or another one across your private network, using the LLM provider's native computer-use tool, not a homegrown imitation of it.
If a human could do it by looking at the screen and clicking, this can do it. No API, no CLI, no browser extension required.
"What's on my screen right now?"
"Click the Export button in that accounting app and save it to my desktop."
"Fill this dialog in for me — I'm done pressing buttons."
This is not the usual one-line,
--app-level pointer at a behavior bundle. It drives a real desktop, so the install has four moving parts that live outside this repository: a version floor on three upstream modules, a model that supports native computer use, a target machine with per-platform prerequisites, and — for remote targets — SSH key auth anduvon the far end.The per-platform prerequisites are hard requirements, not recommendations:
Target Requires Windows WSL2. A plain Windows host — including one reachable over Windows OpenSSH — is not supported. You SSH into the WSL2 side; every action crosses into Win32 via powershell.exeinteropmacOS Two separate TCC grants — Screen Recording (capture) and Accessibility (input). Neither is checked at mount; the tools appear and then fail on first use Linux An X11 session. Wayland is not supported, and not detected — under XWayland the probe can pass anyway → docs/SETUP.md has the exact checks, the proven-vs-rough matrix, the full config reference, and the operational facts (a locked screen cannot be driven, on any platform) that otherwise cost you an afternoon.
Registering this bundle is ordinary Amplifier bundle management:
# 1. Register it (the name "computer-use" comes from this repo's own bundle.md)
amplifier bundle add git+https://github.com/microsoft/amplifier-bundle-computer-use@main#subdirectory=behaviors/computer-use.yaml --app
# 2. Use it for a session
amplifier run --bundle computer-use "What's on my screen right now?"
Confirm it registered correctly with amplifier bundle show computer-use — it should
list tool-computer-use, hook-computer-use, and computer-use:computer-operator.
Working from a local clone instead of GitHub? amplifier bundle add file:///path/to/amplifier-bundle-computer-use instead.
Registering the bundle is not the same as the tools working. See
docs/SETUP.md for the four things that decide whether computer and
desktop actually function once a session starts: upstream module versions, a model with
native computer-use support, a reachable target machine, and (for remote targets) SSH.
Both Anthropic and OpenAI post-train their models on a specific server-side tool definition. Sending a lookalike function tool instead gets you noticeably worse targeting — or, on OpenAI, a hard 400. This bundle sends the real one, per vendor:
Anthropic additionally needs the computer-use-2025-11-24 beta header (derived by the
provider itself). Screenshots come back as genuine base64 image content blocks
inside the tool result, exactly as each vendor's own loop does.
Both providers have driven a real desktop through this bundle. Anthropic:
claude-sonnet-4-5, claude-sonnet-5, claude-opus-5. OpenAI: gpt-5.5, verified
end-to-end against a real remote desktop. A Gemini dialect record exists
(gemini-2.5-computer-use), transcribed from captured traffic — no live end-to-end run
through this bundle is claimed for it.
Model support is capability-gated upstream: provider-openai and
provider-anthropic both expose supports_native_computer_use on ModelCapabilities,
and a model without it cannot use this bundle. See docs/SETUP.md §2 for the exact
rules (OpenAI: minor >= 4 and not -nano; Anthropic: per-family version thresholds
mapping to a specific computer_YYYYMMDD wire type).
Amplifier's orchestrator used to stand between a mounted tool and native computer use:
tool results are collapsed to str before reaching the provider, so a screenshot could
never travel back as an image.
That is fixed at one seam — the provider's complete() call — by
hook-computer-use, which:
- unwraps the orchestrator's
ToolResultenvelope and expands screenshot markers into real image blocks, - keeps only the N most recent screenshots inline so long sessions stay affordable.
(ToolSpec used to be built from name/description/parameters only, so a tool
couldn't declare itself a server-side tool type either — hook-computer-use used to
promote it and inject the beta header itself. That is now handled upstream:
loop-streaming preserves a tool's native_tool_spec through its own ToolSpec
construction, and provider-anthropic derives the required anthropic-beta header
itself. hook-computer-use now only verifies that support is present and refuses to
mount if it isn't — see _fail_if_orchestrator_native_tool_spec_unsupported, and
docs/SETUP.md §1 for the exact commits you need.)
Remove the hook and the tool degrades cleanly to an ordinary function tool. Nothing is monkey-patched on disk; nothing rots when the orchestrator changes.
| Piece | Role |
|---|---|
modules/tool-computer-use |
computer (native action set) + desktop (windows, clipboard) |
modules/hook-computer-use |
The wire-format seam described above |
agents/computer-operator.md |
Operating discipline for driving a live machine |
bridge.ps1 |
WSL2 → Windows execution via powershell.exe + Win32 |
computer — the native action set: screenshot, zoom, cursor_position,
mouse_move, left_click, right_click, middle_click, double_click,
triple_click, left_mouse_down, left_mouse_up, left_click_drag, scroll, key,
hold_key, type, wait, plus screen_info, list_windows, focus_window.
desktop — what the native schema cannot express: list_windows, focus_window,
screen_info, get_clipboard, set_clipboard, list_monitors, select_monitor.
Focus a window before typing into it; use the clipboard to pull exact text out of an app;
use select_monitor to switch which monitor the model sees, mid-session.
On a remote (
ssh://) target, nine of these are not implemented yet.left_mouse_down,left_mouse_up,left_click_drag,scroll,hold_key,desktop.list_windows,desktop.focus_window,desktop.get_clipboard, anddesktop.set_clipboardall raiseBackendError("... over the wire is Phase 2"). Seedocs/SETUP.md§5.
The display is captured at full physical resolution, downscaled to a model-friendly size (default 1280 long edge, aspect ratio preserved), and coordinates the model emits are scaled back to physical pixels. On a 3840×2160 display that is an exact 3× factor.
Anthropic's guidance is explicit: sending screenshots above WXGA is both slower and
less accurate. Raising max_edge is usually the wrong lever.
The keys that matter most. Full reference in docs/SETUP.md §6.
tools:
- module: tool-computer-use
config:
target: ssh://user@host # omit entirely to drive THIS machine
max_edge: 1280 # long edge of the image the model sees
enable_zoom: true
read_only: false # true = screenshots only, all input blocked
# tool_version: auto-resolved from the live model — leave unset
hooks:
- module: hook-computer-use
config:
max_inline_screenshots: 3 # older screenshots collapse to text
# unattended_writes_ok: true # explicit, logged opt-out — see SafetyRemote defaults are stricter than local, because a remote machine is by definition one you are not looking at:
| Key | Local default | Remote (ssh://) default |
|---|---|---|
read_only |
false |
true |
gate_writes |
off | on, whenever read_only is off |
clipboard_read_policy |
allow |
redact |
See docs/SETUP.md — this is the part that is more involved than a
typical bundle. In brief:
- Recent-enough upstream modules —
loop-streaming, andprovider-anthropicand/orprovider-openai, need changes not yet reflected in their package version (all three still declare1.0.0, so there is no version floor to check up front). You do not need to pin anything by commit before you start: the bundle probes the actually-installed code at mount time rather than trusting a version string, and refuses to mount (naming the exact commit to upgrade to) if the orchestrator cannot carry the native tool form. Seedocs/SETUP.md§1 for how the probe works. - A model with
supports_native_computer_use— OpenAI:minor >= 4, not-nano. Anthropic: per-family version thresholds. §2 of the setup doc has the tables. - A target machine — local or remote over SSH — meeting its platform's hard
prerequisites. These are not optional, and two of the three are not caught at mount:
- Windows: WSL2 is required. There is no native-Windows code path.
windows.pyis a "WSL2 -> Windows desktop backend"; every action crosses into Win32 viapowershell.exeinterop. A plain Windows host — including one reachable over Windows OpenSSH — is not supported; the probe reportswslpath not on PATH (not running under WSL2?)and nothing mounts. For a remote Windows target, the SSH server must run inside WSL. Tracked inBACKLOG.md. - macOS: two separate TCC grants — Screen Recording (capture) and Accessibility (input). Granting one does not grant the other, and neither is checked at mount — the tools appear healthy and fail on first use.
- Linux: an X11 session +
python-xlib. Wayland is not supported, and not detected — the probe tests onlyDISPLAY/XTEST, which XWayland satisfies, so it can report available on a Wayland desktop. Unverified territory; use real X11.
- Windows: WSL2 is required. There is no native-Windows code path.
- For remote targets — key-based SSH (
BatchMode=yes; a passphrase-locked key will not prompt, it will fail), a trusted host key, anduvon the far end (mandatory — the connection resolvesuvbefore anything else and fails if it is absent). The target's own per-platform prerequisites above still apply; SSH does not bypass them. - Pillow (installed with the tool module).
No agent installed on the target, and no admin rights, extra service, or new listening port: our files ship down the SSH pipe and are removed at session end — nothing to update or uninstall. That is a claim about our agent, not a claim that the target needs no setup. It still needs everything listed above.
Four distinct mechanisms. Be clear which one you are relying on:
| Control | Kind | Strength |
|---|---|---|
read_only: true |
Code | Enforced. Every mutating action rejected in execute() before anything reaches the desktop. Screenshots still work. |
Write gate (gate_writes) |
Code + human | Enforced. Per-action ask_user approval, default deny, across 15 mutating actions. On by default for a remote, non-read-only target. |
| Presence guard / halt | Code | Enforced and unconditional. Halts before the next write the moment a human is detected at the machine. No config key can disable the halt once a guard exists; the halt is durable across sessions and cleared only by a human running scripts/resume_after_halt.py. |
Stop Conditions in agents/computer-operator.md |
Prompt | Model judgment. Nothing inspects the screen or blocks an action. |
The agent is instructed to stop and ask when it sees credential prompts, CAPTCHAs, destructive confirmations, anything that would send or publish on your behalf, or a screen that twice fails to match expectations. That instruction is followed well in practice — but it is guidance to a model, not a gate. Treat it as such.
unattended_writes_ok is a deliberate opt-out, not a convenience toggle. With no TTY,
the write gate denies rather than crashing on an unanswerable prompt. Setting
unattended_writes_ok: true allows those writes with no human confirmation — always
logged at WARNING, never a default, never inferred from the environment. If prompts are
merely tedious in an interactive session, you want read_only: true instead.
A locked screen cannot be driven — on any platform. macOS and Windows both switch to a
secure session that refuses synthetic input by design. The requirement is an unlocked,
logged-in GUI session, not merely "the screen is on." The tool now detects this and
fails loud; previously it handed a lock-screen capture to the model as if it were the
desktop. docs/SETUP.md §8 has the measurements and the three distinct macOS diagnoses.
On-screen content can try to manipulate the agent. Anything the agent can read, it can
be influenced by: a dialog, a web page, or a document saying "click OK to confirm" or
"enter the password" is indistinguishable from a legitimate UI. This is an inherent risk of
computer use, not a defect in this bundle. Do not point it at untrusted screens
unsupervised, and use read_only: true when you only need it to look.
The clipboard goes to your model provider. desktop.get_clipboard returns the target's
clipboard as tool output, which becomes part of the conversation sent to the API and lands
in durable logs. The default is allow locally and redact (length + digest, never the
text) on a remote target; clipboard_read_policy: block refuses outright. If you have just
copied a password or token, clear the clipboard before letting the agent read it.
Screenshots touch disk briefly. Captures land in %TEMP%\amplifier-computer-use\ on a
Windows target (cleaned after 30 minutes) and in a per-session subdirectory of
~/.amplifier/computer-use/shots/ on the controller (cleaned after 2 hours). Directories
are created 0700 and files 0600.
- macOS
type_textsilently no-ops while returning success — OPEN. On an unlocked Mac,keyworks andtypereturnssuccess: truewhile entering nothing. Localized to the type path; hypothesis (unconfirmed) is that it posts to a specific app rather than the system-wide event tap.key-only flows on macOS are unaffected. Logged inBACKLOG.md. - Nine actions are unimplemented over the remote wire — see Tools above.
- No whole-session end-to-end run of the hook, native promotion, screenshot rewriting
and the write gate all executing together (
BACKLOG.md). - The Windows on-desktop indicator overlay is not built (Linux and macOS announce are).
Full symptom→cause→fix table in docs/SETUP.md §10.
Set a trace path and you get a plain-text record of what the hook actually did:
AMPLIFIER_COMPUTER_USE_TRACE=/tmp/cu-trace.log amplifier run --bundle computer-use "..."MOUNTED max_inline=3
WRAPPED provider=AnthropicProvider module=amplifier_module_provider_anthropic
complete: markers=1 messages_with_blocks=3
No markers= line → screenshots are not reaching the model. A mount-time
ComputerUseNativeToolPassthroughUnsupportedError means the installed loop-streaming
does not yet carry computer's native tool form to the wire on its own — see that
error's message for the exact commit to upgrade to. A provider that fails its capability
probe does not raise; the hook logs which integration points it tried
(_derive_native_tool_betas for Anthropic's dated types, _convert_tools_from_request
for OpenAI's bare computer) and wraps nothing.
python tests/test_wire_format.pyAsserts the real bytes: screenshot markers expanded to image blocks that survive
model_dump() without an API-rejected visibility key, recency window enforced, and
graceful degradation when a screenshot file is gone. (Native tool-spec passthrough is no
longer done by this hook — it's verified at mount time instead; see
tests/test_native_tool_passthrough_guard.py.)
Originally created by @ckrabach617 as a
Windows-from-WSL2 bundle. That original work is preserved in this repository's
git history and its copyright is retained in LICENSE.
This fork extends it to a platform-backend architecture (Windows, macOS, Linux
X11), per-monitor targeting, and remote operation over a private network. See
BACKLOG.md for what is known and not yet done, and docs/ for the
design record.
Note
This project is not currently accepting external contributions, but we're actively working toward opening this up. We value community input and look forward to collaborating in the future. For now, feel free to fork and experiment!
Most contributions require you to agree to a Contributor License Agreement (CLA) declaring that you have the right to, and actually do, grant us the rights to use your contribution. For details, visit Contributor License Agreements.
When you submit a pull request, a CLA bot will automatically determine whether you need to provide a CLA and decorate the PR appropriately (e.g., status check, comment). Simply follow the instructions provided by the bot. You will only need to do this once across all repos using our CLA.
This project has adopted the Microsoft Open Source Code of Conduct. For more information see the Code of Conduct FAQ or contact opencode@microsoft.com with any additional questions or comments.
This project may contain trademarks or logos for projects, products, or services. Authorized use of Microsoft trademarks or logos is subject to and must follow Microsoft's Trademark & Brand Guidelines. Use of Microsoft trademarks or logos in modified versions of this project must not cause confusion or imply Microsoft sponsorship. Any use of third-party trademarks or logos are subject to those third-party's policies.