datuloar/dsh-computer-use-win ↗★ 2

dsh-cu

Windows computer use for the DeepSeek Harness: screenshot, click, type, scroll, read windows and the clipboard, plus an on-screen indicator the human sees and stops with ESC. No dependencies, no driver to install. 适合需要让Agent直接控制Windows系统GUI、执行点击和截图任务的用户。

Package
dsh-cu
Compatibility
Unverified
Harness peer range
>=0.1.0-rc.8 <0.2.0 || ^0.1.1-rc.1 || ^0.1.2-alpha.1 || ^0.1.5-rc.1
Version
1.2.0
License
MIT
Last updated
Sep 18, 2026

Install

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:datuloar/dsh-computer-use-win

dsh-computer-use

CI license platform

Windows computer use for agents: read the screen as text, click controls by name, screenshot, type, scroll, drag — plus an on-screen indicator the human can see and stop with ESC.

The indicator the human sees while an agent drives the machine

A text-only model can drive a GUI with this: dsh-cu tree and dsh-cu read turn a window into a few hundred tokens of roles, names and clickable points, and dsh-cu tap "Save" clicks the one control that matches. Screenshots are there when pixels are the answer, not as the only way in.

Built for the DeepSeek Harness; the CLI works with any agent or script. The CLI is dsh-cu, the skill it registers is dsh-computer-use, the repository is dsh-computer-use-win — three names, one thing.

Nothing to install. No packages, no driver, no SDK: the backend compiles its own C# with the .NET Framework that ships with Windows.

Why this one

  • Cheap enough for a text-only model. tree reads the control tree through UI Automation (native apps, and web pages once the browser's accessibility is woken up), read runs the OCR that ships with Windows, and both print a clickable point per line. A window costs a few hundred tokens instead of a 4 MB screenshot, and no image ever leaves the machine.
  • The human can see it and kill it. A rounded status panel that names the command being run ("click 640,380"), a glow along the screen edge and a ring that follows the real pointer; ESC stops everything. Injected ESC is ignored, so an agent cannot switch off its own indicator.
  • One pointer, always. The ring is drawn around your own cursor, so there is never a second arrow chasing the first one. --hide-cursor is the opt-in for the other model: the drawn marker becomes the only pointer on screen.
  • Input is refused while the indicator is down. move, click, drag, wheel, type, key, keys, clipboard and run answer with a refusal until dsh-cu overlay has put the frame on screen. That gate lives in the tool, not in the prompt.
  • It runs on a stock Windows box. Other computer-use plugins drive an external native driver (cua-driver and friends) that has to be installed first; this one needs only the PowerShell and .NET that Windows already has.
  • It is verified against the real machine. dsh-cu self-test exercises every command and reports pass/fail; environment-dependent checks (Notepad window, focus) warn instead of lying.
  • Works from inside the Harness sandbox. The Harness file sandbox blocks piped stdio for nested processes; the CLI falls back to file-descriptor capture, so dsh-cu still works where a naive spawn(..., {stdio: 'pipe'}) gets EPERM.

What it is not: it has no browser DOM (UI Automation sees what an app exposes, which is less than the DOM and more than pixels), no multi-monitor support, and no per-application allowlist.

Install

WayCommandWhat you get
Skill + CLI + PATHpowershell -ExecutionPolicy Bypass -File install.ps1SKILL.md in $DSH_HOME\skills\dsh-computer-use\ and a dsh-cu shim on the user PATH
DSH profile plugindsh plugin --profile web add github:datuloar/dsh-computer-use-winthe bundle mounts the plugin, which registers the skill from the repository
Nothing at allbin\dsh-cu.cmd or node src\cli.js from a clonethe same CLI, no PATH change

Use one skill delivery path: the copied $DSH_HOME\skills\dsh-computer-use\SKILL.md and the plugin registration carry the same skill name.

Without any of them: node src\cli.js or bin\dsh-cu.cmd .

The filesystem skill root is scanned live, so no restart is needed for the copy; the profile plugin needs dsh web restarted once.

The loop

dsh-cu overlay                       # the human is warned for 8 s, ESC cancels
dsh-cu shot                          # capture the screen, print the path
dsh-cu click 640 380                 # click there
dsh-cu type "hello"                  # type into the focused window
dsh-cu keys ctrl+l                   # key combinations
dsh-cu drag 420 300 980 300          # press, glide, release
dsh-cu shot --region 1200 600 420 240   # read one detail instead of the whole screen
dsh-cu overlay --stop                # stop and clean up
  1. dsh-cu overlay before the first input.
  2. dsh-cu focus (or --title ) the target window.
  3. dsh-cu shot and look at the image.
  4. Act on coordinates taken from that screenshot; confirm a move with dsh-cu pos.
  5. dsh-cu shot again — a step is done only when the second screenshot says so.
  6. dsh-cu overlay --stop when finished.

Commands

dsh-cu shot [file.png] [--region    ]   capture the screen, or a crop of it
dsh-cu read [--region    ] [--lang ] [--json]   the screen as text, with a point per line
dsh-cu tree [--pid N | --title ] [--depth N] [--all] [--json]   the control tree of a window
dsh-cu find  [--pid N | --title ] [--json]   controls whose name contains the text
dsh-cu tap  [--pid N | --title ]   click the one control named like that
dsh-cu move                         move the pointer
dsh-cu click   [left|right|middle|double|triple]   click
dsh-cu drag             press, glide to the second point, release
dsh-cu wheel  [ ] [--horizontal]   scroll, 120 = one notch, at that point when x y are given
dsh-cu type                         type into the focused window
dsh-cu key  [down|up]               Enter, Esc, Tab, Space, Ctrl, Alt, F1..F12, arrows
dsh-cu keys                        ctrl+l, ctrl+shift+t, ...
dsh-cu clipboard [get|set |clear]   read or write the clipboard
dsh-cu pos                                pointer position and foreground window
dsh-cu windows [--json]                   visible windows with pid, title and rectangle
dsh-cu focus 

| --title bring a window to the front, verified dsh-cu wait-window [seconds] wait for a window to appear, then focus it dsh-cu run run a command line — see the safety rules dsh-cu broker [start|stop|status] elevated input broker, asks for UAC dsh-cu overlay [--announce N] [--quiet] [--hide-cursor] on-screen indicator, ESC stops it dsh-cu overlay-state whether the indicator is running dsh-cu display primary screen size and DPI dsh-cu mcp serve the same tools over MCP (stdio) dsh-cu doctor what this machine can do dsh-cu self-test exercise every command dsh-cu overlay --stop stop the indicator, restore the cursor dsh-cu overlay --restore-cursor emergency cursor rescue

read, tree, find and tap are the cheap path; shot writes into %TEMP%\dsh-computer-use\ and prints SAVED ; --region x y w h crops it to that rectangle in screen coordinates. --quiet draws only the pointer marker, for pixel-accurate reads. Arguments are handed to the backend as a JSON array, so typed text and window titles keep their spaces, quotes and leading dashes.

drag presses at the first point, glides to the second in eased steps and releases it, which is what sliders, selections and drag-and-drop targets expect; click also takes middle and triple.

As MCP tools (DeepSeek Harness, Claude Code, opencode, Cursor)

dsh-cu mcp serves every command as native tools over MCP stdio: ui_tree, find, tap, read_text, screenshot (returns the image), click, move, drag, scroll, type, press, windows, focus, wait and indicator. It keeps one backend process warm, so a call takes 10–200 ms instead of the ~550 ms a fresh PowerShell costs, and input tools start the on-screen indicator by themselves the first time.

DeepSeek Harness — add a row to your profile patch (\profiles\web\cordis.patch.yml); the tools then appear as mcp__pc__click, mcp__pc__ui_tree, ...:

- id: mcp-pc
  name: '@deepseek-ai/dsh-mcp-client'
  config:
    serverName: pc
    transport: stdio
    command: node
    args: ['C:\\path\\to\\dsh-computer-use\\src\\mcp.js']
    toolCallTimeoutMs: 120000

Claude Code: claude mcp add pc -- node C:\path\to\dsh-computer-use\src\mcp.js. opencode and Cursor take the same command in their MCP settings.

If the human presses ESC, the indicator records it and every input tool refuses for the next ten minutes with a message telling the agent to stop and ask — an agent cannot quietly restart the indicator it was just stopped with.

Driving a GUI without vision

dsh-cu overlay                       # the human is warned, ESC cancels
dsh-cu focus --title "Settings"      # pick the window
dsh-cu tree                          # roles, names and a clickable point each
dsh-cu tap "Bluetooth"               # click the one control with that name
dsh-cu tree                          # verify: the tree is the new state

tree walks the window with UI Automation and prints one line per control:

TREE Invoice draft (6 elements, depth 6)
  Text "Invoice draft" @1066,643
  Edit "Customer name" @1250,754
  Button "Save" @1075,1013
  Button "Cancel" @1245,1013

find "save" filters that list, and tap "Save" clicks the single match — or refuses and prints the candidates when the name is ambiguous, so a wrong click is never silent. Before clicking, tap checks that the point really belongs to the target window and refuses when something covers it.

Chrome, Edge and other Chromium browsers keep their accessibility tree switched off until a client asks for it; the first tree or find wakes the renderers and waits for the page, later calls are immediate. A browser window then lists its links, buttons and fields the same way, which is enough to drive most web UIs without a DOM. Firefox and Electron apps expose whatever they expose — check with tree before planning a run.

When a window shows no tree (canvas apps, games, remote desktops, a PDF viewer), read falls back to the OCR that ships with Windows:

dsh-cu read --region 1200 600 420 240 --lang en-US
READ 3 lines (en-US, 420x240 at 1200,600)
[1310,648] Customer name
[1288,712] Design system audit
[1276,764] Save

Each line carries the point to click, small regions are upscaled before recognition, and --lang picks an installed language pack (dsh-cu read --lang ru). doctor lists what is installed.

A screenshot is still the right tool when the question is about pixels — colours, layout, an image, a chart. Then shot --region keeps it small, and a vision-capable route (for DeepSeek Harness: the Vision Toolkit plugin) can answer questions about it.

Scrolling follows the pointer, so pass it the page: dsh-cu wheel -600 1280 700 is five notches down at that point, --horizontal scrolls sideways, and key PageDown / keys ctrl+End scroll the focused page. Every scroll moves the coordinates you measured — take a fresh shot before clicking.

What the human sees

The whole screen while an agent works: the status panel, the edge glow and the pointer ring

Countdown before the first inputA click, shown where it landed
Amber panel counting down, ESC cancelsRipple rings on the button the agent clicked
  • The panel names the command that is running; typed text is shown as a character count only.
  • The edge glow breathes while the agent is in control and is solid during the countdown.
  • The ring follows your own pointer and pulses on every click. Screenshots never contain the system cursor, which is why only the ring shows up in the images above.

All images on this page were taken over a neutral demo window, not a real desktop.

Safety model

  • Input is refused while the indicator is down — and again the moment the indicator process dies, so a killed indicator cannot leave the human blind for the seconds the heartbeat stays fresh. Override only deliberately: DSH_CU_ALLOW_NO_INDICATOR=1.
  • The indicator comes up before input is allowed, and the CLI waits out the full announcement (default 8 s) after the frame is confirmed on screen.
  • ESC stops it — a physical press only; injected ESC is ignored.
  • The panel says what is happening. Every command writes a one-line label the indicator shows next to the status text, so the human can follow along; type is reported as a character count rather than the text, so a typed password never appears on screen.
  • overlay --stop kills the indicator by pid file, reloads the system cursors and removes its state files. The indicator also stops itself after five minutes without a command.
  • One pointer, never two. By default nothing replaces your cursor: the indicator draws a ring around it, so the pointer you see is the real one and the ring only says "the agent is here". dsh-cu overlay --hide-cursor is the opt-in for screen sharing and recordings: it blanks the system cursor and draws the marker arrow instead, so there is still exactly one pointer. The tool records that it blanked the cursor, overlay --stop gives it back, and if the indicator is killed the next dsh-cu command restores it and says so. dsh-cu doctor reports cursor: visible / replaced by the marker / hidden, but not by this tool.
  • run and broker start are the escape hatches. broker start raises a UAC prompt and lets run execute a command line elevated, for windows this process cannot reach (UIPI). Both are gated behind the indicator and are meant for explicit human requests — the Harness already has a shell, so the only thing they add is elevation.
  • clipboard get is gated too: it can expose whatever the human last copied.
  • The skill instructs the agent never to press anything destructive without an explicit request.

Environment variables

VariableWhat it does
DSH_CU_ALLOW_NO_INDICATOR=1let input through while the indicator is down — deliberate override, nothing else bypasses the gate
DSH_CU_UI_SCALE=1.5scale the indicator on top of the display DPI, for a 4K panel or weak eyes (0.75–4)
DSH_CU_POWERSHELL=powershellthe default; set it to another PowerShell only if yours lives elsewhere
DSH_CU_POWERSHELL=pwshrun the backend with another PowerShell executable

Running from inside a sandboxed agent session

The Harness file sandbox denies nested processes a pipe, so node-spawns-powershell with stdio: 'pipe' fails with EPERM before PowerShell even starts. The CLI detects that exact failure and re-runs the backend with two file descriptors instead of pipes, so every command keeps working. If you would rather skip Node entirely, call the backend directly — it takes the same commands:

powershell -NoProfile -ExecutionPolicy Bypass -File backends\windows.ps1 shot C:\temp\shot.png

Limits, stated plainly

  • Windows only. doctor and self-test still run elsewhere; every other command refuses.
  • Elevated windows cannot be driven from a non-elevated process (UIPI) — use broker start.
  • The UAC consent dialog and the lock screen are unreachable: they run on the secure desktop, whi