1azybug/dsh-real-time-computer-use ↗★ 1
dsh-real-time-computer-use
Real-time Computer Use for DeepSeek Harness on Windows: full-screen capture, timestamp replay, and mouse/keyboard injection 适合需要在Windows环境下进行自动化屏幕交互和键鼠模拟的任务。
インストール
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:1azybug/dsh-real-time-computer-useドキュメント
README 全文を読む ↗Configuration
Every knob lives in cordis.patch.yml, and the shipped defaults are not a neutral baseline: they are the
configuration one GUI-evaluation setup runs with. Out of the box you get a 16-tool surface; switching the four
optional groups on widens it to 21 tools and re-enables ordinary image reading.
Read those four defaults as answers to four questions, not as arbitrary flags:
| Switch | Default | With the default | Turning it on adds |
|---|---|---|---|
triggerTools | false | nothing watches the screen waiting for a condition | wait_for_change, act_when — poll a region until it changes, then act |
singleImageTools | false | screen reading only ever arrives as the same-frame thumbnail + tiles form | screen_observe, region_observe — hand back a single image |
denyReadImage | true | a tool guard rejects every read_image call — with or without region, for every session, not only screen captures | normal image reading works again |
diffTool | false | no frame differencing | screen_diff — report which grid cells changed |
Why those defaults: on a 2560×1440 desktop a single full-screen image is either downscaled past legibility (a 46 px
control becomes 19 px) or covers one corner, so the plugin standardizes on one full-view form — screen_grid, one
capture turned into a same-frame thumbnail plus 1:1 tiles. Letting read_image back in would restore the
single-image path for any image, not only screen captures. The triggerTools entry is an evaluation constraint
(watching for a visual condition and then acting counts as cheating in that setting), not a technical limit. The
comments in the file record each decision.
The remaining keys are ordinary capture settings:
| Key | Default | Meaning |
|---|---|---|
backend | dxgi | capture backend: dxgi (GPU copy, ~0.42 ms/frame) or gdi (~28.9 ms/frame) |
frameIntervalMs | 16 | capture interval ≈ 62.5 fps — the only place the capture rate is decided (screen_watch has no fps argument). Deliberately far below the 33.333 ms per-frame bound: at 25 ms only 8.3 ms of headroom is left and the once-per-segment encoder pre-build (10–15 ms of concurrent work) crosses it; at 16 ms the headroom is 17.3 ms and the same jitter stays inside. Measured under a full-screen load with a second helper recording: 4/2408 misses at 25 ms vs 1/11217 at 16 ms. Costs ≈79% of one core (vs 52%) and 1.5× the recording size. |
frameCapacity | 1800 | frames retained by the ring buffer; only the jpeg codec uses it — h264 windows are bounded by retain_s and the byte budget |
jpegQuality | 70 | JPEG quality |
codec | h264 | frame storage: h264 (in-memory segments, ~445 MB per 20 min) or jpeg (per-frame files, 6–11 GB) |
captureDir | '' | frame output directory; empty means the helper's own temp directory |
recordDir | '' | recording directory (Windows path). When set, every screen_watch start also encodes the captured frames straight into an mp4 in this directory — one continuous encoder, timestamps taken from the real capture instants, so a fluctuating capture rate does not compress playback. Measured at 2560×1440: ~1.1 MB per 5 s, versus ~631 MB for the same span written as per-frame JPEG. Recording implies h264. Empty means no recording. |
helperPath | '' | helper executable; empty means the bundled helper/CuHelper.exe |
The plugin's own code defaults differ for frameIntervalMs (33), backend (gdi) and codec (jpeg): the patch
file is where a deployment states its choice, and the code keeps the conservative value for deployments that mount the
plugin without it. denyReadImage is true in both.