piecuzwhynot/dsh-external-workers ↗★ 1

dsh-external-workers

Persistent external agent lanes for DeepSeek Harness — delegate long-running tasks to Claude Code, Codex CLI and Antigravity CLI through their own official CLIs and your existing subscription logins, with one durable session per lane and jobs that survive a restart or a context compaction. 适合需要将长耗时任务委托给Claude Code等外部CLI,且需跨重启保持会话的用户。

패키지
dsh-external-workers
호환성
미검증
Harness peer 범위
*
버전
0.3.0
라이선스
MIT
최근 업데이트
2026. 10. 3.

설치

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:piecuzwhynot/dsh-external-workers

Configuration

Everything lives in config.json next to this package; most of it is read on every use, so edits apply without a restart.

Which sessions get the tools (scope)

"scope": {
  "mode": "per-session",
  "excludePresets": ["my-persona-preset"],
  "excludeWorkspaces": ["C:\\Users\\me\\companion-workspace"],
  "excludeSessionIds": ["session-…"]
}
  • per-session (default) registers the tools into each chosen agent's own scope. An excluded session has no registration at all — not a hidden tool, an absent one.
  • off gives them to nobody; global gives them to every session (rarely what you want).
  • If the agent-scoped registry is unavailable the plugin registers nothing and logs an error — it never falls back to a global registration.
  • A session that switches onto an excluded preset mid-life loses the tools again.

Lanes

"lanes": [  { "id": "main", "worker": "claude", "role": "example lane: planning and code review", "model": "opus", "effort": "high", "speed": null },
  { "id": "quick", "worker": "gpt", "role": "example lane: fast edits and scripted chores", "model": "gpt-6-luna", "effort": "medium", "speed": "priority" },
  { "id": "deep", "worker": "gpt", "role": "example lane: the heavy model work", "model": "gpt-6.1-sol", "effort": "xhigh", "speed": "priority" },
  { "id": "research", "worker": "google", "role": "example lane: research and data gathering", "model": "gemini-3.1-pro-high", "effort": "high", "speed": null }
]

A lane is the unit of work: one product + one model (and effort/speed) + one role, with its own persistent session and its own working directory ~/.dsh/workers/ (a lane that inherited a pre-lanes session keeps that session's original directory instead — see the upgrade note below). The same product runs several lanes, and that is the point — Codex and Claude each bind a session to one model and degrade it when that thread is resumed on another, so gpt + gpt-6-luna and gpt + gpt-6.1-sol must not share a session.

  • id must match /^[a-z0-9][a-z0-9_-]{0,31}$/, and it names the lane's own workspace (~/.dsh/workers/) — for a lane that inherited a session, the inherited directory wins.
  • Delete the whole lanes key and the plugin falls back to one lane per product, named after the product — the behaviour before lanes existed.
  • Manage them at runtime with worker_config: action: "set" edits a lane, "add" creates one on a product, "remove" drops one, "show" is the roster.

Upgrading from a pre-lanes version. Records used to be keyed by product name. As soon as the config declares named lanes, the first lane of each product adopts that product's own record, so the session you already have is never thrown away; every further lane of that product starts with a fresh session on purpose. Two details an upgrader should know. The adopted record keeps its session id, its state and its working directory — the directory travels with the session on purpose, because these products key their stored conversations by working directory (Claude Code files transcripts under ~/.claude/projects//), so a session resumed from a different folder may not be found at all. A lane renamed from claude therefore keeps running in ~/.dsh/workers/claude/; only genuinely new lanes get ~/.dsh/workers//. And the adopted record's old role/model/effort/speed are not kept: those four used to be assigned product-wide, so they describe no particular lane — the lane's config, or a later worker_config call, owns them now. Jobs recorded before lanes existed carry no lane id, and readers fall back to the product name; nothing in the ledger is rewritten or lost.

Per worker

"workers": {
  "claude": { "enabled": true, "cliPath": null, "permissionMode": "acceptEdits", "extraArgs": [] },
  "gpt":    { "enabled": true, "cliPath": null, "sandboxMode": "workspace-write", "approvalPolicy": "never", "extraArgs": [] },
  "google": { "enabled": true, "cliPath": null, "extraArgs": [] }
}

This block is CLI behaviour only: enabled, cliPath, and each product's own gate (permissionMode for claude, sandboxMode / approvalPolicy for gpt, extraArgs for all three). role, model, effort and speed are not read from here — they belong to a lane. cliPath pins a binary (default: newest install, then PATH).

Granting a worker more power is a one-line change — see docs/USAGE.md.

Custom products

The three products above are the ones with a hand-written adapter. Any other agent CLI can be added by description, without touching the code — Kimi, Qwen, Grok, Doubao, a local model runner, whatever ships next month. A product is a recipe in config.json:

"products": {
  "kimi": {
    "label": "Kimi CLI",
    "bin": { "names": ["kimi.exe", "kimi"], "roots": ["~/.kimi/bin"], "hint": "install kimi" },
    "args": {
      "base":   ["-p", "{prompt}"],
      "resume": ["--resume", "{session}"],
      "model":  ["--model", "{model}"],
      "extra":  ["--output-format", "json"]
    },
    "output": { "format": "json", "session": "session_id", "text": "result", "model": "model" },
    "models": { "command": ["models"], "format": "lines" }
  }
}

That is all it takes: the product immediately gets a delegate_kimi tool, its own lanes, its own sessions, durable jobs, artifact collection and the panel card. Concretely:

  • bin — how to find the CLI: names are searched on PATH, roots are scanned for the newest match, and workers..cliPath pins one outright.
  • args — templates with placeholders {prompt} {session} {model} {effort} {speed} {cwd} {timeout} {image} {index}. Sections are emitted in the order base, cwd, resume, model, effort, speed, images, prompt, extra, and a section whose placeholder has no value is dropped whole — so a run with no model never leaves a dangling --model behind. A CLI that takes the prompt right after its flag puts both in base: ["-p", "{prompt}"]; one that takes it last uses args.prompt.
  • output — json (one envelope), jsonl (last event wins per field) or text (stdout as the result, plus an optional sessionPattern with a capture group). This is how the bridge learns the session id, which is what makes a lane's thread continue across runs.
  • models — the CLI's own model-list command, so worker_config(action: "models") reports reality instead of guesswork.

A recipe is data, so it is validated, not trusted: a broken one is reported at plugin start and by worker_config, naming the field and the rule. Five lines in a config file is the whole cost of supporting a new CLI — and if you would rather not write it, ask the DSH agent to: it can read the CLI's --help and fill the recipe in for you.

Not only CLI products: API-only models too

The example above drives a CLI. A product that offers only an HTTP API has no command to run, so the plugin ships one: examples/api-lane.mjs turns an OpenAI-compatible endpoint into a lane. That shape covers far more than OpenAI — xAI (Grok), Moonshot (Kimi), DashScope's compatible mode (Qwen), Volcengine Ark (Doubao) and most local servers all speak it.

"products": {
  "grok": {
    "label": "Grok (xAI API)",
    "bin": { "names": ["node.exe", "node"], "roots": [] },
    "args": {
      "base":   ["
/examples/api-lane.mjs", "--history-dir", "{cwd}/.api-history"],
      "prompt": ["-p", "{prompt}"],
      "resume": ["--resume", "{session}"],
      "model":  ["--model", "{model}"]
    },
    "output": { "format": "json", "session": "session_id", "text": "result", "model": "model" }
  }
}

The key never goes in config.json: the script reads API_LANE_API_KEY from the environment (or --key-file ), with API_LANE_BASE_URL and API_LANE_MODEL alongside it.

Two things this buys you, and one it does not:

  • It iterates. The script keeps the conversation on disk per session, so --resume sends the whole thread back — a follow-up builds on everything the worker already saw, exactly like a CLI lane. A plain "call the API once" wrapper would send one message every time; there is a test that fails if that regresses.
  • The packet is inlined. The bridge normally tells a worker to read jobs//packet.md. A raw API model has no file tools, so the script reads that packet and sends it with the prompt, and says so.
  • But an API model has no tools at all. It cannot read your repo, run your tests, or write files into jobs//out/. It answers in text, and the answer is the deliverable. If you need a worker that actually does things in your project, use a product with a real CLI — the API lane is for asking, reviewing, drafting and thinking, not for touching files.

Why a lane beats a plain subagent

DSH's built-in subagent seam is one-shot by design — a fresh process, a fresh thread, one turn. The upstream provider README says it plainly: "no continuation, resume, pooling, progress stream, or product-session persistence." Every call starts from zero, so you re-brief the model every time, and whatever it learned in the previous call is gone.

A lane is the opposite: a worker that iterates.

Plain subagentA lane
What it remembersnothing; each call starts from zeroits own session, kept across calls
Follow-upsre-explain the whole task"continue, but do X instead"
Models it can holdone model per callone model per lane, several lanes per product
If it failsthe turn is gonethe job is on disk: read it, resume it, retry it
If your chat is compactedthe subagent call is gonethe worker session and its jobs are untouched
Working directorythe caller'sits own, per lane
Who runs the modelyour DSH process hosts the requestthe product's own agent, on your own subscription

The practical difference: a subagent is a question you ask. A lane is a colleague you hired — brief them once, then keep handing them the next piece of the same job, and they still remember the first one.

Quota and credits (Codex)

delegate_gpt reads Codex's own usage records before every run, so it can stop before starting work that would hit the wall:

  • At/above quota.warnAtPercent (default 90) on either the 5-hour or the weekly window, the tool does not start. It returns the real numbers and waits for you; continuing needs an explicit allow_quota: true.
  • If the plan allowance is already spent, the next run would be billed to credits — that needs a separate allow_credits: true, so approving a percentage never silently approves money.
  • A run that does spend credits is reported afterwards with the amount, on the job record and in the completion notice.
"quota": { "enabled": true, "warnAtPercent": 90 }

The numbers come from Codex's rollout files ($CODEX_HOME/sessions/**), which every Codex surface writes — this plugin's lanes, the Codex desktop app, and the CLI — so the reading stays current even when the usage happened somewhere else. Nothing is queried over the network and no credential is read.

Only Codex reports this. Claude Code computes its own five-hour/weekly usage but renders it only in its TUI (nothing is cached where a plugin can read it), and the Antigravity CLI logs no numbers at all. For those two the bridge stays silent rather than showing an invented number — and a hard limit still surfaces reactively: the job comes back as retryable with failureKind: quota.

Another practical limit: credits can only be reported after a run, not mid-run. The worker is a single CLI process; nothing can be intercepted while it spends. What the bridge guarantees is that the spend is never silent and that the next task cannot start on credits without your yes.