huyang218/dsh-plugins--packages-tool-retry0

dsh-plugin-tool-retry

Retries transient tool-call failures in DeepSeek Harness — socket resets, rate limits, timeouts — for the tools an operator declares safe to repeat, with backoff and a deadline on every retry.

包名
dsh-plugin-tool-retry
版本
0.1.0
许可证
MIT
最近更新
2026年8月18日

安装

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:huyang218/dsh-plugins#efa02c42cd65f80b544659dfd6dbd00b6f293d41&path:packages/tool-retry

dsh-plugin-tool-retry

English | 中文

Retries a tool call that failed for a reason that might clear — a reset socket, a rate limit, a timeout — for the tools you declare safe to repeat.

dsh already retries the model request (dsh-llm-retry) and enforces tool deadlines (dsh-tool-call-timeout-policy). Nothing retried the tool call itself, so a data source closing one socket mid-scan ended the task, and an agent told only that the data was unavailable was left to improvise.

dsh plugin --profile web add dsh-plugin-tool-retry

It retries nothing until you say what is safe

Repeating a tool that wrote, ordered or sent something does it twice. Nothing in the tool contract says which tools are idempotent — isConcurrencySafe is about overlap, not repetition — so the operator names them, and the default list is empty:

- id: tool-retry
  config:
    retryTools: ['astock_*', 'ainfo_*', 'web_fetch']   # read-only tools only

Patterns are literal names or a trailing * prefix, deliberately not regular expressions: this list decides what may execute twice, and a stray . matching everything is not a mistake worth enabling.

What counts as worth retrying

Failures where the request never got an answer, or the answer said "later": socket resets and fetch failed, 429/5xx, rate-limit messages in either language, timeouts. A wrong argument, a permission denial or a missing file will fail identically forever; retrying those only makes the user wait.

Backoff is exponential and capped, because the failures worth retrying are usually a remote system under load, and a tight loop is how a client turns someone else's brief overload into its own outage.

Two things it will not do

It never turns a failure into a success. An exhausted retry returns the real failure with one line appended — how many attempts were made — so the model reads a persistent outage as persistent rather than as one unlucky call, and reports it instead of trying the same thing again.

It never leaves a retry unstoppable. cordis consumes its waterfall listener list, so a repeated next() reaches the tool body directly and skips every wrapper already shifted off — including the official timeout policy. Each retry therefore runs under a deadline of this plugin's own, fused with the caller's cancellation rather than replacing it. Cancelling during the backoff returns the failure already in hand instead of starting an attempt nobody is waiting for.

Config

KeyDefaultMeaning
retryTools[]Tools safe to repeat; empty means nothing is retried
maxAttempts3Attempts per call, including the first
backoffMs500Wait before the first retry, doubling after
maxBackoffMs8000Cap on the wait
retryDeadlineMs120000Deadline imposed on each retry attempt

Pairs naturally with tool-health, which remembers what has been failing, and tool-usage, which measures what it costs.

License

MIT