zhiheng-zhang-Mera/dsh-restart ↗★ 0
dsh-restart
Safe restart execution for DeepSeek Harness: request validation, checkpoint gating, restart locking, graceful shutdown, crash-loop breaking and an external supervisor. It never decides when to restart. 用于执行安全的进程重启。适合需要防崩溃循环、确保重启过程安全稳定的生产环境。
同名パッケージの別リポジトリ
インストール
$
npx -p @deepseek-ai/dsh dsh plugin --profile web add github:zhiheng-zhang-Mera/dsh-restartドキュメント
README 全文を読む ↗Configuration
Every key is optional; the values below are the shipped defaults, and each is documented
with the risk of changing it in cordis.patch.yml. Types and effects
in full: docs/operations.md.
| Key | Type | Default | Effect |
|---|---|---|---|
enabled | boolean | true | master switch; false refuses every request with DISABLED |
applicationRestart.enabled | boolean | true | whether application restarts may happen |
applicationRestart.minIntervalMs | number | 1200000 | enforced cooldown between application restarts (20 min) |
systemRestart.enabled | boolean | false | whether the system mode may be used |
systemRestart.minIntervalMs | number | 3600000 | enforced cooldown between system restarts (60 min) |
allowedSources | string[] | dsh-health-scheduler, dsh-cli, operator | who may submit; an empty list is a ConfigError |
allowedPriorities | string[] | low, normal, high, emergency | accepted priority vocabulary |
allowSystemReboot | boolean | false | second gate for mode: system; a request must also acknowledge |
safety.checkpointRequired | boolean | true | whether a requested checkpoint must succeed |
safety.duplicateSuppression | boolean | true | whether a replayed requestId returns the previous answer |
safety.crashLoopLimit | number | 3 | unclean starts inside the window that trip the breaker |
safety.crashLoopWindowMs | number | 600000 | rolling breaker window (10 min) |
safety.safeModeOnLoop | boolean | true | whether tripping the breaker also enters safe mode |
safety.shutdownTimeoutMs | number | 90000 | shutdown budget, also used as the checkpoint budget |
safety.allowForceTerminate | boolean | false | declared, validated, never consulted by this release |
safety.allowRestartWithoutSupervisor | boolean | false | whether to exit with nobody to relaunch |
supervisor.heartbeatIntervalMs | number | 5000 | supervisor heartbeat period |
supervisor.heartbeatTimeoutMs | number | 30000 | age after which the supervisor counts as absent |
supervisor.relaunchTimeoutMs | number | 90000 | time a relaunched pid has to be alive |
supervisor.launchCommand | string[] | null | null | relaunch command; null means "reuse the supervisor's argv" |
supervisor.launchArgs | string[] | [] | extra arguments appended to the relaunch |
supervisor.launchCwd | string | null | null | working directory for the relaunch |
supervisor.pollIntervalMs | number | 1000 | pid poll interval |
supervisor.ticketTtlMs | number | 600000 | how long a pending ticket stays valid |
supervisor.detach | boolean | true | whether the supervisor runs detached — declared and validated, read only by spawnSupervisor(), which this package never calls: see the divergence note below. Only the uncalled spawnSupervisor() consults it |
storage.directory | string | null | null | audit-log directory; null = the state directory |
storage.maxLogBytes | number | 4194304 | audit log rotation threshold |
storage.maxRecentAttempts | number | 25 | attempts kept in memory for status |
knownReasonCodes | string[] | nine codes | codes accepted without complaint; unknown ones are logged, not refused |
Invalid documents are refused at load with a dotted path, and the plugin continues on the defaults rather than failing the host's boot. For example:
dsh-restart config: supervisor.heartbeatTimeoutMs must exceed supervisor.heartbeatIntervalMs (60000), received 30000
dsh-restart config: allowedSources must list at least one source; an empty list would refuse every request, including an operator request