xarleyn/dsh-plugins--plugins-dsh-model-safety-gate ↗★ 1

@yadsh/dsh-model-safety-gate

Independent two-layer safety gate around the DeepSeek Harness agent loop: deterministic and model-classifier verdicts for prompts, streamed output, tools, and tool results 适合安全策略严格的环境,需配置 classifier 与模式以控制拦截级别

パッケージ
@yadsh/dsh-model-safety-gate
互換性
未検証
Harness ピア範囲
catalog:dsh
Cordis ピア範囲
catalog:dsh
バージョン
0.1.0
ライセンス
MIT
最終更新
2026/09/13

インストール

$npx -p @deepseek-ai/dsh dsh plugin --profile web add github:xarleyn/dsh-plugins#fd3cdd95651ec4eafe52d510751ba3aba3c9248d&path:plugins/dsh-model-safety-gate

ドキュメント

README 全文を読む ↗

Configuration

All options are optional; defaults are shown.

enabled: true            # master switch for the whole gate
mode: warn               # off | audit | warn | enforce — default decision profile

classifier:
  backend: none          # none | dsh | openai-compatible
  provider: ""           # backend: dsh — DSH provider id for the classifier model
  model: ""              # backend: dsh — model id
  baseURL: ""            # backend: openai-compatible — endpoint base URL
  apiKey: ""             # backend: openai-compatible — API key (kept out of logs)
  timeoutMs: 3000        # classifier request timeout
  maxTokens: 128         # bounded structured response
  temperature: 0
  failureMode: rules-only # closed | open | rules-only | ask — on timeout/error/malformed
  requireLocal: false     # true forbids remote (openai-compatible) endpoints

input:
  enabled: true          # gate user prompts on agent/pre-step
  safetyAction: block    # allow | warn | block for safety verdicts
  qualityAction: warn    # allow | warn | block for quality-only verdicts (block = opt-in)

output:
  enabled: true          # gate main-model streaming output
  mode: buffered         # observe | interrupt | buffered
  text: true             # check the visible-answer channel
  reasoning: true        # check the reasoning channel when the provider streams it
  checkEveryChars: 512   # new quarantined chars between classifier snapshots
  windowChars: 1536      # snapshot window size sent to the classifier
  lookbehindChars: 768   # preceding context included with each window
  minCheckIntervalMs: 250
  maxBufferedChars: 8192 # overflow fails closed in buffered mode

tools:
  enabled: true          # gate tool calls on tools/pre-execute
  semanticClassifier: true

toolResults:
  enabled: true          # scan tool results on tools/post-execute
  classifyUntrustedSources: true

audit:
  enabled: true
  includeRawContent: false # opt-in raw content logging (default: hashes only)

ui:
  enabled: true
  showWarnings: true

allowSessionOverride: true # false forbids per-session downgrade of the global mode

Deployment presets

ProfileInputOutputToolsFailure mode
Personalwarninterruptaskrules-only
Balancedblockbufferedaskrules-only
Strictblockbuffered (text + reasoning)blockclosed, session override disabled

Privacy

If the classifier backend is openai-compatible, prompts, streamed output, and reasoning content are sent to that endpoint. The configuration surface reports this; set classifier.requireLocal: true to forbid remote endpoints entirely.