Hubert-hwk/dsh-loop-doctor--packages-loop-doctor-apply ↗★ 0

@hubert-hwk/dsh-loop-doctor-apply

循环修复的写入层:人工复核台账、自动应用策略门与已采纳修复沉淀。 适合需持久记录并选择性自动应用修复的用户,默认关闭自动应用。

套件
@hubert-hwk/dsh-loop-doctor-apply
相容性
待驗證
Harness 依賴範圍
^0.1.0-rc.8
Cordis 依賴範圍
^4.0.1
版本
0.1.0-rc.5
授權
MIT
最近更新
2026年8月20日

安裝

此插件尚未提供可驗證的 bundle,或相容性檢查未通過。請先閱讀倉庫說明。 閱讀完整 README ↗

@deepseek-ai/dsh-loop-doctor-apply

中文 | English

The write plane of the loop-doctor family — the P2 piece: a durable human-review ledger, a LOOP_DOCTOR_AUTO_APPLY-style policy gate (off by default), and fix-sediment growth from accepted fixes.

It sits on the advisor's loop-doctor/suggestion feed and owns the only two durable decisions the family makes. "Applying" a fix means promoting it into the accepted-fix sediment (accepted-fixes.jsonl); the advisor merges that sediment back into its patterns at startup, so the next time the same detour class appears, the exact validated fix is injected directly. The only writes this package ever performs are to its own ledger and sediment files.

The human-review ledger

Every suggestion lands as pending in a JSONL review ledger (loop-doctor-review.jsonl in the working directory, or ledgerFile):

node examples/review-demo.mjs   # the interactive review walkthrough

Statuses: pending → applied | rejected. A human decides through the exported API (approveFix / rejectFix / createReviewer), a review tool, or by reading the ledger file directly. Ledger rewrites are serialized and atomic (tmp + rename). Every approval — auto or human — is observable: the plugin emits loop-doctor/fix-applied, and createReviewer's onApplied hook fires for human approvals with the same payload. createReviewer also accepts a reload hook (e.g. () => ctx.loopDoctorPatterns.reload()), so a human-approved fix takes effect in the current process exactly like auto-apply.

The auto-apply gate

- id: loop-doctor-apply
  name: '@deepseek-ai/dsh-loop-doctor-apply'
  config:
    ledgerFile: ''          # default ./loop-doctor-review.jsonl
    sedimentFile: ''        # default ./loop-doctor-accepted-fixes.jsonl
    autoApply: false        # LOOP_DOCTOR_AUTO_APPLY — default OFF (human-confirm)
    autoApplyMinConfidence: 0.9

With autoApply: true, a suggestion skips the human queue only when every gate of decideAutoApply passes:

  1. the pattern is verifiedBy: 'replay' (expert-judged fixes never auto-apply);
  2. confidence clears autoApplyMinConfidence (default 0.9 — the cap);
  3. the replay outcome is present, verified, and not simulated (policy-cap counterfactuals never auto-apply — only behavior-modeling findings measured on the exact log may).

Everything else stays pending for a human. Auto-applied fixes enter the sediment with autoApplied: true and emit loop-doctor/fix-applied.

The sediment

accepted-fixes.jsonl is the harness's self-optimization memory. Each line is an AcceptedFix: the accepted pattern snapshot (per-tool context preserved in sub) plus provenance (sourceSessionId, autoApplied, acceptedAt). The advisor merges it into its pattern set at startup and live: this plugin declares the loopDoctorPatterns service (so mounting it without the advisor fails loud) and calls reload() after every sediment write, so an accepted fix takes effect in the current process — no restart. Appends are serialized and the plugin claims each fix id synchronously, so concurrent auto-applies of the same fix never duplicate a line.

P3: apply-then-measure

After an accepted fix, this plugin counts occurrences of the same detour per session (the suggestion feed is the observation stream). The first session observed after acceptance is the baseline; once measureWindowSessions (default 5) post sessions have been finalized, it emits loop-doctor/fix-effectiveness comparing the post average to the baseline:

config:
  measureWindowSessions: 5   # post-acceptance sessions per sample; <1 disables
{ "fixId": "retry-storm-repeat:repeat:grep",
  "signalKind": "retry-storm", "sub": "repeat:grep",
  "baselineOccurrences": 1, "postOccurrences": [0, 0, 0, 0],
  "trend": "improved",
  "detail": "… directional evidence over a small window, not a statistical claim" }

Sessions are finalized at disposal; sequential sessions are also finalized as the next one starts, so a run of clean sessions completes a window while the plugin is alive. The sample is explicitly labeled directional — it is the harness's own evidence that self-optimization is working, not a claim of statistical significance.

See the family README for the safety baseline.

Model Experience

An accepted fix changing the next report

What the model sees

Nothing from the plugin directly. After a fix is accepted (auto or human) and the patterns reload, the next loop_doctor_diagnose report in the same process carries the accepted variant's composite fixPatternId — the only model-visible difference.

Suggestion line before vs after acceptance
- suggestion [retry-storm-repeat] (prompt, 80%): …   ← before
- suggestion [retry-storm-repeat:repeat:grep] (prompt, 80%): …   ← after reload
Token effect

Zero. Ledger, sediment, and effectiveness samples are operator-facing files.

KV Cache effect

None. No model-visible bytes change.

Known Limitations and Deferred Work

  • One review process per ledger. The ledger is not safe for concurrent writers; a single review process (or tool) owns each file.
  • Auto-apply is deliberately narrow. Only replay-verified, non-simulated, ceiling-confidence fixes; widening it is a policy decision, not a TODO.
  • Effectiveness is directional. measureWindowSessions samples are small-window evidence, explicitly labeled "not a statistical claim".
  • No rejection memory. A rejected suggestion can re-enter the ledger in a later session; negative sediment (suppression) is deferred.